A training method, an image generation method and a device for a sample image generation model
By using the training method of sample image generation model in industrial production, and using mask marker fusion and image fusion technology, the problem of difficulty in collecting defect images is solved, and the quality of generated images and the performance of defect detection models are improved.
Patent Information
- Application Number
- CN202510310832.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-03-14
AI Technical Summary
In industrial production, due to the difficulty in collecting defect images, the performance of the defect detection model is affected, and defect-free images are difficult to obtain, which affects the stability of the sample image generation model and the generated image quality.
A training method for sample image generation model is adopted, by obtaining template diagrams, target diagrams, first mark diagrams and second mark diagrams, mask mark fusion is performed, target mark diagrams are generated, and the generator is used to image fusion of target mark diagrams and template diagrams to generate high-quality sample images.
The stability of the sample image generation model and the quality of the generated images are improved, and sample images close to the real defect image can be generated, which enhances the performance of the defect detection model.
Smart Images

Figure CN119832114B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of machine vision, and particularly to a method for training a sample image generation model, an image generation method, and an apparatus therefor. Background Art
[0002] Generally, in industrial production, there is a high requirement for the product yield rate. Using a defect detection model to perform industrial defect detection on product images to determine whether a product is qualified is the main way to ensure the product yield rate. Among them, whether the defect images are sufficient directly affects the performance of the defect detection model, and further affects the product yield rate. The so-called defect images are product images obtained by collecting images of products with product defects and including defects (flaws) in the image content.
[0003] However, due to the strictness of the industrial production environment, it is often not suitable to collect a large number of defect images of defective products. And in some cases, due to the high yield rate of the production line, products are not likely to have defects, which also makes it difficult to collect defect images.
[0004] In the related art, in order to obtain a large number of defect images for training a defect detection model, generally, a sample image generation model is trained to obtain sample images for training the defect detection model.
[0005] However, in the above-mentioned related art, defect-free images are required for training the sample image generation model. However, since defect-free images are difficult to collect, and when there are minor flaws in the defect-free images that do not reach the defect standard but do not affect the product being a good product, it will affect the stability of the sample image generation model, and thus affect the quality of the generated sample images. Summary of the Invention
[0006] The purpose of the embodiments of the present application is to provide a method for training a sample image generation model, an image generation method, and an apparatus therefor. The specific technical solutions are as follows:
[0007] In a first aspect, the embodiments of the present application provide a method for training a sample image generation model, including:
[0008] Obtaining a sample image; wherein, the sample image includes a template image, a target image, a first marked image, and a second marked image; the first marked image is an image obtained by performing a predetermined process on the template image, the second marked image is an image obtained by performing the predetermined process on the target image, and the predetermined process is a process of setting a mask mark for a defect.
[0009] Using the sample images, continue to train the sample image generation model to be trained until the sample image generation model converges; wherein, the sample image generation model includes: a generator and a discriminator; each training process of the sample image generation model includes:
[0010] Perform mask label fusion on the first label map and the second label map to obtain a target label map; wherein, the mask label fusion is used to fuse the mask labels existing in the first label map and the mask labels existing in the second label map into the same image, and represent the mask labels existing in the first label map in the same image in a preset manner;
[0011] Input the target label map and the template map into the generator, so that the generator performs image fusion on the target label map and the template map to obtain a generated image; wherein, the training convergence target of the generator includes: when there are mask labels represented in the preset manner in the target label map, using the image content of the template map to repair the target area; the target area is the image area indicated by the mask labels represented in the preset manner;
[0012] Input the target map and the generated image into the discriminator, so that the discriminator makes a true / false judgment on the generated image based on the target map and the generated image.
[0013] In a second aspect, an embodiment of the present application provides an image generation method, including:
[0014] Obtain a pre-collected real product image as an initial image;
[0015] Obtain a mask label map; wherein, the mask label map is generated by adding mask labels of various types of defects to a blank image with the same size as the initial image, and the mask labels of different types of defects are different;
[0016] Input the initial image and the mask label map into the generator of the pre-trained sample image generation model, so that the generator fuses the initial image and the mask label map to obtain a generated image corresponding to the initial image;
[0017] Wherein, the generated image corresponding to the initial image includes defects of the type indicated by each mask label of the mask label map, and the sample image generation model is trained based on the training method of a sample image generation model provided in the first aspect.
[0018] In a third aspect, an embodiment of the present application provides a training device for a sample image generation model, including:
[0019] A sample image acquisition module for acquiring sample images; wherein, the sample images include a template image, a target image, a first marked image, and a second marked image; the first marked image is an image obtained by performing a predetermined process on the template image, the second marked image is an image obtained by performing the predetermined process on the target image, and the predetermined process is a process of setting mask marks for defects.
[0020] A model training module for training a sample image generation model to be trained using the sample images until the sample image generation model converges; wherein, the sample image generation model includes: a generator and a discriminator.
[0021] Each training process of the sample image generation model includes:
[0022] Performing mask mark fusion on the first marked image and the second marked image to obtain a target marked image; wherein, the mask mark fusion is used to fuse the mask marks existing in the first marked image and the mask marks existing in the second marked image into the same image, and to represent the mask marks existing in the first marked image in a preset manner in the same image.
[0023] Inputting the target marked image and the template image into the generator, so that the generator performs image fusion on the target marked image and the template image to obtain a generated image; wherein, the training convergence target of the generator includes: when there are mask marks represented in the preset manner in the target marked image, using the image content of the template image to repair the target area; the target area is the image area indicated by the mask marks represented in the preset manner.
[0024] Inputting the target image and the generated image into the discriminator, so that the discriminator makes a true / false judgment on the generated image based on the target image and the generated image.
[0025] In a fourth aspect, an embodiment of the present application provides an image generation device, including:
[0026] An image acquisition module for acquiring a pre-collected real product image as an initial image.
[0027] A mark acquisition module for acquiring a mask mark image; wherein, the mask mark image is generated by adding mask marks of various types of defects to a blank image having the same size as the initial image, and the mask marks of different types of defects are different.
[0028] An image generation module, configured to input the initial image and the mask marking map into a generator of a pre-trained sample image generation model, so that the generator fuses the initial image and the mask marking map to obtain a generated image corresponding to the initial image;
[0029] Wherein, the generated image corresponding to the initial image includes: defects of each type indicated by each mask marking in the mask marking map, and the sample image generation model is trained based on the training method of the above sample image generation model.
[0030] In a fifth aspect, an embodiment of the present application provides an electronic device, including:
[0031] A memory, configured to store a computer program;
[0032] A processor, configured to implement any of the training methods of the sample image generation model provided in the first aspect, and / or any of the image generation methods provided in the second aspect when executing the program stored in the memory.
[0033] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, it implements any of the training methods of the sample image generation model provided in the first aspect, and / or any of the image generation methods provided in the second aspect.
[0034] In a seventh aspect, an embodiment of the present application further provides a computer program product containing instructions, which when running on a computer, causes the computer to execute any of the training methods of the sample image generation model provided in the first aspect, and / or any of the image generation methods provided in the second aspect.
[0035] Advantageous effects of the embodiments of the present application:
[0036] As can be seen above, a sample image generation model provided by an embodiment of the present application includes a generator and a discriminator. In the process of training the sample image generation model each time, by fusing the mask markings of the first marking map and the second marking map in the sample image, in the obtained target marking map, while there are mask markings in the first marking map and mask markings in the second marking map, the mask markings existing in the first marking map are characterized in a preset manner. Thus, the generator performs image fusion on the target marking map and the template map, and when there are mask markings characterized in a preset manner in the target marking map, the image content of the template map is used to repair the area indicated by such mask markings. In this way, the generated image can be close to the real defect image (both the background part and the defect part are close), so as to achieve the effect of deceiving the discriminator. Furthermore, the stability of the sample image generation model is improved, and the quality of the generated sample images is further improved.
[0037] Moreover, an image generation method provided by an embodiment of the present application, with the help of the trained sample image generation model, can use a small number of real product images to generate a large number of sample images for defect detection model training. And the generated sample images can not only include the image content in the above initial images, but also include the defects corresponding to each mask marking added when generating the mask marking map. Since the initial image is a real product image, when there are defects in the initial image, the generated image corresponding to the initial image can not only include the real defects of the product, but also include generated defects. And when making the above mask marking map, various types of defects and the mask markings corresponding to various types of defects can be added to the blank image. Thus, the diversity of the generated defects in the generated images is greatly improved, and the generated defects in the generated images can have a good data distribution. In this way, by applying the solution provided by the embodiment of the present application, sufficient defect images can be obtained as the images for training the defect detection model to improve the performance of the defect detection model.
[0038] In addition, in the present application, it is not limited whether the template map has defects, so that in the training process and the image generation process of the above sample image generation model, the template map and the initial image can be defect images or defect-free images. This makes the above training process and image generation process have high flexibility in selecting the template map and the initial image, avoiding the influence on model training and image generation caused by unqualified image selection or difficult image acquisition (such as difficult acquisition of defect-free images).
[0039] Of course, it is not necessary for any product or method implementing the present application to achieve all the above advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other embodiments can also be obtained based on these drawings.
[0041] Figure 1 It is a schematic flowchart of a method for training a sample image generation model provided by an embodiment of the present application;
[0042] Figure 2 It is a schematic flowchart of an image generation method provided by an embodiment of the present application;
[0043] Figure 3(a) is an exemplary training flowchart of a sample image generation model in a specific embodiment provided by an embodiment of the present application;
[0044] Figure 3(b) is another exemplary training flowchart of a sample image generation model in a specific embodiment provided by an embodiment of the present application;
[0045] Figure 4 It is an inference flowchart of a sample image generation model in a specific embodiment provided by an embodiment of the present application;
[0046] Figure 5 It is a schematic structural diagram of a training device for a sample image generation model provided by an embodiment of the present application;
[0047] Figure 6 It is a schematic structural diagram of an image generation device provided by an embodiment of the present application;
[0048] Figure 7 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0049] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art based on the present application belong to the scope of protection of the present application.
[0050] For the convenience of understanding, first, the professional terms and background technology involved in the present application will be specifically described.
[0051] Deep learning is a research direction in the field of Machine Learning (ML). It is introduced into machine learning to make it closer to the original goal - Artificial Intelligence (AI). Among them, deep learning is to learn the internal laws and representation levels of sample data, and the information obtained during these learning processes is very helpful for the interpretation of data such as text, images, and sounds. The ultimate goal of deep learning is to enable machines to have the ability to analyze and learn like humans, and be able to recognize data such as text, images, and sounds. The applicable directions of the embodiments of this application mainly include classification, object detection, semantic segmentation, instance segmentation, and text localization in the field of computer vision.
[0052] Generative Adversarial Network (GAN) is a deep learning neural network model that learns tasks through a game method. The generative adversarial network consists of two parts: a generator and a discriminator. The generator and the discriminator learn through a game during the model training stage. After training is completed, the generator can be used to generate image samples, achieving the effect and purpose of being indistinguishable from the real ones.
[0053] Defect-free image: It refers to an image without any defects. In the defect detection task, it refers to a good product image, also known as an OK image, and is also called a template image or a reference image in related technologies.
[0054] Defective image: It refers to an image with one or more defects. In the defect detection task, it refers to a defective product image, also known as an NG image or a defective image.
[0055] Model training: By setting a certain goal, the deep neural network is continuously made to approach this goal, and the parameters of the deep neural network are adjusted during this process.
[0056] Defect detection: It refers to using deep learning technology to detect and mark the defective parts in industrial products, generally using technologies such as object detection, semantic segmentation, and instance segmentation.
[0057] In industrial production, the requirement for the product qualification rate is relatively high. The industrial defect detection task is the main part to ensure the product qualification rate. Whether the defective images are sufficient directly affects the performance of the defect detection model and the qualification rate of the output products. Due to the strictness of the industrial production environment, it is often not suitable to collect a large amount of data for training the above-mentioned defect detection model. That is, in some cases, due to the difficulty of generating defective samples (i.e., defective images) (the qualification rate of the normal production line is relatively high), it is difficult to collect defective samples; and in some cases, due to the existence of weak defects in the good products that do not meet the defect standards, it is also difficult to collect "pure" defect-free samples. And collecting a large number of defect-free samples and then manually screening may not only affect normal production but also consume labor costs.
[0058] In the related art, among the methods for obtaining a defective image as a sample image for training a defect detection model, there are the following three exemplary methods:
[0059] The first method: By means of data augmentation (not relying on a neural network), the original defects in the collected defective images are transformed and then copied, pasted, or Poisson fused in other regions. However, this method cannot generate defects with sufficient diversity, and its defect features are still based on the defects in the collected defective images. The main change is the positional relationship.
[0060] The second method: Use a generative adversarial network to generate defects, and then copy, paste, or Poisson fuse the defects onto the image. However, this method does not consider generating the defect-free regions in the images using a neural network. Therefore, the pixel values of the defect-free regions are fixed, which is likely to cause overfitting in the training of the model using this generated sample, and there may be a sense of falsehood and a sense of boundary in the defect regions and defect-free regions after pasting.
[0061] The third method: Collect some defect-free images and a small number of defective images. By using the defect-free images as template images and pairing them with the defective images to train a generative adversarial network, defective images are generated based on the defect-free images, thereby generating a certain number of sample images with generated defects. Moreover, since the defect-free regions are also generated by a network model (the generator in the generative adversarial network), the pixel diversity of the defect-free regions is also guaranteed. However, this method requires collecting both defect-free images and defective images at the same time. In some cases, there are some minor defects in the defect-free images that do not meet the defect standards (such as the size not reaching a certain size, the contrast not reaching a certain threshold, etc. The common oral expressions are relatively light dirt, relatively small impurities, relatively light oil stain blocks, etc.). These minor defects are not labeled. If they are used as defect-free images for training, it will affect the stability of the training of the generative adversarial network, thereby affecting the quality of the generated sample images. On the other hand, for a single sample image generated by this method, there are only generated defects and it cannot contain both real defect features and generated defect features, and the diversity of defects in the sample images is insufficient, which is likely to cause a large difference between the data distribution during training and the data distribution in the actual inference scenario.
[0062] As can be seen from the above, in the related art, when solving the technical problem of how to obtain sufficient defective images as sample images for training a defect detection model, it is necessary to collect defect-free images and use the defect-free images as template images, and moreover, it is impossible to generate a single sample image that contains both real defect features and generated defect features.
[0063] That is to say, the following problems exist in the related art: 1. If it is difficult to collect a large number of OK images (defect-free images), or there are only a small number of NG images (defective images), the above-mentioned related art cannot be used. 2. When there are minor defects in the OK images (defect-free images) that do not meet the defect standards but do not affect the product being a good product, it is not conducive to the training effect of the generative adversarial network. If manual re-screening of the OK images is required, the labor cost is relatively high. 3. For a single generated sample image, there is only a generated defect, and it cannot contain both real defect features and generated defect features. The diversity of defects in the sample image is insufficient, and the model using this sample set is prone to overfitting.
[0064] In order to obtain sufficient defective images as sample images for training a defect detection model, so as to improve the performance of the defect detection model, and further ensure the yield rate of products, and solve the problems existing in the above-mentioned related art, first, an embodiment of the present application provides a training method for a sample image generation model to improve the stability of the sample image generation model, thereby improving the quality of the generated sample images.
[0065] Among them, this method can be applied to various industrial production scenarios that require training a defect detection model to detect whether a product is qualified. For example, an automobile manufacturing enterprise trains a defect detection model to detect whether the automobile parts produced are qualified. Another example is a mobile phone manufacturing enterprise that trains a defect detection model to detect whether the mobile phones produced are qualified, etc. In addition, the execution subject of this method can be various electronic devices. The electronic device is specifically used to train and obtain a sample image generation model. In specific applications, the electronic device can be a server, a desktop computer, etc. Moreover, the electronic device can be a single electronic device or an electronic device cluster composed of multiple electronic devices. Based on this, the embodiment of the present application does not specifically limit the application scenario and execution subject of this method.
[0066] A training method for a sample image generation model provided by an embodiment of the present application may include the following steps:
[0067] Obtain sample images; wherein, the sample images include a template image, a target image, a first marked image, and a second marked image; the first marked image is an image obtained by performing a predetermined process on the template image, and the second marked image is an image obtained by performing the predetermined process on the target image. The predetermined process is a process of setting a mask mark for defects;
[0068] Use the sample images to train the sample image generation model to be trained until the sample image generation model converges; wherein, the sample image generation model includes a generator and a discriminator;
[0069] Each training process of the sample image generation model includes:
[0070] Perform mask label fusion on the first labeled image and the second labeled image to obtain a target labeled image; wherein, the mask label fusion is used to fuse the mask labels existing in the first labeled image and the mask labels existing in the second labeled image in the same image, and to characterize the mask labels existing in the first labeled image in the same image in a preset manner;
[0071] Input the target labeled image and the template image into the generator, so that the generator performs image fusion on the target labeled image and the template image to obtain a generated image; wherein, the training convergence objective of the generator includes: when there are mask labels characterized in the preset manner in the target labeled image, using the image content of the template image to perform image repair on the target region; the target region is the image region indicated by the mask labels characterized in the preset manner;
[0072] Input the target image and the generated image into the discriminator, so that the discriminator makes a true / false judgment on the generated image based on the target image and the generated image.
[0073] As can be seen above, a sample image generation model provided in an embodiment of the present application includes a generator and a discriminator. In each training process of the sample image generation model, through the mask label fusion of the first labeled image and the second labeled image in the sample image, in the obtained target labeled image, while there are mask labels in the first labeled image and mask labels in the second labeled image, the mask labels existing in the first labeled image are characterized in a preset manner; thus, through the generator performing image fusion on the target labeled image and the template image, and when there are mask labels characterized in the preset manner in the target labeled image, using the image content of the template image to perform image repair on the region indicated by such mask labels, in this way, the obtained generated image can be close to a real defect image (both the background part and the defect region are close), so as to achieve the effect of deceiving the discriminator, and further, improve the stability of the sample image generation model and further improve the quality of the generated sample images.
[0074] In addition, in the present application, whether the template image has defects is not limited, so that in the training process of the above sample image generation model and the subsequent image generation process, the template image and the initial image can be defect images or defect-free images, which makes the above training process and subsequent image generation process have high flexibility in the selection of the template image and the initial image, and avoids the influence on model training and image generation caused by unqualified image selection or difficult image acquisition (such as difficult acquisition of defect-free images).
[0075] Next, in conjunction with the accompanying drawings, a training method for a sample image generation model provided by an embodiment of the present application will be specifically described.
[0076] Figure 1 FIG. 4 is a schematic flowchart of a training method for a sample image generation model provided by an embodiment of the present application. The method may include the following steps:
[0077] S101: Obtain sample images.
[0078] Among them, the sample images include a template image, a target image, a first marked image, and a second marked image; the first marked image is an image obtained by performing a predetermined process on the template image, and the second marked image is an image obtained by performing a predetermined process on the target image. The predetermined process is a process of setting a mask mark for defects.
[0079] In the present application, a sample image generation model to be trained including a generator and a discriminator may be pre-constructed. Then, training images for training the sample image generation model are obtained as sample images. The number of sample images may be multiple, and each sample image includes a set of images: a template image, a target image, a first marked image, and a second marked image. It can be understood that in order to ensure that there are differences between each sample image, so as to achieve the purpose of effectively training the sample image generation model with different sample images, for any two sample images, at least one of the template image and the target image is different.
[0080] Among them, the template image and the target image are real product images collected in advance, and the target image is a defective image. The template image may be a defective image or a defect-free image, and there are no overlapping defective areas between the template image and the target image.
[0081] The first marked image is an image obtained by performing a process of setting a mask mark for defects on the template image. That is, the mask marked image of the template image obtained by marking the defects of the template image with the mask mark is used as the first marked image. Among them, when the defects of the template image are marked with the mask mark, the first marked image can be automatically generated.
[0082] Optionally, when there are defects in the template image, that is, when the template image is a defective image, for each defect in the template image, a graphic with a mask mark of the same type as the defect covers the area where the defect is located in the first marked image. Then, the covered graphic is the mask mark of the marked defect, and the area in the first marked image where the above mask mark is not added is a blank area.
[0083] Correspondingly, the second marked image is the image obtained by performing the process of setting mask marks for defects on the target image. That is to say, the mask mark image of the target image obtained by performing defect annotation on the target image using the mask marks is used as the second marked image. Among them, when performing defect annotation on the target image using the mask marks, the second marked image can be automatically generated.
[0084] Optionally, for each defect in the target image, in the second marked image, a graphic with the same mask mark as the type of the defect is covered in the area where the defect is located. Then, the covered graphic is the mask mark of the marked defect, and the area in the second marked image where the above mask marks are not added is a blank area.
[0085] Among them, for any defect, the covered graphic can be a graphic with the same shape as the defect and having the mask mark color preset for the type of the defect; it can also be a graphic with the same shape as the defect. Of course, it can also be a graphic having the mask mark color preset for the type of the defect. In this regard, the embodiments of the present application do not make specific limitations. Any method in the related art that can implement setting mask marks for defects can be applied to the present application.
[0086] S102: Use the sample image to train the sample image generation model to be trained until the sample image generation model converges;
[0087] Among them, the sample image generation model includes: a generator and a discriminator;
[0088] Each training process of the sample image generation model may include the following steps:
[0089] Step A1: Perform mask mark fusion on the first marked image and the second marked image to obtain a target marked image;
[0090] Among them, mask mark fusion is used to fuse the mask marks existing in the first marked image and the mask marks existing in the second marked image in the same image, and to represent the mask marks existing in the first marked image in a preset manner in the same image;
[0091] Step A2: Input the target marked image and the template image into the generator so that the generator performs image fusion on the target marked image and the template image to obtain a generated image;
[0092] Among them, the training convergence target of the generator includes: when there are mask marks represented in a preset manner in the target marked image, using the image content of the template image to repair the target area; the target area is the image area indicated by the mask marks represented in a preset manner;
[0093] Step A3: Input the target image and the generated image into the discriminator, so that the discriminator can judge the authenticity of the generated image based on the target image and the generated image.
[0094] In this application, the sample image generation model to be trained can be a generative adversarial network, which includes a generator and a discriminator. Therefore, when inputting the sample image into the sample image generation model for training, the first labeled image and the second labeled image can be first subjected to mask label fusion to obtain a target labeled image. Then, the obtained target labeled image and the template image are input into the generator, so that the generator performs image fusion on the target labeled image and the template image to obtain a generated image. Exemplarily, the mask labels existing in the first labeled image and the mask labels existing in the second labeled image can be fused into the same blank image. Of course, it is not limited to the blank image. For example, it can also be an image with a preset background. Exemplarily, the preset background can be a product line background, etc., which is also reasonable.
[0095] Among them, the training objective of the above generator is that it is expected that the generated image generated by the generator can deceive the discriminator of the sample image generation model to be trained, and it is expected that the discriminator judges the above generated image as true (True). Furthermore, in order to achieve the above training objective, for the above generator, it is expected that the generated image is as similar as possible, or even the same, as the real defect image (that is, the defect image of the product collected).
[0096] To achieve the above purpose, when performing mask label fusion, the mask labels existing in the first labeled image corresponding to the template image included in the target labeled image can be characterized in a preset manner. Thus, the training convergence objective of the generator can include: when there are mask labels characterized in a preset manner in the target labeled image, use the image content of the template image to repair the image area indicated by the mask labels characterized in the preset manner.
[0097] Among them, the essence of the above mask label fusion process can be understood as: fusing the mask labels of the first labeled image and the mask labels of the second labeled image into the same image, then the obtained target labeled image can simultaneously include the mask labels of the first labeled image and the mask labels of the second labeled image. Among them, since the above template image can be a defect image or a defect-free image, and the above target image is a defect image, therefore, the target labeled image must include the mask labels of the second labeled image, and when the above template image is a defect image, the target labeled image can also include the mask labels of the first labeled image; correspondingly, when the above template image is a defect-free image, at this time, the mask labels of the first labeled image can be empty, so that the target labeled image can only include the mask labels of the second labeled image.
[0098] Among them, in order to distinguish the mask marks of the first marked image and the mask marks of the second marked image, and to determine whether the subsequent determination model converges, during the above mask mark fusion process, the mask marks existing in the first marked image are characterized in the same image in a preset manner.
[0099] Optionally, in a specific implementation manner, when both the template image and the target image are defective images, step A1 above, which performs mask mark fusion on the first marked image and the second marked image to obtain a target marked image, includes the following steps:
[0100] Step A11: On the basis that the mask marks existing in the second marked image are retained and the mask marks existing in the first marked image are characterized in a preset manner, the mask marks existing in the first marked image and the mask marks existing in the second marked image are fused into the same image to obtain a target marked image.
[0101] In this specific implementation manner, the mask marks existing in the second marked image are kept unchanged and the mask marks existing in the first marked image are characterized in a preset manner to achieve the purpose of fusing the defects existing in the first marked image and the defects existing in the second marked image. Of course, in other implementation manners, when the mask marks existing in the first marked image are characterized in a preset manner and the positions and types of the defects indicated by the mask marks in the second marked image can still be recognized in the target marked image, during the process of fusing into the target marked image, the mask marks existing in the second marked image can also change, for example: the mark color changes, or the mark morphology changes, which are all feasible.
[0102] By characterizing the mask marks existing in the first marked image in a preset manner in the same image, the trained sample image generation model can be made to have the ability to repair the image area indicated by the mask marks characterized in a preset manner. That is to say, if the mask marks existing in the first marked image are not specially processed (characterized in a preset manner), then the generated image obtained by the sample image generation model will retain the defects in the first marked image. As a result, regardless of whether there are defects in the template image, the generated image will retain the defects. That is, when the template image is a defect-free image, there are no defects at the corresponding positions in the generated image; when the template image is a defective image, the defects will be retained at the corresponding positions in the generated image. This leads to the presence of defects in the generated image other than the defects in the target image. And if there are defects in the image other than the defects in the target image, the image is likely to be determined as a false image.
[0103] Moreover, if the mask marks in the first marked image are not specially processed, then, for the entire training process, it may be difficult for the model to converge due to the existence of these mask marks. In another case, the model may converge, but in one situation, the model may randomly repair the image area indicated by the mask marks in the first marked image. That is to say, if the template image corresponding to the first marked image has defects, in actual use, when obtaining the generated image, these defects may be randomly repaired or deformed, etc., resulting in the inability to retain the real defects. In another case, when there are no defects in the template image corresponding to the first marked image, the model may also add random defects relative to the defects in the second marked image, which will also cause the real defects not to be retained when generating the image.
[0104] Among them, there can be various implementation manners for representing the mask marks existing in the first marked image in a preset manner in the same image.
[0105] Optionally, in a specific implementation manner, the above-mentioned representation of the mask marks existing in the first marked image in a preset manner in the same image may include: adjusting the color value of the area indicated by the mask marks existing in the first marked image in the same image to a specified color value.
[0106] In this specific implementation manner, when performing mask mark fusion, in the same image where the mask marks existing in the first marked image and the mask marks existing in the second marked image are fused, the color value of the area indicated by the mask marks existing in the first marked image can be adjusted to a specified color value. Among them, the above-mentioned specified color value can also be referred to as a reserved color value. Exemplarily, the above-mentioned specified color value can be black.
[0107] In this way, through the above adjustment of the color value, it is convenient to distinguish the mask marks existing in the two marked images in the same image. Thus, when obtaining the generated image by using the generator, by determining the area with the specified color value, the area that needs to be image-repaired can be directly determined, so as to improve the determination efficiency of the area to be image-repaired.
[0108] Optionally, the above-mentioned representation of the mask marks existing in the first marked image in a preset manner in the same image includes: adjusting the area identifier of the area indicated by the mask marks existing in the first marked image in the same image to a predetermined area identifier.
[0109] In this specific implementation manner, when performing mask mark fusion, in the same image where the mask marks existing in the first marked image and the mask marks existing in the second marked image are fused, the area identifier of the area indicated by the mask marks existing in the first marked image can be adjusted to a predetermined area identifier.
[0110] In this way, through the adjustment of the above-mentioned region identifiers, it is convenient to distinguish the mask marks existing in each of the two marked graphs in the same image. Thus, when obtaining the generated image by using the generator, by determining the region belonging to the predetermined region identifier, the region that needs to be image-inpainted can be directly determined, so as to improve the efficiency of determining the region to be image-inpainted.
[0111] It should be noted that the goal of image inpainting is: to ensure that the generated image output by the generator can deceive the discriminator as much as possible. Otherwise, in the case where there are mask marks represented by the above preset method in the generated image, the generated image can be easily judged as False by the discriminator. By performing image inpainting, the adversarial training between the above generator and discriminator can proceed normally and stably. At the same time, the above image inpainting can also increase the learning of structural knowledge by the above generator and discriminator, and improve the quality of the finally obtained image. Moreover, during the training process of the above sample image generation model, using image inpainting can enable the above generator and discriminator to work stably.
[0112] Optionally, in a specific implementation manner, in step A2 above, the generator performs image fusion on the target marked graph and the template graph to obtain the generated image, including:
[0113] Step A21: Generate image content corresponding to the image content of the reference region of the template graph in the specified region of the target marked graph except for the existing mask marks, to obtain an intermediate image;
[0114] Wherein, the reference region is an image region with the same image position as the specified region;
[0115] Step A22: Perform predetermined image processing on the intermediate image to obtain the generated image;
[0116] Wherein, the predetermined image processing includes a first sub-processing and a second sub-processing;
[0117] The first sub-processing includes: in the case where there are mask marks represented by a preset method in the target marked graph, using the image content of the template graph to perform image inpainting on the target region in the intermediate image;
[0118] The second sub-processing includes: according to the morphology of other mask marks not represented by the preset method in the target marked graph, generating defects of the type indicated by the other mask marks in the image region indicated by the other mask marks in the intermediate image.
[0119] In this specific implementation, after receiving the target marker map and the template map, the generator can generate image content corresponding to the image content of the reference area of the template map in a specified area other than the mask markers existing in the locations in the target marker map, obtaining an intermediate image. That is, for the specified area, the generator can generate the image content of the specified area of the target marker map according to the image content of the reference area of the template map.
[0120] Among them, when generating the image content of the specified area of the target marker map, the generated image content should be as close as possible to the image content of the reference area of the template map, so that the finally obtained generated image can be as close as possible to the real defect image. The above image content can be understood as the visual pixels of the image.
[0121] Then, perform predetermined image processing on the obtained intermediate image to obtain a generated image. Among them, the above predetermined image processing includes a first sub-processing and a second sub-processing.
[0122] That is to say, after obtaining the intermediate image, determine whether there are mask markers characterized in a preset manner in the target marker map. If so, use the image content of the template map to repair the target area in the intermediate image.
[0123] And, after obtaining the intermediate image, determine other mask markers in the target marker map that are not characterized in a preset manner, and generate defects of the type indicated by the other mask markers in the image area indicated by the other mask markers in the intermediate image according to the morphology of the other mask markers.
[0124] Among them, considering the mask markers originally belonging to the second marker map, that is, other mask markers in the target marker map that are not characterized in a preset manner, which have a morphology, so information such as the shape and size of the defects to be generated can be determined according to the morphology of the other mask markers, and the type of the generated defects can be determined according to the type indicated by the other mask markers. In this way, the generator can generate defects with corresponding morphology and type in the image area indicated by the other mask markers in the intermediate image according to the morphology of the other mask markers.
[0125] Optionally, the above mask markers can include two attributes: morphology and color. Among them, different colors correspond to different types of mask markers. Therefore, after determining information such as the shape and size of the defects to be generated, the type of the defects to be generated can be determined according to the color of the other mask markers, so that the generator can generate defects of the type with the color of the mask marker in the image area indicated by the other mask markers in the intermediate image according to the morphology of the other mask markers. It should be noted that when the mask markers also include color, the color value of this color is different from the above specified color value.
[0126] Among them, the embodiments of the present application do not limit the execution order of the above first sub - processing and second sub - processing. The first sub - processing can be executed first and then the second sub - processing; or the second sub - processing can be executed first and then the first sub - processing, both of which are reasonable.
[0127] It should be emphasized that the above template image can be a defective image or a non - defective image. When the above template image is a non - defective image, there are no mask marks belonging to the first mark image in the above target mark image, that is, there are no mask marks characterized in a preset manner. Therefore, after generating the image content of the specified area of the above target mark image and generating defects of the corresponding type of the mask mark in the image area indicated by the other mask mark in the intermediate image according to the form of the other mask mark, the generator can complete the image generation and obtain the generated image. At this time, the obtained generated image is a defective image.
[0128] When the above template image is a defective image, there are mask marks belonging to the first mark image in the above target mark image, that is, there are mask marks characterized in a preset manner. Therefore, it is necessary to process the mask marks belonging to the first mark image in the above target mark image. For example, when the above template image is a defective image, there are mask marks characterized in a preset manner in the above target mark image. Therefore, the generator needs to process this type of mask mark.
[0129] Among them, when the above template image is a defective image, in order to achieve the training purpose of the generator, it is expected that the defects of the above generated image are as similar as possible, or even the same, as those of the above target image. Considering that there are no defects with overlapping areas between the above template image and the target image, there are no defects of the template image in the target image. Furthermore, in order to make the defects of the obtained generated image as similar as possible, or even the same, as those of the above target image, there should be no defects in the area indicated by the mask mark characterized in a preset manner in the obtained generated image, but it should be as close as possible to the image content of the above target image. Further, considering that the image content of the specified area of the target mark image is generated based on the template image, when the above template image is a defective image and there are mask marks characterized in a preset manner, for the mask marks belonging to this type in the target mark image, the image content of the template image can be used to repair the image in the area where the mask marks are located.
[0130] In this way, after generating the image content of the specified area of the above target mark image; generating defects of the corresponding type of the mask mark in the image area indicated by the other mask mark in the intermediate image according to the form of the other mask mark; and repairing the image of the mask mark characterized in a preset manner, the generator can complete the image generation and obtain the generated image. At this time, the obtained generated image is a defective image.
[0131] After obtaining the above-generated image, the above target image and the generated image can then be input into the discriminator of the sample image generation model to be trained.
[0132] Among them, the training purpose of the above discriminator is to judge the real defect image as True and the generated defect image as False, that is, the training purpose of the above discriminator is to judge the above target image as True and the above generated image as False.
[0133] Based on this, when obtaining the above-generated image, a True label can be added to the above target image, and a False label can be added to the above-generated image. Thus, the target image and the generated image with labels added are input into the discriminator, and the discriminator can first learn the image features of the real defect image and the generated defect image based on the image content of the obtained target image and generated image, as well as the labels added to the target image and the generated image respectively. Furthermore, a True label can be added to the above-generated image, and the generated image with the True label added is input into the discriminator again, so that the discriminator can first judge the authenticity of the generated image based on the learned image feature knowledge, and determine whether the judgment result of the discriminator is accurate according to the similarity or difference between the obtained judgment result and the True label added to the above-generated image.
[0134] Among them, for the above-generated image with the True label added, when the judgment result of the discriminator is that the generated image is False, it means that the image generated by the above generator fails to deceive the discriminator, and the model parameters of the above generator can be adjusted so that the adjusted generator can generate an image closer to the real defect image (target image); when the judgment result of the discriminator is that the generated image is True, it means that the image generated by the above generator can deceive the discriminator, but it does not meet the training purpose of the discriminator itself. Therefore, the parameters of the above discriminator can be adjusted so that the adjusted discriminator can distinguish the real defect image (target image) and the generated defect image (generated image) as much as possible, judge the real defect image as True, and judge the generated defect image as False.
[0135] As described above, in the embodiment of the present application, the training process of the above sample image generation model includes a generator training stage and a discriminator training stage, and the above generator training stage and discriminator training stage appear alternately and influence each other.
[0136] Among them, in the above generator training stage, the generator generates image content corresponding to the image content of the reference area of the template image in the blank area of the target marker map, and generates defects of the corresponding type of the mask marker in the area where the mask marker is located according to the morphology of the mask marker of the second marker map, obtaining a generated image. The training purpose is to expect the discriminator to judge the above generated image as True. In the above discriminator training stage, the training purpose is to judge the real defect image as True and the generated defect image as False.
[0137] That is, the above generator hopes that the obtained output can deceive the discriminator, while the above discriminator hopes not to be deceived by the output of the generator. Moreover, the judgment result of the above discriminator can affect the adjustment of the model parameters of the generator and the discriminator, and the generated image output by the above generator can affect the judgment result of the discriminator. Therefore, the training purposes of the above generator and discriminator are contrary, and the output results of the two affect each other. In this way, through the alternating training of the above generator and discriminator, the balance of the output results and training purposes of the above generator and discriminator can be achieved, and the model training can be completed.
[0138] In this way, when training the sample image generation model to be trained, the obtained template image, target image, first marker map, and second marker map can be used to alternately train the generator and discriminator of the sample image generation model to be trained, and finally achieve the balance between the output results of the above generator and discriminator, such as Nash equilibrium. Thus, finally, a trained sample image generation model that meets the actual application requirements can be obtained.
[0139] Therefore, after obtaining the judgment result of the above discriminator for judging the authenticity of the above generated image, the loss value of the discriminator can be calculated based on the judgment result of the above discriminator, and the loss value of the generator can be calculated based on the judgment result of the above discriminator, the target image, and the generated image. Furthermore, the model parameters of the discriminator can be adjusted based on the loss value of the discriminator, and the model parameters of the generator can be adjusted based on the loss value of the generator.
[0140] Among them, the loss function used to calculate the loss value of the discriminator and the loss function used to calculate the loss value of the generator can be conventional generative adversarial network loss functions.
[0141] In this way, when the loss value of the above discriminator and the loss value of the above generator both meet the preset loss conditions, it can be considered that the output results and training purposes of the above generator and discriminator reach a balance, completing the training of the generator and the discriminator, that is, completing the training of the sample image generation model to be trained, and obtaining a trained sample image generation model.
[0142] In addition, optionally, in some cases, when the number of iterations of the sample image reaches a set value during the training process, the training is stopped, and the model obtained when the training is stopped is used as the trained sample image generation model.
[0143] As mentioned above, in this specific implementation, the principle of training the generative adversarial network is: the discriminator will judge the difference (loss, loss value) between the predicted image (generated image) and the real image (target image) generated by the generative adversarial network, and the training goal is to make the predicted image (generated image) closer to the real defect image. Among them, when training the generative adversarial network, the defects in the defect image (target image) with defects can be marked with special marks (mask marks), so that the generative adversarial network can repair the defects in the template image. In this way, based on the defect image (target image) and the template image, the predicted defect image (generated image) generated by the generative adversarial network can be close to the background part of the template image and the defect part of the real defect image (target image). In this way, the predicted defect image can be close to the real defect image, achieving the effect of deceiving the discriminator of the generative adversarial network. Therefore, this application does not limit whether the template image has defects, and regardless of whether the template image has defects, the discriminator can confirm the ability of the generative adversarial network to generate defect images without being disturbed by the defects of the template image.
[0144] Based on this, the following problems existing in the prior art can be solved: In order to normally train the generative adversarial network, the template image can only be an OK image and cannot have defects. Otherwise, the predicted image generated by the generative adversarial network will contain the defects in the template image, making it difficult to approach the real image and affecting the work of the discriminator.
[0145] At this point, after completing the training of the above sample image generation model, the image generation method provided by the embodiment of the present application can be used to generate a large number of sample images using the generator in the above trained sample image generation model, and the defect detection model training can be performed using the generated large number of sample images. Finally, the defect detection model obtained by training is used to perform defect detection on the generated product to ensure the yield rate of the generated product.
[0146] As can be seen above, a sample image generation model provided by an embodiment of the present application includes a generator and a discriminator. In the process of training the sample image generation model each time, by fusing the mask markings of the first marking map and the second marking map in the sample image, in the obtained target marking map, while there are mask markings in the first marking map and mask markings in the second marking map, the mask markings existing in the first marking map are characterized in a preset manner; thus, the generator performs image fusion on the target marking map and the template map, and when there are mask markings characterized in a preset manner in the target marking map, the image content of the template map is used to repair the area indicated by such mask markings, so that the obtained generated image can be close to a real defect image (both the background part and the defect area are close), so as to achieve the effect of deceiving the discriminator, and further improve the stability of the sample image generation model and further improve the quality of the generated sample images.
[0147] In addition, in the present application, whether the template map has defects is not limited, so that in the training process of the above sample image generation model and the subsequent image generation process, the template map and the initial image can be defect images or defect-free images, which makes the above training process and subsequent image generation process have high flexibility in selecting the template map and the initial image, and avoids the influence on model training and image generation caused by unqualified image selection or difficult image acquisition (such as difficult acquisition of defect-free images).
[0148] Optionally, in a specific implementation manner, a training method for a sample image generation model provided by an embodiment of the present application may further include the following steps:
[0149] Step B: Collect multiple real product images, and perform predetermined processing on each real product image to obtain a marking map of each real product image;
[0150] Among them, the multiple real product images at least include: defect images.
[0151] Correspondingly, the above step S101, obtaining a sample image, may include the following steps:
[0152] Step C: Obtain a template map and a target map from the multiple real product images, and obtain a marking map of the template map as the first marking map, and obtain a marking map of the target map as the second marking map to obtain a sample image.
[0153] In this specific implementation, multiple real product images can be collected in advance. Moreover, these multiple real product images at least include defective images. That is to say, these multiple real product images can all be defective images, or can be partially defective images and partially non-defective images. Furthermore, predetermined processing can be performed on each real product image to obtain a labeled map of each real product image.
[0154] Among them, for each real product image, when the real product image is a defective image, in the labeled map of the real product image, it can include mask labels for covering each defect in the real product image. The shape of each mask label is the same as the shape of the covered defect, and the type of the mask label can represent the type of the covered defect; while when the real product image is a non-defective image, the labeled map of the real product image can be a blank image.
[0155] Since the construction of the sample image needs to follow certain construction principles: the target image is a defective image, and the template image can be a defective image or a non-defective image, and there are no overlapping defects between the template image and the target image. Therefore, after obtaining multiple real product images, according to this construction principle, a template image and a target image can be obtained from these multiple real product images. Then, a labeled map of the template image can be further obtained as the first labeled map, and a labeled map of the target image can be obtained as the second labeled map, thereby completing the acquisition of the sample image.
[0156] In this specific implementation, the above template image can be a defective image or a non-defective image. Among them, a defective image is a product image obtained by collecting an image of a product with product defects and the image content includes defects (flaws). Correspondingly, a non-defective image refers to a product image obtained by collecting an image of a product and the image content does not include defects. That is, a non-defective image is an image without any flaws. In the defect detection task, it refers to a good product image, also known as an OK image. A defective image is also called a defective image, which is an image with one or more flaws. In the defect detection task, it refers to a defective product image, also known as an NG image. Thus, in this application, since it is not necessary to limit that the above template image must be a defective image or a non-defective image, the difficulty of constructing the sample image can be reduced and the flexibility of using the template image for model training can be improved.
[0157] Moreover, in actual scenarios, in some cases, since good products, i.e., qualified products, may have minor defects that do not meet the defect standards, it is relatively difficult to collect "pure" defect-free images. After collecting a large number of defect-free images, manually screening for "pure" defect-free images not only affects normal production but also incurs labor costs. Therefore, when the above template image is a defect image, it is possible to avoid collecting defect-free images of the product. Thus, it is possible to avoid the impact on the training of the above sample image generation model caused by some defects in the collected defect-free images that do not meet the defect standards, and it is also possible to avoid the labor costs and the impact on product generation caused by manually selecting "pure" defect-free images.
[0158] Optionally, in a specific implementation, step C of obtaining a template image and a target image from multiple real product images may include the following steps:
[0159] Step C1: Randomly select two images from multiple real product images;
[0160] Step C2: When the two selected images include a defect image and a defect-free image, determine the selected defect-free image as the template image and the selected defect image as the target image;
[0161] Step C3: When the two selected images are both defect images, determine whether there is an overlapping area in the defects of the two selected images; if not, execute step C4; if so, execute step C5.
[0162] Step C4: Determine one of the selected images as the template image and the other selected image as the target image;
[0163] Step C5: Ignore the two selected images and return to the above step C1.
[0164] Step C6: When the two selected images are both defect-free images, ignore the two selected images and return to the above step C1.
[0165] In this specific implementation, two images can be randomly selected from the above collected multiple real product images first. Since the above multiple real product images can only include defect images, or can include both defect images and defect-free images, there can be multiple situations for the two selected images, including: selecting two defect images, selecting two defect-free images, and selecting one defect image and one defect-free image. For the above different situations, the situations where the two selected images can be used as the template image and the target image respectively are different.
[0166] Among them, when both of the two extracted images are defect-free images, since neither of the two images can be used as the target image, it is necessary to discard the current extraction result and re-extract images, that is, ignore the two extracted images and return to the step of randomly extracting two images from multiple real product images to re-perform image extraction.
[0167] When the two extracted images include one defective image and one defect-free image, since there are no defects in the defect-free image, there are no defects with overlapping regions in the two extracted images. Therefore, the extracted defective image can be directly determined as the target image, and the extracted defect-free image can be determined as the template image.
[0168] When the two extracted images include two defective images, since both of the two extracted images have defects, and the finally determined target image and template image have no defects with overlapping regions, it is necessary to first determine whether there are overlapping regions in the defects of the two extracted images.
[0169] For example, the respective labeled images corresponding to the two extracted images can be obtained, so as to determine whether there are mask marks in the two labeled images: mask marks that belong to the two labeled images respectively and have overlapping regions. When it exists, it indicates that there are overlapping regions in the defects of the two extracted images, and when it does not exist, it indicates that there are no overlapping regions in the defects of the two extracted images.
[0170] In this way, when it is determined that there are no overlapping regions in the defects of the two extracted images, the two extracted images can be used as the template image and the target image respectively. Thus, one of the extracted images can be determined as the template image, and the other extracted image can be determined as the target image. Among them, the two extracted images can be determined as the template image and the target image respectively in various ways. For example, one of the extracted images can be randomly determined as the template image, and the other extracted image can be determined as the target image. Another example is that the first extracted image can be determined as the template image, and the later extracted image can be determined as the target image, etc.
[0171] Correspondingly, when it is determined that there are overlapping regions in the defects of the two extracted images, the two extracted images cannot be used as the template image and the target image respectively. Therefore, it is necessary to discard the current extraction result and re-extract images, that is, ignore the two extracted images and return to the step of randomly extracting two images from multiple real product images to re-perform image extraction.
[0172] Secondly, the embodiment of the present application provides an image generation method to improve the diversity of generation defects in the generated images, enable the generation defects in the generated images to have a good data distribution, and at the same time, provide sufficient defective images for the training of the defect detection model to improve the performance of the defect detection model.
[0173] Among them, this method can be applied to various industrial production scenarios that require training a defect detection model to detect whether a product is qualified. For example, an automobile manufacturing enterprise trains a defect detection model to detect whether the produced automobile parts are qualified. Another example is a mobile phone manufacturing enterprise that trains a defect detection model to detect whether the produced mobile phones are qualified, etc. In addition, the execution subject of this method can be various electronic devices. Specifically, the electronic device is used to generate images by using the trained sample image generation model. In a specific application, the electronic device can be a server, a desktop computer, etc. Moreover, the electronic device can be a single electronic device or an electronic device cluster composed of multiple electronic devices. Based on this, the embodiment of the present application does not specifically limit the application scenario and execution subject of this method.
[0174] It should be noted that the electronic device that executes the image generation method provided by the embodiment of the present application and the electronic device that executes the above-mentioned training method of a sample image generation model can be the same device or different devices. In this regard, the embodiment of the present application does not make a specific limitation.
[0175] The image generation method provided by the embodiment of the present application may include the following steps:
[0176] Obtain a pre-collected real product image as an initial image;
[0177] Obtain a mask marking map; wherein, the mask marking map is generated by adding mask markings of various types of defects to a blank image with the same size as the initial image, and the mask markings of different types of defects are different;
[0178] Input the initial image and the mask marking map into the generator of the pre-trained sample image generation model, so that the generator fuses the initial image and the mask marking map to obtain a generated image corresponding to the initial image;
[0179] Among them, the generated image corresponding to the initial image includes: defects of the type indicated by each mask marking of the mask marking map, and the sample image generation model is trained based on the above-mentioned training method of a sample image generation model.
[0180] As can be seen above, for an image generation method provided in an embodiment of the present application, by means of a trained sample image generation model, a small number of real product images can be used to generate a large number of sample images for training a defect detection model. Moreover, the generated sample images can not only include the image content in the above initial images, but also include defects of the types corresponding to the respective mask marks added when generating the mask mark map. Since the initial images are real product images, when there are defects in the initial images, the generated images corresponding to the initial images can include not only the real defects of the products, but also generated defects. Moreover, when making the above mask mark map, various types of defects and the mask marks corresponding to the various types of defects can be added to a blank image, thereby greatly increasing the diversity of the generated defects in the generated images and enabling the generated defects in the generated images to have a good data distribution. In this way, by applying the solution provided in the embodiment of the present application, sufficient defect images can be obtained as the images for training the defect detection model to improve the performance of the defect detection model.
[0181] In addition, in the present application, whether the template image has defects is not limited, so that in the training process and the image generation process of the above sample image generation model, the template image and the initial image can be defect images or defect-free images, which makes the above training process and image generation process have high flexibility in selecting the template image and the initial image, and avoids the influence on model training and image generation caused by unqualified image selection or difficult image acquisition (such as difficult acquisition of defect-free images).
[0182] Next, with reference to the accompanying drawings, a specific description will be given of an image generation method provided in an embodiment of the present application.
[0183] Figure 2 The following is a schematic flowchart of an image generation method provided in an embodiment of the present application, and the method may include the following steps:
[0184] S201: Obtain a pre-collected real product image as an initial image.
[0185] In the present application, when training a sample image generation model and using the sample image generation model to generate sample images, a pre-collected real product image can be first obtained as an initial image.
[0186] Among them, the above-mentioned initial image can be a defective image or a defect-free image. As mentioned above, a defective image is a product image obtained by collecting images of products with product defects and including defects (flaws) in the image content. Correspondingly, a defect-free image refers to a product image obtained by collecting images of products and not including defects in the image content. That is, a defect-free image is an image without any flaws, which refers to a good product image, also known as an OK image, in the defect detection task. A defective image, also known as a defective image, is an image with one or more flaws, which refers to a defective product image, also known as an NG image, in the defect detection task. Thus, in the present application, since it is not necessary to limit the above-mentioned initial image to be a defective image or a defect-free image, the flexibility of image generation using the initial image can be improved.
[0187] Moreover, in the actual scenario, in some cases, since good products, that is, qualified products, may have weak flaws that do not meet the defect standard, it is relatively difficult to collect "pure" defect-free images. After collecting a large number of defect-free images and then manually screening "pure" defect-free images, it not only affects normal production but also consumes labor costs. Therefore, when the above-mentioned initial image is a defective image, it is not necessary to collect defect-free images of the product. Thus, it is possible to avoid some flaws that do not meet the defect standard in the collected defect-free images from affecting the training of the above-mentioned sample image generation model and the above-mentioned image generation, and it is also possible to avoid the labor costs and the impact on product generation caused by manually selecting "pure" defect-free images.
[0188] S202: Obtain a mask marking map.
[0189] Among them, the mask marking map is generated by adding mask markings of various types of defects to a blank image with the same size as the initial image, and the mask markings of different types of defects are different.
[0190] In the present application, when generating a sample image using the above-mentioned trained sample image generation model, a mask marking map needs to be prepared, and the mask marking map is generated by adding mask markings of various types of defects to a blank image with the same size as the initial image. That is, a blank image with the same size as the above-mentioned initial image can be obtained first, and then mask markings of various types of defects are added to the blank image, and the mask markings of different types of defects are different.
[0191] Among them, various types of defects that may exist in the product can be determined in advance, and mask marks for each type of defect can be set. In order to distinguish different types of defects, the mask marks for different types of defects are different. For example, the mask mark for scratch-type defects is mark 1, the mask mark for sunken pit-type defects is mark 2, and so on. The process of setting marks for defects in the image, that is, using the mask mark to perform defect marking processing on the image, can be understood as using the mask mark to cover the area where the defect is located in the image. At this time, the shape of the mask mark is the same as the shape of the defect in the image, and the type is the type to which the defect in the preset image belongs.
[0192] Optionally, the mask marks for different types of defects are distinguished by colors. For example, the color of the mask mark for scratch-type defects is red, the color of the mask mark for sunken pit-type defects is yellow, the mask color for bulge-type defects is green, and so on. Generally, the mask mark can be understood as a graphic with a specific shape and color for covering defects in the image. Thus, when using the mask mark to cover the area where the defect is located in the image, the shape of the mask mark is the same as the shape of the defect in the image, and the color is the color of the mask mark of the type to which the defect in the preset image belongs. Optionally, when using the mask mark to label the defect in the image, the edge line of the mask mark is outlined along the edge of the area where the defect is located to obtain the graphic of the mask mark. Furthermore, the corresponding color is filled in the outlined graphic according to the type of the defect.
[0193] As mentioned above, the above mask mark map is obtained by directly adding mask marks to a blank image, and there are no defects in the blank image. Therefore, the mask marks added to the blank image are not used to cover defects in the image, but are used to label the generated defects that are expected to exist in the generated image. That is, for each mask mark in the mask mark map, in the sample image generated based on the mask mark map, there can be a defect with the same position, the same shape, and the type of the mask mark as the mask mark. Among them, the above shape can include information such as shape and size. For example, assuming that the mask color for bulge-type defects is mark 1, when there is a circular mask mark with a radius of size B centered at image position A and marked as 1 in the mask mark map, then in the image generated based on the mask mark map, there will be a bulge-type defect with a radius of size B centered at image position A.
[0194] Based on this, when generating the above mask marking map, information such as the position, shape, and type of the generation defects desired to exist in the target sample image for the model training task can be determined according to the actual application needs, the actual production situation of the product, etc. Thus, according to the information of the generation defects desired to exist in the above target sample image, corresponding mask markings are added to the blank image to obtain the mask marking map. Among them, one or more mask markings of one type of defect can be added to the above mask marking map, or multiple mask markings of multiple types of defects can be added, and the number of mask markings of each type of defect can be one or more. Moreover, mask markings can be added at any position on the above mask marking map, and the shape of each added mask marking can be arbitrary. That is, when generating the above mask marking map, the position, number, shape, and type of the mask markings added to the mask marking map are not restricted. In this way, a large number of mask marking maps with different added mask markings can be generated, thereby greatly improving the richness of the mask marking maps. Furthermore, when using the above mask marking map to generate sample images, a large number of sample images with different defects can also be generated. Therefore, the diversity of the generation defects in the generated sample images can be greatly improved, and the generation defects in the generated sample images can have a good data distribution to ensure that sufficient defect images can be finally obtained as sample images for training the defect detection model to improve the performance of the defect detection model.
[0195] Among them, the above mask marking map can be generated in various ways. For example, using computer peripherals such as a mouse, in a blank image with the same size as the above initial image, mask markings (or the colors of mask markings) of various preset types of defects are used to draw one or more mask markings of one or more types of defects; for another example, mask markings such as curves, ellipses, polygons, etc. are generated in a blank image with the same size as the above initial image through a function.
[0196] S203: Input the initial image and the mask marking map into the generator of the sample image generation model pre-trained, so that the generator fuses the initial image and the mask marking map to obtain the generated image corresponding to the initial image.
[0197] Among them, the generated image corresponding to the initial image includes: defects of the type indicated by each mask marking of the mask marking map, and the sample image generation model is trained based on a training method of a sample image generation model provided in an embodiment of the present application.
[0198] In this application, after obtaining the above-mentioned initial image and mask marking map, the initial image and mask marking map can be input into the generator of the sample image generation model pre-trained, so that the generator fuses the initial image and the mask marking map to obtain the generated image corresponding to the initial image.
[0199] It should be noted that the generated image corresponding to the above-mentioned initial image is the target sample image for the subsequent model training task, for example, the target sample image for the training task of the defect detection model.
[0200] Among them, the image area of the above-mentioned mask marking map is divided into two parts according to whether there is a mask marking. The area without a mask marking is the blank area, and the other part is the area with a mask marking. The image area in the initial image with the same image position as the blank area is used as the target reference area.
[0201] In this way, after the generator obtains the above-mentioned initial image and mask marking map, for the blank area in the mask marking map, the generator can generate image content corresponding to the image content of the target reference area of the initial image in the above-mentioned blank area. That is, for the blank area in the mask marking map, the generator can generate the image content of the blank area according to the image content of the target reference area of the initial image. And the above-mentioned image content can be understood as the visual pixels of the image.
[0202] Among them, in an ideal state, the image content of the generated blank area is consistent with the image content of the target reference area of the initial image. In practical applications, the image content of the generated blank area has a high similarity with the image content of the target reference area of the initial image. And, among the image content of the generated blank area, the significant features in the image content of the target reference area of the initial image can be retained to a great extent. Therefore, when there is a defect in the target reference area of the initial image, the defect features of the defect can be retained to a great extent in the generated image. And since the initial image is a real product image, the defect in the initial image is a real defect. Thus, real defect features can exist in the generated image.
[0203] Furthermore, for each mask marking in the mask marking map, since the mask marking has a shape and a type, the shape, size, etc. of the defect to be generated can be determined according to the shape of the mask marking, and the type of the defect to be generated can be determined according to the type of the mask marking. Thus, according to the shape of each mask marking, a defect of the type indicated by the mask marking is generated in the area where the mask marking is located in the mask marking map.
[0204] Optionally, different types of mask marks can be characterized by different colors. Thus, the type of defect to be generated can be determined according to the color of the mask mark, and defects of the type with the color of the mask mark can be generated in the area where the mask mark is located in the mask mark map according to the shape of each mask mark. For example, assume that the mask color of the convex type of defect is green. When there is a green circular mask mark with the center at image position A and radius B in the mask mark map, then, in the mask mark map, a convex type of defect with the center at image position A and radius B is generated in the area where the green circular mask mark is located. Different types of mask marks can also be characterized by different visual elements. The visual elements can be understood as graphic elements with relatively small sizes. The mask mark can be composed of several visual elements. For example, the mask mark of type A defect is composed of multiple first-type visual elements, and the mask mark of type B defect is composed of multiple second-type visual elements. It should be noted that different colors are used to distinguish different types of mask marks. Different colors can be understood as multiple colors with different hues, such as gray, red, green, etc.; or different variants of the same color. For example, the gray variants can include titanium gray, pink gray, lead gray, etc. with different gray levels, and all of these are reasonable.
[0205] As described above, the mask marks added in the mask mark map represent the defects that are expected to exist in the generated image. Therefore, the defects generated for each mask mark are not the defects originally existing in the initial image, but generated defects. Thus, generated defect features can also exist in the generated image.
[0206] That is, the generated image can not only include real defect features, but also include generated defect features. Thus, the diversity of the generated sample images can be further improved, and the sample images can be ensured to have a good data distribution.
[0207] As can be seen above, an image generation method provided by an embodiment of the present application can generate a large number of sample images by means of a trained sample image generation model using a small number of real product images for defect detection model training. Moreover, the generated sample images can not only include the image content in the above initial images, but also include defects corresponding to the types of each mask mark added when generating the mask mark map. Since the initial image is a real product image, when there are defects in the initial image, the generated image corresponding to the initial image can include not only the real defects of the product but also the generated defects. Moreover, when making the above mask mark map, various types of defects and the corresponding mask marks of various types of defects can be added to a blank image, thereby greatly improving the diversity of the generated defects in the generated image and enabling the generated defects in the generated image to have a good data distribution. In this way, by applying the solution provided by the embodiment of the present application, sufficient defect images can be obtained as images for training the defect detection model to improve the performance of the defect detection model.
[0208] In addition, in the present application, whether the template image has defects is not limited, so that in the training process and the image generation process of the above sample image generation model, the template image and the initial image can be defect images or defect-free images, which makes the above training process and image generation process have high flexibility in selecting the template image and the initial image, avoiding the influence on model training and image generation caused by unqualified image selection or difficult image acquisition (such as difficult acquisition of defect-free images).
[0209] Optionally, in a specific implementation manner, an image generation method provided by an embodiment of the present application may further include the following steps:
[0210] Step D: Add defect labels to the defects in the generated image corresponding to the initial image based on the mask marks of the defects in the initial image and the mask marks existing in the mask mark map, and filter the defects in the generated image corresponding to the initial image that do not meet the preset defect conditions based on the defect labels.
[0211] In this specific implementation manner, since the mark map of each real product image collected can be obtained, and the above initial image is extracted from multiple collected real product images, the mask mark map of the above initial image can be obtained, and the mask marks of the real defects in the initial image can be marked in the mark map of the initial image. Moreover, as mentioned above, the mask marks of the generated defects expected to be generated in the target sample image for model training tasks can be marked in the above mask mark map.
[0212] Based on this, for the target sample image generated by the generator of the above sample image generation model, the forms and types of the generated defects and real defects in the target sample image are both known. Therefore, the defect labels can be added to the defects in the target sample image corresponding to the initial image by using the labeled map and mask labeled map of the initial image.
[0213] Among them, the added defect labels can be represented in various ways. For example, the minimum bounding rectangle of the defect can be used as the label, or the polygon multi-point contour map of the defect can be used as the label, etc.
[0214] Furthermore, according to the differences in scenarios such as product type, product generation requirements, and applicable products of the product, when performing defect detection on the product, the requirements for the forms, types, etc. of the defects labeled in the target sample image are also different. Therefore, before training the defect detection model using the above target sample image, the defect conditions can be preset according to the actual needs of product defect detection as the preset defect conditions to retain the defects that meet the actual needs of product defect detection for use in defect detection model training. For example, setting the defect size, setting the defect shape, setting the image contrast of the area where the defect is located, etc.
[0215] Thus, using the defect labels added to the above target sample image, the defects in the above target sample image are filtered, that is, the defect labels of the defects that do not meet the preset defect conditions in the above target sample image are removed. In this way, when training the defect detection model using the target sample image after the above defect filtering, the defects labeled by the filtered defect labels will not be used as defects for the defect detection model to learn image features, but only as the image areas without defects in the target sample image to learn image features. Thus, the defect detection model will not learn the image features of the defects labeled by the filtered defect labels. Furthermore, the trained defect detection model will not regard the regional defects in the image whose image features are similar to the image features of the defects labeled by the above filtered defect labels as defects.
[0216] To facilitate the understanding of an image generation method provided by an embodiment of the present application and the training method of the sample image generation model as described above, below, specific examples shown in FIGS. 3(a), 3(b) and Figure 4 will be used for specific illustration.
[0217] Among them, Fig. 3(a) is an exemplary training flowchart of a sample image generation model, and Fig. 3(b) is another exemplary flowchart of the sample image generation model. Among them, the real defect map is the target map, the defect map mask label is the mask label map (the second label map) for the target map, the template map mask label is the mask label map (the first label map) for the template map, the fused mask label is the target label map, the generated defect map is the generated image, and the generator (also called the GAN generator) and the discriminator (also called the GAN discriminator) are the generator and the discriminator of the sample image generation model.
[0218] As shown in Fig. 3(a), the real defect map has defects 01, 02, and 03. The defect map mask label has the mask label 011 corresponding to defect 01, the mask label 021 corresponding to defect 02, and the mask label 031 corresponding to defect 03. The template map has defect 04, and the template map mask label has the mask label 041 corresponding to defect 04. In the fused mask label, there are the mask label 012 corresponding to defect 01, the mask label 022 corresponding to defect 02, the mask label 032 corresponding to defect 03, and the mask label 042 corresponding to defect 04 (this mask label 042 has been adjusted to the specified color, black). In the generated defect map, there are the defect 013 corresponding to defect 01, the defect 023 corresponding to defect 02, and the defect 033 corresponding to defect 03.
[0219] As shown in Fig. 3(b), the real defect map has defects 05, 06, and 07. The defect map mask label has the mask label 051 corresponding to defect 05, the mask label 061 corresponding to defect 06, and the mask label 071 corresponding to defect 07. The template map has defect 08, and the template map mask label has the mask label 081 corresponding to defect 08. In the fused mask label, there are the mask label 052 corresponding to defect 05, the mask label 062 corresponding to defect 06, the mask label 072 corresponding to defect 07, and the mask label 082 corresponding to defect 08 (this mask label 082 has been adjusted to the specified color, black). In the generated defect map, there are the defect 053 corresponding to defect 05, the defect 063 corresponding to defect 06, and the defect 073 corresponding to defect 07.
[0220] Figure 4 It is the inference flowchart of the sample image generation model, that is, the generation flowchart of the target sample image (the generated image corresponding to the initial image). Among them, the defect mask label to be generated is the mask label map, the template map is the initial image, the template map mask label is the label map of the initial image, the generated defect map is the target sample image (that is, the generated image corresponding to the initial image), and the generator is the generator of the sample image generation model.
[0221] As Figure 4As shown, the mask marks to be generated for the defect map include mask mark 111, mask mark 121, and mask mark 131; the template map has a defect 14; the mask mark of the template map has a mask mark 141; the generated defect map has defects 112, 122, 132, and 142.
[0222] And, in the specific embodiments shown in FIG. 3(a), FIG. 3(b) and Figure 4 the sample image generation model is a generative adversarial network.
[0223] As shown in FIG. 3(a), FIG. 3(b) and Figure 4 shown, the training process of the sample image generation model may include the following Step1-Step7, and the generation process of the target sample image may include the following Step8-Step12. Specifically:
[0224] Step1: Sample collection. In this specific embodiment, only a small amount of defect images need to be collected, and the real defects in the collected defect images are masked and labeled to obtain the labeled map of each collected defect image. Among them, different types of defects are masked and labeled with different colors (exemplarily, in FIG. 3(a), mask mark 011, mask mark 021, and mask mark 031 in the defect map mask mark are three mask marks with different gray levels, and in FIG. 3(b), mask mark 051, mask mark 061, and mask mark 071 in the defect map mask mark are three mask marks with different gray levels, that is, in FIG. 3(a) and FIG. 3(b), three mask marks are characterized by different gray levels), and the collected defect images can be called defect samples, and the labeled masks can be called mask marks. Among them, in order to improve the training effect of the sample image generation model, when masking and labeling the real defects in the collected defect images, careful labeling is usually required. Therefore, the above mask labeling can also be called mask fine-labeling.
[0225] Step2: Training image pairing. Randomly select two images from the defect images collected in Step1 to form an image pair, one of which is called the template image and the other is called the real defect image (target image), and finally obtain four input quantities, namely the real defect image (target image), the defect map mask mark (the mask mark map or the second mark map about the target image), the template image, and the template image mask mark (the mask mark map or the first mark map about the template image), that is, the four input quantities on the leftmost side shown in FIG. 3(a) and FIG. 3(b).
[0226] Step 3: Process through the filtering module. Input the defect map mask label and the template map mask label into the filtering module to determine whether there are overlapping defective positions in the defect map mask label (target map mask label map) and the template map mask label (template map mask label map). If there is an overlap, repeat Step 2 to randomly extract a new pair of images. If there is no overlap in the defective positions, proceed to Step 4.
[0227] Step 4: Process through the fusion module. Through the fusion module, fuse the defect map mask label and the template map mask label. Specifically: Retain the color pixel values of each mask label in the defect map mask label (exemplarily, in Figure 3(a), the color pixel values of mask label 012 and mask label 011 are the same, the color pixel values of mask label 022 and mask label 021 are the same, and the color pixel values of mask label 032 and mask label 031 are the same; in Figure 3(b), the color pixel values of mask label 052 and mask label 051 are the same, the color pixel values of mask label 062 and mask label 061 are the same, and the color pixel values of mask label 072 and mask label 071 are the same), and convert each mask label in the template map mask label to a specified color, for example, black (exemplarily, the pixel value of mask label 042 in Figure 3(a) and the pixel value of mask label 082 in Figure 3(b) are both black). After fusion, all the mask labels of the defects in the defect map mask label and the template map mask label are fused into a single label map to obtain the fused mask label (target label map). For the defects in the template map, at the position of each defect in the fused mask label (target label map), annotate with a mask label of the specified color. For example, the black triangles in Figure 3(a) and Figure 3(b) are the mask labels of the template map mask label (template map mask label map). For the defects in the real defect map (target map), at the position of each defect in the fused mask label (target label map), annotate with a mask label of the color set for the type of that defect. For example, the curves and the irregular figures formed by the curves in Figure 3(a) and Figure 3(b) are the mask labels in the defect map mask label (target map mask label map) (exemplarily, mask label 011, mask label 021, and mask label 031 in Figure 3(a); mask label 051, mask label 061, and mask label 071 in Figure 3(b)). Among them, the colors set for the mask labels of different types of defects are different from the above-specified color to avoid the reserved color (the reserved color is the specified color, exemplarily the color of the triangles in Figure 3(a) and Figure 3(b)).
[0228] Step 5: Generator Training. The template graph and the fused mask label (target label graph) are input together into the GAN generator of the generative adversarial network for the training of the generative adversarial network. The generation objective of the GAN generator is as follows: in the blank area of the fused mask label, generate visual pixels corresponding to the reference area of the template graph; in the colored mask label part of the fused mask label (target label graph) (the mask label of the defect in the real defect graph), generate corresponding types of defects; and in the black mask label part of the fused mask label (target label graph), perform image inpainting to restore the visual pixels corresponding to the template graph (exemplarily, there are defects 013, 023, and 033 in the generated defect graph in Fig. 3(a), and there are defects 053, 063, and 073 in the generated defect graph in Fig. 3(b)). In this way, generate and output the generated defect graph (generated image), and the training objective of the GAN generator is to expect the generated defect graph (generated image) to deceive the discriminator in Step 6 and expect the discriminator to judge the generated defect graph (generated image) as True.
[0229] It should be noted that in the above Step 5, the purpose of performing image inpainting on the black mask label part is to ensure that the generated defect graph (generated image) output by the generator can deceive the discriminator as much as possible. The area responsible for the black mask label part is easily judged as a fake graph by the discriminator. This design is to enable the adversarial training between the generator and the discriminator to proceed normally and stably, while enhancing the learning of structural knowledge and improving the quality of data generation.
[0230] Step 6: Discriminator Training. The real defect graph (target graph) and the generated defect graph (generated image) output in Step 5 are jointly input into the GAN discriminator. The training objective of the GAN discriminator is to judge the real defect graph (target graph) as True and the generated defect graph (generated image) as False; that is, the discriminator outputs a True / False discrimination result.
[0231] As mentioned above, in the above Step 5 and Step 6, the GAN generator generates visual pixels corresponding to the reference area of the template graph in the blank area of the fused mask label, generates corresponding types of defects in the colored mask label part of the fused mask label, and performs image inpainting in the black mask label part of the fused mask label to restore the visual pixels corresponding to the template graph, obtaining the generated defect graph. Its training objective is for the discriminator to judge the generated defect graph as True. The training objective of the GAN discriminator is to judge the real defect graph as True and the generated defect graph as False.
[0232] Step 7: Perform cyclic iterative training according to Step 2 - Step 6. When the number of iterations reaches the set value, stop the training and output the best model, that is, obtain the trained sample image generation model.
[0233] Step 8: Inference template image selection. Randomly select one image from the defect images collected in Step 1 as the template image (initial image) for target sample image generation, and obtain the labeled image of the template image (initial image) for target sample image generation as the template image mask of the template image (initial image) for target sample image generation.
[0234] Step 9: Defect mask image generation. To generate a sample image, a defect mask label to be generated (mask label image) needs to be prepared. In the training process of the sample image generation model shown in Fig. 3(a) and Fig. 3(b) and Figure 4 in the inference process of the sample image generation model shown, the function of the fused mask label (target label image) in the training process and the defect mask label to be generated (mask label image) in the inference process is the same. When preparing the defect mask label to be generated (mask label image), the color of the defect label mask (mask label) can be pre-agreed according to the type of defect (exemplarily, Figure 4 mask label 111, mask label 121, and mask label 131 in are pre-determined colors). Thus, as long as the type of defect is determined, and then the color of the mask label in the defect mask label to be generated (mask label image) is determined, the defect mask label to be generated (mask label image) can be generated in various ways. For example, using a computer peripheral such as a mouse, draw mask labels of a certain shape in the blank image with the determined colors, which can be one or more, one type or more types. For another example, generate the mask label of the defect through function drawing, such as curves, ellipses, polygons, etc.
[0235] Step 10: Generator inference. The template image (initial image) for target sample image generation and the defect mask label to be generated (mask label image) are input into the GAN generator of the generative adversarial network. Thus, the GAN generator can output a generated defect image (target sample image).
[0236] Among them, the difference between Step 10 and Step 5 in the training process of the above sample image is that the black mask label is missing in the defect mask label to be generated (mask label image). According to Step 5, in Step 10, there is no area in the defect mask label to be generated (mask label image) that needs image repair, but only defects need to be generated at the corresponding positions of the colored mask label part, and visual pixels corresponding to the reference area of the template image need to be generated in the blank area. Therefore, the true defect features of the template image (initial image) for target sample image generation will still be retained and generated, and new generated defects will be additionally generated on this basis. As Figure 4As shown in the generated defect map (target sample image), a single generated image can contain both real defect features and generated defect features, namely, including defect 112, defect 122, defect 132, and defect 142, improving the sample diversity and ensuring a good data distribution of defects in the sample image.
[0237] Step11: Label generation. So far, the known information includes: the mask label of the defect map to be generated (mask label map) and the mask label (label map) of the template map (initial image) for generating the target sample image. Therefore, on the generated defect map (target sample image) output by Step10, the mask labels of all defects are known. Thus, the label information can be automatically generated by the code to add defect labels to each defect on the generated defect map (target sample image), and then, the target sample image with defect labels is output. Among them, the generated defect label can be the smallest circumscribed rectangle or the multi-point polygon contour of the defect. As Figure 4 shown in the rightmost image in, the generated defect label is the smallest circumscribed rectangle.
[0238] Step12: Label filtering. Label filtering is to filter the defect labels of defects that do not meet the defect standards (preset defect conditions) in the sample image according to the preconditions (preset defect conditions) of different projects (actual requirements of product defect detection). For example, setting the defect size, setting the defect contrast, etc.
[0239] By repeating the sample image generation model inference according to Step8 - Step12, a certain number of "generated defect maps" can be generated. The "generated defect maps" are automatically attached with labels, obtaining a certain number of sample images.
[0240] It should be noted that the solution provided in the embodiment of the present application can be compatible with the scenario where the template map in the above training process and inference process is a defect-free image, that is, in the above training process and inference process, if a defect-free image is used as the template map, there is no need to change the solution provided in the embodiment of the present application. Therefore, the minimum condition to be achieved by the solution provided in the embodiment of the present application is: only a small number of defect images are collected.
[0241] As described above, in the solution provided by the embodiments of the present application, a generative adversarial network (a generative adversarial network for paired training of a template graph and a defect image) is used. Among them, the real defect samples (target graphs), the mask labels corresponding to the real defects (target graph mask label graphs), the template samples (template graphs), and the defect mask labels corresponding to the template samples (template graph mask label graphs) are paired and input into the generative adversarial network for training. Moreover, before training the generative adversarial network, it is necessary to perform mask label fusion on the mask labels corresponding to the real defects (target graph mask label graphs) and the defect mask labels corresponding to the template samples (template graph mask label graphs) to obtain the fused mask labels (target label graphs), and the fused mask labels (target label graphs) will participate in the above-mentioned training process of the generative adversarial network. Among them, the above-mentioned mask label fusion process can be regarded as a preprocessing process before model training. Moreover, the above-mentioned template graph is not limited to a defect-free image, and can also be a defective image (defect image). In the model inference stage, the constructed defect mask (mask label graph), the template sample (template graph), and the defect mask corresponding to the template sample (template graph mask label graph) are simultaneously input into the sample image generation model to generate new defects on the template sample (template graph) without changing the visual features of the original defects on the template sample (template graph), thereby generating a large number of defect samples (sample images). The newly generated data can automatically generate defect labels without additional manual annotation. The generated defect labels and defect samples (sample images) can be used for specific defect detection training tasks, such as semantic segmentation, object detection, instance segmentation, etc.
[0242] Among them, the template graph mentioned in the solution provided by the embodiments of the present application supports defective images (defect images). By annotating the mask labels of the template graph and pairing them to input into the generative adversarial network, it is possible to avoid the influence of minor defects that do not reach the standard of defect degree in some defect-free images on the training results, ensuring the quality of the generated data (sample images). At the same time, each generated sample image has both real sample features and generated sample features, and the data diversity of the defects in the sample images is better. In addition, in terms of resource requirements, it avoids the workload of collecting a large number of defect-free images and manually selecting "pure" defect-free images in related methods.
[0243] Corresponding to the above-mentioned training method of a sample image generation model provided by the embodiments of the present application, the embodiments of the present application also provide a training device for a sample image generation model.
[0244] Figure 5 Shown in the structural schematic diagram of a training device for a sample image generation model provided by the embodiments of the present application, as Figure 5 shown, the device includes:
[0245] A sample image acquisition module 510 is configured to acquire sample images. Among them, the sample images include a template image, a target image, a first marked image, and a second marked image. The first marked image is an image obtained by performing a predetermined process on the template image, and the second marked image is an image obtained by performing the predetermined process on the target image. The predetermined process is a process of setting mask marks for defects.
[0246] A model training module 520 is configured to use the sample images to train a sample image generation model to be trained until the sample image generation model converges. Among them, the sample image generation model includes a generator and a discriminator.
[0247] Each training process of the sample image generation model includes:
[0248] Perform mask mark fusion on the first marked image and the second marked image to obtain a target marked image. Among them, the mask mark fusion is used to fuse the mask marks existing in the first marked image and the mask marks existing in the second marked image into the same image, and to represent the mask marks existing in the first marked image in a preset manner in the same image.
[0249] Input the target marked image and the template image into the generator, so that the generator performs image fusion on the target marked image and the template image to obtain a generated image. Among them, the training convergence target of the generator includes: when there are mask marks represented in the preset manner in the target marked image, using the image content of the template image to repair the target area. The target area is the image area indicated by the mask marks represented in the preset manner.
[0250] Input the target image and the generated image into the discriminator, so that the discriminator makes a true / false judgment on the generated image based on the target image and the generated image.
[0251] Optionally, in a specific implementation manner, the generator performs image fusion on the target marked image and the template image to obtain a generated image, including:
[0252] Generate image content corresponding to the reference area of the template image in a specified area other than the existing mask marks in the target marked image to obtain an intermediate image. Among them, the reference area is an image area with the same image position as the specified area.
[0253] Perform a predetermined image process on the intermediate image to obtain a generated image. Among them, the predetermined image process includes a first sub-process and a second sub-process.
[0254] The first sub - processing includes: when there is a mask mark characterized in the preset manner in the target mark map, using the image content of the template map to perform image patching on the target area in the intermediate image;
[0255] The second sub - processing includes: according to the morphology of the mask mark of the second mark map, generating defects of the type indicated by the identifier of the mask mark of the second mark map in the image area indicated by the mask mark of the second mark map in the intermediate image.
[0256] Optionally, in a specific implementation manner, both the template map and the target map are defect images;
[0257] The fusing the mask marks of the first mark map and the second mark map to obtain a target mark map includes:
[0258] On the basis that the mask marks existing in the second mark map are retained and the mask marks existing in the first mark map are characterized in the preset manner, fusing the mask marks existing in the first mark map and the mask marks existing in the second mark map into the same image to obtain a target mark map.
[0259] Optionally, in a specific implementation manner, the device further includes:
[0260] An image annotation module, configured to collect multiple real - product images and perform the predetermined processing on each real - product image to obtain a mark map of each real - product image; wherein, the multiple real - product images at least include: defect images;
[0261] The sample image acquisition module 510 includes:
[0262] An image acquisition sub - module, configured to obtain a template map and a target map from the multiple real - product images, and obtain the mark map of the template map as the first mark map, and obtain the mark map of the target map as the second mark map, so as to obtain a sample image.
[0263] Optionally, in a specific implementation manner, the image acquisition sub - module is specifically configured to:
[0264] Randomly extract two images from the multiple real - product images;
[0265] When the two extracted images include: a defect image and a non - defect image, determining the extracted non - defect image as the template map and determining the extracted defect image as the target map;
[0266] In the case where both of the two extracted images are defective images, determine whether there is an overlapping area in the defects of the two extracted images; if not, determine one of the extracted images as the template image and the other extracted image as the target image; otherwise, ignore the two extracted images and return to the step of randomly extracting two images from the multiple real product images.
[0267] In the case where both of the two extracted images are defect-free images, ignore the two extracted images and return to the step of randomly extracting two images from the multiple real product images.
[0268] Optionally, in a specific implementation manner, the characterizing the mask marks existing in the first marked image in a preset manner in the same image includes:
[0269] Adjust the color value of the area indicated by the mask mark existing in the first marked image in the same image to a specified color value.
[0270] Corresponding to the image generation method provided in the embodiment of the present application, the embodiment of the present application further provides an image generation device. Figure 6 For the structural schematic diagram of an image generation device provided in the embodiment of the present application, as Figure 6 shown, the device includes:
[0271] An image acquisition module 610, configured to acquire a pre-collected real product image as an initial image;
[0272] A mark acquisition module 620, configured to acquire a mask mark image; wherein, the mask mark image is generated by adding mask marks of various types of defects in a blank image having the same size as the initial image, and the mask marks of different types of defects are different;
[0273] An image generation module 630, configured to input the initial image and the mask mark image into a generator of a pre-trained sample image generation model, so that the generator fuses the initial image and the mask mark image to obtain a generated image corresponding to the initial image;
[0274] Wherein, the generated image corresponding to the initial image includes: defects of the type indicated by each mask mark of the mask mark image, and the sample image generation model is trained based on a training method of a sample image generation model provided in the embodiment of the present application.
[0275] Optionally, in a specific implementation manner, the device further includes:
[0276] A label adding module, configured to add defect labels to defects in the generated image corresponding to the initial image based on the mask labels of the defects in the initial image and the mask labels existing in the mask label map, and filter defects that do not meet the preset defect conditions in the generated image corresponding to the initial image based on the defect labels.
[0277] An embodiment of the present application further provides an electronic device, as Figure 7 shown, including:
[0278] A memory 701, configured to store a computer program;
[0279] A processor 702, configured to implement the steps of any of the training methods and / or sample image generation methods of the sample image generation model provided in the embodiments of the present application when executing the program stored on the memory 701.
[0280] And the above-mentioned electronic device may further include a communication bus and / or a communication interface, and the processor 702, the communication interface, and the memory 701 complete communication with each other through the communication bus.
[0281] The communication bus mentioned in the above-mentioned electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0282] The communication interface is used for communication between the above-mentioned electronic device and other devices.
[0283] The memory may include a Random Access Memory (RAM), and may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0284] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0285] In another embodiment provided by the present application, there is also provided a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the training methods and / or sample image generation methods of the sample image generation model provided by the embodiments of the present application are implemented.
[0286] In another embodiment provided by the present application, there is also provided a computer program product containing instructions. When it runs on a computer, the computer is caused to execute the steps of any of the training methods and / or sample image generation methods of the sample image generation model provided by the embodiments of the present application.
[0287] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer, or a data storage device such as a server, data center, etc. that includes one or more integrated available media. The available media may be magnetic media (for example, floppy disk, hard disk, magnetic tape), optical media (for example, DVD), or solid state disk (SSD), etc.
[0288] It should be noted that, in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
[0289] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device embodiments, electronic device embodiments, computer-readable storage medium embodiments and computer program product embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the corresponding parts of the method embodiments for the relevant content.
[0290] The above description is only a preferred embodiment of the present application and is not intended to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application are all included in the protection scope of the present application.
Claims
1. A training method for a sample image generation model, characterized in that: The method comprises: Acquire a sample image; wherein the sample image includes a template image, a target image, a first mark image, and a second mark image; the first mark image is an image obtained after a predetermined process is performed on the template image, and the second mark image is an image obtained after the predetermined process is performed on the target image, and the predetermined process is a process of setting a mask mark for a defect; Using the sample image, training a sample image generation model to be trained until the sample image generation model converges; wherein the sample image generation model includes: a generator and a discriminator; Each training process of the sample image generation model includes: Performing mask mark fusion on the first mark map and the second mark map to obtain a target mark map; wherein the mask mark fusion is used to fuse the mask marks existing in the first mark map and the mask marks existing in the second mark map into the same image, and to represent the mask marks existing in the first mark map in the same image in a preset manner; Inputting the target label image and the template image into the generator, so that the generator performs image fusion on the target label image and the template image to obtain a generated image; The training convergence target of the generator includes: when there is a mask mark represented in the preset manner in the target mark map, using the image content of the template map to perform image repair on the target area; the target area is the image area indicated by the mask mark represented in the preset manner; The target image and the generated image are input to the discriminator, so that the discriminator makes a true or false judgment on the generated image based on the target image and the generated image.
2. The method according to claim 1, characterized in that The generator performs image fusion on the target label image and the template image to obtain a generated image, including: In a designated area other than the existing mask mark in the target mark image, an image content corresponding to the image content of the reference area of the template image is generated to obtain an intermediate image; wherein the reference area is an image area having the same image position as the designated area; Performing a predetermined image processing on the intermediate image to obtain a generated image; wherein the predetermined image processing includes a first sub-processing and a second sub-processing; The first sub-processing includes: when there is a mask mark characterized in the preset manner in the target mark map, using the image content of the template map to perform image repair on the target area in the intermediate image; The second sub-processing includes: generating defects of the type indicated by the other mask marks in the image area indicated by the other mask marks in the intermediate image according to the forms of the other mask marks that are not represented in the preset manner in the target mark map.
3. The method according to claim 1, characterized in that The template image and the target image are both defect images; The step of performing mask label fusion on the first label image and the second label image to obtain a target label image includes: On the basis that the mask marks existing in the second label map are retained and the mask marks existing in the first label map are characterized in the preset manner, the mask marks existing in the first label map and the mask marks existing in the second label map are fused into the same image to obtain a target label map.
4. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: Collect multiple real product images, and perform the predetermined processing on each real product image to obtain a labeling image of each real product image; wherein the multiple real product images at least include: defect images; The obtaining of the sample image comprises: A template image and a target image are obtained from the multiple real product images, and a labeled image of the template image is obtained as a first labeled image, and a labeled image of the target image is obtained as a second labeled image to obtain a sample image.
5. The method according to claim 4, characterized in that The acquiring the template image and the target image from the plurality of real product images comprises: Randomly extract two images from the plurality of real product images; In the case where the two extracted images include: a defective image and a non-defective image, the extracted non-defective image is determined as a template image, and the extracted defective image is determined as a target image; In the case that both the extracted two images are defective images, it is determined whether there is an overlapping area of defects in the two extracted images; if not, one of the extracted images is determined as a template image, and the other extracted image is determined as a target image; otherwise, the two extracted images are ignored, and the process returns to the step of randomly extracting two images from the multiple real product images; In the case that both the extracted two images are defect-free images, the two extracted images are ignored, and the process returns to the step of randomly extracting two images from the multiple real product images.
6. The method according to any one of claims 1 to 3, characterized in that: The mask marks existing in the first mark map are characterized in the same image in a predetermined manner, including: The color value of the area indicated by the mask mark existing in the first mark map in the same image is adjusted to a specified color value.
7. An image generation method, characterized in that: The method comprises: Obtain a pre-collected real product image as an initial image; Acquire a mask mark map; wherein the mask mark map is generated by adding mask marks of various types of defects to a blank image of the same size as the initial image, and the mask marks of different types of defects are different; Inputting the initial image and the mask label map into a generator of a pre-trained sample image generation model, so that the generator fuses the initial image and the mask label map to obtain a generated image corresponding to the initial image; The generated image corresponding to the initial image includes: defects of the type indicated by each mask mark of the mask mark image, and the sample image generation model is trained based on the training method of the sample image generation model according to any one of claims 1-6.
8. The method according to claim 7, characterized in that The method further comprises: Based on the mask marks of the defects in the initial image and the mask marks existing in the mask mark map, defect labels are added to the defects in the generated image corresponding to the initial image, and based on the defect labels, defects that do not meet preset defect conditions in the generated image corresponding to the initial image are filtered.
9. A training device for a sample image generation model, characterized in that: The device comprises: A sample image acquisition module, used to acquire a sample image; wherein the sample image includes a template image, a target image, a first mark image, and a second mark image; the first mark image is an image obtained after a predetermined process is performed on the template image, and the second mark image is an image obtained after the predetermined process is performed on the target image, and the predetermined process is a process of setting a mask mark for a defect; A model training module, used to train a sample image generation model to be trained using the sample image until the sample image generation model converges; wherein the sample image generation model includes: a generator and a discriminator; Each training process of the sample image generation model includes: Performing mask mark fusion on the first mark map and the second mark map to obtain a target mark map; wherein the mask mark fusion is used to fuse the mask marks existing in the first mark map and the mask marks existing in the second mark map into the same image, and to represent the mask marks existing in the first mark map in the same image in a preset manner; The target label image and the template image are input to the generator, so that the generator performs image fusion on the target label image and the template image to obtain a generated image; wherein the training convergence target of the generator includes: when there is a mask mark represented in the preset manner in the target label image, using the image content of the template image to perform image repair on the target area; the target area is the image area indicated by the mask mark represented in the preset manner; The target image and the generated image are input to the discriminator, so that the discriminator makes a true or false judgment on the generated image based on the target image and the generated image.
10. An image generating device, characterized in that: The device comprises: An image acquisition module is used to acquire a real product image collected in advance as an initial image; A marking acquisition module, used to acquire a mask marking map; wherein the mask marking map is generated by adding mask markings of various types of defects to a blank image of the same size as the initial image, and the mask markings of different types of defects are different; An image generation module, used for inputting the initial image and the mask label map into a generator of a pre-trained sample image generation model, so that the generator fuses the initial image and the mask label map to obtain a generated image corresponding to the initial image; The generated image corresponding to the initial image includes: defects of the type indicated by each mask mark of the mask mark image, and the sample image generation model is trained based on the training method of the sample image generation model according to any one of claims 1-6.
11. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, for implementing any of the methods described in claims 1-6 and / or claims 7-8 when executing a program stored in a memory.
12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of claims 1 to 6 and / or claims 7 to 8 is implemented.
Citation Information
Patent Citations
Image fusion method and system based on generative adversarial network, and storage medium
CN111754446A
Defect image generation method and device, electronic equipment and storage medium
CN114972268A