Small sample defect detection method and system based on generative adversarial network
By using generative adversarial networks for sample synthesis and refinement, the problem of data scarcity and imbalance in defect detection in industrial manufacturing is solved, achieving efficient and robust defect detection.
Patent Information
- Application Number
- CN202511084994.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies for defect detection in industrial manufacturing suffer from data imbalance and sample scarcity, making it difficult for traditional models to adapt to rapid iteration, resulting in low recall rates and poor robustness to changes in lighting and viewing angles.
Generative adversarial networks (GANs) are used for sample synthesis. Combined with multi-channel image acquisition and preprocessing, the GAN generates synthetic samples and mixes them with real samples to train a defect segmentation model. Defect detection is performed using DeepLabV3+ and MobileNetV2 network architectures. Level Set algorithm and Gaussian filtering are introduced for mask refinement.
It significantly improves the recall and accuracy of defect detection, enables the detection of minute defects in complex environments, enhances the segmentation performance of image segmentation models, and provides better technical applications.
Smart Images

Figure CN120976146A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of industrial defect detection, and particularly relates to a small sample defect detection method and system based on a generative adversarial network. BACKGROUND
[0002] With the development of industrial intelligentization and automatic generation technology, industrial manufacturing quality detection technology appears, which ensures the reliability and safety of products by quality detection on industrial manufacturing parts. However, through process optimization and quality control, the modern production line makes the normal product ratio usually more than 99%, and the defect sample becomes an extreme minority class. The defect morphology is affected by multiple variables such as material characteristics, processing technology, environmental factors, and is characterized by irregular geometric shape, multi-scale feature and complex texture anomaly, which is difficult to cover all defect types by limited samples.
[0003] In the traditional technology, the convolutional neural network relying on a large number of labeled defect samples is easy to identify normal samples as the dominant optimization target in the class imbalance scene, resulting in low defect detection recall rate. The manually designed features are difficult to capture the semantic information of defects, and the robustness to light changes and angle differences is poor.
[0004] The current intelligent manufacturing industry has full inspection demand for industrial manufacturing products, and new materials and new processes give birth to new defect morphology. The traditional model needs to be re-labeled and trained for months due to the lack of corresponding training data, and it is difficult to adapt to the rapid iteration rhythm of industry. The demand of intelligent manufacturing for detection efficiency and accuracy forms a sharp contradiction with the data dilemma of high imbalance and difficult acquisition of defect samples. SUMMARY
[0005] Therefore, it is necessary to provide a small sample defect detection method and system based on a generative adversarial network, which can break through the data imbalance and sample scarcity dilemma to detect defects of industrial manufacturing.
[0006] In a first aspect, the application provides a small sample defect detection method based on a generative adversarial network, comprising:
[0007] Performing multi-channel image acquisition on a target industrial part and preprocessing to obtain a part image set;
[0008] Inputting the part image set into a defect segmentation model to obtain a defect segmentation mask and a defect category labeling result;
[0009] Performing thinning processing on the defect segmentation mask to obtain a real defect mask;
[0010] Obtaining a defect detection report of the target industrial part according to the real defect mask and the defect category labeling result; the defect detection report includes the position coordinates of each defect and the corresponding defect category.
[0011] In one embodiment, the defect segmentation model is obtained as follows:
[0012] Multi-channel image acquisition and preprocessing are performed on the defective target industrial parts to obtain a defect image set;
[0013] Based on generative adversarial networks, a synthetic image set is obtained by synthesizing samples from a defective image set. The generative adversarial network includes a StyleGAN2 generator, a PatchGAN discriminator, and a DiffAugment enhancement structure.
[0014] Synthetic sample pseudo-labels are generated from the synthetic image set using a pre-trained segmentation model to obtain synthetic sample pairs;
[0015] A fusion training set is constructed based on the defect image set and synthetic sample pairs, and a network architecture based on DeepLabV3+ and MobileNetV2 is trained to obtain a defect segmentation model.
[0016] In one embodiment, multi-channel image acquisition is performed on the defective target industrial component, and preprocessing is carried out to obtain a defect image set, including:
[0017] Image acquisition of the target component is performed using photometric stereo method to obtain channel images; the channel images include curvature map, texture map and distance map;
[0018] Gaussian filtering is applied to the channel images, and image normalization and contrast stretching are then applied to obtain the filtered and enhanced image set.
[0019] For each image in the filtered and enhanced image set, multiple preset enhancement strategies are executed in a random combination to obtain a defective image set; the enhancement strategies include random rotation, flipping, scaling, brightness perturbation, and adding mild Gaussian noise.
[0020] In one embodiment, a pre-trained segmentation model is used to generate synthetic sample pseudo-labels from the synthetic image set to obtain synthetic sample pairs, including:
[0021] The pre-trained DeepLabV3+ network is used to perform preliminary segmentation on the synthesized image to obtain a coarse mask;
[0022] The Level Set algorithm is used to refine the contour of the coarse mask to obtain the refined mask.
[0023] Morphological methods are used to fill the holes in the refined mask to obtain a defect mask;
[0024] Synthetic sample pairs are constructed using defect masks and corresponding synthetic images.
[0025] In one embodiment, a fused training set is constructed based on a defect image set and synthetic sample pairs, and a network architecture based on DeepLabV3+ and MobileNetV2 is trained to obtain a defect segmentation model, including:
[0026] A subset of defect images is annotated with real defect masks to obtain real sample pairs; each real sample pair includes a real defect image and its corresponding real defect mask.
[0027] Real sample pairs and synthetic sample pairs are mixed in a preset ratio and then enhanced to obtain a fused training set;
[0028] According to the preset training objectives, the DeepLabV3+ backbone semantic segmentation architecture is trained based on the fused training set to obtain the first model parameters;
[0029] Based on preset defect category labels, a MobileNetV2 classification network is trained according to the connected defect regions to obtain the second model parameters; the connected defect regions are obtained by segmenting and fusing images in the training set using the DeepLabV3+ backbone semantic segmentation architecture;
[0030] The defect segmentation model is obtained based on the first model parameters and the second model parameters.
[0031] In one embodiment, a set of component images is input into a defect segmentation model to obtain a defect segmentation mask and defect category annotation results, including:
[0032] The images in the component image set are segmented to obtain the defect segmentation mask of the corresponding image; the defect segmentation mask includes multiple connected defect regions;
[0033] Defects are identified and labeled in connected defect regions to obtain defect category labeling results.
[0034] In one embodiment, the defect segmentation mask is refined to obtain a true defect mask, including:
[0035] The Level Set algorithm is used to refine the mask edges of the defect segmentation mask to obtain an enhanced mask.
[0036] Gaussian filtering is used to detect and enhance the boundary stability of the mask, thereby obtaining the boundary strength;
[0037] If the boundary strength is lower than the preset threshold, the enhanced mask is etched to obtain the real defect mask;
[0038] If the boundary strength is higher than the preset threshold, the enhanced mask is determined to be a real defect mask.
[0039] Secondly, this application also provides a small-sample defect detection system based on generative adversarial networks, comprising:
[0040] The image acquisition module is used to acquire multi-channel images of the target industrial parts and perform preprocessing to obtain a set of part images;
[0041] The defect detection module is used to input the component image set into the defect segmentation model to obtain the defect segmentation mask and defect category labeling results;
[0042] The mask processing module is used to refine the defect segmentation mask to obtain the real defect mask;
[0043] The inspection report module is used to generate a defect inspection report for the target industrial component based on the actual defect mask and defect category labeling results.
[0044] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of any of the above-described generative adversarial network-based few-sample defect detection methods.
[0045] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the above-described generative adversarial network-based few-sample defect detection methods.
[0046] The aforementioned few-sample defect detection method and system based on generative adversarial networks significantly improves the perception of defect features in input images by fusing image signals with different physical properties through multi-channel image acquisition, providing a richer foundation of structural and textural information for the model. The training process of the defect segmentation model overcomes the limitations of few-sample defects, simultaneously outputting location and category information, thus possessing stronger result description capabilities and being suitable for more complex quality analysis or defect statistics scenarios. By introducing a post-processing refinement mechanism for the mask boundary, including edge strength judgment and optional erosion processing, the realism and stability of the final defect mask are improved, reducing the possibility of misjudgment caused by false edges. The results are summarized into a defect detection report containing spatial location and semantic information, with clearly defined fields and structure, making it easy to use as an industrial data interface and supporting integration into factory quality inspection, maintenance, or traceability systems. Attached Figure Description
[0047] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0048] Figure 1 This is a flowchart illustrating the small-sample defect detection method based on generative adversarial networks of the present invention.
[0049] Figure 2 This is a schematic diagram of the training process of the defect segmentation model of the present invention;
[0050] Figure 3 This is a flowchart illustrating the steps of step S203.
[0051] Figure 4 This is a structural diagram of the small-sample defect detection system based on generative adversarial networks of the present invention. Detailed Implementation
[0052] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0053] In one embodiment, such as Figure 1 As shown, a few-sample defect detection method based on generative adversarial networks is provided. This embodiment illustrates the application of this method to a terminal. It is understood that this method can also be applied to a server, and further to a system including both a terminal and a server, and implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0054] S101. Multi-channel image acquisition is performed on the target industrial component, and preprocessing is carried out to obtain a component image set.
[0055] In a schematic representation, an industrial vision system acquires multiple image channels of the same target workpiece based on different physical properties or lighting conditions. These include a curvature map reflecting the surface geometry, a texture map showing surface texture variations, and a range map obtained using 3D imaging technology. These represent different visual dimensions of the workpiece, enabling the capture of richer defect features. Optionally, a multi-source imaging system can acquire multi-channel data and then fuse the three-channel images into a standard RGB format image.
[0056] Furthermore, preprocessing operations are performed on the image data. Specifically, image preprocessing mainly includes two aspects: noise suppression and structure enhancement. Gaussian filtering can effectively smooth the image background while preserving the boundary features of defective edge regions, thereby reducing the impact of invalid textures in the image on subsequent model interference. In addition, normalization and contrast enhancement strategies can be combined to ensure that the brightness and dynamic range of different images remain consistent, improving the training stability of the network.
[0057] S102. Input the component image set into the defect segmentation model to obtain the defect segmentation mask and defect category labeling results.
[0058] Indicatively, the defect segmentation model consists of a segmentation architecture and a classification architecture. The segmentation architecture employs DeepLabv3+ (a semantic segmentation model), Pyramid Attention Network (PAN), or similar dilated convolutional structures. After sufficient training on an industrial sample training set, it can perform pixel-level segmentation prediction on the input image and output a defect region mask. Furthermore, a lightweight classification architecture is introduced to perform feature discrimination on the segmented suspected defect regions. Architectures such as MobileNetV2 (mobile neural network) and EfficientNet (efficient neural network) are used to classify each defect, identifying whether it is a crack, dent, contamination, corrosion, or other different types of defects, thus forming a complete closed loop for defect localization and identification.
[0059] The defect segmentation model is trained by fusing real samples with defect samples synthesized by a generative adversarial network during the pre-training stage. By introducing high-quality synthetic defect images and accurate pseudo-annotation masks, it overcomes the risk of overfitting in small sample scenarios and achieves better generalization ability when real data is insufficient.
[0060] S103. Refine the defect segmentation mask to obtain the real defect mask.
[0061] Semantic segmentation networks often suffer from information loss during downsampling and feature fusion, leading to jagged or offset defect edge predictions. To illustrate this, an enhanced Level Set evolution algorithm is introduced to refine the contours of the initially generated mask. The Level Set method is a boundary fitting algorithm based on minimizing an energy function. It iteratively evolves based on the edge gradient information of the input mask to obtain a closed curve that more closely matches the actual boundary. Specifically, using the segmentation mask as the initial region, and combining image gradient and boundary strength control parameters, a finite number of evolution rounds are performed to achieve precise refinement of the defect contour, thereby improving the edge position and defect closure integrity of the true defect mask.
[0062] S104. Obtain a defect detection report for the target industrial component based on the actual defect mask and defect category labeling results; the defect detection report includes the location coordinates of each defect and its corresponding defect category.
[0063] Based on the refined real-world defect mask and corresponding defect category information, a structured defect detection report is generated. For example, the defect detection report can include not only the location coordinates of each defect region in the image, but also its identified defect category label and confidence score. Optionally, the defect detection report can be output in JSON, CSV, or PDF format according to application requirements, supporting integration with existing Quality Management Systems (QMS) or Manufacturing Execution Systems (MES) for scenario deployment.
[0064] In the aforementioned few-sample defect detection method based on generative adversarial networks (GANs), multi-channel image acquisition and preprocessing enhance the defect representation dimension, enabling the defect segmentation model to perceive the surface state of the target workpiece from multiple dimensions such as morphological structure, texture differences, and spatial depth, significantly improving the perceptibility of both minute and complex defects. Simultaneously, the Gaussian filtering and image enhancement strategies introduced in the preprocessing stage effectively suppress background noise and unify image style, providing a cleaner and more consistent image feature space for subsequent model input. The defect segmentation model adopted originates from an architecture that integrates real and synthetic data for joint training. During the training phase, it fully utilizes defect images generated by the GAN and pseudo-labels to form synthetic sample pairs, overcoming the problem of insufficient model training under traditional data scarcity conditions, allowing the model to maintain high recall and IoU accuracy during real-world deployment. When processing images, the defect segmentation model simultaneously outputs defect segmentation masks and defect category annotations, providing not only the location range of each defect but also clarifying its semantic category, constructing a collaborative inference channel from localization to semantic classification, ensuring both spatial accuracy and semantic integrity. Based on the original segmentation results, the Level Set algorithm is introduced for boundary evolution, and combined with Gaussian filtering to determine boundary strength. This effectively eliminates false alarm regions and weak response regions in the model's edge prediction, reducing the risk of false detections and missed detections, and enhancing the visualization quality of the mask. The defect detection report constructed based on the mask and classification results can provide clear location coordinates and category labels for each defect, which is beneficial for result tracking.
[0065] In one embodiment, such as Figure 2 As shown, the defect segmentation model is obtained in the following way:
[0066] S201. Multi-channel image acquisition is performed on the target industrial component with defects, and preprocessing is carried out to obtain a defect image set.
[0067] To illustrate, in scenarios with small sample sizes, priority is given to acquiring real images containing defects. If the initial defect samples are extremely few, manual or semi-automatic segmentation methods are allowed to mark the central region of the image as the defect area, serving as the basis for subsequent sample synthesis. To ensure the structural richness and completeness of feature representation, a multi-channel image acquisition method is also used to collect defect samples, for example. Specifically, photometric stereo imaging technology is used to acquire images of the target component. The same workpiece surface is illuminated by multiple light sources with different angles and intensities to reconstruct its local curvature, texture distribution, and depth information, accurately acquiring defect information while enriching the number of real samples.
[0068] S202. Based on the generative adversarial network, samples are synthesized from the defective image set to obtain the synthesized image set; the generative adversarial network includes the StyleGAN2 generator, the PatchGAN discriminator, and the DiffAugment enhancement structure.
[0069] Indicatively, the Generative Adversarial Network (GAN) uses StyleGAN2 as the generator architecture and the standard PatchGAN structure as the discriminator, possessing the ability to discriminate the consistency of local image structures. Furthermore, during training, DiffAugment (a differentiable enhancement algorithm) is introduced, making the image enhancement operation differentiable on both the generator and discriminator input paths, thereby improving training stability and the diversity of synthesized images in small-sample scenarios. During training, adversarial training is performed between the generator and discriminator using the Adam optimizer, with a learning rate lr = 2e-4, gradient β1 = 0.5, batch size set to 16, and training epochs ranging from 10,000 to 20,000 steps, using FID (Fréchet Inception Distance) as the convergence metric. After training, the GAN is used to generate a batch of synthesized images with a style consistent with real defect images.
[0070] S203. Use a pre-trained segmentation model to generate pseudo-labels for synthetic samples from the synthetic image set to obtain synthetic sample pairs.
[0071] As an illustration, a pre-trained semantic segmentation model can be used, such as a DeepLabv3+ model pre-trained on a similar data domain, or an initial network trained with a very small number of real samples. Pixel-level segmentation is performed on each synthetic image to obtain an initial mask, which yields the defect distribution of the synthetic samples, forming a synthetic image and its pseudo-label mask pair, thus constituting a synthetic sample pair.
[0072] S204. Construct a fusion training set based on the defect image set and synthetic sample pairs, and train a network architecture based on DeepLabV3+ and MobileNetV2 to obtain a defect segmentation model.
[0073] A defect image set was fused with synthetic sample pairs to train the backbone defect segmentation model. Illustratively, a two-stage collaborative network architecture based on DeepLabv3+ and MobileNetV2 was employed. DeepLabv3+ served as the backbone semantic segmentation model, responsible for extracting image spatial structure and contextual information, and predicting defect region masks. MobileNetV2 acted as a lightweight image classifier, classifying each defect region cropped from the segmentation results. During training, the segmentation model used a combined loss function of Dice Loss and Binary Cross Entropy (BCE) to jointly optimize mask overlap accuracy and pixel classification accuracy; the classification model used a cross-entropy loss function for optimization. The Adam optimizer was used during training, with a learning rate of 1e-4, batch size of 8–16, and 100–200 training epochs, depending on the convergence of the validation set metrics.
[0074] After training, the model will be able to perform pixel-level defect localization and category recognition on any input industrial image, forming a defect segmentation model. The training method of this model has the advantage of making full use of a small number of real samples and improving performance with the help of synthetic data.
[0075] In one embodiment, multi-channel image acquisition is performed on the defective target industrial component, and preprocessing is carried out to obtain a defect image set, including:
[0076] S21. Use photometric stereo method to acquire images of the target component to obtain channel images; the channel images include curvature map, texture map and distance map.
[0077] In a schematic manner, photometric stereoscopic imaging technology is used to acquire images of the target component. The same workpiece surface is illuminated by multiple light sources with different angles and intensities to reconstruct its local curvature, texture distribution, and depth information.
[0078] S22. Perform Gaussian filtering on the channel images, and apply image normalization and contrast stretching to obtain the filtered and enhanced image set.
[0079] Indicatively, Gaussian filtering is used to smooth the image, reducing the impact of random noise while preserving edge and texture features. For example, a 5×5 kernel with a standard deviation σ = 1.0 is typically chosen. Further, image normalization is applied to compress pixel values to the [0,1] or [-1,1] range to fit the neural network input specification, and contrast stretching is performed on the image to enhance the grayscale difference between defective and normal areas.
[0080] S23. For each image in the filtered and enhanced image set, execute multiple preset enhancement strategies in a random combination to obtain a defective image set; the enhancement strategies include random rotation, flipping, scaling, brightness perturbation, and adding mild Gaussian noise.
[0081] To illustrate, in order to improve the model's generalization ability and expand the training data scale, various conventional enhancement strategies are applied to the preprocessed image set, including but not limited to image rotation, horizontal and vertical flipping, random scaling, brightness perturbation, and Gaussian noise injection. The enhancement operations are applied to each image in a random combination; for example, each image generates 9 enhanced versions, expanding the initial defect image set from N images to 10N images, forming a complete defect image set for training the generative model.
[0082] In one embodiment, such as Figure 3 As shown, a pre-trained segmentation model is used to generate pseudo-labels for synthetic samples from the synthetic image set, resulting in synthetic sample pairs, including:
[0083] S301. Use the pre-trained DeepLabV3+ network to perform preliminary segmentation on the synthesized image to obtain a coarse mask.
[0084] Indicatively, a pre-trained DeepLabV3+ network is used to perform preliminary semantic segmentation on the synthetic image. The pre-trained DeepLabV3+ model can be derived from an existing dataset in the same industrial inspection field or pre-trained on a small number of real samples. Its purpose is to generate a rough defect mask for the synthetic image. Since the synthetic image itself is generated by the StyleGAN2 model, it possesses a certain degree of image realism and structural features. The pre-trained DeepLabV3+ segmentation network can output preliminary defect boundaries based on the image surface texture and region consistency.
[0085] S302. Use the Level Set algorithm to refine the contour of the coarse mask to obtain the refined mask.
[0086] Indicatively, the Level Set algorithm uses a coarse mask as the initial contour function and combines image grayscale, edge intensity, and gradient direction as evolutionary drivers, performing a finite number of iterations. During the evolution process, the boundary gradually converges from a broken, jagged state to a smoother, more continuous curve shape, thus forming a refined mask with a stable structure and consistent shape.
[0087] S303. Use morphology to fill the holes in the refined mask to obtain a defect mask.
[0088] Since Level Set evolution may create holes or empty spots in local strong edge regions, in order to ensure the connectivity and geometric integrity of the mask, image morphological operations are further used to fill holes and correct edges in the thinned mask, including morphological processing such as region dilation, closing operation and area filtering. This can effectively remove small-sized pseudo-predicted regions and fill the holes inside the segmentation, so that the defect region has a closed structure and a clear outer contour. This makes the defect mask not only fit the actual shape of the defect region on the edge, but also more meaningful in terms of structural connectivity and semantic expression.
[0089] S304. Construct synthetic sample pairs using defect masks and corresponding synthetic images.
[0090] By pairing the corrected defect mask with its corresponding synthetic image, standard-format synthetic sample pairs are constructed. Each sample pair includes a synthetic image and a structurally complete defect mask. The synthetic sample pairs are structurally highly consistent with real manually labeled samples and can be seamlessly incorporated into the training set for segmentation model training. For example, introducing 30%–40% of synthetic sample pairs into the total sample size can effectively improve the robustness and generalization ability of the model, especially in scenarios with rare defects or very few categories.
[0091] In one embodiment, a fused training set is constructed based on a defect image set and synthetic sample pairs, and a network architecture based on DeepLabV3+ and MobileNetV2 is trained to obtain a defect segmentation model, including:
[0092] S31. Label a portion of the defect image set with real defect masks to obtain real sample pairs; the real sample pairs include real defect images and their corresponding real defect masks.
[0093] As an illustration, a subset of real defect images are manually labeled to establish highly reliable real sample pairs. These real sample pairs can provide initial stability and supervision strength during training, playing a crucial role, especially in the early stages of model training. Optionally, each real sample pair includes each real image and its corresponding pixel-level defect mask. The labeling method can employ multi-annotator consensus or expert review mechanisms to ensure labeling quality.
[0094] S32. Mix real sample pairs and synthetic sample pairs according to a preset ratio and perform enhancement processing to obtain a fused training set.
[0095] Real sample pairs are mixed with the aforementioned synthetic sample pairs at a preset ratio. For example, the proportion of real samples is controlled at 60%–70%, and the proportion of synthetic samples is controlled at 30%–40%. The mixed sample set will then undergo uniform cropping, normalization, and image enhancement processing to prevent sample distribution shifts from affecting training. Image enhancement strategies include operations such as rotation, flipping, translation, and blurring to ensure that the model maintains its discriminative ability under deformation, occlusion, and blurring conditions.
[0096] S33. According to the preset training objective, train the DeepLabV3+ backbone semantic segmentation architecture based on the fused training set to obtain the first model parameters.
[0097] Intuitively, a DeepLabV3+ backbone semantic segmentation network is trained using a fused training set. The network architecture includes an encoder using an Xception or ResNet101 backbone and a decoder employing a combination of multi-scale dilated convolutions and upsampling structures to achieve pixel-level defect mask output. The training loss function is a composite of Dice Loss and BCE to balance the few-shot class problem and pixel overlap. The optimizer uses Adam, with an initial learning rate of 1e-4, batch size of 8–16, and 100–200 training epochs. The convergence criteria are the validation set IoU (Intersection over Union) and mDice (Multi-class Dice Coefficient).
[0098] S34. Based on the preset defect category labels, train the MobileNetV2 classification network according to the connected defect regions to obtain the second model parameters; the connected defect regions are obtained by segmenting and fusing images in the training set using the DeepLabV3+ backbone semantic segmentation architecture.
[0099] The connected defect regions obtained from the images in the training set segmented and fused using the DeepLabV3+ backbone semantic segmentation architecture are cropped and normalized to a fixed size, and then input into a predefined MobileNetV2 lightweight classification network. For example, the classification network uses preset defect category labels such as cracks, pores, pits, and corrosion as the supervision target, and is trained using the cross-entropy loss function. Transfer tuning can be completed in approximately 100 iterations. Optionally, if no labels are available, a two-class coarse classification strategy can be used to label connected defect regions as either defects or non-defects.
[0100] S35. Obtain the defect segmentation model based on the first model parameters and the second model parameters.
[0101] Based on the first model parameters of the semantic segmentation model and the second model parameters of the defect classification model, they are integrated into a complete defect segmentation model, which has both pixel-level segmentation capability and region-level classification capability. It can perform end-to-end defect recognition tasks on any input industrial image and output structured defect location, contour and category information to meet the application needs of intelligent quality inspection, defect screening and other applications.
[0102] In one embodiment, a set of component images is input into a defect segmentation model to obtain a defect segmentation mask and defect category annotation results, including:
[0103] S41. Segment the images in the component image set to obtain the defect segmentation mask of the corresponding image; the defect segmentation mask includes multiple connected defect regions.
[0104] Indicatively, after the defect segmentation model receives the image set of the component to be tested, it can achieve pixel-level prediction of potential defect regions in the image based on the dilated convolution and multi-scale context awareness capabilities of the DeepLabV3+ backbone architecture. The output defect segmentation mask is a binary image of the same size as the input image, where foreground pixels are marked as defect regions and background pixels represent normal parts. Since defects often occur locally, the segmentation mask generally contains multiple spatially connected regions, called connected defect regions. Structurally, they usually appear as spots, stripes, or cracks, characterizing possible flaws in the target component.
[0105] S42. Identify and label the connected defect regions to obtain the defect category labeling results.
[0106] Furthermore, each connected defect region is cropped from the original image and standardized to a fixed-size image, which is then fed into the trained MobileNetV2 classification model. This model boasts advantages such as lightweight design and high concurrency, enabling it to quickly analyze the texture, edge morphology, and local contrast features of the input image to determine its defect category, including cracks, corrosion, foreign objects, and bubbles. The classification model outputs the category label and confidence value for each region, thereby generating defect category annotation results corresponding to the segmentation mask.
[0107] In one embodiment, the defect segmentation mask is refined to obtain a true defect mask, including:
[0108] S51. Use the Level Set algorithm to refine the mask edges of the defect segmentation mask to obtain an enhanced mask.
[0109] This illustration demonstrates how the Level Set contour evolution algorithm is used to refine the edges of a defect segmentation mask. The algorithm uses the segmentation mask as the initial boundary and constructs an energy function based on the image gradient field to drive the curve evolution process. Through multiple iterations, the edge contour gradually evolves from a jagged, burr-like state to a smooth, closed, continuous boundary, making it particularly suitable for defect areas with incomplete or occluded contours in industrial images.
[0110] S52. Gaussian filtering is used to detect and enhance the boundary stability of the mask, and the boundary strength is obtained.
[0111] To illustrate, a Gaussian filter map is introduced as a benchmark to perform stability testing on the boundary strength of the enhancement mask. Specifically, a Gaussian filter operation is reapplied to the image processed by the Level Set, and the edge gradient magnitude is calculated after generating a smooth image. For each connected region boundary, the corresponding edge strength distribution on the Gaussian filter map is extracted, and the average edge magnitude is calculated as the stability score of the defect boundary.
[0112] S53. If the boundary strength is lower than the preset threshold, the enhancement mask is etched to obtain the real defect mask.
[0113] S54. If the boundary strength is higher than the preset threshold, then the enhanced mask is determined to be the real defect mask.
[0114] The process involves judgment and branching based on boundary stability scores and preset stability thresholds. If the detected boundary strength is below the threshold, it indicates that the region may be a false alarm, a false edge, or a region with high prediction uncertainty, and an erosion operation is performed on the enhancement mask. Erosion is a typical image morphology operation that can remove isolated pixels, shrink boundaries, and eliminate disconnected small regions, thereby removing unstable or unrealistic prediction components. The eroded mask is the structurally optimized true defect mask. If the boundary strength is above the threshold, the current enhancement mask is considered to have sufficient edge confidence and geometric stability and can be directly identified as a true defect mask without additional erosion processing.
[0115] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0116] Based on the same inventive concept, this application also provides a generative adversarial network-based small-sample defect detection system for implementing the above-described generative adversarial network-based small-sample defect detection method. The solution provided by this system is similar to the implementation described in the above method. Therefore, the specific limitations in one or more embodiments of the generative adversarial network-based small-sample defect detection system provided below can be found in the limitations of the generative adversarial network-based small-sample defect detection method described above, and will not be repeated here.
[0117] In one exemplary embodiment, such as Figure 4 As shown, a few-sample defect detection system based on generative adversarial networks is provided, including:
[0118] The image acquisition module 401 is used to acquire multi-channel images of the target industrial component and perform preprocessing to obtain a component image set;
[0119] The defect detection module 402 is used to input the component image set into the defect segmentation model to obtain the defect segmentation mask and defect category labeling results;
[0120] The mask processing module 403 is used to refine the defect segmentation mask to obtain a real defect mask;
[0121] The inspection report module 404 is used to generate a defect inspection report for the target industrial component based on the actual defect mask and defect category labeling results.
[0122] In one embodiment, it further includes:
[0123] The image acquisition module 401 is also used to acquire multi-channel images of the target industrial parts with defects and to preprocess them to obtain a defect image set;
[0124] The sample synthesis module is used to synthesize samples from a set of defective images based on a generative adversarial network to obtain a synthesized image set.
[0125] The defect labeling module is used to generate synthetic sample pseudo-labels for the synthetic image set using a pre-trained segmentation model, thus obtaining synthetic sample pairs.
[0126] The model training module is used to construct a fused training set based on the defect image set and synthetic sample pairs, and train a network architecture based on DeepLabV3+ and MobileNetV2 to obtain a defect segmentation model.
[0127] In one embodiment, it further includes:
[0128] The image acquisition module 401 is also used to acquire images of the target component using photometric stereo method to obtain channel images; the channel images include curvature maps, texture maps and distance maps;
[0129] The image processing module is used to perform Gaussian filtering on the channel images and apply image normalization and contrast stretching to obtain a filtered and enhanced image set.
[0130] The image processing module is also used to execute multiple preset enhancement strategies on each image of the filtered and enhanced image set in a random combination to obtain a defective image set; the enhancement strategies include random rotation, flipping, scaling, brightness perturbation, and adding mild Gaussian noise.
[0131] In one embodiment, it further includes:
[0132] The image segmentation module is used to perform preliminary segmentation of the synthesized image using a pre-trained DeepLabV3+ network to obtain a coarse mask;
[0133] The mask processing module 403 is also used to refine the contour of the rough mask using the Level Set algorithm to obtain a refined mask;
[0134] The mask processing module 403 is also used to fill the mask holes in the refined mask using morphology to obtain a defect mask;
[0135] The sample synthesis module is also used to construct synthetic sample pairs using defect masks and corresponding synthetic images.
[0136] In one embodiment, it further includes:
[0137] The training set construction module is used to annotate a portion of the defect image set with real defect masks to obtain real sample pairs; the real sample pairs include real defect images and their corresponding real defect masks.
[0138] The training set construction module is also used to mix real sample pairs and synthetic sample pairs according to a preset ratio and perform enhancement processing to obtain a fused training set;
[0139] The model training module is also used to train the DeepLabV3+ backbone semantic segmentation architecture according to the preset training objectives and the fused training set to obtain the first model parameters;
[0140] The model training module is also used to train the MobileNetV2 classification network based on the preset defect category labels and the connected defect regions to obtain the second model parameters; the connected defect regions are obtained by segmenting and fusing images in the training set using the DeepLabV3+ backbone semantic segmentation architecture;
[0141] The model training module is also used to obtain a defect segmentation model based on the first model parameters and the second model parameters.
[0142] In one embodiment, it further includes:
[0143] The segmentation module is used to segment the images in the component image set to obtain the defect segmentation mask of the corresponding image; the defect segmentation mask includes multiple connected defect regions;
[0144] The identification module is used to identify and label connected defect regions, and obtain defect category labeling results.
[0145] In one embodiment, it further includes:
[0146] The mask processing module 403 is also used to refine the mask edges of the defect segmentation mask using the Level Set algorithm to obtain an enhanced mask;
[0147] The mask processing module 403 is also used to enhance the boundary stability of the mask by using Gaussian filtering to obtain the boundary strength;
[0148] The mask processing module 403 is also used to perform etch processing on the enhanced mask if the boundary strength is lower than a preset threshold to obtain a real defect mask;
[0149] The mask processing module 403 is also used to determine that the enhanced mask is a real defect mask if the boundary strength is higher than a preset threshold.
[0150] In one embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps in the above method embodiments.
[0151] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.
[0152] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0153] The above-described embodiments are merely illustrative of several implementation methods of the embodiments of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the embodiments of this application, and these modifications and improvements all fall within the protection scope of the embodiments of this application.
Claims
1. A small-sample defect detection method based on generative adversarial networks, characterized in that, The method includes: Multi-channel image acquisition and preprocessing are performed on the target industrial component to obtain a component image set; The component image set is input into the defect segmentation model to obtain the defect segmentation mask and defect category labeling results; The defect segmentation mask is refined to obtain a true defect mask; A defect detection report for the target industrial component is obtained based on the actual defect mask and the defect category labeling results; the defect detection report includes the location coordinates of each defect and its corresponding defect category.
2. The method according to claim 1, characterized in that, The defect segmentation model is obtained in the following way: Multi-channel image acquisition and preprocessing are performed on the defective target industrial component to obtain a defect image set; Based on a generative adversarial network, a sample is synthesized from the defective image set to obtain a synthesized image set; the generative adversarial network includes a StyleGAN2 generator, a PatchGAN discriminator, and a DiffAugment enhancement structure. Using a pre-trained segmentation model, synthetic sample pseudo-labels are generated from the synthetic image set to obtain synthetic sample pairs; A fusion training set is constructed based on the defect image set and the synthetic sample pairs, and a network architecture based on DeepLabV3+ and MobileNetV2 is trained to obtain the defect segmentation model.
3. The method according to claim 2, characterized in that, The process involves acquiring multi-channel images of the defective target industrial component and preprocessing them to obtain a defect image set, including: The target component is imaged using a photometric stereo method to obtain a channel image; the channel image includes a curvature map, a texture map, and a distance map. The channel image is subjected to Gaussian filtering, and image normalization and contrast stretching are applied to obtain the filtered and enhanced image set. For each image in the filtered and enhanced image set, multiple preset enhancement strategies are executed in a random combination to obtain the defective image set; the enhancement strategies include random rotation, flipping, scaling, brightness perturbation, and adding mild Gaussian noise.
4. The method according to claim 2, characterized in that, The process of generating synthetic sample pseudo-labels from the synthetic image set using a pre-trained segmentation model to obtain synthetic sample pairs includes: The synthesized image is initially segmented using a pre-trained DeepLabV3+ network to obtain a coarse mask; The coarse mask is refined using the Level Set algorithm to obtain a refined mask. The morphological method is used to fill the mask holes in the refined mask to obtain a defect mask; The synthetic sample pair is constructed using the defect mask and the corresponding synthetic image.
5. The method according to claim 2, characterized in that, The step of constructing a fusion training set based on the defect image set and the synthetic sample pairs, and training a network architecture based on DeepLabV3+ and MobileNetV2 to obtain the defect segmentation model includes: A portion of the defect image set is annotated with a real defect mask to obtain real sample pairs; the real sample pairs include real defect images and their corresponding real defect masks. The real sample pairs and the synthetic sample pairs are mixed in a preset ratio and then enhanced to obtain the fused training set. According to the preset training objective, the DeepLabV3+ backbone semantic segmentation architecture is trained based on the fused training set to obtain the first model parameters; Based on preset defect category labels, a MobileNetV2 classification network is trained according to the connected defect regions to obtain the second model parameters; the connected defect regions are obtained by segmenting and fusing images in the training set using the DeepLabV3+ backbone semantic segmentation architecture; The defect segmentation model is obtained based on the first model parameters and the second model parameters.
6. The method according to claim 2, characterized in that, The step of inputting the component image set into the defect segmentation model to obtain the defect segmentation mask and defect category labeling results includes: The images in the component image set are segmented to obtain the defect segmentation mask for the corresponding image; the defect segmentation mask includes multiple connected defect regions; The connected defect region is identified and labeled to obtain the defect category labeling result.
7. The method according to claim 6, characterized in that, The refinement process of the defect segmentation mask to obtain the true defect mask includes: The Level Set algorithm is used to refine the mask edges of the defect segmentation mask to obtain an enhanced mask. The boundary stability of the enhanced mask is detected using Gaussian filtering to obtain the boundary strength; If the boundary strength is lower than a preset threshold, the enhanced mask is etched to obtain the real defect mask. If the boundary strength is higher than a preset threshold, then the enhanced mask is determined to be the real defect mask.
8. A small-sample defect detection system based on generative adversarial networks, characterized in that, The system includes: The image acquisition module is used to acquire multi-channel images of the target industrial parts and perform preprocessing to obtain a set of part images; The defect detection module is used to input the component image set into the defect segmentation model to obtain the defect segmentation mask and defect category labeling results; The mask processing module is used to refine the defect segmentation mask to obtain a real defect mask; The inspection report module is used to generate a defect inspection report for the target industrial component based on the actual defect mask and the defect category labeling results.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.