Defect sample generation method and device

By querying the original image and segmenting mask expansion operations of defective products in industrial manufacturing, defect diffusion editing images are generated, which solves the problem of scarcity of defect samples and achieves the effect of generating high-quality defect samples on the basis of limited samples.

CN120219877APending Publication Date: 2025-06-27GUANGDONG AOPUTE TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510276754.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the field of industrial manufacturing, defective products are rare and the production process is highly optimized, making it difficult to collect sufficient defect samples to train deep learning models.

Method used

By obtaining the prompt image set of querying the original image set and the defective image, the segmentation mask of the defective image is obtained using image query, the expansion operation is performed, and the defect diffusion editing image is generated based on the start and end points of multiple rays, thereby generating high-quality defect samples.

Benefits of technology

Generate large batches of high-quality defect samples based on limited defect samples to meet the training needs of deep learning models and improve the accuracy and efficiency of defect detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219877A_ABST
    Figure CN120219877A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a defect sample generation method and device. In the embodiment of the invention, a query original image set and a prompt small image set of a defect image can be obtained, and image query is performed on the query original image set based on each prompt small image in the prompt small image set. Obtaining a first segmentation mask of a defect image of each query original image in the query original image set and a corresponding defect category; performing expansion operation on the segmentation mask of the defect image of each query original image in the query original image set, and generating a second segmentation mask of the expanded defect image of each query original image in the query original image set; and generating a defect diffusion editing image corresponding to each query original image in the query original image set based on a starting point and an end point determined by a plurality of rays uniformly emitted in a plurality of specified directions in the first segmentation mask and the second segmentation mask, and the second segmentation mask, and generating a defect sample based on the defect diffusion editing image and the corresponding defect category corresponding to each query original image in the query original image set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of deep learning technology, and in particular, to a method and device for defective samples. Background Art

[0002] In the field of industrial manufacturing, defect detection, as a key technology, is crucial for monitoring the production process of industrial products and ensuring product quality, and is used to achieve defect identification and precise positioning of industrial products. However, in the actual production environment, defective products are often relatively rare, and coupled with the highly optimized production process, it is difficult to collect sufficient defective samples to train deep learning models for defect detection.

[0003] Therefore, how to generate a large number of high-quality defective samples based on limited defective samples to meet the training of relevant models has become a formidable challenge in industrial defect detection. Summary of the Invention

[0004] Multiple aspects of this application provide a method and device for generating defective samples, which are used to generate a large number of high-quality defective samples based on limited defective samples to meet the training of relevant models.

[0005] An embodiment of this application provides a method for generating defective samples, including: Obtaining a query original image set and a hint thumbnail image set of defective images, and performing image queries on the query original image set based on each hint thumbnail in the hint thumbnail image set to obtain a first segmentation mask of the defective image of each query original image in the query original image set and the corresponding defective category; Performing a dilation operation on the segmentation mask of the defective image of each query original image in the query original image set to generate a second segmentation mask of the expanded defective image of each query original image in the query original image set; Generating a defective diffusion editing image corresponding to each query original image in the query original image set based on the starting and ending points determined by multiple rays uniformly emitted in multiple specified directions in the first segmentation mask and the second segmentation mask, and the second segmentation mask, and generating defective samples based on the defective diffusion editing image corresponding to each query original image in the query original image set and the corresponding defective category.

[0006] An embodiment of this application also provides a device for generating defective samples, including: A mask acquisition module, configured to obtain a query original image set and a hint thumbnail image set of defective images, and perform image queries on the query original image set based on each hint thumbnail in the hint thumbnail image set to obtain a first segmentation mask of the defective image of each query original image in the query original image set and the corresponding defective category; A mask generation module, configured to perform a dilation operation on the segmentation masks of the defective images of each query original image in the query original image set, and generate second segmentation masks of the expanded defective images of each query original image in the query original image set; A sample generation module, configured to generate a defective diffusion edited image and a corresponding defective category corresponding to each query original image in the query original image set based on the starting points and ending points determined by multiple rays uniformly emitted in multiple specified directions in the first segmentation mask and the second segmentation mask, and the second segmentation mask, and generate defective samples based on the defective diffusion edited images corresponding to each query original image in the query original image set.

[0007] An embodiment of the present application further provides an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. The processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps in the defective sample generation method as described above are executed.

[0008] An embodiment of the present application further provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor is caused to implement the steps in the defective sample generation method provided by the embodiment of the present application.

[0009] The defective sample generation method provided by the embodiment of the present application can obtain a query original image set and a hint thumbnail image set of defective images, and perform an image query from the query original image set based on each hint thumbnail in the hint thumbnail image set to obtain the first segmentation mask and the corresponding defective category of the defective image of each query original image in the query original image set. Then, a dilation operation is performed on the segmentation mask of the defective image of each query original image in the query original image set to generate a second segmentation mask of the expanded defective image of each query original image in the query original image set. Finally, based on the starting points and ending points determined by multiple rays uniformly emitted in multiple specified directions in the first segmentation mask and the second segmentation mask, and the second segmentation mask, a defective diffusion edited image corresponding to each query original image in the query original image set is generated, and defective samples are generated based on the defective diffusion edited images corresponding to each query original image in the query original image set. On the one hand, by using the query original image that matches the hint thumbnail, the first segmentation mask and the corresponding defective category of the defective image similar to the hint thumbnail in the query original image are determined. On the other hand, the defective area is edited by performing a dilation operation and defective diffusion editing on the first segmentation mask in sequence, and new defective samples that do not deviate from the defective category of the original query original image are generated based on the edited defective images, which can realize generating a large number of high-quality defective samples based on a limited number of defective samples to meet the training of related models. Description of the Drawings

[0010] The accompanying drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings: Figure 1 It is a schematic flowchart of a method for generating defect samples provided for an exemplary embodiment of the present application; Figure 2 It is a schematic structural diagram of an image segmentation network in the method for generating defect samples provided for an embodiment of the present application; Figure 3 It is a schematic diagram of region filtering for a segmentation mask in the method for generating defect samples provided for an embodiment of the present application; Figure 4 It is a schematic diagram of determining a starting point for a segmentation mask and expanding based on the starting point in the method for generating defect samples provided for an embodiment of the present application; Figure 5 It is a schematic flowchart of point tracking in a method for generating defect samples provided for an embodiment of the present application; Figure 6 It is a schematic diagram of determining a segmentation mask for the generated defect samples in the defect detection method provided for an exemplary embodiment of the present application; Figure 7 It is a schematic diagram of the effect of restoring the generated defect samples to the original image in the method for generating defect samples provided for an exemplary embodiment of the present application; Figure 8 It is a schematic structural diagram of a device for generating defect samples provided for an exemplary embodiment of the present application; Figure 9 It is a schematic structural diagram of an electronic device provided for an exemplary embodiment of the present application. Detailed implementation manners

[0011] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0012] As described in the background art, in order to generate a large number of high-quality defect samples based on a limited number of defect samples, in the related art, when training a deep learning model, in the face of limited training samples, data augmentation methods are generally used to perform transformation operations such as rotation, scaling, cropping, and flipping on the existing training samples to generate more training samples. At the same time, the related art also uses methods such as adding noise, adjusting brightness and contrast to expand the dataset of training samples and improve the generalization ability of the deep learning model. On this basis, the related art also uses the method of transfer learning, using the pre-trained model obtained by training on a large dataset, transferring knowledge to the current deep learning training task, using the weights of the pre-trained model as the initial weights of the model in this deep learning training task, and then fine-tuning the network parameters of the pre-trained model. However, in industrial defect detection, the image backgrounds of defect samples are generally similar. The data augmentation method cannot effectively enhance the local area (such as the defective area) of the defect samples in a targeted manner.

[0013] In recent years, with the continuous development of generative large models, using generative technology to generate high-quality defect samples has become a popular research direction. For example, the interactive image editing method DragGAN based on Generative Adversarial Networks (GAN) enables users to directly drag specific points in the image through the "point dragging" technology, so as to achieve pixel-level geometric editing and realize realistic and interactive image editing.

[0014] However, since DragGAN relies on a pre-trained GAN model, its generalization ability is limited, and it is difficult to achieve ideal editing effects when dealing with complex multi-object scenes or images with diverse styles. To overcome these limitations, DragDiffusion introduces a generative method based on diffusion models, which significantly improves the generality and accuracy of the editing process by optimizing the latent features of a single time step and combining identity-preserving fine-tuning technology. However, when using DragDiffusion to generate defect samples, it usually requires manual editing of the defect areas one by one, which greatly affects the generation speed of the defect samples.

[0015] In view of this, embodiments of the present application propose a method and device for generating defective samples, which can obtain a query original image set and a hint thumbnail image set of defective images, and perform image queries on the query original image set based on each hint thumbnail in the hint thumbnail image set to obtain the first segmentation mask and corresponding defect category of the defective images of each query original image in the query original image set; perform a dilation operation on the segmentation masks of the defective images of each query original image in the query original image set to generate a second segmentation mask of the expanded defective images of each query original image in the query original image set; generate a defective diffusion edited image corresponding to each query original image in the query original image set based on the starting and ending points determined by multiple rays uniformly emitted in multiple specified directions in the first segmentation mask and the second segmentation mask, and generate defective samples based on the defective diffusion edited images corresponding to each query original image in the query original image set, which can generate a large number of high-quality defective samples on the basis of limited defective samples to meet the training of relevant models.

[0016] The technical solutions provided by the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0017] Figure 1 It is a schematic flowchart of a method for generating defective samples provided by an exemplary embodiment of the present application. As Figure 1 shown, the method includes: Step 110, obtain a query original image set and a hint thumbnail image set of defective images, and perform image queries on the query original image set based on each hint thumbnail in the hint thumbnail image set to obtain the first segmentation mask and corresponding defect category of the defective images of each query original image in the query original image set.

[0018] Among them, each query original image in the query original image set is an original image that may contain a defective image. It should be noted that at the initial stage, the positions and shapes of the defective images included in each query original image are not determined. The hint thumbnail image set of defective images is a pre-collected set of defective images containing different defect types, corresponding defect categories, and segmentation masks of defect regions.

[0019] As an exemplary embodiment, performing image queries on the query original image set based on each hint thumbnail in the hint thumbnail image set to obtain the first segmentation mask and corresponding defect category of the defective images of each query original image in the query original image set includes: Perform image queries on the query original image set based on each hint thumbnail in the hint thumbnail image set through an image retrieval model to obtain the initial segmentation mask and corresponding defect category of the defective images of each query original image in the query original image set; Perform noise filtering operations on the initial segmentation masks of the defective images of each query original image in the query original image set to obtain the first segmentation masks of the defective images of each query original image in the query original image set.

[0020] Among them, the defect category corresponding to the defective image of each query original image in the query original image set is determined by the corresponding hint thumbnail. Specifically, it can be determined based on the defect category corresponding to the hint thumbnail that matches each query original image.

[0021] Exemplarily, the image retrieval model can be implemented based on the existing fast matching retrieval algorithm. This fast matching retrieval algorithm has powerful feature matching and retrieval capabilities, such as Figure 2 In step1 of the matching and retrieval stage, it is the main framework of the image retrieval model constructed based on the fast matching retrieval algorithm. Its input mainly consists of the hint thumbnail image set Det_Object of the defective image and the query original image set Det_Set. According to each hint thumbnail in the hint thumbnail image set Det_Object of the defective image, this fast matching retrieval model can quickly match and search for the target objects (i.e., the corresponding query original images) similar to each hint thumbnail in the hint thumbnail image set Det_Object of the defective image from the query original image set Det_Set, and segment out the initial segmentation mask Mask containing the defective area and its corresponding defect category (this defect category is determined by the defect category in the hint thumbnail similar to the query original image) in the query original images similar to each hint thumbnail.

[0022] Specifically, the image retrieval model is used to take the feature vector representation of the specified dimension of the input image by the defect feature encoding network as the hint information, and retrieve the defect location and defect category of the input image from the input image. The image retrieval model can include three modules, specifically the image encoding module, the defect hint feature encoding module, and the feature decoding module. Among them, the image encoding module aims to map the input image to be segmented into the image feature space. The defect hint feature encoding module is used to extract features from the input hint thumbnail Det_Object to obtain the feature vector of the input hint thumbnail Det_Object, and use the feature vector of the input hint thumbnail Det_Object as the hint information of the image retrieval model and input it into the image retrieval model. The feature decoding module mainly has two functions. One is to integrate the two feature vectors embedding respectively output by the image encoding module and the defect hint feature encoding module of the image retrieval model, and then decode the initial segmentation mask Mask (i.e., the image retrieval result) from the feature map of these two feature vectors embedding.

[0023] Therefore, for a given query original image set Det_Set, first, it is necessary to identify and obtain the hint thumbnail image set Det_Object of defective images. Then, input the query original image set Det_Set and the hint thumbnail image set Det_Object of defective images into the image retrieval model for inference, and the initial segmentation masks of the defective regions in each query original image in the query original image set Det_Set can be obtained. Finally, due to the possible small-area noise points in the partial segmentation results obtained in the above steps, noise filtering methods such as area screening can be used to filter out the small-area noise points in the initial segmentation masks to obtain the first segmentation mask Image Mask of the defective images of each query original image in the final relatively accurate query original image set.

[0024] Figure 3 This is a schematic diagram of region filtering for the segmentation mask in the defective sample generation method provided by an embodiment of this application. Among them, (a) is an example of a query original image, (b) is the initial segmentation mask obtained based on the query original image (a). It can be seen that there are small-area noise points near the initial segmentation mask in (b). (c) is the first segmentation mask of the defective image of the query original image obtained after filtering out the noise points near the initial segmentation mask in (b) by the method of area screening. Among them, the white area in (c) is the segmentation result, that is, the first segmentation mask.

[0025] Step 120: Perform a dilation operation on the segmentation masks of the defective images of each query original image in the query original image set to generate the second segmentation mask of the expanded defective images of each query original image in the query original image set.

[0026] Among them, performing a dilation operation on the segmentation masks of the defective images of each query original image in the query original image set can specifically perform a dilation operation on the first segmentation masks of the defective images of each query original image in the query original image set to generate the second segmentation masks of the expanded defective images of each query original image in the query original image set. That is to say, the second segmentation masks of the defective images of each query original image are generated based on their corresponding first segmentation masks.

[0027] Exemplarily, perform a dilation operation on the first segmentation mask Imagemask of the defective images of each query original image in the query original image set. First, the size of the first segmentation mask Image mask can be calculated. Then, one-tenth of the short side size of the first segmentation mask Image mask is used as the dilation coefficient. Finally, use the dilation operator of the square structuring element to perform dilation processing on the first segmentation mask Image mask. Among them, the dilation operator is a mathematical morphology operation. By performing a pixel-by-pixel "coverage" operation on the first segmentation mask Image mask and the square structuring element, the boundary of the first segmentation mask Image mask is expanded outward to obtain the second segmentation mask.

[0028] Step 130, generate a defective diffusion editing image corresponding to each query original image in the query original image set based on the starting points and ending points determined by multiple rays uniformly emitted in multiple specified directions in the first segmentation mask and the second segmentation mask, and the second segmentation mask, and generate defective samples based on the defective diffusion editing images corresponding to each query original image in the query original image set.

[0029] Among them, the multiple specified directions may include N specified directions in practical applications. The N specified directions can be N directions that diverge uniformly outward, where N is any integer greater than or equal to 3 and less than or equal to 8. Among them, the angle between any two adjacent directions is the same. For example, when N is 3, the 3 specified directions can be any 3 directions that diverge uniformly outward, where the angle between two adjacent directions is 120°. When N is 8, the 8 specified directions may include due north, due south, due west, due east, north-south, southwest, east-west, and northeast, or the 8 specified directions can also be any 8 directions that diverge uniformly outward, where the angle between two adjacent directions is 45°.

[0030] Optionally, in order to make the generated defects more abundant, the embodiment of the present application can use the center point of the first segmentation mask as the starting point, uniformly emit multiple rays in multiple specified directions of the starting point, and use the intersection points between these rays and the edges of the first segmentation mask and the second segmentation mask as the starting points for defective image editing. Specifically, generating a defective diffusion editing image corresponding to each query original image in the query original image set based on the starting points and ending points determined by multiple rays uniformly emitted in multiple specified directions in the first segmentation mask and the second segmentation mask, and the second segmentation mask, includes: Use the center point of the first segmentation mask as the starting point, and uniformly emit multiple rays in multiple specified directions of the starting point. The number of rays is the same as the number of specified directions; For each ray, the intersection point of each ray with the edge of the first segmentation mask is used as the starting point of the ray in the corresponding specified direction, and the intersection point of each ray with the edge of the second segmentation mask is used as the ending point of the ray in the corresponding specified direction; Crop the image of the region containing the second segmentation mask as the edited input image; Based on the starting point and ending point of each ray in the corresponding specified direction, the edited input image, and the corresponding second segmentation mask, use a pre-trained image diffusion editing model to generate defect diffusion edited images corresponding to each query original image in the query original image set.

[0031] Among them, taking the center point of the first segmentation mask as the starting point, the lengths of multiple rays evenly emitted in multiple specified directions from the starting point, the angle between two adjacent rays, and the ray lengths can be the same. To generate richer defects, when the number of multiple specified directions is 8, among the 8 starting point pairs determined based on the intersection points of these 8 rays with the edge of the first segmentation mask and the edge of the second segmentation mask respectively, randomly select 1-8 starting point pairs as the editing starting points and input them into the DragDiffusion model for image editing. Among them, the intersection points of each ray with the edge of the first segmentation mask and the edge of the second segmentation mask can be used as a starting point pair.

[0032] Among them, the edited input image is obtained in the following way: perform region cropping with a 1.1-fold frame expanded by the expanded mask expansion_mask (it should be understood that this 1.1-fold is only an exemplary description and should not constitute a limitation on the size of the edited input image, as long as the size of the edited input image is slightly larger than the second segmentation mask), save the coordinates of the cropped region, and use the cropped region object as the edited input image of the DragDiffusion model for image editing.

[0033] Exemplarily, the acquisition of the editing points of the starting point and ending point of each ray in the corresponding specified direction is mainly achieved based on the DragDiffusion model. Among them, the DragDiffusion model is an image editing tool based on the diffusion model, aiming to achieve interactive image editing through the method of point dragging. Figure 2The second framework in [the model] is the main framework of the model, which mainly includes three parts: editing point acquisition, Diffusion model, and feature optimization. The generation of the image can be achieved through the input image mask regions (the first segmentation mask and the second segmentation mask) and the automatically generated editing point pairs (the starting point pairs determined based on the intersection points between the rays emitted in multiple specified directions and the first segmentation mask and the second segmentation mask). First, the editing point pairs are obtained according to the initialization editing point pair acquisition rule. Secondly, the image is input into the DragDiffusion network for fine-tuning training to make the network better adapt to the current image. Finally, the feature optimization training is realized by using feature supervision and point tracking techniques, and the generated defect diffusion edited image Gen_Set is obtained.

[0034] Optionally, after generating the defect diffusion edited image, it is also necessary to determine the defect region and defect category in the image as the defect label of the image to generate a complete defect sample. Specifically, based on querying the defect diffusion edited images corresponding to each query original image in the original image set, defect samples are generated, including: Determine the defect region and defect category of the defect diffusion edited image corresponding to each query original image in the query original image set; Based on the defect diffusion edited images corresponding to each query original image in the query original image set, and the defect region and defect category of the defect diffusion edited images corresponding to each query original image in the query original image set, defect samples are generated.

[0035] Optionally, to avoid the defect region of the defect diffusion edited image corresponding to each query original image determined by the image segmentation model being larger than the region of the defect diffusion edited image corresponding to each query original image, the embodiments of the present application may use the points within a specified range based on the distance from the ray to the end point of the ray in the corresponding specified direction as the hint negative points of the ray and input them into the image segmentation model. Specifically, determining the defect region and defect category of the defect diffusion edited image corresponding to each query original image in the query original image set includes: For each ray, based on the specified range based on the distance from the ray to the end point of the ray in the corresponding specified direction, determine the hint negative points of the ray; Through the image segmentation model, using the starting point, end point, and hint negative points of each ray in the corresponding specified direction as hint information, determine the segmentation mask of the defect region and the defect category of the defect diffusion edited image corresponding to each query original image in the query original image set; Based on the segmentation mask of the defect region and the defect category of the defect diffusion edited image corresponding to each query original image in the query original image set, determine the defect region and defect category of the defect diffusion edited image corresponding to each query original image in the query original image set.

[0036] As an example, for each ray, the points at a specified number of pixels (the value of the specified number can be the downsampling multiple of the encoder of the image segmentation model) from the end point of the ray in the corresponding specified direction can be determined as the hint negative points of the ray.

[0037] Exemplarily, the calculation method for generating the start and end point pairs for editing based on the first segmentation mask Image mask of the defective images of each query original image in the query original image set Det_Set may include the following steps: First, perform a dilation operation on the first segmentation mask Image mask of the defective images of each query original image to generate an expanded second segmentation mask expansion_mask. Then, starting from the center of the first segmentation mask Image mask, emit rays in eight directions evenly outward. For each ray, calculate its intersection points with the first segmentation mask Image mask and the expanded second segmentation mask expansion_mask respectively to determine the intersection point set. Specifically, the intersection points in the same ray direction can be processed in pairs to form corresponding point pairs, that is, the start and end point set.

[0038] Optionally, in order to make the generated defects more abundant, the embodiments of the present application can also edit the start and end points of each ray in the corresponding specified direction. Specifically, through the pre-trained image diffusion editing model, based on the start and end points of each ray in the corresponding specified direction, the editing input image, and the corresponding second segmentation mask, generate the defective diffusion edited images corresponding to each query original image in the query original image set, including: Edit the start and end points of each ray in the corresponding specified direction through the pre-trained image diffusion editing model to obtain the edited start and end points of each ray in the corresponding specified direction; Generate the defective diffusion edited images corresponding to each query original image in the query original image set through the pre-trained image diffusion editing model based on the edited start and end points of each ray in the corresponding specified direction, the editing input image, and the corresponding second segmentation mask.

[0039] To make the generated defects more abundant, the embodiments of the present application can randomly select 1-8 point pairs from the start-end point set and input them into the DragDiffusion model for image editing to edit the selected start and end points. At the same time, in the direction of the extension line of the ray where the randomly selected start and end points are located, a point is determined at an optional threshold distance from the end point (usually the downsampling multiple of the encoder of the image segmentation model, i.e., 16 pixels) as the prompt negative point of the image segmentation model. As an example, the area can be cropped with a 1.1 times frame expanded by the expanded mask expansion_mask at the same time, and the coordinates of the cropped area are saved, and the object of the cropped area is used as the edited input image of the DragDiffusion model for image editing.

[0040] Next, the obtained edited start-end point pairs, the edited input image, and its second segmentation mask expansion_mask are input into the DragDiffusion model for feature extraction and fine-tuning training. The image input first obtains the Image_Embedding feature through the feature extraction module inside the DragDiffusion model. Subsequently, the DragDiffusion model is fine-tuned to enhance the adaptability of the DragDiffusion model to the extracted features.

[0041] Based on feature-based motion supervision, precise control of specific editing points in the image can be achieved by optimizing the feature vectors, thus ensuring the smooth coherence of semantic information during the image editing process. Among them, the schematic diagram of the feature supervision process is as Figure 4 shown, where the red dot is the editing start point and the blue dot is the editing end point. During the feature supervision process, the optimization of the feature vectors is achieved. Figure 4 What is shown in is a stage in the optimization process, specifically manifested as the update of the image features according to the direction of the editing point pair. Specifically, the embodiments of the present application can calculate the corresponding direction scalar by analyzing the direction from the editing start point to the editing end point. Based on the direction pointed by this direction scalar, a pixel matrix region with a radius of 1 pixel point is defined with the editing start point as the center as the current feature. At the same time, according to the direction scalar, the nearest point is found on the editing start point, and another pixel matrix region with a radius of 1 pixel point is defined with it as the center as the reference feature. By calculating the average value of the current feature and the reference feature, the input for the next feature supervision is obtained. This process ensures the continuity and smoothness of the features during the editing process, that is, it ensures the smooth coherence of semantic information during the image editing process.

[0042] Since the update of motion supervision changes the Image_Embedding features, which may lead to changes in the positions of the editing start and end points. Therefore, after optimizing the intermediate feature vectors during the diffusion process, the corresponding processing point positions can be updated through point tracking technology. The schematic diagram of the point tracking process is as shown in Figure 5 Figure Figure 5 , where the red dot is the editing start point and the blue dot is the editing end point. During the point tracking process, based on the results of feature supervision, the change of the features can be calculated, and the adjusted position of the editing start point can be determined, so as to realize the dynamic movement of the editing start point towards the editing end point. This process is manifested as a tracking behavior. Since the update of feature supervision changes the Image_Embedding features, the features corresponding to the editing start point also change. Therefore, it is necessary to use point tracking technology to update the position of the editing start point to ensure the synchronization of each step of the update, so as to realize the accuracy of the editing process. Specifically, in the next feature stage after feature supervision, the point tracking technology searches for the point most similar to the original feature within the range of 3 pixel points around the editing start point to determine the next tracking point, so as to realize the precise synchronization of the editing process.

[0043] Among them, the image segmentation model can specifically be the SAM model (Segment Anything Model), which aims to achieve fast segmentation of any object through simple prompts. The SAM model is a prompt-based model that can segment any image without any annotation. Its input can be an image and a prompt. The prompt can be a point, a box, text, or a mask, used to indicate the target to be segmented, and the output is a segmentation mask.

[0044] Optionally, since defect editing is only performed on the defect areas in the query original image, after editing, in order to retain the features of the complete defect samples, the edited defect image can also be mapped back to the original query original image. Specifically, based on the defect diffusion edited images corresponding to each query original image in the query original image set, as well as the defect areas and defect categories of the defect diffusion edited images corresponding to each query original image in the query original image set, defect samples are generated, including: Mapping the segmentation masks and defect categories of the defect diffusion edited images corresponding to each query original image in the query original image set to each query original image in the query original image set to obtain the defect areas and defect categories of each query original image in the query original image set; Generating defect samples based on the defect areas and defect categories of each query original image in the query original image set.

[0045] Exemplarily, generating defect samples can be realized based on the SAM model. Figure 2The first framework in [[ID=]] is the main framework of the model, which mainly includes three parts: the encoder, decoder, and visual prompt of the image. By simply using prompt hints such as points, boxes, or masks, high-precision segmentation of the target area can be achieved, and the segmentation accuracy exceeds that of manual annotation. Therefore, in order to further obtain higher-precision segmentation results, first, the above-mentioned editing start and end points and negative points are used as prompts, and the Prompt Token vector is obtained through the Prompt Encoder. Similarly, the generated defective diffusion editing image Gen_Set passes through the Image Encoder to obtain the image encoding Image Embedding, and then the Prompt Token and the image encoding Image Embedding are input into the decoder of SAM for inference, so as to obtain the segmentation mask corresponding to Gen_Set, as Figure 6 is a schematic diagram of the effect generated for the defective image, where Figure 6 (a)is the generated defective image, Figure 6 (b)is the segmentation mask of the defective area in the generated defective image.

[0046] Optionally, similar to Figure 3 , small-area noise points that may exist in the generated data segmentation mask obtained in the above steps are further filtered to obtain a higher-precision defective mask Det_Mask.

[0047] Exemplarily, based on the coordinates of the image cropped by the area containing the second segmentation mask, the generated defective image and its label are restored to the original image, and the defective data and its corresponding label generated based on the query original image set Det_set can be obtained, as Figure 7 is a schematic diagram of the effect of mapping the generated defective image and its label to the original image.

[0048] Finally, in order to ensure that all the labels of the generated defects meet the requirements, the labels of the generated defects can be further manually confirmed.

[0049] The defect sample generation method provided by the embodiments of this application can obtain a query original image set and a hint thumbnail image set of defect images, and perform image queries on the query original image set based on each hint thumbnail in the hint thumbnail image set to obtain the first segmentation mask and the corresponding defect category of the defect images of each query original image in the query original image set. Then, perform a dilation operation on the segmentation masks of the defect images of each query original image in the query original image set to generate the second segmentation mask of the extended defect images of each query original image in the query original image set. Finally, based on the starting and ending points determined by multiple rays uniformly emitted in multiple specified directions in the first segmentation mask and the second segmentation mask, and the second segmentation mask, generate the defect diffusion editing images corresponding to each query original image in the query original image set, and generate defect samples based on the defect diffusion editing images corresponding to each query original image in the query original image set. On the one hand, use the query original images that match the hint thumbnails to determine the first segmentation mask and the corresponding defect category of the defect images in the query original images that are similar to the hint thumbnails. On the other hand, through dilation operations and defect diffusion editing on the first segmentation mask in sequence to achieve the editing of the defect area, and generate new defect samples that do not deviate from the defect categories of the original query original images based on the edited defect images, which can realize generating a large number of high-quality defect samples based on a limited number of defect samples to meet the training of relevant models.

[0050] Figure 8 FIG. 4 is a schematic structural diagram of a defect sample generation device 800 provided by an exemplary embodiment of this application. As Figure 8 shown, the device 800 includes: a mask acquisition module 810, a mask generation module 820, and a sample generation module 830, where: The mask acquisition module 810 is configured to obtain a query original image set and a hint thumbnail image set of defect images, and perform image queries on the query original image set based on each hint thumbnail in the hint thumbnail image set to obtain the first segmentation mask and the corresponding defect category of the defect images of each query original image in the query original image set; The mask generation module 820 is configured to perform a dilation operation on the segmentation masks of the defect images of each query original image in the query original image set to generate the second segmentation mask of the extended defect images of each query original image in the query original image set; The sample generation module 830 is configured to generate the defect diffusion editing images and the corresponding defect categories corresponding to each query original image in the query original image set based on the starting and ending points determined by multiple rays uniformly emitted in multiple specified directions in the first segmentation mask and the second segmentation mask, and the second segmentation mask, and generate defect samples based on the defect diffusion editing images corresponding to each query original image in the query original image set.

[0051] Optionally, the sample generation module 830 is specifically configured to: Taking the center point of the first segmentation mask as the starting point, uniformly emit a plurality of rays in a plurality of specified directions from the starting point, and the number of rays is the same as the number of specified directions; For each ray, taking the intersection point of each ray and the edge of the first segmentation mask as the starting point of the ray in the corresponding specified direction, and taking the intersection point of each ray and the edge of the second segmentation mask as the end point of the ray in the corresponding specified direction; Cropping the image of the region containing the second segmentation mask as the edited input image; Based on the starting point and end point of each ray in the corresponding specified direction, the edited input image, and the corresponding second segmentation mask, the defect diffusion edited image corresponding to each query original image in the query original image set is generated by a pre-trained image diffusion editing model.

[0052] Optionally, the sample generation module 830 is specifically configured to: Determine the defect regions and defect categories of the defect diffusion edited images corresponding to the query original images in the query original image set; Based on the defect diffusion edited images corresponding to the query original images in the query original image set, and the defect regions and defect categories of the defect diffusion edited images corresponding to the query original images in the query original image set, defect samples are generated.

[0053] Optionally, the sample generation module 830 is specifically configured to: For each ray, based on a specified range from the end point of the ray in the corresponding specified direction, determine the prompt negative point of the ray; Using an image segmentation model, with the starting point, end point, and prompt negative point of each ray in the corresponding specified direction as prompt information, determine the segmentation mask of the defect region of the defect diffusion edited image corresponding to each query original image in the query original image set; Based on the segmentation mask and defect category of the defect region of the defect diffusion edited image corresponding to each query original image in the query original image set, determine the defect region and defect category of the defect diffusion edited image corresponding to each query original image in the query original image set.

[0054] Optionally, the sample generation module 830 is specifically configured to: Edit the starting point and end point of each ray in the corresponding specified direction through a pre-trained image diffusion editing model to obtain the edited starting point and end point of each ray in the corresponding specified direction; Based on the edited start and end points of each ray in the corresponding specified direction, the edited input image, and the corresponding second segmentation mask, a pre-trained image diffusion editing model generates the defect diffusion edited images corresponding to each query original image in the query original image set.

[0055] Optionally, the mask acquisition module 810 is specifically configured to: Through an image retrieval model, based on each prompt thumbnail in the prompt thumbnail image set, perform image queries on the query original image set to obtain the initial segmentation mask and corresponding defect category of the defect image of each query original image in the query original image set; Perform a noise filtering operation on the initial segmentation mask of the defect image of each query original image in the query original image set to obtain the first segmentation mask of the defect image of each query original image in the query original image set.

[0056] Optionally, the sample generation module 830 is specifically configured to: Map the segmentation mask and defect category of the defect diffusion edited image corresponding to each query original image in the query original image set to each query original image in the query original image set to obtain the defect region and defect category of each query original image in the query original image set; Generate defect samples based on the defect regions and defect categories of each query original image in the query original image set.

[0057] The defect sample generation device 800 can implement Figures 1 - 7 the method of the method embodiment, and specifically, reference can be made to Figures 1 - 7 the defect sample generation method shown in the embodiment, which will not be elaborated here.

[0058] Figure 9 This is a schematic structural diagram of an electronic device provided by an exemplary embodiment of the present application. As Figure 9 shown, the device includes: a memory 91 and a processor 92.

[0059] The memory 91 is used to store computer programs and can be configured to store various other data to support operations on the computing device. Examples of these data include instructions for any application program or method for operating on the computing device, contact data, phone book data, messages, images, videos, etc.

[0060] A processor 92, coupled to a memory 91, is configured to execute a computer program in the memory 91 for: obtaining a query original image set and a hint thumbnail image set of defective images, and performing an image query on the query original image set based on each hint thumbnail in the hint thumbnail image set to obtain a first segmentation mask of the defective images of each query original image in the query original image set and the corresponding defective category; performing a dilation operation on the segmentation masks of the defective images of each query original image in the query original image set to generate a second segmentation mask of the extended defective images of each query original image in the query original image set; generating a defective diffusion editing image corresponding to each query original image in the query original image set based on the starting points and ending points determined by multiple rays uniformly emitted in multiple specified directions in the first segmentation mask and the second segmentation mask, and the second segmentation mask, and generating defective samples based on the defective diffusion editing images corresponding to each query original image in the query original image set and the corresponding defective categories.

[0061] Further, as Figure 9 shown, the electronic device further includes: a communication component 93, a display 94, a power supply component 95, an audio component 96, and other components. Figure 9 Only some components are schematically shown in Figure 9 and it does not mean that the electronic device only includes Figure 9 the components shown. Additionally, according to different implementation forms of the traffic playback device, Figure 9 the components within the dashed box in Figure 9 are optional components rather than essential components. For example, when the electronic device is implemented as a terminal device such as a smart phone, a tablet computer, or a desktop computer, it may include

[0062] the components within the dashed box in

[0063] the above; when the electronic device is implemented as a server - side device such as a conventional server, a cloud server, a data center, or a server array, it may not include Figure 9 the components within the dashed box in Figure 9 above.

[0062] Correspondingly, an embodiment of the present application further provides a computer - readable storage medium storing a computer program, and when the computer program is executed by a processor, it causes the processor to be able to implement the steps in the above - mentioned defective sample generation method embodiment.

[0063] The above Figure 9 mentioned communication component is configured to facilitate communication between the device where the communication component is located and other devices in a wired or wireless manner. The device where the communication component is located can access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component receives broadcast signals or broadcast - related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component may further include a near - field communication (NFC) module, radio - frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra - wideband (UWB) technology, Bluetooth (BT) technology, etc.

[0064] The above-mentioned Figure 9 The memory in the above can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0065] The above-mentioned Figure 9 The display in the above includes a screen, and the screen can include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operations.

[0066] The above-mentioned Figure 9 The power supply component in the above provides power for various components of the device where the power supply component is located. The power supply component can include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device where the power supply component is located.

[0067] The above-mentioned Figure 9 The audio component in the above can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC), and when the device where the audio component is located is in an operation mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive external audio signals. The received audio signals can be further stored in the memory or sent via the communication component. In some embodiments, the audio component further includes a speaker for outputting audio signals.

[0068] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0069] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, as well as the combination of flows and / or blocks in the flowchart and / or block diagram. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 means for implementing the functions specified in one block or multiple blocks.

[0070] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 means for implementing the functions specified in one block or multiple blocks.

[0071] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 means for implementing the functions specified in one block or multiple blocks.

[0072] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0073] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.

[0074] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0075] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0076] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.

Claims

1. A defect sample generation method, characterized in that: include: Acquire a query original image set and a prompt thumbnail image set of defect images, and perform an image query on the query original image set based on each prompt thumbnail in the prompt thumbnail image set to acquire a first segmentation mask and a corresponding defect category of each defect image of the query original image in the query original image set; Performing an expansion operation on the segmentation mask of the defect image of each query original image in the query original image set to generate a second segmentation mask of the expanded defect image of each query original image in the query original image set; Based on the starting point and end point determined by multiple rays uniformly emitted in multiple specified directions in the first segmentation mask and the second segmentation mask, and the second segmentation mask, a defect diffusion edited image corresponding to each query original image in the query original image set is generated, and based on the defect diffusion edited image corresponding to each query original image in the query original image set and the corresponding defect category, a defect sample is generated.

2. The method according to claim 1, characterized in that Based on the starting points and end points determined by the plurality of rays uniformly emitted in the plurality of specified directions in the first segmentation mask and the second segmentation mask, and the second segmentation mask, a defect diffusion edited image corresponding to each query original image in the query original image set is generated, including: Taking the center point of the first segmentation mask as a starting point, uniformly emitting multiple rays in multiple specified directions of the starting point, wherein the number of the rays is consistent with the number of the specified directions; For each ray, taking the intersection point of each ray with the edge of the first segmentation mask as the starting point of the ray in the corresponding specified direction, and taking the intersection point of each ray with the edge of the second segmentation mask as the end point of the ray in the corresponding specified direction; Crop the image of the region containing the second segmentation mask as the editing input image; A defect diffusion editing image corresponding to each query original image in the query original image set is generated through a pre-trained image diffusion editing model based on the starting point and end point of each ray in the corresponding specified direction, the editing input image and the corresponding second segmentation mask.

3. The method according to claim 2, characterized in that Generating defect samples based on the defect diffusion edited images and corresponding defect categories corresponding to each query original image in the query original image set includes: Determine the defect area and defect category of the defect diffusion edited image corresponding to each query original image in the query original image set; Defect samples are generated based on the defect diffusion edited images corresponding to each query original image in the query original image set and the defect areas and defect categories of the defect diffusion edited images corresponding to each query original image in the query original image set.

4. The method according to claim 3, characterized in that The step of determining the defect area and defect category of the defect diffusion edited image corresponding to each query original image in the query original image set comprises: For each ray, determining a prompt negative point of the ray based on a specified range from an end point of the ray in a corresponding specified direction; Determine the segmentation mask of the defect area of ​​the defect diffusion edited image corresponding to each query original image in the query original image set by using the image segmentation model with the starting point, the end point and the prompt negative point of each ray in the corresponding specified direction as prompt information; Based on the segmentation masks and defect categories of the defect areas of the defect diffusion edited images corresponding to each query original image in the query original image set, the defect areas and defect categories of the defect diffusion edited images corresponding to each query original image in the query original image set are determined.

5. The method according to claim 2, characterized in that Generate a defect diffusion edited image corresponding to each query original image in the query original image set based on the starting point and the end point of each ray in the corresponding specified direction, the edited input image and the corresponding second segmentation mask through a pre-trained image diffusion edited model, including: The starting point and the end point of each ray in the corresponding specified direction are edited by using a pre-trained image diffusion editing model to obtain the edited starting point and the end point of each ray in the corresponding specified direction; A defect diffusion edited image corresponding to each query original image in the query original image set is generated through a pre-trained image diffusion editing model based on the edited starting and ending points of each ray in the corresponding specified direction, the edited input image and the corresponding second segmentation mask.

6. The method according to claim 1, characterized in that Based on each hint thumbnail in the hint thumbnail image set, an image query is performed from the query original image set to obtain a first segmentation mask and a corresponding defect category of each query original image in the query original image set, including: Performing an image query from the query original image set based on each hint thumbnail in the hint thumbnail image set through an image retrieval model, and obtaining an initial segmentation mask and a corresponding defect category of each query original image in the query original image set; A noise filtering operation is performed on the initial segmentation mask of the defect image of each query original image in the query original image set to obtain a first segmentation mask of the defect image of each query original image in the query original image set.

7. The method according to claim 4, characterized in that Generating defect samples based on the defect diffusion edited images corresponding to each query original image in the query original image set and the defect areas and defect categories of the defect diffusion edited images corresponding to each query original image in the query original image set includes: Mapping the segmentation mask and defect category of the defect diffusion edited image corresponding to each query original image in the query original image image set to each query original image in the query original image image set to obtain the defect area and defect category of each query original image in the query original image image set; Defect samples are generated based on the defect areas and defect categories of each query original image in the query original image set.

8. A defect sample generating device, characterized in that: include: A mask acquisition module, used to acquire a query original image set and a prompt thumbnail image set of defect images, and perform an image query from the query original image set based on each prompt thumbnail in the prompt thumbnail image set to acquire a first segmentation mask and a corresponding defect category of each defect image of the query original image in the query original image set; A mask generation module, configured to perform an expansion operation on the segmentation mask of the defect image of each query original image in the query original image set to generate a second segmentation mask of the expanded defect image of each query original image in the query original image set; A sample generation module is used to generate defect diffusion edited images and corresponding defect categories corresponding to each query original image in the query original image set based on the starting point and end point determined by multiple rays uniformly emitted in multiple specified directions in the first segmentation mask and the second segmentation mask, and to generate defect samples based on the defect diffusion edited images corresponding to each query original image in the query original image set.

9. An electronic device, characterized in that: include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, the processor and the memory communicate via a bus, and when the machine-readable instructions are executed by the processor, the steps in any one of the methods described in claims 1 to 7 are performed.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the processor is caused to implement the steps in the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Defect image generation method and device, equipment and storage medium

    CN121120867A

  • Methods, apparatus, devices and storage media for generating defect images

    CN121120867B