Defect dividing method, defect dividing apparatus, and defect dividing system

By fine-tuning and distilling the initial prompt segmentation model and replacing the prompt encoder with an image prompt generator, the problem that the existing model cannot segment unknown defects and high computational volume in industrial defect segmentation is solved, and the industrial defect segmentation effect with high precision and low computational volume is achieved.

CN120219736APending Publication Date: 2025-06-27BEIJING LUSTER LIGHTTECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510209093.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing prompt segmentation model cannot segment unknown defects in industrial defect segmentation scenarios, and it is very computationally expensive in low computing scenarios and is difficult to apply.

Method used

By fine-tuning the initial prompt segmentation model using the industrial segmentation data set, a first prompt segmentation model suitable for industrial scenarios is obtained; then the first prompt segmentation model is lightened by staged distillation to obtain the second prompt segmentation model; finally, the prompt encoder in the second prompt segmentation model is replaced with an image prompt generator, and the third prompt segmentation model is obtained. This model supports visual images as input and has the ability to segment out unknown defects.

Benefits of technology

The ability to segment unknown defects in industrial defect segmentation scenarios is realized, and the calculation amount is reduced in low computing power scenarios, improving segmentation accuracy and generalization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219736A_ABST
    Figure CN120219736A_ABST
Patent Text Reader

Abstract

The invention discloses a defect segmentation method, a defect segmentation device and a defect segmentation system. The defect segmentation method provided by the embodiment of the invention comprises the following steps: performing fine adjustment on an initial prompt segmentation model by utilizing an industrial segmentation data set to obtain a first prompt segmentation model; distilling the first prompt segmentation model in a staged distillation mode to obtain a second prompt segmentation model which is lighter than the first prompt segmentation model; a second prompt encoder in the second prompt segmentation model is replaced with an image prompt generator, a third prompt segmentation model is obtained, and the image prompt generator supports a prompt image as input; and performing defect segmentation on the target image based on the prompt image through a third prompt segmentation model to obtain a target segmentation result. Therefore, the image prompt generator supports the prompt image as input, so that the third prompt segmentation model has the capability of segmenting unknown defects, and has relatively high generalization and relatively high segmentation precision in an industrial defect segmentation scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of machine vision technology, and particularly to a defect segmentation method, a defect segmentation device, and a defect segmentation system. Background Art

[0002] A vision foundation model refers to a foundation model in the field of computer vision that has been pre-trained on a large amount of data and has general vision understanding capabilities. These models can adapt to a variety of downstream tasks and can achieve optimal or competitive results without having to be trained from scratch. In recent years, thanks to the booming development of vision foundation models, prompt segmentation models have made great breakthroughs and achieved quite high accuracy in many natural scenes. However, current prompt segmentation models do not have the ability to segment unknown defects in industrial defect segmentation scenarios. Summary of the Invention

[0003] Embodiments of this application provide a defect segmentation method, a defect segmentation device, a defect segmentation system, and a computer-readable storage medium to solve at least one of the above-mentioned technical problems.

[0004] The defect segmentation method according to the embodiments of this application includes:

[0005] Fine-tuning an initial prompt segmentation model using an industrial segmentation dataset to obtain a first prompt segmentation model;

[0006] Distilling the first prompt segmentation model through a staged distillation method to obtain a second prompt segmentation model that is lightweight relative to the first prompt segmentation model;

[0007] Replacing a second prompt encoder in the second prompt segmentation model with an image prompt generator to obtain a third prompt segmentation model, where the image prompt generator supports a prompt image as input;

[0008] Performing defect segmentation on a target image based on the prompt image through the third prompt segmentation model to obtain a target segmentation result.

[0009] In some embodiments, the step of fine-tuning an initial prompt segmentation model using an industrial segmentation dataset to obtain a first prompt segmentation model includes:

[0010] Using the industrial segmentation dataset as a training set and performing fine-tuning on the initial prompt segmentation model using the LoRA fine-tuning method to obtain the first prompt segmentation model.

[0011] In some embodiments, the step of using the industrial segmentation dataset as a training set and performing fine-tuning on the initial prompt segmentation model using the LoRA fine-tuning method to obtain the first prompt segmentation model includes:

[0012] Construct the initial prompt segmentation model, where the initial prompt segmentation model includes an initial image encoder, an initial prompt encoder, and an initial decoder, and the initial image encoder uses ViT-L;

[0013] Freeze the initial prompt encoder and the initial decoder;

[0014] Freeze the encoder parameters in ViT-L except for the attention mechanism parameters;

[0015] Use the industrial segmentation dataset as the training set and train the attention mechanism parameters in ViT-L using the LoRA fine-tuning method to obtain the first prompt segmentation model.

[0016] In some embodiments, the first prompt segmentation model includes a first image encoder, a first prompt encoder, and a first decoder. Distill the first prompt segmentation model through a staged distillation method to obtain a second prompt segmentation model that is lighter than the first prompt segmentation model, including:

[0017] During the first-stage distillation process, distill the first image encoder;

[0018] During the second-stage distillation process, distill the first image encoder and the first decoder to determine a second image encoder and a second decoder;

[0019] Determine the second prompt segmentation model according to the second image encoder, the first prompt encoder, and the second decoder.

[0020] In some embodiments, during the first-stage distillation process, distilling the first image encoder includes:

[0021] Perform an alignment operation on the image embedding vectors output by the first image encoder and the image embedding vectors output by the second image encoder;

[0022] During the second-stage distillation process, distilling the first image encoder and the first decoder to determine a second image encoder and a second decoder includes:

[0023] Perform an alignment operation on the image embedding vectors output by the first image encoder and the image embedding vectors output by the second image encoder, and perform an alignment operation on the segmentation results output by the first decoder and the segmentation results output by the second decoder to determine the second image encoder and the second decoder.

[0024] In some embodiments, the defect segmentation of the target image based on the prompt image by the third prompt segmentation model to obtain the target segmentation result includes:

[0025] Performing defect segmentation on the target image based on the prompt image by the third prompt segmentation model, and outputting an initial segmentation result;

[0026] Judging whether there is any missed detection of defects according to the initial segmentation result;

[0027] When there is no missed detection of defects, taking the initial segmentation result as the target segmentation result;

[0028] The defect segmentation method further includes:

[0029] When there is a missed detection of defects, adding the missed defect image to the prompt image.

[0030] In some embodiments, after adding the missed defect image to the prompt image when there is a missed detection of defects, the defect segmentation method further includes:

[0031] Judging whether the number of the missed defect images reaches a predetermined number;

[0032] When the number of the missed defect images reaches the predetermined number, taking the third prompt segmentation model as the initial prompt segmentation model, and returning to the step of fine-tuning the initial prompt segmentation model with the industrial segmentation dataset to obtain the first prompt segmentation model.

[0033] In some embodiments, the prompt segmentation model is a Segment Anything Model.

[0034] The defect segmentation device according to the embodiments of the present application includes:

[0035] A fine-tuning module, configured to fine-tune an initial prompt segmentation model with an industrial segmentation dataset to obtain a first prompt segmentation model;

[0036] A distillation module, configured to distill the first prompt segmentation model by a staged distillation method to obtain a second prompt segmentation model that is lightweight relative to the first prompt segmentation model;

[0037] A replacement module, configured to replace the second prompt encoder in the second prompt segmentation model with an image prompt generator to obtain a third prompt segmentation model, where the image prompt generator supports a prompt image as input;

[0038] A segmentation module, configured to perform defect segmentation on a target image based on the prompt image by the third prompt segmentation model to obtain a target segmentation result.

[0039] The defect segmentation system according to the embodiments of the present application, the defect segmentation system includes one or more processors and a memory, the memory stores a computer program, and when the computer program is executed by the processor, the defect segmentation method according to any of the above embodiments is implemented.

[0040] The computer-readable storage medium according to the embodiments of the present application, on which a computer program is stored, and when the program is executed by a processor, the defect segmentation method according to any of the above embodiments is implemented.

[0041] For the defect segmentation method, defect segmentation device, defect segmentation system, and computer-readable storage medium according to the embodiments of the present application, first, fine-tune the initial prompt segmentation model using an industrial segmentation dataset to obtain a first prompt segmentation model, so that the first prompt segmentation model can adapt to the image domain of the industrial scenario to improve the segmentation accuracy; second, distill the first prompt segmentation model through a staged distillation method to obtain a second prompt segmentation model that is lighter than the first prompt segmentation model to reduce the computational amount, so that it can be applied to industrial low-computing-power scenarios; third, replace the second prompt encoder in the second prompt segmentation model with an image prompt generator to obtain a third prompt segmentation model, where the image prompt generator supports a prompt image as input, so that the third prompt segmentation model has the ability to segment unknown defects, and when defect-segmenting a target image based on the prompt image through the third prompt segmentation model, it has strong generalization and high segmentation accuracy.

[0042] The additional aspects and advantages of the embodiments of the present application will be partially given in the following description, partially become apparent from the following description, or be understood through the practice of the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, where:

[0044] Figure 1 is one of the flow diagrams of the defect segmentation method according to some embodiments of the present application;

[0045] Figure 2 is the flow diagram of the construction of the third prompt segmentation model according to some embodiments of the present application;

[0046] Figure 3 is the flow diagram of the application of the third prompt segmentation model according to some embodiments of the present application;

[0047] Figure 4 is the second flow diagram of the defect segmentation method according to some embodiments of the present application;

[0048] Figure 5It is the third flowchart diagram of the defect segmentation method according to some embodiments of the present application;

[0049] Figure 6 It is the fourth flowchart diagram of the defect segmentation method according to some embodiments of the present application;

[0050] Figure 7 It is the fifth flowchart diagram of the defect segmentation method according to some embodiments of the present application;

[0051] Figure 8 It is the sixth flowchart diagram of the defect segmentation method according to some embodiments of the present application;

[0052] Figure 9 It is the seventh flowchart diagram of the defect segmentation method according to some embodiments of the present application;

[0053] Figure 10 It is the module diagram of the defect segmentation device according to some embodiments of the present application;

[0054] Figure 11 It is the module diagram of the defect segmentation system according to some embodiments of the present application;

[0055] Figure 12 It is the connection state diagram of the computer-readable storage medium and the processor according to some embodiments of the present application.

[0056] Explanation of reference numerals:

[0057] Initial prompt segmentation model 101, initial image encoder 1011, initial prompt encoder 1012, initial decoder 1013, first prompt segmentation model 102, first image encoder 1021, first prompt encoder 1022, first decoder 1023, second prompt segmentation model 103, second image encoder 1031, second prompt encoder 1032, second decoder 1033, third prompt segmentation model 104, third image encoder 1041, image prompt generator 1042, third decoder 1043, defect segmentation device 100, fine-tuning module 10, distillation module 20, replacement module 30, segmentation module 40, defect segmentation system 200, processor 210, memory 220, computer-readable storage medium 300, computer program 310, processor 320. Detailed implementation manners

[0058] The following further describes the embodiments of the present application with reference to the accompanying drawings. The same or similar reference numerals in the drawings denote the same or similar elements or elements having the same or similar functions throughout. Additionally, the embodiments of the present application described below with reference to the accompanying drawings are exemplary only for explaining the embodiments of the present application and should not be construed as limiting the present application.

[0059] A vision foundation model refers to a foundation model in the field of computer vision that has been pre-trained on large-scale data and has general visual understanding capabilities. These models can adapt to a variety of downstream tasks and achieve optimal or competitive results without having to be trained from scratch. In recent years, thanks to the booming development of vision foundation models, prompt segmentation models have made great breakthroughs, achieving quite high accuracy in numerous natural scenes and even obtaining competitive results in scenes not included in the training set.

[0060] It has been found through research that prompt segmentation models can be subdivided into three different methods according to the way of obtaining prompts: The first method is that the prompt consists of information provided by humans such as points, boxes, and text. For example, for the Segment Anything Model (SAM), under large-scale natural scene data, it is trained by combining information such as the point coordinates and circumscribed rectangles of the target area to be segmented. During inference, only by manually clicking a point or drawing a box, the Segment Anything Model can output the mask of the corresponding target. The second method is that the prompt is obtained based on the target image, also known as the self-prompt method, which is often used for panoramic segmentation. The third method is that the prompt consists of visual images, where the visual image is a prompt image and the mask of the object to be segmented, which is often applied to scenarios with a small number of samples, that is, few-shot scenarios.

[0061] Among the above three different methods, for the first method, information such as points and boxes of the target area to be segmented needs to be provided manually as a prompt to obtain the mask of the target, so it is not applicable to the general industrial defect segmentation scenario. For the second method, although it does not require a prompt provided by humans as input, if it is applied to the general industrial defect segmentation scenario, its generalization ability will be greatly restricted. This is because the premise for the self-prompt model to have strong generalization abilities such as cross-domain segmentation is that it needs to not learn semantic features of specific categories and specific regions, but only perform region segmentation based on gray-scale differences, that is, panoramic segmentation. However, the general industrial defect segmentation scenario requires specific abnormal regions to be segmented. Therefore, applying the self-prompt method to the industrial defect segmentation scenario will weaken its ability. For the third method, based on the visual prompt input, it can interact with the target image to obtain the preliminary position information and semantic features of the target to be segmented, and then the mask of the target to be segmented can be obtained after decoding. This method has the ability to segment unknown classes and has strong generalization, but it is trained based on large-scale data of natural scenes. Therefore, it performs poorly in industrial scenes with semantic differences. In addition, currently it does not support mixed categories as prompts to segment the masks of different classes at one time. If directly applied to the industrial defect segmentation scenario, multiple inferences or multiple batches of inferences are required. Therefore, it does not meet the low latency and low hardware conditions of industrial scenes.

[0062] In view of this, an industrial defect segmentation method supporting a hint function is proposed in an embodiment of the present application to solve at least one of the above-mentioned technical problems. The embodiment of the present application mainly includes a process of training an initial hint segmentation model 101 using an industrial segmentation dataset to obtain a hint segmentation model suitable for industrial defect segmentation scenarios, and a process of enabling the hint segmentation model to support visual images as hints to segment the mask of the corresponding target.

[0063] Please refer to Figures 1 to 3 , the defect segmentation method of the embodiment of the present application includes:

[0064] 010: Fine-tune the initial hint segmentation model 101 using an industrial segmentation dataset to obtain a first hint segmentation model 102;

[0065] 020: Distill the first hint segmentation model 102 through a staged distillation method to obtain a second hint segmentation model 103 that is lighter than the first hint segmentation model 102;

[0066] 030: Replace the second hint encoder 1032 in the second hint segmentation model 103 with an image hint generator 1042 to obtain a third hint segmentation model 104, where the image hint generator 1042 supports hint images as inputs;

[0067] 040: Perform defect segmentation on the target image based on the hint image through the third hint segmentation model 104 to obtain a target segmentation result.

[0068] For the defect segmentation method of the embodiment of the present application, firstly, fine-tune the initial hint segmentation model 101 using an industrial segmentation dataset to obtain a first hint segmentation model 102, so that the first hint segmentation model 102 can adapt to the image domain of industrial scenarios to improve the segmentation accuracy; secondly, distill the first hint segmentation model 102 through a staged distillation method to obtain a second hint segmentation model 103 that is lighter than the first hint segmentation model 102 to reduce the computational amount, so that it can be applied to industrial low-computing-power scenarios; thirdly, replace the second hint encoder 1032 in the second hint segmentation model 103 with an image hint generator 1042 to obtain a third hint segmentation model 104, where the image hint generator 1042 supports hint images as inputs, so that the third hint segmentation model 104 has the ability to segment unknown defects, and has strong generalization and high segmentation accuracy when performing defect segmentation on the target image based on the hint image through the third hint segmentation model 104.

[0069] Specifically, a large-scale industrial segmentation dataset containing various types of industrial defect annotations can be collected in advance. Using the industrial segmentation dataset as the training set to fine-tune the initial prompt segmentation model 101, a first prompt segmentation model 102 is obtained, enabling the first prompt segmentation model 102 to adapt to the image domain of industrial scenarios, that is, enabling the first prompt segmentation model 102 to effectively process the image features in industrial scenarios.

[0070] Among them, industrial scenarios include but are not limited to industrial production environments such as manufacturing, automated production lines, and energy facilities. The image domain refers to the performance characteristics of image data in a specific environment, such as the resolution, color space, lighting conditions, object shape, and texture of the image. Since in industrial scenarios, image data often has some unique characteristics, such as complex backgrounds, diverse objects, uneven lighting, etc., a first prompt segmentation model 102 that adapts to the image domain of industrial scenarios is required to accurately perform object segmentation on these complex images. Fine-tuning the initial prompt segmentation model 101 using the industrial segmentation dataset allows the initial prompt segmentation model 101 to learn on the large-scale industrial segmentation dataset and gradually adapt to the image domain of industrial scenarios. During the fine-tuning process, the initial prompt segmentation model 101 continuously adjusts its internal parameters, thereby improving the segmentation accuracy.

[0071] Then, the first prompt segmentation model 102 is distilled through a staged distillation method to obtain a second prompt segmentation model 103 that is lightweight compared to the first prompt segmentation model 102. Specifically, refer to the staged distillation algorithm of EdgeSAM, and adopt a two-stage distillation method with the first prompt segmentation model 102 as the teacher model to distill a lightweight student model downward, that is, the second prompt segmentation model 103, to reduce the computational load and thus be applicable to industrial low-computing-power scenarios.

[0072] Next, the second prompt encoder 1032 in the second prompt segmentation model 103 is replaced with an image prompt generator 1042 to obtain a third prompt segmentation model 104. Among them, the image prompt generator 1042 supports a prompt image as input. The prompt encoder, that is Figure 2 the prompt encoder in, only supports information such as points, boxes, and text as prompts. Replacing the original second prompt encoder 1032 that only supports information such as points, boxes, and text as prompts with the image prompt generator 1042 enables the third prompt segmentation model 104 to support visual images as prompts, thus possessing the segmentation ability for unknown defects. In the implementation manner of this application, the role of the image prompt generator 1042 is to match the target area in the target image that is similar in category to the corresponding area of the prompt image mask. This process simulates the process of artificially giving prompt information such as points, boxes, and text as position priors.

[0073] When performing defect segmentation, the third prompt segmentation model 104 performs defect segmentation on the target image based on the prompt image to obtain the target segmentation result. Since the third prompt segmentation model 104 has the ability to segment unknown defects, when performing defect segmentation on the target image based on the prompt image through the third prompt segmentation model 104, it has strong generalization and high segmentation accuracy. Among them, the prompt image can be provided manually, and the number of prompt images can be multiple to support mixed categories as prompts. When the industrial segmentation dataset does not contain images of a certain type of defect, the defect of this type in the target image can be segmented by giving a prompt image containing this type of defect. For example, when the industrial segmentation dataset does not contain images of type A defects, the type A defects in the target image can be segmented by giving a prompt image containing type A defects. It has been found through research that although traditional image matching techniques can achieve similar defect segmentation functions, they need to formulate specific parameter configuration schemes for diverse industrial defect segmentation scenarios, and their final segmentation accuracy is not high.

[0074] It should be noted that the prompt segmentation models in the above initial prompt segmentation model 101, first prompt segmentation model 102, second prompt segmentation model 103, and third prompt segmentation model 104 can all be segmentation - everything models. The segmentation - everything model has a powerful generalization ability and can adapt to different types of images. In addition, the segmentation - everything model also has high precision and can accurately identify and segment objects in the image.

[0075] Please refer to Figure 2 and Figure 4 In some embodiments, fine - tuning the initial prompt segmentation model 101 using the industrial segmentation dataset to obtain the first prompt segmentation model 102 (i.e., 010) includes:

[0076] 011: Using the industrial segmentation dataset as the training set, and performing fine - tuning on the initial prompt segmentation model 101 in the LoRA fine - tuning manner to obtain the first prompt segmentation model 102.

[0077] In the embodiments of the present application, the initial prompt segmentation model 101 is fine - tuned in the LoRA (Low - Rank Adaptation of Large Language Models) fine - tuning manner to obtain the first prompt segmentation model 102, so that the first prompt segmentation model 102 can adapt to the image domain of the industrial scenario. Compared with the adapter fine - tuning manner, the LoRA fine - tuning manner fine - tunes the model parameters through a low - rank matrix, can not increase additional parameters during model inference, greatly reduces the calculation cost, and does not cause inference delay.

[0078] Please refer to Figure 2and Figure 5 In some embodiments, an industrial segmentation dataset is used as a training set, and the initial prompt segmentation model 101 is fine-tuned in a LoRA fine-tuning manner to obtain a first prompt segmentation model 102 (i.e., 011), including:

[0079] 0111: Construct an initial prompt segmentation model 101, where the initial prompt segmentation model 101 includes an initial image encoder 1011, an initial prompt encoder 1012, and an initial decoder 1013, and the initial image encoder 1011 uses ViT-L;

[0080] 0112: Freeze the initial prompt encoder 1012 and the initial decoder 1013;

[0081] 0113: Freeze the encoder parameters in ViT-L except for the attention mechanism parameters;

[0082] 0114: Use the industrial segmentation dataset as a training set and train the attention mechanism parameters in ViT-L in a LoRA fine-tuning manner to obtain a first prompt segmentation model 102.

[0083] Specifically, as Figure 2 shown, the initial prompt segmentation model 101 includes an initial image encoder 1011, an initial prompt encoder 1012, and an initial decoder 1013, while the first prompt segmentation model 102 includes a first image encoder 1021, a first prompt encoder 1022, and a first decoder 1023. The image encoder is the Figure 2 image encoder in Figure 2 and the decoder is the

[0084] mask decoder in

[0085] When performing fine-tuning, the initial prompt encoder 1012 and the initial decoder 1013 are frozen. It can be understood that the initial prompt encoder 1012 is used to process prompt information such as points, boxes, and text provided by the user, and the initial decoder 1013 is used to generate the final segmentation mask based on the image embedding vector output by the initial image encoder 1011 and the prompt information. Since the processing methods of the initial prompt encoder 1012 and the initial decoder 1013 are relatively general and already have the above capabilities, when performing fine-tuning, the parameters of the initial prompt encoder 1012 and the initial decoder 1013 are frozen to reduce the consumption of computing resources.

[0086] In the embodiments of the present application, an industrial segmentation dataset is used as the training set, and only the queries (Q) and keys (K) in the attention mechanism of each stage of ViT-L are trained using the LoRA fine-tuning method to achieve fine-tuning of the model without significantly increasing the model size. In addition, except for the above Q and K, the remaining encoder parameters in ViT-L are frozen to reduce the parameters to be updated during fine-tuning, speed up the training speed, and reduce the risk of overfitting.

[0087] In summary, the embodiments of the present application use the initial prompt segmentation model 101 with ViT-L as the initial image encoder 1011 as the basis for fine-tuning, and adopt the strategy of freezing most parameters and only performing LoRA fine-tuning on key parameters, which can achieve efficient and effective fine-tuning. After fine-tuning, the first prompt segmentation model 102 can not only maintain high accuracy in natural scenes, but also show high accuracy in industrial scenes where the image domain is very different from natural scenes.

[0088] It should be noted that after fine-tuning, the first image encoder 1021 can be obtained by fine-tuning the initial image encoder 1011, while the initial prompt encoder 1012 can be directly used as the first prompt encoder 1022, and the initial decoder 1013 can be directly used as the first decoder 1023. The first prompt segmentation model 102 can be determined according to the first image encoder 1021, the first prompt encoder 1022, and the first decoder 1023.

[0089] Please refer to Figure 2 and Figure 6 , in some embodiments, the first prompt segmentation model 102 includes a first image encoder 1021, a first prompt encoder 1022, and a first decoder 1023. The first prompt segmentation model 102 is distilled through a staged distillation method to obtain a second prompt segmentation model 103 (i.e., 020) that is lighter than the first prompt segmentation model 102, including:

[0090] 021: In the first-stage distillation process, the first image encoder 1021 is distilled;

[0091] 022: During the second-stage distillation process, distill the first image encoder 1021 and the first decoder 1023 to determine the second image encoder 1031 and the second decoder 1033;

[0092] 023: Determine the second prompt segmentation model 103 according to the second image encoder 1031, the first prompt encoder 1022, and the second decoder 1033.

[0093] Specifically, as Figure 2 shown, the first prompt segmentation model 102 includes the first image encoder 1021, the first prompt encoder 1022, and the first decoder 1023, while the second prompt segmentation model 103 includes the second image encoder 1031, the second prompt encoder 1032, and the second decoder 1033.

[0094] In the first-stage distillation, only the first image encoder 1021 is distilled. For example, using ViT-L as the teacher model, distill a lightweight pure convolutional structure model as the student model downward. The second-stage distillation is based on the first-stage distillation. Load the encoder weights of the first image encoder 1021 trained in the first-stage distillation, and introduce the first decoder 1023 to simultaneously distill the first image encoder 1021 and the first decoder 1023, and finally distill to obtain the second image encoder 1031 and the second decoder 1033. Among them, the second image encoder 1031 is more lightweight than the first image encoder 1021. The second decoder 1033 can be more lightweight than the first decoder 1023, or can be not more lightweight than the first decoder 1023, and there is no limitation here.

[0095] In the embodiments of the present application, through the two-stage distillation method, the student model can gradually learn the feature representation ability of the teacher model. The first-stage distillation process can transfer the knowledge of the first image encoder 1021 to the second image encoder 1031, and only distill the first image encoder 1021, reducing the training complexity and computational resource requirements. The second-stage distillation process introduces the first decoder 1023, which can further optimize the performance of the first image encoder 1021 and ensure the consistency of the final segmentation result by adjusting the first decoder 1023.

[0096] It should be noted that after two-stage distillation, the second image encoder 1031 and the second decoder 1033 can be distilled from the first image encoder 1021 and the first decoder 1023, and the first prompt encoder 1022 can be directly used as the second prompt encoder 1032. The second prompt segmentation model 103 can be determined according to the second image encoder 1031, the second prompt encoder 1032, and the second decoder 1033.

[0097] Please refer to Figure 2 and Figure 7 , in some embodiments, during the first-stage distillation process, distilling the first image encoder 1021 (i.e., 021) includes:

[0098] 0211: Performing an alignment operation on the image embedding vectors output by the first image encoder 1021 and the image embedding vectors output by the second image encoder 1031;

[0099] During the second-stage distillation process, distilling the first image encoder 1021 and the first decoder 1023 to determine the second image encoder 1031 and the second decoder 1033 (i.e., 022) includes:

[0100] 0221: Performing an alignment operation on the image embedding vectors output by the first image encoder 1021 and the image embedding vectors output by the second image encoder 1031, and performing an alignment operation on the segmentation results output by the first decoder 1023 and the segmentation results output by the second decoder 1033 to determine the second image encoder 1031 and the second decoder 1033.

[0101] Specifically, the image embedding vector is the Figure 2 image embedding in, and the segmentation result can be a segmentation mask.

[0102] During the first-stage distillation process, the way to perform the alignment operation can be: calculating the feature distillation loss (such as MSE loss) between the image embedding vectors output by the first image encoder 1021 and the image embedding vectors output by the second image encoder 1031 to align the image embedding vectors. Among them, when the feature distillation loss is smaller, it indicates that the similarity between the image embedding vectors output by the first image encoder 1021 and the image embedding vectors output by the second image encoder 1031 is higher. By minimizing the difference between the two, the image embedding vectors can be aligned.

[0103] During the second-stage distillation process, the alignment operation can be performed in the following ways: Calculate the feature distillation loss (such as MSE loss) between the image embedding vectors output by the first image encoder 1021 and the image embedding vectors output by the second image encoder 1031 to align the image embedding vectors; in addition, calculate the segmentation loss (such as Dice loss or cross-entropy loss) between the segmentation results output by the first decoder 1023 and the segmentation results output by the second decoder 1033 to align the segmentation results. Among them, when the feature distillation loss is smaller, it indicates that the similarity between the image embedding vectors output by the first image encoder 1021 and the image embedding vectors output by the second image encoder 1031 is higher. By minimizing the difference between the two, the image embedding vectors can be aligned. When the segmentation loss is smaller, it indicates that the similarity between the segmentation results output by the first decoder 1023 and the segmentation results output by the second decoder 1033 is higher. By minimizing the difference between the two, the segmentation results can be aligned.

[0104] In this way, it can be ensured that the lightweight second prompt segmentation model 103 can be maximally close to the first prompt segmentation model 102 in terms of feature extraction and segmentation tasks.

[0105] Please refer to Figure 2 , the second prompt segmentation model 103 includes a second image encoder 1031, a second prompt encoder 1032, and a second decoder 1033, while the third prompt segmentation model 104 includes a third image encoder 1041, an image prompt generator 1042, and a third decoder 1043. After obtaining the second prompt segmentation model 103, replace the second prompt encoder 1032 in the second prompt segmentation model 103 with the image prompt generator 1042, and the second image encoder 1031 can be directly used as the third image encoder 1041, and the second decoder 1033 can be directly used as the third decoder 1043. The third prompt segmentation model 104 can be determined according to the third image encoder 1041, the image prompt generator 1042, and the third decoder 1043.

[0106] Please refer to Figure 3 and Figure 8 , in some embodiments, the third prompt segmentation model 104 is used to perform defect segmentation on the target image based on the prompt image to obtain the target segmentation result (i.e., 040), including:

[0107] 041: Use the third prompt segmentation model 104 to perform defect segmentation on the target image based on the prompt image and output the initial segmentation result;

[0108] 042: Determine whether there is a missed defect detection according to the initial segmentation result;

[0109] 043: When there is no missed defect detection, use the initial segmentation result as the target segmentation result;

[0110] The defect segmentation method further includes:

[0111] 050: When there is a missed detection of a defect, add the image of the missed-detected defect to the prompt image.

[0112] Specifically, as Figure 2 and Figure 3 shown, the target image can be input into the third prompt segmentation model 104 by the third image encoder 1041, and the prompt image can be input into the third prompt segmentation model 104 by the image prompt generator 1042. In addition, the prompt mask corresponding to the prompt image (i.e., Figure 2 the mask in

[0113] can also be input into the third prompt segmentation model 104. The third image encoder 1041 outputs an image embedding vector according to the target image. After performing a convolution operation on the prompt mask, it is added to the image embedding vector and input into the third decoder 1043; at the same time, the prompt image is also input into the third decoder 1043 by the image prompt generator 1042. Finally, the third decoder 1043 outputs a segmentation result, and the segmentation result can be the segmentation mask of the target to be segmented. The segmentation result at this time can be used as the initial segmentation result. Figure 2 According to the initial segmentation result, it can be judged whether there is a missed detection of a defect in the target image, and this process can be judged manually. When there is no missed detection of a defect, it means that the segmentation result is accurate, and then the initial segmentation result can be used as the target segmentation result, that is, the final segmentation result. When there is a missed detection of a defect, it means that the segmentation result is inaccurate. At this time, add the image of the missed-detected defect to the prompt image, that is,

[0114] the prompt image input into the image prompt generator 1042 in

[0115] as an expansion of the prompt image. In this way, when performing defect segmentation on the target image subsequently, the third prompt segmentation model 104 can use the image of the missed-detected defect as a prompt to identify the defects of the corresponding category in the target image. Figure 3 and Figure 9 Through the above solution, when a missed detection occurs on-site, the third prompt segmentation model 104 supports using the image of the missed-detected defect as a prompt to detect more than 90% of similar missed-detected defects without retraining the third prompt segmentation model 104 (that is, without re-executing steps 010 to 030). In this way, the model iteration cycle of on-site projects can be greatly reduced, and the number of interventions by algorithm personnel can be reduced.

[0116] 060: Judge whether the number of images of the missed-detected defects reaches a predetermined number;

[0117] 070: When the number of undetected defect images reaches a predetermined number, use the third prompt segmentation model 104 as the initial prompt segmentation model 101, and return the step of fine-tuning the initial prompt segmentation model 101 using the industrial segmentation dataset to obtain the first prompt segmentation model 102.

[0118] Specifically, the predetermined number can be set in advance manually. When the number of undetected defect images still reaches the predetermined number by adding undetected defect images to the prompt images, then use the third prompt segmentation model 104 as the initial prompt segmentation model 101 and perform model training again (i.e., re-execute steps 010 - 030) to update the third prompt segmentation model 104. In this way, when performing defect segmentation through the third prompt segmentation model 104 subsequently, it can have a relatively high segmentation accuracy.

[0119] In summary, the industrial defect segmentation method with a support prompt function according to the embodiments of the present application can solve the problem of large human consumption caused by a long and frequent model iteration cycle in the industrial defect segmentation scenario, and has strong generalization ability, and can segment scene defects not seen by the model with a small number of prompt images.

[0120] Since the model has the ability to segment unknown defects, in the initial stage of the project when the sample size is small, compared with traditional methods, it has a higher detection accuracy; in the middle and late stages of the project, since more than 90% of the undetected defects can be solved by only adding prompt images without re-training the model, compared with traditional methods, it can reduce the project iteration cycle and the number of times algorithm personnel intervene, thereby achieving cost reduction and efficiency improvement.

[0121] It is experimentally obtained that the third prompt segmentation model 104 based on deep learning according to the embodiments of the present application: (1) While maintaining its high accuracy and strong generalization ability in natural scenes, it still has a competitive segmentation accuracy for industrial scenes with completely different image domains. (2) In the training and inference stages, the video memory resources used are both less than 12G, which has sufficient feasibility for on-site implementation of industrial detection projects and can be effectively applied on-site. (3) When no prompt images are added, the detection ability can reach more than 90% of the detection ability of a traditional segmentation model trained using 100 pieces of data as the training set. Therefore, it can play a great promoting role in the cold start stage of the project when the real samples are extremely few or even non-existent.

[0122] Please refer to Figure 2 、 Figure 3 and Figure 10, the defect segmentation device 100 of the embodiments of the present application includes a fine-tuning module 10, a distillation module 20, a replacement module 30, and a segmentation module 40. The defect segmentation method of the embodiments of the present application can be implemented by the defect segmentation device 100 of the embodiments of the present application. For example, the fine-tuning module 10 is used to fine-tune the initial prompt segmentation model 101 using an industrial segmentation dataset to obtain a first prompt segmentation model 102. The distillation module 20 is used to distill the first prompt segmentation model 102 through a staged distillation method to obtain a second prompt segmentation model 103 that is lightweight relative to the first prompt segmentation model 102. The replacement module 30 is used to replace the second prompt encoder 1032 in the second prompt segmentation model 103 with an image prompt generator 1042 to obtain a third prompt segmentation model 104, where the image prompt generator 1042 supports a prompt image as input. The segmentation module 40 is used to perform defect segmentation on the target image based on the prompt image through the third prompt segmentation model 104 to obtain a target segmentation result.

[0123] In some embodiments, the fine-tuning module 10 is specifically configured to use the industrial segmentation dataset as a training set and fine-tune the initial prompt segmentation model 101 using the LoRA fine-tuning method to obtain a first prompt segmentation model 102.

[0124] In some embodiments, the fine-tuning module 10 is specifically configured to: construct an initial prompt segmentation model 101, where the initial prompt segmentation model 101 includes an initial image encoder 1011, an initial prompt encoder 1012, and an initial decoder 1013, and the initial image encoder 1011 uses ViT-L; freeze the initial prompt encoder 1012 and the initial decoder 1013; freeze the encoder parameters in ViT-L except for the attention mechanism parameters; use the industrial segmentation dataset as a training set and train the attention mechanism parameters in ViT-L using the LoRA fine-tuning method to obtain a first prompt segmentation model 102.

[0125] In some embodiments, the first prompt segmentation model 102 includes a first image encoder 1021, a first prompt encoder 1022, and a first decoder 1023. The distillation module 20 is specifically configured to: during the first-stage distillation process, distill the first image encoder 1021; during the second-stage distillation process, distill the first image encoder 1021 and the first decoder 1023 to determine a second image encoder 1031 and a second decoder 1033; determine the second prompt segmentation model 103 according to the second image encoder 1031, the first prompt encoder 1022, and the second decoder 1033.

[0126] In some embodiments, the distillation module 20 is specifically configured to: perform an alignment operation on the image embedding vectors output by the first image encoder 1021 and the image embedding vectors output by the second image encoder 1031; perform an alignment operation on the image embedding vectors output by the first image encoder 1021 and the image embedding vectors output by the second image encoder 1031, and perform an alignment operation on the segmentation results output by the first decoder 1023 and the segmentation results output by the second decoder 1033, to determine the second image encoder 1031 and the second decoder 1033.

[0127] In some embodiments, the defect segmentation device 100 includes an adding module. The segmentation module 40 is specifically configured to: perform defect segmentation on the target image based on the prompt image through the third prompt segmentation model 104, and output an initial segmentation result; determine whether there is any missed detection of defects according to the initial segmentation result; when there is no missed detection of defects, use the initial segmentation result as the target segmentation result. The adding module is configured to add the missed detection defect images to the prompt image when there is a missed detection of defects.

[0128] In some embodiments, the defect segmentation device 100 further includes a judgment module and a return module. The judgment module is configured to judge whether the number of the missed detection defect images reaches a predetermined number. The return module is configured to, when the number of the missed detection defect images reaches the predetermined number, use the third prompt segmentation model 104 as the initial prompt segmentation model 101, and return the step of fine-tuning the initial prompt segmentation model 101 using the industrial segmentation dataset to obtain the first prompt segmentation model 102.

[0129] In some embodiments, the prompt segmentation model is a Segment Anything Model.

[0130] It should be noted that the explanations of the defect segmentation method in the foregoing embodiments are equally applicable to the defect segmentation device 100 in the embodiments of the present application, and will not be elaborated herein.

[0131] Please refer to Figure 11 , the defect segmentation system 200 in the embodiments of the present application includes one or more processors 210 and a memory 220, and the memory 220 stores a computer program. When the computer program is executed by the processor 210, the defect segmentation method in any of the foregoing embodiments is implemented.

[0132] For example, when the computer program is executed by the processor 210, the following defect segmentation method is implemented:

[0133] 010: Fine-tune the initial prompt segmentation model 101 using the industrial segmentation dataset to obtain the first prompt segmentation model 102;

[0134] 020: Distill the first prompt segmentation model 102 through staged distillation to obtain a second prompt segmentation model 103 that is lighter than the first prompt segmentation model 102;

[0135] 030: Replace the second prompt encoder 1032 in the second prompt segmentation model 103 with an image prompt generator 1042 to obtain a third prompt segmentation model 104, where the image prompt generator 1042 supports a prompt image as input;

[0136] 040: Based on the prompt image, perform defect segmentation on the target image through the third prompt segmentation model 104 to obtain a target segmentation result.

[0137] For another example, when the computer program is executed by the processor 210, the following defect segmentation method is implemented:

[0138] 0111: Construct an initial prompt segmentation model 101, where the initial prompt segmentation model 101 includes an initial image encoder 1011, an initial prompt encoder 1012, and an initial decoder 1013, and the initial image encoder 1011 uses ViT-L;

[0139] 0112: Freeze the initial prompt encoder 1012 and the initial decoder 1013;

[0140] 0113: Freeze the encoder parameters in ViT-L except for the attention mechanism parameters;

[0141] 0114: Use the industrial segmentation dataset as the training set and train the attention mechanism parameters in ViT-L using the LoRA fine-tuning method to obtain the first prompt segmentation model 102.

[0142] It should be noted that the explanations of the defect segmentation method in the foregoing embodiments also apply to the defect segmentation system 200 of the embodiments of the present application, and will not be elaborated here.

[0143] Please refer to Figure 12 , a computer-readable storage medium 300 of the embodiments of the present application, on which a computer program 310 is stored. When the program is executed by the processor 320, the defect segmentation method of any of the above embodiments is implemented.

[0144] For example, when the program is executed by the processor 320, the following defect segmentation method is implemented:

[0145] 010: Fine-tune the initial prompt segmentation model 101 using the industrial segmentation dataset to obtain the first prompt segmentation model 102;

[0146] 020: Distill the first prompt segmentation model 102 through staged distillation to obtain a second prompt segmentation model 103 that is lighter than the first prompt segmentation model 102;

[0147] 030: Replace the second prompt encoder 1032 in the second prompt segmentation model 103 with an image prompt generator 1042 to obtain a third prompt segmentation model 104, where the image prompt generator 1042 supports a prompt image as input;

[0148] 040: Perform defect segmentation on the target image based on the prompt image through the third prompt segmentation model 104 to obtain a target segmentation result.

[0149] For another example, when the program is executed by the processor 320, the following defect segmentation method is implemented:

[0150] 0111: Construct an initial prompt segmentation model 101, where the initial prompt segmentation model 101 includes an initial image encoder 1011, an initial prompt encoder 1012, and an initial decoder 1013, and the initial image encoder 1011 uses ViT-L;

[0151] 0112: Freeze the initial prompt encoder 1012 and the initial decoder 1013;

[0152] 0113: Freeze the encoder parameters in ViT-L except for the attention mechanism parameters;

[0153] 0114: Use the industrial segmentation dataset as the training set and train the attention mechanism parameters in ViT-L using the LoRA fine-tuning method to obtain the first prompt segmentation model 102.

[0154] It should be noted that the explanations of the defect segmentation method in the foregoing embodiments are equally applicable to the computer-readable storage medium 300 of the embodiments of the present application, and will not be elaborated here.

[0155] In summary, for the defect segmentation method, defect segmentation device 100, defect segmentation system 200, and computer-readable storage medium 300 according to the embodiments of the present application, first, the initial prompt segmentation model 101 is fine-tuned using an industrial segmentation dataset to obtain a first prompt segmentation model 102, so that the first prompt segmentation model 102 can adapt to the image domain of industrial scenarios to improve the segmentation accuracy; second, the first prompt segmentation model 102 is distilled through a staged distillation method to obtain a second prompt segmentation model 103 that is lightweight relative to the first prompt segmentation model 102 to reduce the computational load, so that it can be applied to industrial low-computing power scenarios; third, the second prompt encoder 1032 in the second prompt segmentation model 103 is replaced with an image prompt generator 1042 to obtain a third prompt segmentation model 104, where the image prompt generator 1042 supports a prompt image as input, so that the third prompt segmentation model 104 has the ability to segment unknown defects, and has strong generalization and high segmentation accuracy when defect-segmenting a target image based on the prompt image through the third prompt segmentation model 104.

[0156] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0157] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or more executable instructions for implementing a specific logical function or process, and the scope of the preferred embodiments of the present application includes additional implementations, where the functions can be executed in a manner that is not shown or discussed, including in a substantially simultaneous manner or in a reverse order according to the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application belong.

[0158] The logic and / or steps represented in the flowchart or otherwise described herein can, for example, be considered as a definitional sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable storage medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in conjunction with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a computer-readable storage medium can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable storage medium include the following: an electrical connection part with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable storage medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other appropriate processing as necessary, and then stored in a computer memory.

[0159] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well-known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGA), field programmable gate arrays (FPGA), etc.

[0160] Those of ordinary skill in the art can understand that all or part of the steps carried out in the method of the above embodiments can be completed by instructing relevant hardware through a program. The said program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments. In addition, in each of the embodiments of the present application, the functional units can be integrated into a processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. If the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disc, etc.

[0161] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application. The scope of the present application is defined by the claims and their equivalents.

Claims

1. A defect segmentation method, characterized in that: include: Using the industrial segmentation dataset, the initial prompt segmentation model is fine-tuned to obtain the first prompt segmentation model; Distilling the first prompt segmentation model in a staged distillation manner to obtain a second prompt segmentation model that is lighter than the first prompt segmentation model; Replacing the second hint encoder in the second hint segmentation model with an image hint generator to obtain a third hint segmentation model, wherein the image hint generator supports hint images as input; Defect segmentation is performed on the target image based on the prompt image using the third prompt segmentation model to obtain a target segmentation result.

2. The defect segmentation method according to claim 1, characterized in that: The method of fine-tuning the initial prompt segmentation model using the industrial segmentation dataset to obtain the first prompt segmentation model includes: The industrial segmentation dataset is used as a training set, and the initial prompt segmentation model is fine-tuned using the LoRA fine-tuning method to obtain the first prompt segmentation model.

3. The defect segmentation method according to claim 2, characterized in that: The method of using the industrial segmentation dataset as a training set and fine-tuning the initial prompt segmentation model using a LoRA fine-tuning method to obtain the first prompt segmentation model includes: Constructing the initial hint segmentation model, wherein the initial hint segmentation model includes an initial image encoder, an initial hint encoder and an initial decoder, and the initial image encoder adopts ViT-L; freezing the initial hint encoder and the initial decoder; Freeze the encoder parameters in ViT-L except the attention mechanism parameters; The industrial segmentation dataset is used as a training set, and the attention mechanism parameters in the ViT-L are trained using the LoRA fine-tuning method to obtain the first prompt segmentation model.

4. The defect segmentation method according to claim 1, characterized in that: The first prompt segmentation model includes a first image encoder, a first prompt encoder and a first decoder, and the first prompt segmentation model is distilled by a staged distillation method to obtain a second prompt segmentation model that is lighter than the first prompt segmentation model, including: In the first stage distillation process, the first image encoder is distilled; In the second stage distillation process, the first image encoder and the first decoder are distilled to determine a second image encoder and a second decoder; The second hint segmentation model is determined based on the second image encoder, the first hint encoder and the second decoder.

5. The defect segmentation method according to claim 4, characterized in that: The first image encoder is distilled in the first stage distillation process, including: Performing an alignment operation on the image embedding vector output by the first image encoder and the image embedding vector output by the second image encoder; In the second stage distillation process, distilling the first image encoder and the first decoder to determine a second image encoder and a second decoder includes: An alignment operation is performed on the image embedding vector output by the first image encoder and the image embedding vector output by the second image encoder, and an alignment operation is performed on the segmentation result output by the first decoder and the segmentation result output by the second decoder to determine the second image encoder and the second decoder.

6. The defect segmentation method according to claim 1, characterized in that: The step of performing defect segmentation on the target image based on the prompt image by using the third prompt segmentation model to obtain a target segmentation result includes: Perform defect segmentation on the target image based on the prompt image by using the third prompt segmentation model, and output an initial segmentation result; Determining whether there are any defects missed according to the initial segmentation result; When there is no defect missed detection, taking the initial segmentation result as the target segmentation result; The defect segmentation method further comprises: When there are defects that are missed, the missed defect image is added to the prompt image.

7. The defect segmentation method according to claim 6, characterized in that: When there is a defect missed detection, after adding the missed defect image to the prompt image, the defect segmentation method further includes: Determining whether the number of the missed-detection defect images reaches a predetermined number; When the number of the missed defect images reaches the predetermined number, the third prompt segmentation model is used as the initial prompt segmentation model, and the process returns to the step of fine-tuning the initial prompt segmentation model using the industrial segmentation dataset to obtain the first prompt segmentation model.

8. The defect segmentation method according to any one of claims 1 to 7, characterized in that: The suggested segmentation model is a segmentation everything model.

9. A defect segmentation device, characterized in that: include: A fine-tuning module, used for fine-tuning the initial prompt segmentation model using the industrial segmentation dataset to obtain a first prompt segmentation model; a distillation module, configured to distill the first prompt segmentation model in a staged distillation manner to obtain a second prompt segmentation model that is lighter than the first prompt segmentation model; a replacement module, configured to replace the second prompt encoder in the second prompt segmentation model with an image prompt generator to obtain a third prompt segmentation model, wherein the image prompt generator supports a prompt image as an input; The segmentation module is used to perform defect segmentation on the target image based on the prompt image through the third prompt segmentation model to obtain a target segmentation result.

10. A defect segmentation system, characterized in that: The defect segmentation system includes one or more processors and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the defect segmentation method according to any one of claims 1 to 8 is implemented.