Non-prompt gray fabric defect detection method and device based on segmentation of all large models

By splitting the combination of all large model (SAM) generator and decoder, the problem of grey cloth defect detection under silent conditions is solved, and efficient and robust grey cloth defect detection is achieved to adapt to various grey cloth scenarios.

CN120298313APending Publication Date: 2025-07-11HUAZHONG UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510287757.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The prior art is difficult to achieve stable detection of grey fabric defects under no prompt conditions, especially when the defects in warp knitted fabrics are highly similar to the background, and the robustness is insufficient.

Method used

Segment Anything Model (SAM) is used to set the defect mask prompt generator, and the image encoder and mask decoder are used to detect grey fabric defects under silent conditions. Combined with the significant area generator and the differentiable binary module to generate efficient mask prompts, and supervised training is used to use binary cross entropy and soft Dice loss.

Benefits of technology

The stable detection of grey fabric defects under silent conditions is achieved, which improves the adaptability and generalization performance of the model, and significantly improves the detection efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298313A_ABST
    Figure CN120298313A_ABST
Patent Text Reader

Abstract

The invention belongs to the related technical field of warp knitting gray fabric defect detection, and discloses a no-prompt gray fabric defect detection method and equipment based on a segmentation emery model, and the method comprises the steps: inputting image data of a to-be-detected gray fabric into the segmentation emery model to achieve gray fabric defect detection; the model comprises an image encoder, a mask decoder, a prompt encoder and a defect mask prompt generator; the image encoder is used for encoding input gray fabric image data to be detected to obtain image feature codes and encoder intermediate layer features; the defect mask prompt generator is used for processing the features of the middle layer of the encoder to obtain a salient region map, and then obtaining a final mask prompt; the prompt encoder is used for encoding the final mask prompt to obtain a prompt code; and the mask decoder is used for decoding the image feature code and the prompt code to obtain a predicted mask image, so that defect detection is realized. According to the invention, stable detection of gray fabric defects is realized without prompt.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field related to the detection of warp knitted fabric defects, and more specifically, relates to a method and device for detecting warp knitted fabric defects without prompts based on the Segment Anything Model (SAM). Background Art

[0002] The accurate detection and segmentation of surface defects on warp knitted fabrics are crucial for the production efficiency and quality control of the textile industry. However, due to the complex texture of fabrics and the diverse forms of defects in industrial scenarios, traditional methods based on manual or rules have obvious deficiencies in detection efficiency and robustness, making it difficult to meet actual requirements. Therefore, deep learning methods have gradually become the mainstream, which can utilize global information and local detail features to improve detection performance.

[0003] Among the existing methods, Chinese Patent CN118691626A relies on users to provide point or box prompts and cannot adapt to the scenario of detection without prompts, resulting in poor universality; Patent CN118691815A performs instance segmentation by fine-tuning SAM, but its prompt generation depends on the region prompt network. Moreover, the defects of warp knitted fabrics are highly similar to the background, and some defects are difficult to be recognized by conventional object detection networks, leading to insufficient robustness. Therefore, there is an urgent need for an intelligent detection method that can stably and efficiently complete the segmentation of warp knitted fabric defects without prompts. Summary of the Invention

[0004] In view of the above defects or improvement requirements of the prior art, the present invention provides a method and device for detecting warp knitted fabric defects without prompts based on the Segment Anything Model (SAM), aiming to solve the problem of how to stably detect warp knitted fabric defects without prompts.

[0005] To achieve the above object, according to one aspect of the present invention, there is provided a method for detecting warp knitted fabric defects without prompts based on the Segment Anything Model (SAM), the method comprising the following steps:

[0006] Inputting the image data of the warp knitted fabric to be detected into the Segment Anything Model (SAM) to achieve the detection of warp knitted fabric defects; wherein, the Segment Anything Model (SAM) includes an image encoder, a mask decoder, a prompt encoder, and a defect mask prompt generator;

[0007] The image encoder is configured to encode the input image data of the warp knitted fabric to be detected to obtain an image feature encoding and an encoder intermediate layer feature, and transmit the image feature encoding and the encoder intermediate layer feature to the mask decoder and the defect mask prompt generator respectively;

[0008] The defect mask prompt generator is configured to process the encoder intermediate layer feature to obtain a significant region map, and then obtain a final mask prompt;

[0009] The prompt encoder is used to encode the final mask prompt from the defect mask prompt generator to obtain a prompt encoding, and transmit the prompt encoding to the mask decoder;

[0010] The mask decoder is used to decode the image feature encoding and the prompt encoding to obtain a predicted mask image, realizing defect detection.

[0011] Further, the defect mask prompt generation module includes a salient region generator and a differentiable binarization module. The salient region generator is used to process the intermediate layer features of the encoder to obtain a salient region map, and transmit the salient region image to the differentiable binarization module; the differentiable binarization module performs binarization processing on the salient region image to obtain a final mask prompt.

[0012] Further, the salient region generator consists of four layers of attention module Block’ 1~4 , one layer of intermediate processing layer neck’, one layer of global attention module saliency_attention and one layer of channel fusion post-processing layer conv_out.

[0013] Further, Block’ i is a Transformer attention module, which learns defect features through the attention mechanism to distinguish the texture of fabric defects; neck’ contains two convolutional layers to unify the output sizes of the salient region generator and the image encoder; saliency_attention uses the global attention mechanism to extract the significant numerical regions in the feature map, that is, the saliency map S saliency ; conv_out is used for the generation of the final single-channel salient region, and the corresponding expression is:

[0014]

[0015] where, W conv *S saliency +b conv is a 1×1 convolution, which fuses the original input channels into a single channel; Performs batch normalization operation on the convolved features; ReLU(·), applies the ReLU activation function to the output of batch normalization for activation.

[0016] Further, the intermediate layer feature S9 of the image encoder is input into the salient region generator. The intermediate layer feature S9 extracts the image feature S’ through Block’ 1~4 . The image feature S’ and the feature S 12 of the last layer are fused to obtain the feature S”, and the feature fusion expression is:

[0017] S” = 0.5 * S’ + 0.5 * S 12

[0018] where the "+" operation is the addition of matrix elements; after the fusion is completed, saliency_attention is used to process S”, and a saliency map S of 256×64×64 is obtained. saliency , S saliency Finally, a single-channel saliency region image S’ of size 1×64×64 is obtained through conv_out processing. saliency .

[0019] Furthermore, Block’ of the saliency region generator 1~4 and neck’ have exactly the same structure as Block 9~12 and neck in the image encoder, and share the initial weights.

[0020] Furthermore, during the training of the Segment Anything large model, binary cross-entropy loss and soft Dice loss are used for joint supervised training. The expression of the binary cross-entropy loss is:

[0021]

[0022] where CE(·) represents the cross-entropy loss, predictions’ represents the prediction after activation by the sigmoid activation function, and gts represents the corresponding mask ground truth; the expression of the soft Dice loss is:

[0023]

[0024] where C represents the number of classes, and ∈ is the smoothing term; the expression of the overall loss function is:

[0025] loss = loss BCEseg + loss dice .

[0026] Furthermore, during the training of the Segment Anything large model, the weights of the image encoder and the prompt encoder are frozen, and only the weight parameters of the defect mask prompt generator and the mask decoder are updated.

[0027] The present invention also provides a prompt-free fabric defect detection system based on the Segment Anything large model. The system includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it executes the above-mentioned prompt-free fabric defect detection method based on the Segment Anything large model.

[0028] The present invention also provides a computer-readable storage medium, which stores machine-executable instructions. When the machine-executable instructions are called and executed by a processor, the machine-executable instructions prompt the processor to implement the above-mentioned method for detecting defects in grey cloth without prompting based on segmenting all large models.

[0029] In general, compared with the prior art, the above technical solution conceived by the present invention, the method and device for detecting defects of grey cloth without prompting based on segmenting all large models provided by the present invention mainly have the following beneficial effects:

[0030] 1. A defect mask prompt generator is set in the Segment Anything Model (SAM), which is used to process the encoder intermediate layer features to obtain a significant area map, and then obtain the final mask prompt. The final mask prompt is the prompt required by the mask decoder, does not rely on the area prompt, and thus realizes stable detection of grey cloth defects without prompts, overcoming the problem that high-precision point and frame prompts cannot be obtained in the grey cloth weaving industrial environment, and has good universality.

[0031] 2. The present invention makes full use of the advantages of the SAM model in feature extraction and generalization, and only trains the defect mask prompt generator to generate efficient prompts for the defect locations of grey cloth. This not only greatly accelerates the model training and reasoning process, but also significantly improves the model's adaptability and generalization performance to various grey cloth defect scenarios.

[0032] 3. The present invention successfully applies the segmentation of all large models to industrial scenarios, thus improving its practicality. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 It is a schematic diagram of the framework of a Segment Anything Model (SAM) adopted by a method for detecting defects of grey cloth without prompting based on the Segment Anything Model provided by the present invention;

[0034] Figure 2 yes Figure 1 Schematic diagram of the structure of the salient region generator in the segmentation model in . DETAILED DESCRIPTION

[0035] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0036] Please refer to Figure 1 and Figure 2 , the present invention provides a method for detecting fabric defects without prompts based on the Segment Anything Model. The detection method can efficiently detect fabric surface defects and includes the following steps: inputting the image data of the fabric to be detected into the Segment Anything Model to detect fabric defects; wherein, the Segment Anything Model includes an image encoder, a mask decoder, a prompt encoder, and a defect mask prompt generator;

[0037] The image encoder is used to encode the input image data of the fabric to be detected to obtain image feature encoding and encoder intermediate layer features, and transmit the image feature encoding and encoder intermediate layer features to the mask decoder and the defect mask prompt generator respectively;

[0038] The defect mask prompt generator is used to process the encoder intermediate layer features to obtain a saliency map, and then obtain the final mask prompt;

[0039] The prompt encoder is used to encode the final mask prompt from the defect mask prompt generator to obtain prompt encoding, and transmit the prompt encoding to the mask decoder;

[0040] The mask decoder is used to decode the image feature encoding and prompt encoding to obtain a predicted mask image, realizing defect detection.

[0041] The defect mask prompt generation module includes a saliency region generator and a differentiable binarization module. The saliency region generator is used to process the encoder intermediate layer features to obtain a saliency map, and transmit the saliency region image to the differentiable binarization module. The differentiable binarization module performs binarization processing on the saliency region image to obtain the final mask prompt.

[0042] The saliency region generator consists of four layers of attention module Block’ 1~4 , one layer of intermediate processing layer neck’, one layer of global attention module saliency_attention, and one layer of channel fusion post-processing layer conv_out. Block’ i is a Transformer attention module, which learns defect features through the attention mechanism to distinguish fabric defect textures; neck’ contains two convolutional layers to unify the output sizes of the saliency region generator and the image encoder; saliency_attention has a similar structure to Block’ i , but the number of heads in the multi-head attention mechanism in it is different from that of Block’ iChange 12 to 8, and saliency_attention uses the global attention mechanism to extract the significant numerical regions in the feature map, namely the saliency map S saliency ; conv_out is used to generate the final single-channel significant region, and the corresponding expression is:

[0043]

[0044] where W conv *S saliency +b conv is a 1×1 convolution that fuses the original input channels into a single channel; Perform batch normalization on the convolved features; ReLU(·) applies the ReLU activation function to the output of batch normalization for activation.

[0045] The image I in the image data is successively input into the 12-layer Transformer attention module Block through the image block operation 1~12 . The image I passes through Block i and generates the intermediate layer feature S i . Then the feature S 12 of the last layer is transmitted to an intermediate processing layer neck. The intermediate processing layer neck contains two convolutional layers, and its output is the final image feature encoding S embedding of the image I.

[0046] Input the intermediate layer feature S9 of the image encoder into the significant region generator. The intermediate layer feature S9 extracts the image feature S' through Block' 1~4 . The image feature S' is fused with the feature S 12 of the last layer to obtain the feature S". The feature fusion expression is:

[0047] S" = 0.5 * S' + 0.5 * S 12

[0048] where the "+" operation is the matrix element addition. After the fusion is completed, saliency_attention is used to process S" to obtain a saliency map S of 256×64×64 saliency , S saliency finally passes through conv_out processing to obtain a single-channel significant region image S' of size 1×64×64 saliency .

[0049] Among them, the Block' 1~4 of the significant region generator and neck' have the same structure as Block 9~12 and neck in the image encoder, and share the initial weights.

[0050] S’ saliency After preprocessing, it is input into the differentiable binarization module to obtain the final mask prompt. The binarization module can learn the optimal parameters during the training of the SAM model. The preprocessing includes: image size transformation, S’ saliency The size of the saliency image is 1×64×64. The image is upsampled to the size of the original image I of 1×H×W through interpolation to obtain the saliency image S” saliency Subsequently, a two-dimensional Gaussian kernel with a size of 7×7 is generated, and its expression is:

[0051]

[0052] where x and y are the distances from the corresponding positions in the kernel to the center, σ is the standard deviation, and then the Gaussian kernel is normalized:

[0053] Finally, the Gaussian kernel function is obtained. The Gaussian kernel function is used to perform Gaussian filtering on S” saliency to remove redundant noise and smooth the saliency region to obtain the saliency image heat. The mean μ and standard deviation σ of the saliency image heat are calculated, and the dynamic threshold threshold for binarization segmentation is obtained using μ and σ. The corresponding expression is:

[0054] threshold = μ + γ·σ + δ

[0055] where γ and δ are two learnable parameters in the model, which make the dynamic threshold reach the optimal during the training process. The differentiable binarization segmentation of the saliency image heat is realized using the dynamic threshold to obtain the final mask prompt binarized_mask of the binarized saliency region:

[0056] binarized_mask = Sigmoid((heat - threshold)*t)

[0057]

[0058] where t is a learnable parameter in the model, which makes the binarization effect of the attention image reach the best during the training process.

[0059] The final mask prompt of the binarized saliency region is input into the prompt encoder of SAM to generate the final prompt encoding.

[0060] The final prompt encoding and the image feature encoding S embedding are jointly input into the mask decoder of SAM, and finally the defect mask predictions are generated. During the training process of the SAM model, binary cross-entropy loss and soft Dice loss are used to jointly supervise the training. The expression of the binary cross-entropy loss is:

[0061]

[0062] Among them, CE(·) represents the cross-entropy loss, predictions’ represents the prediction after activation by the sigmoid activation function, and gts represents the corresponding masked ground truth; the expression of the soft Dice loss is:

[0063]

[0064] Among them, C represents the number of categories, and ∈ is a smoothing term to avoid the denominator being zero and affecting the calculation of the loss function. The expression of the overall loss function is:

[0065] loss = loss BCEseg + loss dice

[0066] During training, the SAM mask decoder will separately restore the predicted masks of each image in a training batch to their original image sizes and store them in a tuple, rather than a unified input size, to ensure better training results; when calculating the loss function, all prediction results and ground truth labels will be converted into one-dimensional vectors, and the prediction results and ground truth labels of the same batch will be concatenated along the first dimension for calculation.

[0067] During the model training process, the weights of the image encoder and prompt encoder of SAM are frozen, and only the weight parameters of the salient region generator, differentiable binarization module, and SAM mask decoder are updated.

[0068] The present invention also provides a SAM-based unprompted fabric defect detection system, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it executes the above-mentioned SAM-based unprompted fabric defect detection method.

[0069] The present invention also provides a computer-readable storage medium, which stores machine-executable instructions. When the machine-executable instructions are called and executed by the processor, the machine-executable instructions cause the processor to implement the above-mentioned SAM-based unprompted fabric defect detection method.

[0070] It is easy for those skilled in the art to understand that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for detecting fabric defects without prompts based on the Segment Anything Model, characterized in that, The method includes the following steps: Input the image data of the to-be-detected blank fabric into the Segment-Anything large model to achieve defect detection of the blank fabric; wherein, the Segment-Anything large model includes an image encoder, a mask decoder, a prompt encoder, and a defect mask prompt generator; The image encoder is used to encode the input image data of the to-be-detected blank fabric to obtain an image feature encoding and intermediate layer features of the encoder, and transmit the image feature encoding and the intermediate layer features of the encoder to the mask decoder and the defect mask prompt generator respectively; The defect mask prompt generator is used to process the intermediate layer features of the encoder to obtain a saliency region map, and then obtain a final mask prompt; The prompt encoder is used to encode the final mask prompt from the defect mask prompt generator to obtain a prompt encoding, and transmit the prompt encoding to the mask decoder; The mask decoder is used to decode the image feature encoding and the prompt encoding to obtain a predicted mask image, achieving defect detection.

2. The method for detecting fabric defects without prompts based on the Split Everything Model according to claim 1, characterized in that: The defect mask prompt generator includes a saliency region generator and a differentiable binarization module. The saliency region generator is used to process the intermediate layer features of the encoder to obtain a saliency region map, and transmit the saliency region image to the differentiable binarization module; The differentiable binarization module performs binarization processing on the saliency region image to obtain a final mask prompt.

3. The method for detecting fabric defects without prompts based on the segmented all - in - one large model according to claim 2, wherein: The significant region generator consists of four layers of attention modules Block’ 1~4 , one layer of intermediate processing layer neck’, one layer of global attention module saliency_attention, and one layer of channel fusion post-processing layer conv_out.

4. The method for detecting fabric defects without prompts based on the segmented all - in - one large model according to claim 3, wherein: Block’ i is the Transformer attention module, which learns defect features through the attention mechanism to distinguish the texture of fabric defects; neck’ contains two convolutional layers to unify the output sizes of the saliency region generator and the image encoder; saliency_attention uses the global attention mechanism to extract the significant numerical regions in the feature map, namely the saliency map S saliency ; conv_out is used to generate the final single-channel saliency region, and the corresponding expression is: Among them, W conv *S saliency +b conv is a 1×1 convolution that fuses the original input channels into a single channel; perform batch normalization on the convolved features; ReLU(·), apply the ReLU activation function to activate the output of batch normalization.

5. The method for detecting fabric defects without prompts based on the Split Everything Model according to claim 4, wherein: Input the intermediate layer feature S9 of the image encoder into the salient region generator. The intermediate layer feature S9 passes through Block’ 1~4 to extract the image feature S’. The image feature S’ and the feature S of the last layer 12 are feature-fused to obtain the feature S”. The feature fusion expression is: S” = 0.5 * S’ + 0.5 * S 12 Among them, the "+" operation is the addition of matrix elements; after the fusion is completed, saliency_attention is used to process S", and a saliency map S is obtained saliency , S saliency Finally, a single-channel saliency region image S' is obtained through the processing of conv_out saliency .

6. The method for detecting fabric defects without prompts based on the split all - in - one large model according to claim 3, characterized in that: Block of the significant region generator 1~4 and neck' are exactly the same as the Block 9~12 and neck in the image encoder in terms of structure and share the initial weights.

7. The method for detecting fabric defects without prompts based on the split all - in - one large model according to any one of claims 1 - 6, characterized in that: When training the Segment-Anything large model, binary cross-entropy loss and soft Dice loss are used for joint supervised training. The expression of the binary cross-entropy loss is: wherein, CE(·) represents the cross-entropy loss, predictions’ represents the prediction after being activated by the sigmoid activation function, and gts represents the corresponding mask ground truth; the expression of the soft Dice loss is: wherein, C represents the number of classes, and ∈ is a smoothing term; the expression of the overall loss function is: loss=loss BCEseg +loss dice 。 8. The method for detecting fabric defects without prompts based on the segmented all-large model according to claim 7, wherein: During the training process of the Segment-Anything large model, the weights of the image encoder and the prompt encoder are frozen, and only the weight parameters of the defect mask prompt generator and the mask decoder are updated.

9. A zero-shot fabric defect detection system based on the Segment Anything Model, characterized in that: The system includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, it executes the method for detecting blank fabric defects without prompts based on the Segment-Anything large model according to any one of claims 1-8.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores machine-executable instructions. When the machine-executable instructions are called and executed by the processor, the machine-executable instructions cause the processor to implement the method for detecting blank fabric defects without prompts based on the Segment-Anything large model according to any one of claims 1-8.

Citation Information

Patent Citations

  • Image segmentation method and device, equipment and medium

    CN118691626A

  • Remote sensing image high-quality automatic instance segmentation method based on SAM large model fine tuning

    CN118691815A