A method and system for industrial surface defect detection based on large models
By introducing the rank decomposition matrix layer and the learnable feature fusion module into the large model, the problem that general large models cannot adapt to industrial surface defect detection is solved, and efficient and accurate defect detection is achieved.
Patent Information
- Application Number
- CN202411257019.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-09
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2044-09-09
AI Technical Summary
Existing general large models cannot be directly used for industrial surface defect detection due to the lack of defect image data and their semantic labels, which makes them unable to adapt to downstream tasks in specific scenarios.
An industrial surface defect detection model SAMDSS based on the visual large model SAM is adopted. The large model is fine-tuned by introducing a rank decomposition matrix layer, and a learnable feature fusion module is designed. The image encoder, channel fusion module and mask decoder are combined to perform defect detection.
It significantly improves the defect recognition capability in different industrial scenarios, improves detection accuracy and efficiency, and adapts to the needs of industrial surface defect detection.
Smart Images

Figure CN119205667B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of industrial surface defect detection, and in particular to a large model-based industrial surface defect detection method and system. Background Art
[0002] Defect detection aims to identify abnormal areas in mechanical components or products, such as indentations, dents, or unevenness on metal surfaces. These surface defects not only affect the product's appearance but can also severely impact its serviceability. Therefore, surface defect detection during the manufacturing process is crucial. Manual visual inspection is a traditional method for industrial product quality control. However, this method is often labor-intensive, inefficient, and prone to human error. In contrast, machine vision (including image processing and deep learning) is becoming increasingly popular in realizing intelligent manufacturing. These methods are more efficient than traditional methods and can effectively replace manual visual inspection.
[0003] The emergence of large models now enables researchers to solve a variety of problems using a unified framework. These include large models for vision, such as SAM and SegGPT. These models demonstrate excellent generalization capabilities, making them advantageous in practical applications. However, most large vision models are based on natural scenes rather than industrial scenes. This means that due to the lack of defect image data and its corresponding semantic labels, they cannot be directly applied to solve industrial surface defect detection problems. How to fully realize the potential of general large models and adapt them to downstream tasks in specific scenarios is an urgent problem to be solved. Summary of the Invention
[0004] In response to the problems and needs raised above, this solution proposes a large-scale model-based industrial surface defect detection method and system. By adopting the following technical features, it can achieve the above technical objectives and bring about many other technical effects.
[0005] An object of the present invention is to provide an industrial surface defect detection method based on a large model, comprising the following steps:
[0006] S10: Establish an industrial surface defect detection model SAMDSS based on the visual large model SAM. The industrial surface defect detection model SAMDSS includes: an image encoder, a channel fusion module, a prompt encoder and a mask decoder. The image encoder includes a Transformer layer and a rank decomposition matrix layer connected thereto, and is configured to extract image features; the channel fusion module is configured to process and fuse image features separately through multiple pooling strategies; the mask decoder is configured to receive the fused features of the channel fusion module and the default embedded features of the prompt encoder, and gradually restore spatial information to generate a high-resolution defect segmentation mask;
[0007] S20: Create a defect dataset and divide it into training samples and test samples;
[0008] S30: inputting the training sample into the industrial surface defect detection model SAMDSS to train the industrial surface defect detection model SAMDSS;
[0009] S40: Use the trained industrial surface defect detection model SAMDSS to infer the images in the test sample to measure its accuracy in industrial surface defect detection.
[0010] In addition, the large-model-based industrial surface defect detection method and device according to the present invention may also have the following technical features:
[0011] In one example of the present invention, the rank decomposition matrix layers of the image encoder include two, one of which is connected in parallel to the Q matrix of the self-attention mechanism unit of the Transformer layer and summed with the output features of the Q matrix; the other is connected in parallel to the K matrix of the self-attention mechanism unit of the Transformer layer and summed with the output features of the K matrix; the image encoder is configured to maintain rich prior knowledge while adapting to defect detection requirements through low-rank matrix decomposition.
[0012] In one example of the present invention, the expression of the rank decomposition matrix layer is:
[0013] W LoRA =W+ΔW
[0014] ΔW=BA
[0015] Where W represents the matrix from the Transformer layer for calculating the self-attention mechanism, B and A represent low-dimensional matrices, matrix B is initialized to all zeros, and matrix A is initialized to random values with a normal distribution.
[0016] In one example of the present invention, in step S10, the channel fusion module includes: an average pooling layer, a maximum pooling layer and a global average pooling layer; the channel fusion module processes the image features respectively through a plurality of pooling strategies and performs fusion, including the following steps: performing matrix multiplication operations on the image features in the vertical direction after the average pooling layer and in the horizontal direction after the maximum pooling layer to obtain a first output feature equal to the original size, processing the image features through the global average pooling layer to obtain a second output feature, and then fusing the first output feature processed by the activation function and the second output feature processed by the one-dimensional convolution to output.
[0017] In one example of the present invention, in step S30, when training SAMDSS, the maximum number of iterations is set to 20,000, and CE loss and Dice loss are used as loss functions. The loss function formula is expressed as follows:
[0018] L=(L CE +L Dice ) / 2
[0019] Where, L CE is the CE loss, L Dice Dice loss.
[0020] In one example of the present invention, in step S40, the performance indicators used are IoU and F1 to measure the accuracy of industrial surface defect detection;
[0021] The indicator IoU is the ratio of the intersection and union of the predicted segmentation and the true segmentation, which is expressed as follows:
[0022]
[0023] In the formula, target represents the true value and prediction represents the predicted value.
[0024] The indicator F1 is expressed as follows:
[0025]
[0026] Where TP is the number of true positives, FP is the number of false positives, and FN is the number of false negatives.
[0027] Another object of the present invention is to provide an industrial surface defect detection system based on a large model, comprising:
[0028] A model building module is configured to establish an industrial surface defect detection model SAMDSS based on a large visual model SAM. The industrial surface defect detection model SAMDSS includes: an image encoder, a channel fusion module, a prompt encoder, and a mask decoder. The image encoder includes a Transformer layer and a rank decomposition matrix layer connected thereto, and is configured to extract image features; the channel fusion module is configured to process and fuse image features separately through multiple pooling strategies; the mask decoder is configured to receive the fused features of the channel fusion module and the default embedded features of the prompt encoder, and gradually restore spatial information to generate a high-resolution defect segmentation mask;
[0029] A data set partitioning module is configured to establish a defect data set and divide it into training samples and test samples;
[0030] A data training module is configured to input training samples into the industrial surface defect detection model SAMDSS to train the industrial surface defect detection model SAMDSS;
[0031] The defect detection module is configured to use the trained industrial surface defect detection model SAMDSS to infer images in the test sample to measure its accuracy in industrial surface defect detection.
[0032] In one example of the present invention, the rank decomposition matrix layers of the image encoder include two, one of which is connected in parallel to the Q matrix of the self-attention mechanism unit of the Transformer layer and summed with the output features of the Q matrix; the other is connected in parallel to the K matrix of the self-attention mechanism unit of the Transformer layer and summed with the output features of the K matrix; the image encoder is configured to maintain rich prior knowledge while adapting to defect detection requirements through low-rank matrix decomposition.
[0033] In one example of the present invention, the expression of the rank decomposition matrix layer is:
[0034] W LoRA =W+ΔW
[0035] ΔW=BA
[0036] Where W represents the matrix from the Transformer layer for calculating the self-attention mechanism, B and A represent low-dimensional matrices, matrix B is initialized to all zeros, and matrix A is initialized to random values with a normal distribution.
[0037] In one example of the present invention, the channel fusion module includes: an average pooling layer, a maximum pooling layer and a global average pooling layer, wherein the average pooling layer and the maximum pooling layer perform matrix multiplication operations to obtain a first output feature equal to the original size, the first output feature is connected to an activation function, the global average pooling layer is connected to a one-dimensional convolution layer, and the activation function is connected to the one-dimensional convolution layer through an addition fusion operation.
[0038] The beneficial effects of the present invention are as follows:
[0039] The present invention fine-tunes the large model by introducing a rank decomposition matrix layer, so that it can adapt to the needs of industrial surface defect detection; the present invention designs a learnable feature fusion module to fuse global context information with local features, thereby improving the model's ability to perceive and understand defects; compared with existing mainstream defect image detection methods, the present invention demonstrates multi-faceted technological innovation and performance improvement, significantly improving the defect recognition ability in different industrial scenarios.
[0040] Hereinafter, the best embodiment of the present invention will be described in more detail with reference to the accompanying drawings so that the features and advantages of the present invention can be easily understood. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings of the embodiments of the present invention. The drawings are only used to illustrate some embodiments of the present invention, but not to limit all embodiments of the present invention thereto.
[0042] Figure 1 2 is a schematic structural diagram of a large model-based industrial surface defect detection method according to an embodiment of the present invention;
[0043] Figure 2 2. It is a structural diagram of the self-attention mechanism unit of LoRA in the Transformer layer according to an embodiment of the present invention;
[0044] Figure 3 2 is a schematic structural diagram of a channel fusion module according to an embodiment of the present invention;
[0045] Figure 4 This is the data enhancement result according to an embodiment of the present invention;
[0046] Figure 5 A comparison chart of experimental results of the detection method of the present invention and the detection method in the prior art;
[0047] Figure 6 This is a result diagram of detecting other types of defects using the detection method of the present invention. DETAILED DESCRIPTION
[0048] In order to make the purpose, technical solution and advantages of the technical solution of the present invention clearer, the technical solution of the embodiment of the present invention will be clearly and completely described below in conjunction with the drawings of specific embodiments of the present invention. The same figure marks in the drawings represent the same parts. It should be noted that the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the described embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0049] Unless otherwise defined, the technical or scientific terms used herein shall have the ordinary meaning understood by persons of ordinary skill in the field to which the invention belongs. The words "first", "second" and similar terms used in the patent application specification and claims of the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. Similarly, words such as "a" or "an" do not necessarily indicate a quantity limitation. Words such as "include" or "comprising" mean that the elements or objects preceding the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Words such as "connected" or "connected" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0050] According to a first aspect of the present invention, a large model-based industrial surface defect detection method is Figure 1 As shown, the following steps are included:
[0051] S10: Establish an industrial surface defect detection model SAMDSS based on the visual large model SAM. The industrial surface defect detection model SAMDSS includes: an image encoder, a channel fusion module, a prompt encoder and a mask decoder. The image encoder includes a Transformer layer and a rank decomposition matrix layer connected thereto, which is configured for extracting image features; the channel fusion module is configured to process and fuse image features separately through multiple pooling strategies; the mask decoder is configured to receive the fused features of the channel fusion module and the default embedded features of the prompt encoder, and gradually restore the spatial information to generate a high-resolution defect segmentation mask; it should be noted that the prompt encoder is used for feature input as a feature of the native model, and the mask decoder is used for feature output.
[0052] S20: Create a defect dataset and divide it into training samples and test samples;
[0053] S30: inputting the training sample into the industrial surface defect detection model SAMDSS to train the industrial surface defect detection model SAMDSS;
[0054] S40: Use the trained industrial surface defect detection model SAMDSS to infer the images in the test sample to measure its accuracy in industrial surface defect detection.
[0055] The present invention fine-tunes the large model by introducing a rank decomposition matrix layer, so that it can adapt to the needs of industrial surface defect detection; the present invention designs a learnable feature fusion module to fuse global context information with local features, thereby improving the model's ability to perceive and understand defects; compared with existing mainstream defect image detection methods, the present invention demonstrates multi-faceted technological innovation and performance improvement, significantly improving the defect recognition ability in different industrial scenarios.
[0056] In one example of the present invention, the rank decomposition matrix layers of the image encoder include two, one of which is connected in parallel to the Q matrix of the self-attention mechanism unit of the Transformer layer and summed with the output features of the Q matrix; the other is connected in parallel to the K matrix of the self-attention mechanism unit of the Transformer layer and summed with the output features of the K matrix; the image encoder is configured to use low-rank matrix decomposition (LoRA) to enable the model to maintain rich prior knowledge while adapting to defect detection requirements.
[0057] Specifically, if Figure 2 As shown in the figure, the self-attention mechanism unit of the Transformer layer includes a Q matrix, a K matrix and a V matrix connected in parallel with each other, wherein a rank decomposition matrix layer is connected in parallel to the bypass of the Q matrix and the K matrix respectively. The Q matrix and the rank decomposition matrix layer are summed and multiplied with the V matrix. The result is then input into the activation function together with the sum of the Q matrix and the rank decomposition matrix layer. The features output by the activation function are multiplied with the sum of the K matrix and the rank decomposition matrix to obtain the output features.
[0058] In short, the image encoder fine-tunes the large model by introducing LoRA under the existing Transformer layer architecture, greatly improving the accuracy of image feature extraction and making it adaptable to the needs of industrial surface defect detection.
[0059] In one example of the present invention, the expression of the rank decomposition matrix layer is:
[0060] W LoRA =W+ΔW
[0061] ΔW=BA
[0062] Where W represents the matrix from the Transformer layer for calculating the self-attention mechanism, B and A represent low-dimensional matrices, matrix B is initialized to all zeros, and matrix A is initialized to random values with a normal distribution.
[0063] In one example of the present invention, in step S10, as Figure 3As shown, the channel fusion module includes: an average pooling layer, a maximum pooling layer and a global average pooling layer; the channel fusion module processes the image features separately through multiple pooling strategies and performs fusion, including the following steps: performing matrix multiplication operations on the image features in the vertical direction after the average pooling layer and the horizontal direction after the maximum pooling layer to obtain a first output feature equal to the original size, processing the image features through the global average pooling layer to obtain a second output feature, and then fusing the first output feature processed by the activation function and the second output feature processed by the one-dimensional convolution for output.
[0064] In one example of the present invention, in step S30, when training SAMDSS, the maximum number of iterations is set to 20,000, and CE loss and Dice loss are used as loss functions. The loss function formula is expressed as follows:
[0065] L=(L CE +L Dice ) / 2
[0066] Where, L CE is the CE loss, L Dice Dice loss.
[0067] In one example of the present invention, in step S40, the performance indicators used are IoU and F1 to measure the accuracy of industrial surface defect detection;
[0068] The indicator IoU is the ratio of the intersection and union of the predicted segmentation and the true segmentation, which is expressed as follows:
[0069]
[0070] In the formula, tar get represents the true value, prediction represents the predicted value;
[0071] The indicator F1 is expressed as follows:
[0072]
[0073] Where TP is the number of true positives, FP is the number of false positives, and FN is the number of false negatives.
[0074] According to a second aspect of the present invention, a large model-based industrial surface defect detection system comprises:
[0075] A model building module is configured to establish an industrial surface defect detection model SAMDSS based on a large visual model SAM. The industrial surface defect detection model SAMDSS includes: an image encoder, a channel fusion module, a prompt encoder, and a mask decoder. The image encoder includes a Transformer layer and a rank decomposition matrix layer connected thereto, and is configured to extract image features; the channel fusion module is configured to process and fuse image features separately through multiple pooling strategies; the mask decoder is configured to receive the fused features of the channel fusion module and the default embedded features of the prompt encoder, and gradually restore spatial information to generate a high-resolution defect segmentation mask;
[0076] A data set partitioning module is configured to establish a defect data set and divide it into training samples and test samples;
[0077] A data training module is configured to input training samples into the industrial surface defect detection model SAMDSS to train the industrial surface defect detection model SAMDSS;
[0078] The defect detection module is configured to use the trained industrial surface defect detection model SAMDSS to infer images in the test sample to measure its accuracy in industrial surface defect detection.
[0079] In one example of the present invention, the rank decomposition matrix layers of the image encoder include two, one of which is connected in parallel to the Q matrix of the self-attention mechanism unit of the Transformer layer and summed with the output features of the Q matrix; the other is connected in parallel to the K matrix of the self-attention mechanism unit of the Transformer layer and summed with the output features of the K matrix; the image encoder is configured to maintain rich prior knowledge while adapting to defect detection requirements through low-rank matrix decomposition.
[0080] Specifically, if Figure 2 As shown in the figure, the self-attention mechanism unit of the Transformer layer includes a Q matrix, a K matrix and a V matrix connected in parallel with each other, wherein a rank decomposition matrix layer (LoRA) is connected in parallel to the bypass of the Q matrix and the K matrix respectively. The Q matrix and the rank decomposition matrix layer are summed and multiplied with the V matrix. The result is then input into the activation function together with the sum of the Q matrix and the rank decomposition matrix layer. The features output by the activation function are multiplied with the sum of the K matrix and the rank decomposition matrix to obtain the output features.
[0081] In short, the image encoder fine-tunes the large model by introducing LoRA under the existing Transformer layer architecture, greatly improving the accuracy of image feature extraction and making it adaptable to the needs of industrial surface defect detection.
[0082] In one example of the present invention, the expression of the rank decomposition matrix layer is:
[0083] W LoRA =W+ΔW
[0084] ΔW=BA
[0085] Where W represents the matrix from the Transformer layer for calculating the self-attention mechanism, B and A represent low-dimensional matrices, matrix B is initialized to all zeros, and matrix A is initialized to random values with a normal distribution.
[0086] In one example of the present invention, the channel fusion module includes: an average pooling layer, a maximum pooling layer and a global average pooling layer, wherein the average pooling layer and the maximum pooling layer perform matrix multiplication operations to obtain a first output feature equal to the original size, the first output feature is connected to an activation function, the global average pooling layer is connected to a one-dimensional convolution layer, and the activation function is connected to the one-dimensional convolution layer through an addition fusion operation.
[0087] The beneficial effects of the present invention are as follows:
[0088] The present invention fine-tunes the large model by introducing a rank decomposition matrix layer, so that it can adapt to the needs of industrial surface defect detection; the present invention designs a learnable feature fusion module to fuse global context information with local features, thereby improving the model's ability to perceive and understand defects; compared with existing mainstream defect image detection methods, the present invention demonstrates multi-faceted technological innovation and performance improvement, significantly improving the defect recognition ability in different industrial scenarios.
[0089] Specific cases
[0090] Take the surface defects of silicon steel strip under the interference of oil pollution as an example to illustrate:
[0091] Step S10: Establish a SAM-based industrial surface defect detection model (SAMDSS), introducing a trainable rank decomposition matrix into the SAM image encoder. Design a learnable channel fusion module between the image encoder and decoder, and use multiple pooling techniques (including maximum pooling, average pooling, and global pooling) to fuse global features.
[0092] like Figure 1 As shown in , the original weight parameters in the image encoder are frozen, and LoRA is introduced in the bypass of the Transformer layer (the self-attention mechanism unit of the Transformer layer) in the image encoder. Specifically, LoRA is located in the Q and K projection layers of the Transformer self-attention layer, as shown in Figure 2As shown. The image encoder part only updates a small part of the parameters of the LoRA layer, and the formula is as follows:
[0093] W LoRA =W+ΔW
[0094] ΔW=BA
[0095] Where W represents the matrix from the Transformer self-attention mechanism, B and A represent low-dimensional matrices, B is initialized to all zeros, and A is initialized to random values with a normal distribution;
[0096] Furthermore, the multi-scale information fusion decoder is introduced into SAM, including:
[0097] like Figure 3 As shown, for the image encoder output feature X in , passing through the maximum pooling layer, average pooling layer, and global average pooling layer respectively. Matrix multiplication is performed on the features in the vertical direction after the average pooling layer and the horizontal direction after the maximum pooling layer to obtain features of the same size as the original size. This is then fused with the features after the global average pooling layer to obtain a more comprehensive and integrated feature representation. The formula is as follows:
[0098] X out =sigmoid(Avgpool(X in )×Maxpool(X in )+Gapool(X in ))
[0099] Where, X in Output features for the image encoder, X out Output features of the channel fusion module, Avgpool is the average pooling layer, Maxpool is the maximum pooling layer, Gapool is the global average pooling layer, and sigmoid is the activation function.
[0100] Step S20: Create a defect dataset and divide it into training samples and test samples;
[0101] To construct a defect detection dataset, the present invention collected surface defect data of silicon steel strip, rail defects, and magnetic tile defects under oil contamination. Data augmentation methods were then used to increase the sample size of the data to generate a processed dataset, which was then divided into a training set and a test set.
[0102] It should be noted that if Figure 4As shown, to address the challenge of insufficient data, this paper introduces multiple techniques in image processing to simulate the generation of noisy data. Specifically, random Gaussian noise and Gaussian blur techniques are used to create images that simulate real noise. By varying the image contrast, brightness, and color saturation, and introducing optical distortion effects, defect image samples are generated under diverse environmental conditions. The dataset is randomly divided into training and test sets with an 8:2 ratio.
[0103] Step S30: SAMDSS is trained based on the training samples, wherein a weighted loss function of CE loss and Dice loss is used, and the parameters of the network are automatically updated using the AdamW optimizer.
[0104] Furthermore, when training SAMDSS, CE loss and Dice loss are used as loss functions, and AdamW is used as the optimizer. The loss function formula is expressed as follows:
[0105] L=(L CE +L Dice ) / 2
[0106] Furthermore, the batch size is set to 4; the two momentum parameters β1, β2 and weight decay of the AdamW optimizer are set to 0.9, 0.999 and 0.1; and the maximum number of iterations is set to 20,000.
[0107] During model training, data features are processed by the image encoder, further integrated by the channel fusion unit, and finally resolved by the decoder. A loss function measures the difference between the model-generated sequence and the target sequence. The AdamW algorithm adjusts model parameters based on the gradient information fed back by the loss function, continuously optimizing the model's prediction accuracy.
[0108] Step S40: Use the trained SAMDSS to infer the images in the test sample, calculate the performance index of the model on the test sample and perform model evaluation.
[0109] Furthermore, the validation set is used as the input of the model to evaluate the trained SAM and measure its accuracy in the defect detection task; wherein the performance indicators used are F1 and IoU.
[0110] Furthermore, the performance indicator IoU used is the ratio of the intersection and union of the predicted segmentation and the true segmentation, and the formula is expressed as follows:
[0111]
[0112] In the formula, target represents the true value and prediction represents the predicted value.
[0113] The F1 used here combines the two indicators of Precision and Recall, and the formula is as follows:
[0114]
[0115] Where TP is the number of true positives, FP is the number of false positives, and FN is the number of false negatives.
[0116] In order to fully verify the effectiveness of the design method, the results of introducing other modules between the image encoder and decoder are compared with the channel fusion module proposed by the present invention, as shown in Table 1 below:
[0117] Table 1: Experimental comparison table
[0118]
[0119]
[0120] like Figure 5 As shown in the figure, this paper evaluates the performance of various semantic segmentation models in defect detection, including DeeplabV3+, UNet, Segformer, etc. The comparison of evaluation index values is shown in the following table:
[0121] Table 2: Comparison of evaluation index values
[0122]
[0123] like Figure 6 As shown in the figure, the present invention fine-tunes the large model with a small number of defective images, and can still show good generalization performance in scenes not described in the dataset.
[0124] Compared to existing technologies, this method demonstrates greater accuracy and efficiency. It has significant application value in identifying product defects, quality control, and defect analysis. This detection method provides the manufacturing industry with a more precise and efficient detection tool, significantly improving defect identification capabilities during production to ensure product quality and safety.
[0125] The above describes in detail an exemplary implementation of a large-scale model-based industrial surface defect detection method and system proposed by the present invention with reference to preferred embodiments. However, those skilled in the art will understand that, without departing from the concept of the present invention, various modifications and variations can be made to the above-mentioned specific embodiments, and various technical features and structures proposed by the present invention can be combined in various ways without exceeding the scope of protection of the present invention, which is determined by the appended claims.
Claims
1. A large-scale model-based industrial surface defect detection method, characterized in that: The steps include: S10: Establish an industrial surface defect detection model SAMDSS based on the visual large model SAM. The industrial surface defect detection model SAMDSS includes: an image encoder, a channel fusion module, a prompt encoder and a mask decoder. The image encoder includes a Transformer layer and a rank decomposition matrix layer connected thereto, and is configured to extract image features; the channel fusion module is configured to process and fuse image features separately through multiple pooling strategies; the mask decoder is configured to receive the fused features of the channel fusion module and the default embedded features of the prompt encoder, and gradually restore spatial information to generate a high-resolution defect segmentation mask; S20: Create a defect dataset and divide it into training samples and test samples; S30: inputting the training sample into the industrial surface defect detection model SAMDSS to train the industrial surface defect detection model SAMDSS; S40: Use the trained industrial surface defect detection model SAMDSS to infer the images in the test sample to measure its accuracy in industrial surface defect detection.
2. The large model-based industrial surface defect detection method according to claim 1, characterized in that: The image encoder includes two rank decomposition matrix layers, one of which is connected in parallel to the Q matrix of the self-attention mechanism unit of the Transformer layer and summed with the output features of the Q matrix; the other rank decomposition matrix layer is connected in parallel to the K matrix of the self-attention mechanism unit of the Transformer layer and summed with the output features of the K matrix; the image encoder is configured to maintain rich prior knowledge of the model while adapting to defect detection requirements through low-rank matrix decomposition.
3. The large model-based industrial surface defect detection method according to claim 2, characterized in that: The expression of the rank decomposition matrix layer is: IN LoRA =W+ΔW ΔW=BA Where W represents the matrix from the Transformer layer for calculating the self-attention mechanism, B and A represent low-dimensional matrices, matrix B is initialized to all zeros, and matrix A is initialized to random values with a normal distribution.
4. The large model-based industrial surface defect detection method according to claim 1, characterized in that: In step S10, the channel fusion module includes: an average pooling layer, a maximum pooling layer and a global average pooling layer; the channel fusion module processes the image features respectively through multiple pooling strategies and performs fusion, including the following steps: performing matrix multiplication operations on the image features in the vertical direction after the average pooling layer and in the horizontal direction after the maximum pooling layer to obtain a first output feature equal to the original size, processing the image features through the global average pooling layer to obtain a second output feature, and then fusing the first output feature processed by the activation function and the second output feature processed by the one-dimensional convolution to output.
5. The large model-based industrial surface defect detection method according to claim 1, characterized in that: In step S30, when training SAMDSS, the maximum number of iterations is set to 20,000, and CE loss and Dice loss are used as loss functions. The loss function formula is expressed as follows: L=(L CE +L Dice ) / 2 Where, L CE is the CE loss, L Dice Dice loss.
6. The large model-based industrial surface defect detection method according to claim 1, characterized in that: In step S40, the performance indicators used are IoU and F1 to measure the accuracy of industrial surface defect detection; The indicator IoU is the ratio of the intersection and union of the predicted segmentation and the true segmentation, which is expressed as follows: In the formula, target represents the true value, and prediction represents the predicted value; The indicator F1 is expressed as follows: Where TP is the number of true positives, FP is the number of false positives, and FN is the number of false negatives.
7. An industrial surface defect detection system based on a large model, characterized in that: include: A model building module is configured to establish an industrial surface defect detection model SAMDSS based on a large visual model SAM. The industrial surface defect detection model SAMDSS includes: an image encoder, a channel fusion module, a prompt encoder, and a mask decoder. The image encoder includes a Transformer layer and a rank decomposition matrix layer connected thereto, and is configured to extract image features; the channel fusion module is configured to process and fuse image features separately through multiple pooling strategies; the mask decoder is configured to receive the fused features of the channel fusion module and the default embedded features of the prompt encoder, and gradually restore spatial information to generate a high-resolution defect segmentation mask; A data set partitioning module is configured to establish a defect data set and divide it into training samples and test samples; A data training module is configured to input training samples into the industrial surface defect detection model SAMDSS to train the industrial surface defect detection model SAMDSS; The defect detection module is configured to use the trained industrial surface defect detection model SAMDSS to infer images in the test sample to measure its accuracy in industrial surface defect detection.
8. The large model-based industrial surface defect detection system according to claim 7, characterized in that: The image encoder includes two rank decomposition matrix layers, one of which is connected in parallel to the Q matrix of the self-attention mechanism unit of the Transformer layer and summed with the output features of the Q matrix; the other rank decomposition matrix layer is connected in parallel to the K matrix of the self-attention mechanism unit of the Transformer layer and summed with the output features of the K matrix; the image encoder is configured to maintain rich prior knowledge of the model while adapting to defect detection requirements through low-rank matrix decomposition.
9. The large model-based industrial surface defect detection system according to claim 8, characterized in that: The expression of the rank decomposition matrix layer is: IN LoRA =W+ΔW ΔW=BA Where W represents the matrix from the Transformer layer for calculating the self-attention mechanism, B and A represent low-dimensional matrices, matrix B is initialized to all zeros, and matrix A is initialized to random values with a normal distribution.
10. The large model-based industrial surface defect detection system according to claim 7, characterized in that: The channel fusion module includes: an average pooling layer, a maximum pooling layer and a global average pooling layer, wherein the average pooling layer and the maximum pooling layer perform matrix multiplication operations to obtain a first output feature equal to the original size, the first output feature is connected to an activation function, the global average pooling layer is connected to a one-dimensional convolution layer, and the activation function is connected to the one-dimensional convolution layer through an addition fusion operation.
Citation Information
Patent Citations
Method and device for determining defects of substation equipment
CN113869437A
Defect detection meta-model construction method, defect detection method, equipment and medium
CN117495786A