Dinov3 and SAM-based few-sample industrial defect target detection method and system

By combining Dinov3 and SAM models, a defect reference feature library is generated and high-confidence detection is performed, which solves the detection dilemma caused by the scarcity of defect samples, enables rapid deployment and efficient data accumulation, improves detection accuracy and stability, and lays the foundation for subsequent model optimization.

CN120953269AActive Publication Date: 2025-11-14TROY INFORMATION TECHNOLOGY CO LTD

Patent Information

Application Number
CN202511468203.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-15
Publication Date
2025-11-14
Estimated Expiration
2045-10-15

AI Technical Summary

Technical Problem

In the early stages of new product trial production, production line process changes, or the emergence of new defect patterns, the number of defect samples is extremely small, which makes it impossible to effectively train traditional detection models or result in extremely poor performance, creating a dilemma of 'having defects but being unable to reliably detect them'.

Method used

By combining Dinov3 and SAM models, a method is adopted to generate a defect reference feature library, calculate pixel-wise similarity, filter neighborhoods, filter using spatial attention mechanism, and perform fine segmentation. Combined with semantic matching, texture consistency, and shape regularity evaluation, it can achieve rapid deployment and high-confidence detection, and automatically store the detection results to accumulate labeled data.

Benefits of technology

Achieving rapid and effective detection in the early stages when defect samples are scarce, automatically accumulating high-quality labeled data, improving detection accuracy and stability, reducing false positives and false negatives, and promoting subsequent model optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953269A_ABST
    Figure CN120953269A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence and industrial visual inspection, and discloses a Dinov3 and SAM-based few-sample industrial defect target detection method and system, and the method comprises the steps: building a reference feature library: extracting the features of a few defect reference images through a Dinov3 model, and generating a category prototype through weighted aggregation; generating a similarity response diagram: calculating the pixel-by-pixel similarity of the image to be detected and the reference prototype, and performing context enhancement filtering; generating a segmentation prompt: screening a salient region in combination with a space attention mechanism, and generating a geometric prompt required by the SAM model; fine segmentation is executed; an accurate mask of the defect instance is generated by using an SAM model; and confidence evaluation: multi-dimensional scoring is carried out, and a dynamic threshold value is adopted to screen results. Rapid deployment can be realized without fine adjustment of the model, the problem of data shortage in the initial stage is effectively solved, data is continuously accumulated through automatic detection, a foundation is laid for training a better special model, and the method is suitable for scenes such as new product import or new defect discovery.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and industrial visual inspection technology, and in particular to a method and system for detecting industrial defects based on Dinov3 and SAM with a small sample size. Background Technology

[0002] In intelligent manufacturing and automated quality inspection systems, deep learning models have become a core technology for defect detection. However, their successful application heavily relies on supervised training using large-scale, high-quality labeled datasets. In the initial stages of new product trial production, production line process changes, or the emergence of new defect patterns, the number of defect samples is extremely small, or even zero. This leads to traditional detection models being unable to be effectively trained or performing poorly, creating a dilemma of "defects present but unreliable detection." Therefore, there is an urgent need for a technical solution that can be rapidly deployed and effectively detected in the initial stages when defect samples are extremely scarce.

[0003] Recently, the self-supervised learning model DINOv3 has demonstrated powerful zero-shot transfer capabilities in visual representation learning through large-scale unsupervised pre-training. SAM, as a general segmentation foundation model, can generate high-quality instance masks based on geometric cues without requiring fine-tuning for specific categories. Combining these two models enables a fast and effective technical solution, not only for immediate quality monitoring but, more importantly, for providing a data foundation for subsequently building high-quality labeled datasets and training better specialized models. Summary of the Invention

[0004] This invention provides a method and system for detecting industrial defects using a limited number of samples based on Dinov3 and SAM. It aims to solve the problem of insufficient defect samples in the early stages of industrial projects, which makes it impossible to train an effective detection model. This enables rapid deployment and accurate detection, and accumulates high-quality labeled data for subsequent model iterations.

[0005] To achieve the above objectives, this invention provides a method for detecting industrial defects using a small sample size based on Dinov3 and SAM, the method comprising the following steps: S1: Obtain reference images for different industrial defect categories, extract reference image features using a pre-trained Dinov3 model, perform weighted aggregation of reference image features for each industrial defect category, generate category prototype feature vectors, and construct a defect reference feature library. S2: Use the Dinov3 model to extract features from the image to be tested, calculate pixel-by-pixel similarity between the features of the image to be tested and the category prototype feature vectors in the defect reference feature library, generate an initial response map, and use a neighborhood filtering algorithm to enhance it to obtain an enhanced response map. S3: Based on the enhanced response map, a spatial attention mechanism is introduced to filter significant response regions and generate prompt bounding boxes or prompt key points corresponding to the enhanced response map. S4: Input the cue bounding box or cue key points into SAM, perform fine segmentation on the significant response regions in the enhanced response map, and output the defect instance mask; S5: Combining semantic matching strength, regional texture consistency, and shape regularity, calculate the comprehensive confidence score of each segmentation result, use dynamic thresholds to filter high-confidence results, and output the final detection result.

[0006] Optionally, in step S1, when generating the category prototype feature vector by weighted aggregation of the reference image features for each industrial defect category, the weights used in the weighted aggregation are configured to be determined based on the average cosine similarity between each reference image feature and other reference image features of the same category. The expressions for the weights used in weighted aggregation are as follows: In the formula, For the first Weights of features in a reference image. The average cosine similarity between this feature and the other n-1 similar reference features. For temperature coefficient, The total number of reference images for this defect category; This is the aggregated category prototype feature vector. For the first Features of a reference image.

[0007] Optionally, in step S2, the neighborhood filtering algorithm is used to enhance the response map and obtain an enhanced response map. This is configured as follows: the neighborhood filtering algorithm is used to calculate the similarity between the response values ​​of each spatial location in the initial response map and its local neighborhood, and the original response is weighted and smoothed. The expression for the filtered initial response map is as follows: in, This is the initial response diagram. This is the enhanced response diagram after filtering. For coordinates A local window centered on the center. For Gaussian kernel function, This is the kernel width parameter.

[0008] Optionally, in step S3, when introducing a spatial attention mechanism for salient response region filtering, the expression for the attention weight map generated by the spatial attention mechanism is as follows: In the formula, For attention weights, To enhance the response map, For learnable location encoding, This indicates concatenation along the channel dimension. This represents a 3×3 convolution operation. The sigmoid activation function is used, and the significant response region is formed by... get, This indicates element-wise multiplication.

[0009] Optionally, in step S5, the comprehensive confidence score of each segmentation result is calculated by combining semantic matching strength, region texture consistency, and shape regularity. The specific expression is as follows: In the formula, To calculate the overall confidence score, The normalized semantic matching strength. This is a texture consistency score calculated based on local gray-level variance. For shape regularity scoring based on contour compactness calculation, , , The weight coefficients are non-negative and satisfy the following conditions: .

[0010] Optionally, the normalized semantic matching strength Texture consistency score and shape regularity score The expression is as follows: In the formula, For mask Enhanced response map within the region The average value, The attenuation coefficient is... For mask Gray-level variance in the original gray-level image For the area, For the perimeter of the outline, The 1st step after performing fine-grained segmentation on the salient response region Candidate regions, For mask Enhanced response map within the region The Middle The average value of each candidate region.

[0011] Optionally, in step S5, the expression for the dynamic threshold is specifically as follows: in, For dynamic thresholds, and These are the mean and standard deviation of the overall confidence scores for all candidate regions in the current image under test, respectively. These are the sensitivity adjustment parameters.

[0012] Optionally, a dynamic threshold is used to filter high-confidence results, and the final detection result is output, configured as follows: when the overall confidence score of the segmentation result is... Above the dynamic threshold If the result is positive, retain it; otherwise, delete it as a low-confidence result.

[0013] Optionally, the method further includes: S6: Automatically store the high confidence results output in step S5 and their corresponding association information into the defect sample database to form a structured labeled dataset for subsequent model training; The associated information stored in the defect sample database includes: the original image to be tested, the mask of the detected defect instance, the defect category label, the detection timestamp, and the comprehensive confidence score.

[0014] Furthermore, to achieve the above objectives, the present invention also provides a few-sample industrial defect target detection system based on Dinov3 and SAM, comprising: The module is used to acquire reference images for different industrial defect categories, extract reference image features through a pre-trained Dinov3 model, perform weighted aggregation of reference image features for each industrial defect category, generate category prototype feature vectors, and build a defect reference feature library. The enhancement module is used to extract features of the image under test using the Dinov3 model, calculate pixel-by-pixel similarity between the features of the image under test and the category prototype feature vectors in the defect reference feature library, generate an initial response map, and enhance it using a neighborhood filtering algorithm to obtain an enhanced response map. The generation module is used to introduce a spatial attention mechanism to filter significant response regions based on the enhanced response map and generate prompt bounding boxes or prompt key points corresponding to the enhanced response map. The segmentation module is used to input the cue bounding box or cue key points into SAM, perform fine segmentation on the significant response regions in the enhanced response map, and output defect instance masks. The filtering module combines semantic matching strength, regional texture consistency, and shape regularity to calculate the comprehensive confidence score of each segmentation result, uses dynamic thresholds to filter high-confidence results, and outputs the final detection result.

[0015] The beneficial effects of this invention are as follows: Solving the problem of scarce initial samples: In the early stages of a project, when the number of defective samples is insufficient to train a conventional supervised model, rapid deployment and effective detection can be achieved using only a small number of reference images, filling the gap in quality monitoring. Efficient data accumulation: Through automated detection and high-confidence result screening, high-quality, structured defect annotation data is continuously generated, which significantly accelerates the construction process of the dataset required for subsequent dedicated model training; Enhance system robustness: By incorporating a powerful feature extraction model and a general segmentation base model, along with weighted feature aggregation, context enhancement, and multi-dimensional confidence assessment, the system effectively improves detection accuracy and stability in complex industrial environments, and reduces false alarms and false negatives. Facilitating model iteration: It not only provides detection capabilities with fewer samples, but more importantly, it establishes a solid data foundation for training specialized supervised models with better performance, forming a virtuous cycle of data accumulation and model optimization. Attached Figure Description

[0016] Figure 1 This is a flowchart illustrating the few-sample industrial defect target detection method based on Dinov3 and SAM according to an embodiment of the present invention. Figure 2 This is a schematic diagram illustrating the principle of the few-sample industrial defect target detection method based on Dinov3 and SAM according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the architecture of the few-sample industrial defect target detection method based on Dinov3 and SAM according to an embodiment of the present invention; Figure 4 This is a flowchart illustrating the feature extraction process in the few-sample industrial defect target detection method based on Dinov3 and SAM according to an embodiment of the present invention. Figure 5 This is a schematic diagram of the structure of a few-sample industrial defect target detection system based on Dinov3 and SAM according to an embodiment of the present invention. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0018] This invention provides a method for detecting industrial defects using a small sample size based on Dinov3 and SAM, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the few-sample industrial defect target detection method based on Dinov3 and SAM according to an embodiment of the present invention.

[0019] In this embodiment, a method for detecting industrial defects using a small sample size based on Dinov3 and SAM includes the following steps: S1: Obtain reference images for different industrial defect categories, extract reference image features using a pre-trained Dinov3 model, perform weighted aggregation of reference image features for each industrial defect category, generate category prototype feature vectors, and construct a defect reference feature library. S2: Use the Dinov3 model to extract features from the image to be tested, calculate pixel-by-pixel similarity between the features of the image to be tested and the category prototype feature vectors in the defect reference feature library, generate an initial response map, and use a neighborhood filtering algorithm to enhance it to obtain an enhanced response map. S3: Based on the enhanced response map, a spatial attention mechanism is introduced to filter significant response regions and generate prompt bounding boxes or prompt key points corresponding to the enhanced response map. S4: Input the cue bounding box or cue key points into SAM, perform fine segmentation on the significant response regions in the enhanced response map, and output the defect instance mask; S5: Combining semantic matching strength, regional texture consistency, and shape regularity, calculate the comprehensive confidence score of each segmentation result, use dynamic thresholds to filter high-confidence results, and output the final detection result.

[0020] It should be noted that deep learning models have become a core technology for defect detection in intelligent manufacturing and automated quality inspection systems. However, their successful application heavily relies on supervised training using large-scale, high-quality labeled datasets. In the initial stages of new product trial production, production line process changes, or the emergence of new defect patterns, the number of defect samples is extremely small, or even zero. This leads to traditional detection models being unable to be effectively trained or performing poorly, creating a dilemma of "defects existing but unreliable detection." Therefore, there is an urgent need for a technical solution that can be rapidly deployed and effectively detected in the initial stages when defect samples are extremely scarce.

[0021] To address the aforementioned issues, this embodiment constructs a reference feature library: features from a small number of defect reference images are extracted using the Dinov3 model, and category prototypes are generated through weighted aggregation; a similarity response map is generated: pixel-wise similarity between the test image and the reference prototype is calculated, and context-enhanced filtering is performed; segmentation cues are generated: salient regions are selected using a spatial attention mechanism to generate geometric cues required by the SAM model; refined segmentation is performed: precise masks of defect instances are generated using the SAM model; confidence assessment is conducted: multi-dimensional scoring is performed based on semantic matching, texture consistency, and shape regularity, and results are filtered using dynamic thresholds; data accumulation is achieved: high-confidence detection results are automatically stored to form a structured labeled dataset. Therefore, rapid deployment can be achieved without model fine-tuning, effectively addressing the initial data scarcity problem, and data is continuously accumulated through automated detection, laying the foundation for training better specialized models. This approach is suitable for scenarios such as new product introduction or new defect discovery.

[0022] To more clearly explain the present invention, a specific example of a few-sample industrial defect target detection method based on Dinov3 and SAM is provided below, such as... Figure 2-4 As shown, it includes the following execution process: Step S1: Construct a reference feature library.

[0023] Acquire a small number of reference image samples for each type of industrial defect. In this embodiment, for each defect category, collect 3-10 reference images with a resolution of 1280x1280 pixels. Each image contains a typical defect instance and is equipped with a corresponding mask.

[0024] Each reference image is input into a pre-trained Dinov3 model with frozen weights, and its output feature maps are extracted. For each reference image According to its mask Calculate the average feature vector of the scratched area as the prototype for this instance: in Indicates mask The set of non-zero pixel coordinates.

[0025] Weighted aggregation is performed on multiple instance prototypes of the same defect category. First, the first... An instance prototype Mean cosine similarity with other similar instance prototypes: in The total number of reference images for this category. This is the cosine similarity function.

[0026] Based on similarity Calculate weights : in The value is 2 for the temperature coefficient in this embodiment.

[0027] Aggregate to generate the category prototype feature vector for this defect category: Will Store it in the defect feature library for subsequent matching.

[0028] Step S2: Generate a semantic similarity response map.

[0029] Acquire the image to be tested The image size is 1280x1280 pixels. It is then input into the Dinov3 model to extract feature maps. .Will Each spatial location eigenvectors Compared with the category prototype constructed in step S1 Perform cosine similarity calculation to generate an initial semantic response map. Step S3: Generate segmentation hints.

[0030] For the initial response diagram Perform feature aggregation operation: The response map was smoothed using a Gaussian kernel-weighted average. For each location... In its 3x3 neighborhood Internal calculation: in, For Gaussian kernel function, The kernel width parameter is set to 0.1. This operation effectively suppresses isolated noise responses and enhances the connectivity of continuous defect regions.

[0031] Will With learnable two-dimensional sinusoidal position coding Concatenate along the channel dimension, then input a 3x3 convolutional layer followed by... Activation function to generate spatial attention weight map: Weighting the response graph: The enhanced semantic response map is obtained. .

[0032] Enhance response map Upsample to the original image resolution. Apply threshold segmentation (threshold set to 0.5) to the upsampled response map to obtain a binary saliency map. Identify salient regions through connected component analysis, and calculate the minimum bounding rectangle of each connected component as a bounding box cue input to the SAM model.

[0033] Step S4: Perform refined instance segmentation.

[0034] Input the bounding box prompts generated in step S3 into the SAM model (e.g., the second-generation SMA2 model under the SAM series models), perform fine segmentation on the candidate regions, and output the binary mask of the defect instances. .

[0035] Step S5: Confidence assessment and result optimization.

[0036] Calculate the overall confidence score for each segmentation result. : semantic matching strength Calculate the mask Enhanced response map within the region average And perform Min-Max normalization in all candidate regions of the current image: Texture consistency score Calculate the mask Gray-level variance in the original gray-level image And convert it into a consistency score: Where the attenuation coefficient The value is 0.1.

[0037] Shape regularity score Calculate the mask Contour compactness: in For the area, This is the perimeter of the outline.

[0038] The three scores are weighted and summed to obtain the overall confidence score: Among them, the weighting coefficient , , ,satisfy .

[0039] A dynamic thresholding strategy is used to filter the results. All candidate regions in the current image are calculated. mean with standard deviation Set dynamic thresholds: Sensitivity adjustment parameters The value is 1. If If the result is positive, the test result will be retained; otherwise, it will be considered a low-confidence result and deleted.

[0040] Step S6: Data accumulation.

[0041] The high-confidence detection results selected in step S5, along with their original image to be tested, defect category label, detection timestamp, and overall confidence score, are combined. The samples are stored in a defect sample library to form a structured pre-labeled dataset, which is used for the training and optimization of subsequent supervised learning models.

[0042] It should be noted that this invention proposes a few-sample industrial defect target detection method based on Dinov3 and SAM, which is applied in the early stages of new product introduction, production line process change or new defect type discovery. When the number of defect samples is insufficient to train a conventional supervised learning model, it serves as a transitional detection scheme and is used to accumulate training data.

[0043] Therefore, this invention proposes a few-sample industrial defect target detection method based on Dinov3 and SAM, which has the following technical advantages: Solving the problem of scarce initial samples: In the early stages of a project, when the number of defective samples is insufficient to train a conventional supervised model, rapid deployment and effective detection can be achieved using only a small number of reference images, filling the gap in quality monitoring. Efficient data accumulation: Through automated detection and high-confidence result screening, high-quality, structured defect annotation data is continuously generated, which significantly accelerates the construction process of the dataset required for subsequent dedicated model training; Enhance system robustness: By incorporating a powerful feature extraction model and a general segmentation base model, along with weighted feature aggregation, context enhancement, and multi-dimensional confidence assessment, the system effectively improves detection accuracy and stability in complex industrial environments, and reduces false alarms and false negatives. Facilitating model iteration: It not only provides detection capabilities with fewer samples, but more importantly, it establishes a solid data foundation for training specialized supervised models with better performance, forming a virtuous cycle of data accumulation and model optimization.

[0044] Reference Figure 5 , Figure 5 This is a schematic diagram of the structure of a few-sample industrial defect target detection system based on Dinov3 and SAM according to an embodiment of the present invention.

[0045] like Figure 5As shown, the few-sample industrial defect target detection system based on Dinov3 and SAM proposed in this embodiment of the invention includes: Module 10 is used to acquire reference images of different industrial defect categories, extract reference image features through a pre-trained Dinov3 model, perform weighted aggregation of reference image features for each industrial defect category, generate category prototype feature vectors, and construct a defect reference feature library. Enhancement module 20 is used to extract features of the image under test using the Dinov3 model, calculate pixel-by-pixel similarity between the features of the image under test and the category prototype feature vectors in the defect reference feature library, generate an initial response map, and enhance it using a neighborhood filtering algorithm to obtain an enhanced response map. The generation module 30 is used to introduce a spatial attention mechanism to filter significant response regions based on the enhanced response map and generate prompt bounding boxes or prompt key points corresponding to the enhanced response map. The segmentation module 40 is used to input the cue bounding box or cue key point into SAM, perform fine segmentation on the significant response region in the enhanced response map, and output a defect instance mask; The filtering module 50 is used to combine semantic matching strength, regional texture consistency and shape regularity to calculate the comprehensive confidence score of each segmentation result, use dynamic threshold to filter high confidence results, and output the final detection result.

[0046] Other embodiments or specific implementations of the Dinov3 and SAM-based few-sample industrial defect target detection system of the present invention can be referred to the above-described method embodiments, and will not be repeated here.

[0047] It is understood that in the description of this specification, references to terms such as "one embodiment," "another embodiment," "other embodiments," or "first embodiment to Nth embodiment," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0048] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.

[0049] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.

Claims

1. A method for detecting industrial defects using a small sample size based on Dinov3 and SAM, characterized in that, The method includes the following steps: S1: Obtain reference images for different industrial defect categories, extract reference image features using a pre-trained Dinov3 model, perform weighted aggregation of reference image features for each industrial defect category, generate category prototype feature vectors, and construct a defect reference feature library. S2: Use the Dinov3 model to extract features from the image to be tested, calculate pixel-by-pixel similarity between the features of the image to be tested and the category prototype feature vectors in the defect reference feature library, generate an initial response map, and use a neighborhood filtering algorithm to enhance it to obtain an enhanced response map. S3: Based on the enhanced response map, a spatial attention mechanism is introduced to filter significant response regions and generate prompt bounding boxes or prompt key points corresponding to the enhanced response map. S4: Input the cue bounding box or cue key points into SAM, perform fine segmentation on the significant response regions in the enhanced response map, and output the defect instance mask; S5: Combining semantic matching strength, regional texture consistency, and shape regularity, calculate the comprehensive confidence score of each segmentation result, use dynamic thresholds to filter high-confidence results, and output the final detection result.

2. The method for detecting industrial defects using a small sample size based on Dinov3 and SAM as described in claim 1, characterized in that, In step S1, when generating the category prototype feature vector by weighted aggregation of the reference image features for each industrial defect category, the weights used in the weighted aggregation are configured to be determined based on the average cosine similarity between each reference image feature and other reference image features of the same category. The expressions for the weights used in weighted aggregation are as follows: In the formula, For the first Weights of features in a reference image. The average cosine similarity between this feature and the other n-1 similar reference features. For temperature coefficient, The total number of reference images for this defect category; This is the aggregated category prototype feature vector. For the first Features of a reference image.

3. The method for detecting industrial defects using a small sample size based on Dinov3 and SAM as described in claim 1, characterized in that, In step S2, a neighborhood filtering algorithm is used for enhancement to obtain an enhanced response map. The configuration is as follows: the neighborhood filtering algorithm is used to calculate the similarity between the response values ​​of each spatial location in the initial response map and its local neighborhood, and the original response is weighted and smoothed. The expression for the filtered initial response map is as follows: in, This is the initial response diagram. This is the enhanced response diagram after filtering. For coordinates A local window centered on the center. For Gaussian kernel function, This is the kernel width parameter.

4. The method for detecting industrial defects using a small sample size based on Dinov3 and SAM as described in claim 3, characterized in that, In step S3, when introducing a spatial attention mechanism for salient response region filtering, the expression for the attention weight map generated by the spatial attention mechanism is as follows: In the formula, For attention weights, To enhance the response map, For learnable location encoding, This indicates concatenation along the channel dimension. This represents a 3×3 convolution operation. The sigmoid activation function is used, and the significant response region is formed by... get, This indicates element-wise multiplication.

5. The method for detecting industrial defects using a small sample size based on Dinov3 and SAM as described in claim 1, characterized in that, In step S5, the comprehensive confidence score of each segmentation result is calculated by combining semantic matching strength, region texture consistency, and shape regularity. The specific expression is as follows: In the formula, To calculate the overall confidence score, The normalized semantic matching strength. This is a texture consistency score calculated based on local gray-level variance. For shape regularity scoring based on contour compactness calculation, , , The weight coefficients are non-negative and satisfy the following conditions: .

6. The method for detecting industrial defects using a small sample size based on Dinov3 and SAM as described in claim 5, characterized in that, Normalized semantic matching strength Texture consistency score and shape regularity score The expression is as follows: In the formula, For mask Enhanced response map within the region The average value, The attenuation coefficient is... For mask Gray-level variance in the original gray-level image For the area, For the perimeter of the outline, The 1st step after performing fine-grained segmentation on the salient response region Candidate regions, For mask Enhanced response map within the region The Middle The average value of each candidate region.

7. The method for detecting industrial defects using a small sample size based on Dinov3 and SAM as described in claim 1, characterized in that, In step S5, the expression for the dynamic threshold is as follows: in, For dynamic thresholds, and These are the mean and standard deviation of the overall confidence scores for all candidate regions in the current image under test, respectively. These are the sensitivity adjustment parameters.

8. The method for detecting industrial defects using a small sample size based on Dinov3 and SAM as described in claim 1, characterized in that, A dynamic threshold is used to filter high-confidence results, and the final detection result is output, which is configured as: the overall confidence score of the segmentation result. Above the dynamic threshold If the result is positive, retain it; otherwise, delete it as a low-confidence result.

9. The method for detecting industrial defects using a small sample size based on Dinov3 and SAM as described in claim 1, characterized in that, The method further includes: S6: Automatically store the high confidence results output in step S5 and their corresponding association information into the defect sample database to form a structured labeled dataset for subsequent model training; The associated information stored in the defect sample database includes: the original image to be tested, the mask of the detected defect instance, the defect category label, the detection timestamp, and the comprehensive confidence score.

10. A few-sample industrial defect target detection system based on Dinov3 and SAM, characterized in that, The system includes: The module is used to acquire reference images for different industrial defect categories, extract reference image features through a pre-trained Dinov3 model, perform weighted aggregation of reference image features for each industrial defect category, generate category prototype feature vectors, and build a defect reference feature library. The enhancement module is used to extract features of the image under test using the Dinov3 model, calculate pixel-by-pixel similarity between the features of the image under test and the category prototype feature vectors in the defect reference feature library, generate an initial response map, and enhance it using a neighborhood filtering algorithm to obtain an enhanced response map. The generation module is used to introduce a spatial attention mechanism to filter significant response regions based on the enhanced response map and generate prompt bounding boxes or prompt key points corresponding to the enhanced response map. The segmentation module is used to input the cue bounding box or cue key points into SAM, perform fine segmentation on the significant response regions in the enhanced response map, and output defect instance masks. The filtering module combines semantic matching strength, regional texture consistency, and shape regularity to calculate the comprehensive confidence score of each segmentation result, uses dynamic thresholds to filter high-confidence results, and outputs the final detection result.

Citation Information

Patent Citations

  • MicroLED direct display module appearance defect detection method

    CN117830268A

  • Small-sample industrial anomaly detection method based on cross-modal adaptive interaction

    CN119989247A

  • Automatic defect labeling method and device, electronic equipment and readable storage medium

    CN120198913A

  • Image recognition system for cable joint forming defect detection

    CN120543555A

  • PCB surface defect detection method and system based on double-layer SAM model collaboration

    CN120747085A

Cited By

  • Surface defect detection method and equipment based on DINOv3 model and medium

    CN121190466A

  • Industrial image detection method based on DINOv3

    CN121414723A

  • Disguised target detection image segmentation method based on DINOv3 and SAM2

    CN122090069A

  • Image segmentation method for camouflaged target detection based on DINOv3 and SAM2

    CN122090069B