Product surface defect identification method, system, equipment and medium
By using a self-supervised pre-trained DINOv2 model and a multi-layer feature fusion network in product surface defect recognition, the problems of sample scarcity and multi-scale object recognition are solved, and high-precision and high-efficiency defect recognition are achieved.
Patent Information
- Application Number
- CN202510158752.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art has challenges such as scarcity of samples, imbalance in training sample distribution and multi-scale target recognition in product surface defect recognition, which is difficult to meet the needs of industrial production for high precision and high efficiency.
The self-supervised pre-trained DINOv2 model is used to extract visual semantic features from the product surface image and input them into a multi-layer feature fusion network. Through feature extraction and fusion at different levels, the robustness and generalization ability of the model are improved, and finally a linear classifier is used for defect recognition.
It improves the accuracy and robustness of product surface defect recognition, can effectively handle features of different scales and complexities, and improves the reliability of identification results.
Smart Images

Figure CN120182183A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of product surface defect recognition, and particularly to a product surface defect recognition method, system, device and medium. Background Art
[0002] The recognition of product surface defects in the manufacturing industry is developing rapidly towards intelligence and automation, and its importance is becoming increasingly prominent. With the deep integration of the new generation of information technology and the manufacturing industry, the manufacturing industry is shifting from quantity expansion to quality improvement, and surface defect recognition has become a key link in improving product quality and enhancing product competitiveness. Surface defects not only affect the aesthetics and comfort of products, but may also damage product performance. Therefore, its detection covers multiple links in production, from the intermediate process to the final factory inspection, which are all important steps to ensure product quality. In addition, surface defect recognition also plays a crucial role in reducing production costs, improving production efficiency, and preventing potential economic losses. Against the background of the continuous upgrading of consumption levels, the development of surface defect detection technology has a significant impact on meeting market demands, enhancing corporate image and competitiveness.
[0003] Traditional surface defect recognition technologies mainly rely on image processing algorithms, such as threshold segmentation, edge detection, and clustering. The advantages of these methods are relatively low computational resource requirements and suitability for simple defect recognition tasks. However, their disadvantages are that they rely heavily on subjective factors, have low accuracy, poor real-time performance, low efficiency, and are difficult to meet the requirements of high precision and high efficiency in industrial production. Product surface defect recognition technologies based on deep learning have demonstrated powerful learning capabilities and automatic feature extraction capabilities. These technologies can automatically discover and learn relevant features from data, achieve advanced performance in tasks such as image recognition and natural language processing, and have advantages such as high accuracy, scalability, and flexibility. However, there are also some disadvantages, such as high computational requirements, the need for a large amount of labeled data, poor interpretability, overfitting, and the black box nature.
[0004] Therefore, there is an urgent need to design a new product surface defect recognition scheme to address challenges such as sample scarcity, unbalanced training sample distribution, and multi-scale target recognition faced in industrial defect recognition. Summary of the Invention
[0005] The purpose of the present invention is to provide a product surface defect recognition method, system, device and medium to overcome the defects of the above-mentioned existing technologies, which can effectively alleviate problems such as sample scarcity, unbalanced training sample distribution, and multi-scale target recognition in the field of product surface defect recognition, and has the advantages of high accuracy and high robustness.
[0006] The purpose of the present invention can be achieved by the following technical solutions:
[0007] According to a first aspect of the present invention, there is provided a method for identifying surface defects of a product, comprising:
[0008] S1. Obtain the surface image of the product to be identified and perform standardized preprocessing;
[0009] S2. Extract visual semantic features from the surface image of the product by using a pre-trained DINOv2 model with self-supervision;
[0010] S3. Input the extracted visual semantic features into a multi-layer feature fusion network to obtain multi-layer fusion features;
[0011] S4. Input the output of the multi-layer fusion feature network into a classifier to output the identified defect categories.
[0012] Preferably, the types of surface defects of the product include central defects, rings, edge local defects, edge ring defects, local defects, near-full defects, random defects, and scratches.
[0013] Preferably, the step of extracting visual semantic features from the surface image of the product by using a pre-trained DINOv2 model with self-supervision specifically includes:
[0014] Process the surface image of the product through a Transformer architecture to output a feature dictionary; wherein the feature dictionary contains normalized block tokens, and through the self-attention mechanism and feed-forward network inside the Transformer architecture, each block token is fused to form a global feature representation as the visual semantic feature.
[0015] Preferably, the basic architecture of the multi-layer feature fusion network includes multiple stage modules, and each stage module includes a convolutional layer for feature extraction, a max-pooling layer for reducing the dimension of the feature map, and a max-pooling layer for reducing the dimension of the feature map, and batch normalization and an activation function with a rectified linear unit are performed on the output of each convolutional layer.
[0016] Preferably, the step of inputting the extracted visual semantic features into a multi-layer feature fusion network to obtain multi-layer fusion features specifically includes:
[0017] Input the visual semantic features, extract the last layer of defect image feature maps from each stage to obtain a first feature group containing features of different scales;
[0018] Perform max-pooling operations on the features of different scales in the first feature group respectively, and perform feature splicing after aligning the dimensions;
[0019] Convert the input multiple dimensions into a single dimension through a feature unfolding operation to obtain the final multi-layer fusion features.
[0020] Preferably, inputting the output of the multi-layer fusion feature network into a classifier to output the identified defect category specifically includes: using the Softmax function as the output layer of the classifier, converting the multi-layer fusion features into a probability distribution, and selecting the category with the highest probability as the predicted defect category.
[0021] Preferably, update the model training parameters through backpropagation of the cross-entropy loss function, and update the network weights according to the difference between the predicted probability distribution and the true label.
[0022] According to a second aspect of the invention, there is provided a system adopting the product surface defect recognition method as described above, including:
[0023] A preprocessing module for obtaining the product surface image to be recognized and performing standard preprocessing;
[0024] A feature extraction module for extracting visual semantic features from the product surface image by using the pre-trained DINOv2 model with self-supervised learning;
[0025] A feature fusion module for inputting the extracted visual semantic features into a multi-layer feature fusion network to obtain multi-layer fusion features;
[0026] An identification and output module for inputting the output of the multi-layer fusion feature network into a classifier to output the identified defect category.
[0027] According to a third aspect of the present invention, there is provided an electronic device including a memory and a processor, where a computer program is stored on the memory, and when the processor executes the program, it implements the method described in any one of the above.
[0028] According to a fourth aspect of the present invention, there is provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method described in any one of the above.
[0029] Compared with the prior art, the present invention has the following beneficial effects:
[0030] (1) The pre-trained DINOv2 model with self-supervised learning of the present invention extracts visual semantic features from the product surface image, uses self-supervised learning to extract visual semantic features with high robustness and generalization, and then inputs the extracted self-supervised features into a multi-layer feature fusion network, combining the general features of self-supervision and the fine features with high precision. Through feature extraction and fusion at different levels, the robustness of the model to features of different scales and complexities is improved. Finally, a linear classifier is used for classification to output the defect recognition result, and the recognition accuracy and reliability are higher.
[0031] (2) The multi-level feature fusion network fuses features at different levels in stages, gradually refines the defect representation through cascaded connections, splices and unfolds the features at different levels to ensure the alignment of features at different levels in terms of space and semantics, avoids information loss, realizes the fusion of multi-scale information, and enhances the feature expression ability; through multi-level feature extraction, global feature fusion and end-to-end optimization, it realizes the efficient detection of defect images; it lies in multi-scale feature fusion and global context awareness, enhancing feature robustness and improving classification performance. Brief Description of the Drawings
[0032] Figure 1 is the flowchart of the product surface defect recognition method of the present invention;
[0033] Figure 2 is the model architecture diagram of the product surface defect recognition method provided by the present invention. Detailed Embodiments
[0034] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0035] Embodiment
[0036] As Figure 1 shown, this embodiment provides a product surface defect recognition method, and this method includes the following steps:
[0037] S1. Obtain the product surface image to be recognized and perform standardized preprocessing;
[0038] S2. Use the pre-trained DINOv2 model with self-supervision to extract visual semantic features from the product surface image;
[0039] S3. Input the extracted visual semantic features into the multi-level feature fusion network to obtain multi-level fusion features, specifically including: input the visual semantic features, extract the last layer of defect image feature maps from each stage to obtain the first feature group containing features of different scales; perform max-pooling operations on the features of different scales in the first feature group respectively, and perform feature splicing after aligning the dimensions; convert the input multiple dimensions into a single dimension through feature unfolding operation to obtain the final multi-level fusion features;
[0040] S4. Input the output of the multi-level fusion feature network into the classifier to output the recognized defect categories.
[0041] Next, the method of this embodiment will be introduced in detail.
[0042] Experimental environment: The hardware configuration used in the experiment is an Intel(R) Core(TM) i7-12700H processor and a GTX 3050 graphics card. The software environment is CUDA 11.3 and cuDNN 8.0, and the development environment is Windows 11. The multi-domain feature fusion network model is completed through Pycharm and the open-source deep learning framework Pytorch 1.12.1.
[0043] In this embodiment, the WM-811K wafer surface image dataset collected during the actual semiconductor production process is used for experimental verification. The WM-811K dataset contains a total of 811,457 wafer surface images with marked defect patterns, including 9 wafer patterns, covering 8 defect patterns such as central defects, rings, edge local defects, edge ring defects, local defects, near-full defects, random defects, scratches, and the normal wafer surface pattern. 24,580 wafer images are randomly selected from the WM-811K dataset to form the data used in this embodiment. The images are preprocessed to a size of 128×128 pixels and divided into a training set and a test set in an 8:2 ratio for each type of data as the network input. The training set is used for training and the network parameters are adjusted, and finally the test set is used to evaluate the classification effect.
[0044] 1. The self-supervised pre-trained DINOv2 model
[0045] DINOv2 is a vision Transformer model based on self-supervised learning, used to learn meaningful representations of images in an unsupervised situation. It extracts general visual features from a large number of unlabeled images through self-supervised learning. These features are very useful for a variety of downstream computer vision tasks, such as image classification, object detection, and semantic segmentation. In addition, DINOv2 provides high-performance visual features that can be directly combined with a classifier without fine-tuning and has good cross-domain performance.
[0046] Based on DINOv2, visual feature extraction is performed on the input product surface defect images. The feature extraction process can be expressed as:
[0047] L = XW L , M = XW M , N = XW N , (1)
[0048] where L, M, and M are the query matrix, key matrix, and value matrix respectively, and W L , W M , W N are the corresponding weight matrices, and X is the input feature matrix.
[0049] In DINOV2, a key component is the self-attention mechanism, which allows the model to dynamically focus on different regions when processing images. The calculation of the attention mechanism can be expressed as:
[0050]
[0051] where Attention(L,M,N) represents the attention function. In the attention mechanism, Softmax is used to convert the dot product of the query and the key into weights, and these weights represent the relevance of each element in the input sequence to the current query, and d K represents the dimension of the key matrix.
[0052] The preprocessed product surface defect image tensor is input into the DINOv2 model. The model processes the image through its Transformer architecture and outputs a feature dictionary, which contains normalized patch tokens. These patch tokens are the feature representations after the image is segmented into small patches. Through the self-attention mechanism and the feed-forward network inside its architecture, the individual patch tokens are fused to form a global feature representation. Finally, the model outputs a normalized feature vector as the self-supervised visual semantic feature.
[0053] 2. Multi-layer Feature Fusion Network
[0054] The extracted self-supervised features are input into the multi-layer feature fusion network to further refine and integrate the features. The basic architecture of the multi-layer feature fusion network includes multiple stage modules. Each stage module contains a convolutional layer for feature extraction, a max pooling layer for reducing the dimensionality of the feature map, and a max pooling layer for reducing the dimensionality of the feature map. Batch normalization and an activation function with a rectified linear unit are performed on the output of each convolutional layer.
[0055] In this embodiment, as Figure 2As shown in the figure, the basic architecture of the multi-layer feature fusion network consists of five stages. Each stage contains a convolutional layer for feature extraction and a max pooling layer for feature map dimensionality reduction. In addition, batch normalization (BN) and an activation function with a rectified linear unit (ReLU) are performed on the output of each convolutional layer to increase the network's non-linearity. Before feature fusion, defect image features at different levels are extracted from the network. The last layer of the defect image feature map is extracted from each stage of the network to obtain a set of features at different scales, denoted as {C1, C2, C3, C4, C5}. Then, the feature dimensions are aligned through max pooling operations (Maxpooling) respectively. The process is shown in Equation (3) to obtain {E1, E2, E3, E4, E5}. Then, these five feature maps are concatenated, and the process is shown in Equation (4). Finally, through a feature flattening operation (Flatten), multiple dimensions of the input tensor are converted into a single dimension to obtain the final multi-layer feature fusion feature.
[0056]
[0057] Where: P(i, j) is the value of the pooled feature map at position (i, j), and P(m, n) is the value of the original input feature map at position (m, n). M i,j represents a local area of the pooling window on the input feature map, and this area is determined by the size and stride of the pooling window.
[0058] Concat(E1, E2,...E5) = stack(E1, E2,...E5, dim = d) (4)
[0059] Where d is the specified concatenation dimension.
[0060] 3. Classifier
[0061] The output of the multi-layer fusion feature network is passed through a classifier to output the specific category of the defect. The Softmax function is used as the output layer for this defect recognition, and the multi-layer feature fusion feature output of the network is converted into a probability distribution, where each element represents the probability of the corresponding category. After obtaining the probability distribution, the category with the maximum probability is selected as the prediction result of the neural network, that is, the category of the defect. The specific process is expressed as:
[0062]
[0063] Where, Z i is the original output of the i-th category, and n is the total number of categories.
[0064] The main network layer, classification layer, and feature splicing and unfolding layer in the multi-layer feature fusion network are all trainable layers. The weights and biases of the corresponding parameters will be updated. The model training parameters are updated through backpropagation using the cross-entropy loss function to measure the difference between the predicted probability distribution and the true label, and the network weights are updated accordingly. The process can be expressed as follows:
[0065]
[0066] Among them, n is the total number of categories, Y represents the true label, and Yμ is the probability distribution predicted by the model, which is output by the Softmax function.
[0067] Considering the training cost and stability of the DINOv2 model, this application chooses to directly use the original model to extract features. As a large vision model, DINOv2 has been pre-trained on a large amount of general data and can extract rich visual features, making it applicable to various scenarios. Directly using the original model to extract features can avoid the high cost of modification and re-training, and at the same time utilize its strong generalization ability to avoid the uncertainties that may be introduced after modification.
[0068] The multi-level feature fusion network fuses different-level features (shallow texture + middle-level semantics + deep global information) in stages, gradually refines the defect representation through cascade connection, splices and unfolds the features of different levels to ensure that the features of different levels are aligned in space and semantics, avoid information loss, realize the fusion of multi-scale information, and enhance the expression ability of features. The model is concise and efficient, and through multi-level feature extraction, global feature fusion, and end-to-end optimization, it realizes the efficient detection of defect images. Its innovation lies in multi-scale feature fusion and global context awareness, and its advantage lies in enhancing feature robustness and improving classification performance.
[0069] This embodiment also provides a product surface defect recognition system, including:
[0070] A preprocessing module for obtaining the product surface image to be recognized and performing standard preprocessing;
[0071] A feature extraction module for extracting visual semantic features from the product surface image using the pre-trained DINOv2 model with self-supervision;
[0072] A feature fusion module for inputting the extracted visual semantic features into the multi-layer feature fusion network to obtain multi-layer fusion features;
[0073] An identification output module for inputting the output of the multi-layer fusion feature network into a classifier and outputting the identified defect category.
[0074] The electronic device of the present invention includes a central processing unit (CPU), which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) or computer program instructions loaded from a storage unit into a random access memory (RAM). In the RAM, various programs and data required for device operation can also be stored. The CPU, ROM, and RAM are connected to each other via a bus. An input / output (I / O) interface is also connected to the bus.
[0075] Multiple components in the device are connected to the I / O interface, including: an input unit, such as a keyboard, a mouse, etc.; an output unit, such as various types of displays, speakers, etc.; a storage unit, such as a magnetic disk, an optical disc, etc.; and a communication unit, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit allows the device to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0076] The processing unit executes the various methods and processes described above, such as method S1 - S4. For example, in some embodiments, method S1 - S4 can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device via the ROM and / or the communication unit. When the computer program is loaded into the RAM and executed by the CPU, one or more steps of method S1 - S4 described above can be executed. Alternatively, in other embodiments, the CPU can be configured to execute method S1 - S4 by any other suitable means (e.g., by means of firmware).
[0077] The functions described above herein can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip systems (SOC), complex programmable logic devices (CPLD), and so on.
[0078] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.
[0079] In the context of the present invention, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0080] As described above, only the specific embodiments of the present invention are provided, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily conceive of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A method for identifying surface defects of a product, characterized in that: include: S1. Obtain the surface image of the product to be identified and perform standardized preprocessing; S2, extract visual semantic features from product surface images using the self-supervised pre-trained DINOv2 model; S3, input the extracted visual semantic features into a multi-layer feature fusion network to obtain multi-layer fusion features; S4. Input the output of the multi-layer fusion feature network into the classifier and output the identified defect category.
2. A product surface defect identification method according to claim 1, characterized in that: The types of product surface defects include center defects, rings, edge local defects, edge ring defects, local defects, nearly full defects, random defects and scratches.
3. A product surface defect identification method according to claim 1, characterized in that: The self-supervised pre-trained DINOv2 model is used to extract visual semantic features from product surface images, specifically including: The product surface image is processed through the Transformer architecture to output a feature dictionary. The feature dictionary contains normalized block labels. Through the self-attention mechanism and feedforward network inside the Transformer architecture, each block label is fused to form a global feature representation as a visual semantic feature.
4. A product surface defect identification method according to claim 1, characterized in that: The infrastructure of the multi-layer feature fusion network includes multiple stage modules, each of which includes a convolution layer for feature extraction and a maximum pooling layer for feature map dimensionality reduction, and a maximum pooling layer for feature map dimensionality reduction. The output of each convolution layer is batch normalized and an activation function operation with a rectified linear unit is performed.
5. A product surface defect identification method according to claim 4, characterized in that: The extracted visual semantic features are input into a multi-layer feature fusion network to obtain multi-layer fusion features, specifically including: Input visual semantic features, extract the last layer of defect image feature maps from each stage, and obtain the first feature group containing features of different scales; Perform maximum pooling operations on the features of different scales in the first feature group, align the dimensions, and then perform feature concatenation; The feature expansion operation is used to convert multiple input dimensions into a single dimension to obtain the final multi-layer fusion features.
6. A product surface defect identification method according to claim 1, characterized in that: The output of the multi-layer fusion feature network is input into the classifier to output the identified defect category, which specifically includes: using the Softmax function as the output layer of the classifier, converting the multi-layer fusion features into a probability distribution, and selecting the category with the maximum probability as the predicted defect category.
7. A product surface defect identification method according to claim 6, characterized in that: The model training parameters are updated through back propagation of the cross entropy loss function, and the network weights are updated according to the difference between the predicted probability distribution and the true label.
8. A system using the product surface defect recognition method according to claim 1, characterized in that: include: A preprocessing module, used to obtain the surface image of the product to be identified and perform standardized preprocessing; Feature extraction module, which is used to extract visual semantic features from product surface images using the self-supervised pre-trained DINOv2 model; A feature fusion module is used to input the extracted visual semantic features into a multi-layer feature fusion network to obtain multi-layer fusion features; The recognition output module is used to input the output of the multi-layer fusion feature network into the classifier and output the identified defect category.
9. An electronic device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the program, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Surface defect detection method and equipment based on DINOv3 model and medium
CN121190466A
Part surface defect identification method and device
CN121937445A
A method and apparatus for identifying surface defects in parts.
CN121937445B