Wafer defect detection method based on YOLOv8 and RT-DETR

By combining YOLOv8 and RT-DETR models, the activation function and loss function are optimized, and the distributed focus loss and IoU soft label mechanism are introduced, which solves the problems of low detection efficiency in crystal oscillator surface defect detection and high computing resources in deep learning applications, achieving efficient and accurate defect detection and improvement of algorithm generalization capabilities.

CN120147240APending Publication Date: 2025-06-13CHINA JILIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510200601.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing crystal oscillator surface defect detection methods have problems such as low detection efficiency, inconsistent detection results and difficulty in identifying defects in various forms. The application of deep learning technology in this field faces the problems of complex data set construction and high computing resources and time costs.

Method used

The dual-model fusion method based on YOLOv8 and RT-DETR is adopted to optimize the activation function and loss function through data preprocessing, model training and detection precision positioning, and introduce distributed focus loss and IoU soft label mechanisms to improve the robustness and detection accuracy of the model.

Benefits of technology

It realizes efficient and accurate detection of surface defects of crystal oscillator, significantly improves detection accuracy and speed, enhances the generalization ability of the algorithm, and reduces the computing burden of the model, making it suitable for embedded platform deployment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147240A_ABST
    Figure CN120147240A_ABST
Patent Text Reader

Abstract

The invention discloses a crystal oscillator surface defect detection method based on YOLOv8 and RT-DETR. The method comprises the following steps: S1, data acquisition: acquiring a large number of crystal oscillator surface images by using a high-resolution and high-speed camera; s2, data preprocessing: preprocessing the collected image data, including image denoising, graying and image enhancement operations, and manually marking a defect area by using Label; after labeling is completed, the data set is divided into a training set, a verification set and a test set according to the proportions of 70%, 20% and 10%; s3, training a model: performing model training in a YOLOv8 deep learning framework and an RT-DETR deep learning framework by utilizing the processed data set, the labeled data and the pre-training weight file, and optimizing model performance by replacing and improving an activation function to obtain a trained weight file; and S4, detecting and accurately positioning defects. According to the method, the combination of the YOLOv8 and the RT-DETR model is proposed for the first time by using double-model fusion, and the detection precision and speed are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of machine vision, and particularly relates to a wafer defect detection method based on YOLOv8 and RT-DETR for detecting surface defects of crystal oscillator wafers. Background Art

[0002] As a key component in the fields of vacuum thin film deposition and high-precision metal electroplating, the detection of surface defects of crystal oscillator wafers is crucial for ensuring the reliability and performance of the final products. Traditional methods for detecting surface defects of crystal oscillator wafers mainly rely on manual visual inspection, but this method has low detection efficiency, and the subjective judgment of inspectors will affect the consistency and reliability of the detection results. With the expansion of production scale and the improvement of product quality requirements, the manual detection method can no longer meet the high standards of modern industrial production.

[0003] To overcome the limitations of manual detection, machine vision technology has been widely used in the detection of surface defects of crystal oscillator wafers in recent years. Machine vision technology, with its high objectivity, high efficiency, and high reliability, can effectively identify tiny defects that are difficult to detect by manual inspection. However, there are still some deficiencies in existing machine vision detection methods. For example, although some detection methods based on traditional digital image processing technology can achieve certain detection effects, they rely on manually selected features, which limits their adaptability and efficiency. At the same time, these methods often have poor effects when detecting defects with diverse shapes and difficult to identify, such as white dots, white stripes, and scratches.

[0004] To further improve the accuracy and efficiency of detecting surface defects of crystal oscillator wafers, deep learning technology has been introduced into this field. Deep learning models, with their powerful automatic feature extraction capabilities, have achieved remarkable results in the detection of surface defects of various materials. Among them, the YOLO series of models have shown excellent performance in object detection tasks with high precision and fast detection speed. The RT-DETR model combines the advantages of the Transformer architecture and demonstrates powerful performance in object detection and segmentation tasks.

[0005] However, applying deep learning technology to the detection of surface defects of crystal oscillator wafers still faces some challenges. For example, the construction and preprocessing of the surface defect dataset of crystal oscillator wafers are relatively complex, and refined annotation and enhancement need to be carried out for different types of defects. In addition, the training and optimization of deep learning models also require a large amount of computing resources and time costs. Therefore, how to effectively utilize deep learning technology for the detection of surface defects of crystal oscillator wafers has become a hot and difficult issue in current research. Summary of the Invention

[0006] To solve the above technical problems existing in the prior art, the present invention provides a wafer defect detection method based on YOLOv8 and RT-DETR, aiming to combine the advantages of the two deep learning models, improve the accuracy and efficiency of surface defect detection of crystal oscillator wafers, achieve efficient and accurate detection of surface defects of crystal oscillator wafers, and at the same time enhance the generalization ability of the algorithm.

[0007] The technical solution adopted by the present invention is as follows:

[0008] A method for detecting surface defects of crystal oscillator wafers based on YOLOv8 and RT-DETR, characterized in that it specifically includes the following steps:

[0009] S1. Data collection:

[0010] Use a high-resolution and high-speed camera to obtain a large number of images of the surface of crystal oscillator wafers. The images include the overall view of the surface of the crystal oscillator wafer and the annotation data of the defect area;

[0011] S2. Data preprocessing:

[0012] Preprocess the collected image data, including image denoising, grayscale conversion, and image enhancement operations. Manually annotate the defect areas using Labelme; after the annotation is completed, divide the data set into a training set, a validation set, and a test set at a ratio of 70%, 20%, and 10% respectively;

[0013] S3. Model training:

[0014] Use the processed data set, annotation data, and pre-trained weight file to perform model training in the YOLOv8 and RT-DETR deep learning frameworks respectively. Optimize the model performance by replacing and improving the activation function to obtain a trained weight file;

[0015] S4. Detect and accurately locate defects:

[0016] S41. Input the image of the crystal oscillator wafer to be detected into the loaded YOLOv8 model, and use the fast detection ability of the YOLOv8 model to identify the potential defect areas on the surface of the crystal oscillator wafer to be detected, and output the corresponding anchor boxes and confidence levels;

[0017] S42. Optimize the anchor box settings: By adjusting the number, size, and ratio of the anchor boxes, make them better adapt to the different shapes and scales of the target objects; analyze the target sizes in the training set through the clustering algorithm, and select appropriate anchor box sizes to improve the detection accuracy of small targets, large targets, and objects of different shapes, and reduce the situation of false detection and missed detection;

[0018] S43. Input the identified potential defect area into the RT-DETR model, classify and precisely locate the potential defect area using the precise positioning ability of the RT-DETR model, and output the final detection result.

[0019] Further, in step S2, the preprocessed data also includes cropping, rotating, and scaling the acquired image to adapt to crystal oscillator wafers of different sizes and shapes.

[0020] Further, in step S3, the replacement and improvement of the activation function specifically include:

[0021] Replace the activation function in the YOLOv8 and RT-DETR models from the Silu function to the ReLU function; remove the Sigmoid function in the model to accelerate the model's calculation speed and facilitate deployment on various embedded platforms; the Silu function and the ReLU function are specifically expressed as

[0022] Silu: f(x) =x*sigmoid(x) (1)

[0023]

[0024] Further, in step S4, optimizing the anchor box settings includes:

[0025] Adjust the model structure and parameter configuration according to the actual detection effect;

[0026] Among them, the parameter configuration specifically includes training parameters such as the learning rate, batch size, and number of training epochs during yolo training.

[0027] Introduce the Distribution Focal Loss (DFL) to optimize the probability distribution of the two positions closest to the label in the form of cross-entropy, enabling the model to more accurately focus on the target position and its adjacent areas, thereby improving the robustness and detection accuracy of the model in scenarios with occlusions and moving targets; the Distribution Focal Loss (DFL) is specifically expressed as follows:

[0028] The Distribution Focal Loss (DFL) is specifically expressed as shown below:

[0029] DFL(S i ,S i+1 )=-((y i+1 -y)log(S i )+(y-y i )log(S i+1 )) (3)

[0030] Among them, S i refers to the i-th distribution value in the model prediction distribution, and S i+1 refers to the (i + 1)-th distribution value in the model prediction distribution;

[0031] y i refers to the label of the target category corresponding to the i-th predicted value in the true value, y i+1 refers to the label of the target category corresponding to the (i + 1)-th predicted value in the true value.

[0032] A surface defect detection device for crystal oscillator chips based on YOLOv8 and RT-DETR, characterized in that different hardware devices are used for the deployment of the algorithm platform, including:

[0033] An optical camera for capturing high-resolution surface images of crystal oscillator chips and transmitting the captured surface images of crystal oscillator chips to the algorithm deployment platform;

[0034] The algorithm deployment platform, which is built-in with an AI inference module and executable code, supports the execution of any of the above algorithms, and detects, classifies, and locates surface defects of crystal oscillator chips in real time.

[0035] Compared with the prior art, the beneficial effects of the present invention are reflected in:

[0036] 1. The present invention uses dual-model fusion and for the first time proposes to combine the YOLOv8 and RT-DETR models, making full use of the advantages of both to improve the detection accuracy and speed.

[0037] 2. The present invention optimizes the activation function and loss function, significantly reducing the model calculation burden and realizing efficient deployment on the embedded platform.

[0038] 3. The present invention introduces the DFL loss and IoU soft label mechanism, showing excellent performance in complex scenarios with multiple targets.

[0039] 4. Through various data augmentation and network improvement strategies, the robustness of the model in diverse defect scenarios is significantly enhanced. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 is the flowchart of the surface defect detection of the crystal oscillator chip of the present invention.

[0041] Figure 2 is the application schematic diagram of the YOLOv8 and RT-DETR models of the present invention in the defect detection of crystal oscillator chips.

[0042] Figure 3 is the network model diagram of the YOLOv8 of the present invention.

[0043] Figure 4 is the network model diagram of the RT-DETR of the present invention.

[0044] Figure 5 is the comparison of the detection results of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0045] The present invention will be further described below in conjunction with specific embodiments:

[0046] Embodiment 1:

[0047] Refer to Figure 1 , collect 1000 surface defect images of crystal oscillator wafers, and construct a crystal oscillator wafer defect image dataset.

[0048] Preprocess the dataset, including image contrast adjustment, noise removal, and normalization.

[0049] Select the YOLOv8 model for training, set training parameters such as learning rate and number of iterations, and obtain a trained YOLOv8 model.

[0050] Input the crystal oscillator wafer image to be detected into the trained YOLOv8 model for defect detection.

[0051] Identify and classify the defects on the surface of the crystal oscillator wafer according to the output information, and record the detection results.

[0052] Adjust and optimize the model parameters according to the actual detection effect to improve the detection accuracy and efficiency.

[0053] Embodiment 2:

[0054] Refer to Figures 2 to 4 , similar to Embodiment 1, but select the RT-DETR model for training, and compare the performance differences between YOLOv8 and RT-DETR in crystal oscillator wafer defect detection. Through comparative analysis, assign weights to the training effects of the YOLOv8 and RT-DETR models, and calculate the weighted average value, which is used as the final defect accuracy result.

[0055] 1. Dataset construction and preprocessing:

[0056] Collect a large number of surface images of crystal oscillator wafers, perform normalization processing, image inversion, cropping, blurring and other operations on the dataset for data augmentation to improve the generalization ability of the model. After that, manually label the defect areas using Labelme. After the labeling is completed, divide the dataset into a training set, a validation set, and a test set at a ratio of 70%, 20%, and 10% respectively.

[0057] 2. Model selection and training: Select YOLOv8 and RT-DETR as the object detection models.

[0058] i) In this model training, the native network structure of yolov8 is selected. However, considering the applicability in the actual application scenario, we choose to replace the activation function from the silu function with the relu function. After removing the sigmoid function, the operation speed of the model can be accelerated, which is convenient for the deployment of the model on various embedded platforms

[0059] Silu: f(x) = x * sigmoid(x) (1)

[0060]

[0061] As a real-time detection model, YOLOv8 uses binary cross-entropy as the loss function to determine the confidence level as a certain class:

[0062]

[0063] This calculation method can calculate the entropy of both being and not being this class and add them together, so that the losses calculated by this loss function in the above two cases will not be 0.

[0064] ii) In the training of RT-DETR, similarly, to speed up the operation speed of the model, we choose to replace its activation function from Silu to Mish

[0065] Mish: f(x) = x tanh(ln(1 + e x )) (5)

[0066] It provides the model with better ability to fit complex data distributions. Compared with YOLO, the IoU soft label is added to the loss function of RT-DETR, which means using the IoU between the predicted bounding box and the GT as the label for class prediction.

[0067]

[0068] 3. Defect detection:

[0069] Load the model to detect the test images. In YOLO, we choose CIOU + DFL, which can output the distances from the left, top, right, and bottom borders of the predicted target bounding box to the center point of the target, facilitating better positioning of the defect location. In the model, we also added the DFL loss. DFL optimizes the probabilities of the two positions, one left and one right, closest to the label y in the form of cross-entropy, so that the network can focus on the distribution of the target location and its adjacent areas faster. This loss of learning the positions around the label can enhance the generalization ability of the model in complex situations such as occlusion and moving objects.

[0070] DFL(S i , S i+1 ) = -((y i+1 - y) log(S i ) + (y - y i ) log(S i+1 )) (3)

[0071] RT-DETR uses the native decoder part and has excellent performance, so it is not modified.

[0072] 4. Algorithm optimization:

[0073] According to the actual detection effect, adjust and optimize the model parameters to improve the detection accuracy and efficiency.

[0074] The network structures of YOLO and RT-DETR can be modified, such as CNN or ResNet, to adapt to the multi-device deployment environment.

[0075] Introduce data augmentation technology. Through rotation, scaling, and cropping steps, further improve the generalization ability of the model.

[0076] It can be clearly seen from the results of groups 5(a) and 5(b) that RT-DETR is superior to YOLOv8 in both the number and confidence of defect detections. Especially in the "white" category, RT-DETR can detect more defects and with higher confidence. Further comparison Figure 5 (c) and Figure 5 (d) groups. The result graph shows that YOLOv8 not only has misdetection problems in the "bright" defect category but also has difficulties in distinguishing between "bright" and "white" defects. In contrast, RT DETR significantly improves the detection rate, reduces missed detections and misdetections, and enhances the accuracy and reliability of detection. Figure 5 (e) group results show that RT-DETR has better performance in capturing smaller or more subtle defect features, indicating its progress in reducing the missed detection rate. The result analysis shows that in the surface defect detection of crystal oscillator chips, RT-DETR is superior to YOLOv8 in terms of missed detection rate, misdetection rate, and overall performance. RT-DETR demonstrates its reliability and efficiency as a defect detection tool by increasing the defect detection rate, reducing missed detections and misdetections, and more accurately locating the defect positions.

Claims

1. A method for detecting surface defects of a crystal oscillator based on YOLOv8 and RT-DETR, characterized in that: The specific steps include: S1. Collect data: Use high-resolution and high-speed cameras to obtain a large number of images of the crystal surface, including the overall view of the crystal surface and the annotation data of the defect area; S2. Preprocessing data: The collected image data is preprocessed, including image denoising, grayscale conversion, and image enhancement operations, and the defect areas are manually annotated using Labelme. After the annotation is completed, the data set is divided into a training set, a validation set, and a test set at a ratio of 70%, 20%, and 10%, respectively; S3. Training model: Using the processed data set, labeled data, and pre-trained weight files, the model is trained in the YOLOv8 and RT-DETR deep learning frameworks, respectively. The model performance is optimized by replacing and improving the activation function to obtain the trained weight files. S4. Detection and precise positioning of defects: S41, input the image of the crystal oscillator to be detected into the loaded YOLOv8 model, use the fast detection capability of the YOLOv8 model to identify the potential defect area on the surface of the crystal oscillator to be detected, and output the corresponding anchor frame and confidence; S42. Optimize anchor frame settings: adjust the number, size, and proportion of anchor frames to better adapt to the different forms and scales of target objects; use clustering algorithms to analyze the target sizes in the training set and select appropriate anchor frame sizes to improve the detection accuracy of small targets, large targets, and objects of different shapes, and reduce false detections and missed detections; S43, input the potential defect areas identified above into the RT-DETR model, use the precise positioning capability of the RT-DETR model to classify and precisely locate the potential defect areas, and output the final detection results.

2. According to a method for detecting surface defects of a crystal oscillator based on YOLOv8 and RT-DETR according to claim 1, it is characterized in that: In step S2, preprocessing the data also includes cropping, rotating, and scaling the acquired image to adapt to crystal oscillators of different sizes and shapes.

3. According to a method for detecting surface defects of a crystal oscillator based on YOLOv8 and RT-DETR according to claim 1, it is characterized in that: In step S3, the replacement and improvement of the activation function specifically includes: The activation function in the YOLOv8 and RT-DETR models is replaced by the ReLU function from the Silu function; the Sigmoid function in the model is removed to speed up the calculation of the model and facilitate deployment on a variety of embedded platforms; the Silu function and the ReLU function are specifically expressed as Silu:f(x)=x*sigmoid(x) (1) 4. According to a method for detecting surface defects of a crystal oscillator based on YOLOv8 and RT-DETR according to claim 1, it is characterized in that: In step S4, optimizing the anchor box setting includes: Adjust the model structure and parameter configuration according to the actual detection results; The distribution focus loss DFL is introduced to optimize the probability distribution of the positions on both sides closest to the label in the form of cross entropy, so that the model can focus on the target position and its neighboring area more accurately, thereby improving the robustness and detection accuracy of the model in occlusion and moving target scenes; the specific expression of the distribution focus loss DFL is as follows: DFL(S i ,S i+1 )=-((y i+1 -y)log(S i )+(y-y i )log(S i+1 )) (3) Among them, S i Refers to the i-th distribution value in the model prediction distribution, S i+1 Refers to the i+1th distribution value in the model prediction distribution; y i Refers to the label of the target category corresponding to the i-th predicted value in the true value, y i+1 Refers to the label of the target category corresponding to the i+1th predicted value in the true value.

5. A crystal oscillator surface defect detection device based on YOLOv8 and RT-DETR, characterized in that: Use different hardware devices to deploy the algorithm platform, including: An optical camera is used to capture high-resolution images of the surface of the crystal oscillator and transmit the captured images to the algorithm deployment platform; The algorithm deployment platform has a built-in AI reasoning module and executable code, supports the execution of any algorithm in claims 1 to 4, and detects, classifies and locates surface defects of crystal oscillators in real time.