Compressed image target detection method based on enhancement and knowledge distillation

Through the compression artifact removal network and knowledge distillation method, a collaborative learning framework between teacher model and student model is built, which solves the problem of degradation of accuracy caused by high storage and transmission costs in silicon wafer detection and image compression, and improves detection accuracy and robustness.

CN120339579APending Publication Date: 2025-07-18TAIYUAN UNIVERSITY OF SCIENCE AND TECHNOLOGY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510406477.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art has problems with high storage requirements and transmission costs in silicon wafer defect detection. At the same time, image compression leads to a decrease in detection accuracy, and image distortion of the training set affects the generalization ability of the model.

Method used

The compression artifact removal network and knowledge distillation method are used to build a collaborative learning framework between teacher model and student model, and improve detection performance through logits distillation and feature distillation.

Benefits of technology

This significantly improves the accuracy and robustness of defect detection under compressed images, optimizes image quality, and enhances the recognition ability of the detection network and the reliability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339579A_ABST
    Figure CN120339579A_ABST
Patent Text Reader

Abstract

The invention discloses a compressed image target detection method based on enhancement and knowledge distillation, and belongs to the technical field of target detection methods, and the method comprises the following specific steps: S1, collecting original silicon wafer data, constructing a first data set based on the original silicon wafer data, and training to obtain a teacher model, obtaining silicon wafer data subjected to compression processing under different quality factors, and constructing a silicon wafer data set under different quality factors; s2, processing the compressed data of different quality factors by adopting a compression artifact removal network to obtain an intermediate feature map; s3, constructing a collaborative learning framework of the teacher model and the student model, and guiding the student model to train in a manner of combining logits distillation and feature distillation on compressed data; and S4, performing target detection on original silicon wafer data by using the trained model. According to the method, the teacher model and the student model are constructed by combining logits distillation and feature distillation, so that the defect detection performance under the compressed image is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of target detection methods, and particularly relates to a compressed image target detection method based on enhancement and knowledge distillation. Background Art

[0002] With the rapid development of the photovoltaic industry, due to the brittle characteristics of silicon wafers during the manufacturing process, they are prone to defects such as broken pieces and hidden cracks due to collisions and squeezes, which directly affects the yield of components. To solve this problem, a defect detection system based on deep learning has become the mainstream solution. However, in practical applications, a series of challenges still exist.

[0003] First of all, in order to meet the high-precision detection requirements, high-resolution cameras are usually used to collect images of silicon wafers. Although this method can obtain more detailed image information and help improve the detection accuracy, manufacturers generally require storing production images for several months for traceability. This leads to a huge data storage requirement, which further increases the burden on the system.

[0004] Secondly, the defect detection model needs to be continuously updated and optimized in practical applications, and this depends on a large amount of data for training. Since the training process usually requires long-term calculations, the training is usually carried out on high-performance servers so that the data captured from multiple visual acquisition devices can be integrated together to further train a more accurate model. However, the training and testing are usually carried out at two physical locations, namely the server and the visual acquisition device, which significantly increases the transmission cost. To solve the storage and transmission pressure, the industry generally adopts the JPEG compression method to store images. Although this method effectively reduces the storage and transmission costs, the compression process will cause the loss of high-frequency features in the images, thereby affecting the accuracy of online defect detection. This results in a conflict between the compression efficiency of the images and the integrity of the features, and it is difficult to meet the requirements for efficient detection in industrial applications.

[0005] These problems not only exist in the field of silicon wafer detection but are also significant in the field of machine vision. With the wide application of machine vision technology, data-driven methods have become the core of system performance optimization. The performance of machine vision systems depends heavily on high-quality training data, which are used to simulate various situations that may be encountered in practical applications. However, in actual operation, unknown image distortions may lead to a gap between the source domain and the target domain, affecting the performance of the system, especially when facing image degradation problems. For example, problems such as motion blur, noise, fogging, and compression distortion will have a significant impact on image processing, target detection, and image classification tasks.

[0006] To solve this problem, several new data augmentation methods have been proposed by researchers. The GSES method generates training data with composite distortions by fusing up to 15 types of distortions through kernel density estimation (KDE), avoiding the limitations of traditional methods that train multiple distortion models separately. The DeepCorrect method, on the other hand, identifies and corrects the most noise - vulnerable filters by evaluating the impact of image distortion on the activation of pre - trained convolutional filters. It applies small convolutional layers and residual connections to the outputs of the most damaged filters while keeping the outputs of other filters unchanged, thereby improving the robustness of deep neural networks (DNNs) to distorted images.

[0007] Although existing methods have made progress in the problem of image degradation in the test set, most research focuses on improving robustness when the test set is distorted. No one has paid attention to the situation where the training set consists of degraded images and the test set consists of original images. Image distortion in the training set may cause the model to fail to effectively generalize to original high - quality images, thus affecting the performance in practical applications. Summary of the Invention

[0008] To solve the above - mentioned technical problems, the present invention proposes a compressed image object detection method based on enhancement and knowledge distillation. The method of the present invention uses wafer data processed by compression at different quality factors as the training set, and sets up a compression artifact removal network for this purpose. It constructs a teacher model and a student model by combining logits distillation and feature distillation, improving the defect detection performance under compressed images.

[0009] The technical solution protected by the present invention is as follows: A compressed image object detection method based on enhancement and knowledge distillation, which is specifically carried out according to the following steps:

[0010] Step S1: Collect original wafer data Construct a first data set based on the original wafer data and train a teacher model Obtain wafer data processed by compression at different quality factors q Construct wafer data sets under different quality factors;

[0011] Step S2: Use the compression artifact removal network to process the compressed data Y of different quality factors q q to obtain an intermediate feature map F q ∈R C×H×W . This process is expressed as:

[0012]

[0013] represents the compression artifact removal network, R C×H×W represents the shape of the intermediate feature map, R represents the set of real numbers, and C, H, and W respectively represent the number of channels, width, and height of the feature map;

[0014] Step S3: Construct a teacher model and a student model for collaborative learning framework, and use a combination of logits distillation and feature distillation on the compressed data to guide the training of the student model;

[0015] Step S4: Use the trained model to perform object detection on the original silicon wafer data.

[0016] Furthermore, the specific process of logits distillation in step S3 is as follows:

[0017] During the learning process, the parameters of the teacher model are not updated. For the intermediate feature map F q , it passes through the teacher model T and the student model respectively, which are expressed as:

[0018]

[0019] represents the logits output of F q passing through the teacher model, representing the classification probability and the bounding box regression value respectively, represents the logits output of passing through the student model, representing the classification probability and the bounding box regression value respectively;

[0020] The output process of the student model imitating the teacher model is expressed as:

[0021]

[0022] where σ(·) is the Sigmoid function, M represents the number of anchor points, N represents the number of target categories, represents the logits loss.

[0023] Furthermore, the specific process of feature distillation in step S3 is as follows:

[0024] During the learning process, the parameters of the teacher model are not updated. For the intermediate feature map F q , it passes through the teacher model and the student model respectively. The k-th feature maps of the teacher and the student are represented as T k ∈R C×H×W and S k ∈R C×H×W (k = 1,..,K), and at the same time, the k-th random mask is generated, which is expressed as:

[0025]

[0026] where is a random number in the range (0, 1), x and y are the horizontal and vertical coordinates of the feature map respectively, ζ = 0.65, representing the mask ratio, represents the generated random mask;

[0027] Then use the corresponding mask to cover the student's feature map and generate the teacher's feature map with the remaining pixels. This process is expressed as:

[0028] S′ = G(S·M) = (S·M)·W (6)

[0029] G(F k ) = ReLU(BN(Conv*F k )) (7)

[0030] where G represents the mapping layer, Conv represents the 3×3 convolutional layer, ReLU(·) represents the activation layer, BN(·) represents normalization, W represents the weight of the mapping layer, and F k represents the feature after mask processing;

[0031] The student model mimics the intermediate features of the teacher model, expressed as:

[0032]

[0033] K represents the sum of the distillation layers, C, H, and W represent the shape of the feature map, and S and T represent the features of the student and the teacher respectively.

[0034] Furthermore, train the student model with compressed data of different quality factors q, and the total training loss is expressed as:

[0035]

[0036] where represents the loss of the YOLOv8 network, represents the loss of logits distillation, represents the loss of feature distillation, represents the loss of the decompression artifact network, and α, β, and γ are the corresponding hyperparameters;

[0037] where:

[0038]

[0039]

[0040] is the class loss in the Yolov8 network, is the bounding box regression loss in Yolov8, where η1 and η1 are corresponding hyperparameters, and F i represents the data generated by the compression artifact removal network, and Y i represents the original data, m represents the number of images, v represents the distilled loss decay term, a represents the minimum value of the decay term, set to 0.01, b represents the maximum value of the decay term, set to 1, and T max represents the maximum number of batches controlling the decay process, set to 10, and n i represents the current iteration number.

[0041] The present invention has the following advantages compared with the prior art:

[0042] 1. The present invention proposes an end-to-end network structure combining compression artifact removal and defect detection. By introducing the loss function of the artifact removal network into the detection network, the image quality is optimized and the defect detection performance is improved. The compression artifact removal network effectively removes JPEG compression artifacts, restores subtle defect features, enhances the recognition ability of the detection network for key defects, and thus significantly improves the accuracy and reliability of online detection.

[0043] 2. Through Logit distillation and feature distillation, the present invention constructs a collaborative learning framework for the teacher model and the student model, improving the defect detection performance under compressed images. The student model learns the feature characterization and information extraction capabilities of the teacher model through knowledge distillation, compensates for the performance loss caused by image compression, significantly enhances the detection accuracy and model robustness, and provides an efficient and reliable solution for online detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] The present invention will be further described in detail below with reference to the accompanying drawings.

[0045] Figure 1 is a schematic diagram of the framework of the detection method of the present invention.

[0046] Figure 2 Comparison chart of the data of the present invention and the data of lossy compression under different YOLOv8 network models.

[0047] Figure 3 is an ablation experiment diagram of the training state of the EARN network of the present invention.

[0048] Figure 4 is an experimental data diagram of the present invention comparing different knowledge distillation methods.

[0049] Figure 5 is a visualization chart of the comparison of the detection results of the present invention and the original data. DETAILED DESCRIPTION OF THE INVENTION

[0050] To make the objectives, features, and advantages of the present invention clearly understandable, the following provides a detailed description of the specific implementation manners of the present invention in conjunction with the accompanying drawings.

[0051] As Figure 1 shown, the compression image object detection method based on enhancement and knowledge distillation is specifically carried out according to the following steps:

[0052] Step S1: Collect the original silicon wafer data Construct the first data set based on the original silicon wafer data and train to obtain the teacher model Obtain the silicon wafer data processed by compression under different quality factors q Construct the silicon wafer data sets under different quality factors;

[0053] Step S2: Use the compression artifact removal network to process the compressed data Y of different quality factors q q to obtain the intermediate feature map F q ∈R C×H×W This process is expressed as:

[0054]

[0055] represents the compression artifact removal network, R C×H×W represents the shape of the intermediate feature map, R represents the set of real numbers, and C, H, and W respectively represent the number of channels, width, and height of the feature map;

[0056] Step S3: Construct the collaborative learning framework of the teacher model and the student model to guide the training of the student model by combining logits distillation and feature distillation on the compressed data. The specific processes of the two different distillation methods are as follows:

[0057] The specific process of logits distillation is as follows:

[0058] During the learning process, the parameters of the teacher model are not updated. For the intermediate feature map F q , respectively pass through the teacher model and the student model It is expressed as:

[0059]

[0060] represents the logits output of F q passing through the teacher model, respectively representing the classification probability and the bounding box regression value, represents the logits output passing through the student model, respectively representing the classification probability and the bounding box regression value.

[0061] The output process of the student model mimicking the teacher model is expressed as:

[0062]

[0063] where σ(·) is the Sigmoid function, M represents the number of anchor points, N represents the number of target categories, represents the logits loss.

[0064] The specific process of feature distillation is as follows:

[0065] During the learning process, the parameters of the teacher model are not updated. For the intermediate feature map F q , after passing through the teacher model T and the student model respectively, the k-th feature maps of the teacher and the student are expressed as T k ∈R C×H×W and S k ∈R C×H×W (k = 1,.., K). At the same time, the k-th random mask is generated, expressed as:

[0066]

[0067] where is a random number in the range of (0, 1), x and y are the horizontal and vertical coordinates of the feature map respectively, ζ = 0.65, representing the mask ratio, represents the generated random mask;

[0068] Then, the corresponding mask is used to cover the feature map of the student, and the remaining pixels are used to generate the feature map of the teacher. This process is expressed as:

[0069] S′ = G(S·M) = (S·M)·W (6)

[0070] G(F k ) = ReLU(BN(Conv*F k )) (7)

[0071] where G represents the mapping layer, Conv represents the 3×3 convolutional layer, ReLU(·) represents the activation layer, BN(·) represents normalization, W represents the weight of the mapping layer, and F k represents the feature after mask processing;

[0072] The intermediate features of the student model mimicking the teacher model are expressed as:

[0073]

[0074] K represents the sum of the distillation layers, C, H, and W represent the shape of the feature map, and S and T represent the features of the student and the teacher respectively.

[0075] Total training loss of the network framework Expressed as:

[0076]

[0077] Wherein, represents the loss of the YOLOv8 network, represents the loss of logits distillation, represents the loss of feature distillation, represents the loss of the de-compression artifact network, and α, β, γ are the corresponding hyperparameters;

[0078] Wherein:

[0079]

[0080] is the class loss in the Yolov8 network, is the bounding box regression loss in Yolov8, and η1 and η1 are the corresponding hyperparameters, F i represents the data generated by the compression artifact removal network, Y i represents the original data, m represents the number of images, v represents the distillation loss decay term, which is used to control the influence of the knowledge distillation loss on the total loss, a represents the minimum value of the decay term, which is set to 0.01, b represents the maximum value of the decay term, which is set to 1, T max represents the maximum number of batches to control the decay process, which is set to 10, n i represents the current iteration number.

[0081] Step S4: Use the trained model to perform object detection on the original silicon wafer data.

[0082] Next, a simulation experiment is conducted on the compression image object detection method of the present invention based on enhancement and knowledge distillation.

[0083] Experimental settings

[0084] The method of the present invention is evaluated based on a self-built industrial-grade silicon wafer defect dataset, carried out on an open-source platform based on Pytorch, the system is a 64-bit Ubuntu operating system, the graphics card uses Nvidia RTX2080ti, the maximum video memory of the selected graphics card is 11GB, and the computing power is 7.5. This study systematically evaluates the performance robustness of the scheme in the JPEG compression domain. The experimental setting discrete quality factor parameter space Q = {10, 20, 30, 40, 50, 60}, and compares the performance of data with different quality factors as the training set on the original data.

[0085] The compression artifact removal network uses the output of the first stage of EARN, and the object detection network uses YOLOv8. At the same time, the teacher model uses YOLOv8m, while the student model uses YOLOv8s. We use the ADAMW optimizer, set the initial learning rate to 0.001, and the Batch size to 4.

[0086] Compared with the original data, Figure 2 shows the performance of different YOLOv8 models (YOLOv8n, YOLOv8s, and YOLOv8m) at different compression rates (q values). Each column in the table corresponds to a different compression rate (q = 10, q = 20, q = 30, q = 40, q = 50, q = 60), representing the performance of the model at different compression levels. The performance of each model varies at different compression rates. Among them, YOLOv8m performs the best at lower compression rates, with the highest accuracy, reaching 88.2% and 90.2%, while YOLOv8n and YOLOv8s are relatively lower. In addition, the table also lists the performance of each model on the original data as a comparison benchmark. YOLOv8m achieved the best performance of 92.3% on the original data. It can be seen from the table that as the compression rate increases, the accuracy of all models decreases, but the performance differences of different models in processing compressed images are more obvious. Therefore, the teacher model uses YOLOv8m and the student model uses YOLOv8s. Under the framework of the present invention, for lossy compressed wafer data, they achieved 2.5%, 2%, 2.2%, 2.3%, 2%, and 2.1% respectively. When the training set is distorted images and the test set is original images, the teacher model, by learning the high-level semantic information in the distorted data, has strong robustness and can provide a smoother decision boundary for the student model. The student model, by learning the predictions of the teacher model on the distorted images, can improve its generalization ability for the original images, reduce the interference caused by distortion, and finally also show good performance on the test set of the original images.

[0087] Figure 3 This table shows the performance of the EARN network at different compression rates and compares the effects of the compression artifact removal network and the object detection network in the frozen and training states. In the frozen state of the compression artifact removal network, the accuracy of the EARN model is relatively low. Especially at high compression rates, the performance of the untrained network decreases. While in the training state, the accuracy of the EARN model increases significantly, especially at higher compression rates, where it can effectively remove artifacts and restore image details, thus improving the robustness of object detection. In addition, the object detection network always shows stable performance in the training state, and its accuracy is further improved by combining with the trained EARN network.

[0088] Figure 4Shows the performance comparison of different methods at different compression rates. The results show that for logits distillation, in Logit distillation, the present invention not only focuses on the hard labels of the training set, but uses the unnormalized logits generated by the teacher model as the target to train the student model. These logits contain the decision-making information of the teacher model on the given image, especially the relationship information between classes, and this information is more useful than hard labels in some cases. In the case of image distortion, the teacher model can still provide richer decision-making information based on the knowledge learned during its training process, and this information may help the student model better understand the distorted image.

[0089] For feature distillation, the MGD method performs the best at all compression rates. Especially at high compression rates, the accuracies reach 91.3% and 91.5% respectively, significantly exceeding other methods, indicating that this method has stronger robustness and superior performance under high compression. The reason is that image distortion will affect the integrity of image features in some cases, making it difficult for standard detection or classification models to recognize or recover lost details. However, through random masking and generation mechanisms, the student model can better learn the non-linear relationships of the teacher model's features. Even after image distortion, the student model can still effectively generate feature maps approaching those of the teacher model. This learning method not only focuses on global image features but also aligns local features through generation and masking processes, thus enhancing the robustness of the student model. In contrast, the MIMIC and Logits methods perform better at lower compression, but as the compression rate increases, the accuracy decreases. Although the Logits method can still maintain a relatively high accuracy of 91.1% and 91.8% at high compression rates. The CWD and L1 methods perform relatively weakly.

[0090] The embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the purpose of the present invention.

Claims

1. A compressed image object detection method based on enhancement and knowledge distillation, characterized in that: The specific steps are as follows: Step S1: Collect the original silicon wafer data Construct the first data set based on the original silicon wafer data and train a teacher model Obtain the silicon wafer data after compression processing under different quality factors q Construct silicon wafer data sets under different quality factors; Step S2: Use a compression artifact removal network to process the compressed data Y with different quality factors q q to obtain an intermediate feature map F q ∈R C×H×W . This process is expressed as: Denote the compression artifact removal network as R C×H×W where is the shape of the intermediate feature map, R represents the set of real numbers, and C, H, and W represent the number of channels, width, and height of the feature map, respectively; Step S3: Construct a teacher model and the collaborative learning framework with the student model S, and guide the training of the student model by combining logits distillation and feature distillation on the compressed data Y q ; Step S4: Perform object detection on the original silicon wafer data using the trained model.

2. The method for compressed image object detection based on enhancement and knowledge distillation according to claim 1, wherein: The specific process of logits distillation in step S3 is as follows: During the learning process, the teacher model parameters are not updated. For the intermediate feature map F q , respectively pass through the teacher model student model It is expressed as: Denote F q After the logits output of the teacher model, which respectively represent the classification probability and the bounding box regression value, Denote the logits output after passing through the student model, which respectively represent the classification probability and the bounding box regression value; The process of the student model imitating the output of the teacher model is expressed as: where σ(·) is the Sigmoid function, M represents the number of anchor points, and N represents the number of target categories, represents the logits loss.

3. The method for detecting compressed image targets based on enhancement and knowledge distillation according to claim 2, wherein: The specific process of feature distillation in step S3 is as follows: During the learning process, the teacher model parameters are not updated. For the intermediate feature map F q , after passing through the teacher model and the student model respectively, the k-th feature maps of the teacher and the student are denoted as T k ∈R C×H×W and S k ∈R C×H×W (k = 1,..,K), and at the same time, the k-th random mask is generated, denoted as: Among them is a random number within the range of (0, 1), where x and y are the horizontal and vertical coordinates of the feature map respectively, ζ = 0.65 represents the mask ratio, represents the generated random mask; Then use the corresponding mask to cover the feature map of the student, and generate the feature map of the teacher with the remaining pixels. This process is expressed as: S′ = G(S·M) = (S·M)·W (6) G(F k ) = ReLU(BN(Conv * F k )) (7) Among them, G represents the mapping layer, Conv represents the 3×3 convolutional layer, ReLU(·) represents the activation layer, BN(·) represents normalization, W represents the weight of the mapping layer, and F k represents the feature after masking; The student model imitates the intermediate features of the teacher model, which is expressed as: K represents the sum of the distillation layers, C, H, W represent the shape of the feature map, and S and T represent the features of the student and the teacher respectively.

4. The method for detecting compressed image targets based on enhancement and knowledge distillation according to claim 3, wherein: Training the student model with compressed data of different quality factors q, the total training loss is expressed as: Among them, represents the loss of the YOLOv8 network, represents the loss of logits distillation, represents the loss of feature distillation, represents the loss of the decompression artifact network, and α, β, γ are the corresponding hyperparameters respectively; Where: is the class loss in the Yolov8 network, is the bounding box regression loss in Yolov8, where η1 and η1 are corresponding hyperparameters, and F i represents the data generated by the compression artifact removal network, and Y i represents the original data, m represents the number of images, v represents the distilled loss decay term, a represents the minimum value of the decay term, set to 0.01, b represents the maximum value of the decay term, set to 1, and T max represents the maximum number of batches controlling the decay process, set to 10, and n i represents the current iteration number.