Light guide plate defect detection method based on multi-scale feature fusion and knowledge distillation

Through the multi-scale feature fusion and knowledge distillation method, combined with the RT-DETR-LGP model, MSEIE module and ADPN module, the problems of large difficulty in feature extraction and scale differences in light guide plate defect detection are solved, and efficient small and medium-sized target feature extraction and detection accuracy are achieved.

CN120219310APending Publication Date: 2025-06-27ZHEJIANG SCI-TECH UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510281759.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing light guide defect detection methods have problems such as difficult feature extraction, scale differences affecting model performance, and traditional object detection models are difficult to extract small and medium-sized target features.

Method used

Using multi-scale feature fusion and knowledge distillation methods, the RT-DETR-LGP model combines the MSEIE module and the ADPN module to extract features of different scales and fusions, combining CWKD-based self-distillation strategy to improve detection accuracy and generalization.

Benefits of technology

It effectively solves the difficulty of extracting small defect features in light guide plates, reduces the impact of scale differences on model performance, and improves detection accuracy and generalization without increasing the model size, achieving a balance of accuracy and speed of surface defect detection of light guide plates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219310A_ABST
    Figure CN120219310A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of light guide plate detection, and particularly discloses a light guide plate defect detection method based on multi-scale feature fusion and knowledge distillation, which comprises the following steps: collecting a light guide plate image, preprocessing the light guide plate image, and inputting the preprocessed light guide plate image into an offline trained multi-scale feature fusion model; the multi-scale feature fusion model is improved by taking an RT-DETR model as a baseline network, a backbone network comprises CNN networks with four hierarchical structures, and an ADPN module is adopted in an encoder part to replace an original FPN-PAN module. The offline network training comprises a pre-training stage and a distillation training stage. The multi-scale feature fusion model effectively solves the problem of difficulty in feature extraction of small defects in the light guide plate, reduces the influence of scale differences among various defects in the image, further improves the detection precision and generalization while not increasing the size of the model in combination with a self-distillation strategy, and improves the detection accuracy. And the balance between the light guide plate surface defect detection precision and the detection speed is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of light guide plate detection and deep learning, and specifically relates to a light guide plate defect detection method based on multi-scale feature fusion and knowledge distillation. Background Art

[0002] A light guide plate is an optical component used to evenly distribute light sources, capable of converting point-like or line-like light sources into planar uniform light output, and is widely used in devices such as liquid crystal displays, LED lighting, and projection systems. However, during the production process of the light guide plate, due to factors such as raw materials, equipment usage, processing technology, and manual operation, various processing defects or flaws will inevitably appear on its surface. Therefore, defect detection of the light guide plate can ensure the yield rate of products and effectively prevent the assembly of defective light guide plates into displays, causing greater waste of resources.

[0003] Traditional light guide plate defect detection methods usually involve workers visually judging and classifying in an optometry room, which is time-consuming, laborious, and prone to errors. Moreover, factors such as differences in optometry standards among different workers and environmental conditions will also affect this process, resulting in potential errors in subsequent detection and classification, and affecting the production quality and efficiency of the light guide plate. To solve the low efficiency problem of manual detection, an automatic light guide plate surface defect detection method based on a target detection network is gradually replacing manual detection. However, there are the following challenges in the light guide plate surface defect detection based on the target detection network:

[0004] a) The texture background of the light guide plate image is complex, with a similar contrast to some defects, and redundant confusing features affect feature extraction, requiring more training time to distinguish actual defects from background noise.

[0005] b) There are many types of light guide plate defects, with different sizes and large scale differences in feature patterns, which hinder effective detection.

[0006] c) Small target defects account for a large proportion in the light guide plate defects, and traditional target detection models are difficult to extract effective features.

[0007] In recent years, transformer-based end-to-end detectors (DETRs) have made remarkable progress. The Real-Time Detection Transformer (RT-DETR) is based on the idea of DETR (without the NMS framework) and introduces a conv-based backbone and an efficient hybrid encoder to achieve real-time speed. However, in the light guide plate detection scenario, RT-DETR uses a CNN network (such as ResNet18) as the backbone network, and its feature extraction ability is relatively weak. When dealing with complex light guide plate defect detection tasks, it cannot capture sufficient semantic information and fine-grained features, resulting in performance degradation. Especially in light guide plate defects, there are many small and medium-sized defects, and ResNet18 cannot provide enough receptive fields, easily missing and misdetecting small defects. Summary of the Invention

[0008] The technical problem to be solved by the present invention is to provide a method for detecting light guide plate defects with multi-scale feature fusion and knowledge distillation, so as to efficiently and lightweight identify and extract light guide plate defects using a multi-scale feature fusion model trained with a self-distillation strategy based on CWKD.

[0009] To solve the above technical problem, the present invention provides a method for detecting light guide plate defects with multi-scale feature fusion and knowledge distillation, including collecting light guide plate images, and then inputting the preprocessed light guide plate images into the offline-trained multi-scale feature fusion model RT-DETR-LGP in the upper computer.

[0010] The multi-scale feature fusion model RT-DETR-LGP is improved based on the RT-DETR model as the baseline network, and includes a backbone network, an encoder, and a decoder. The backbone network contains a CNN network with 4 hierarchical structures, and outputs image features of three different scales: feature map F2, feature map F3, and feature map F4. In the encoder, feature map F4 undergoes decoupled single-scale feature interaction to obtain feature map S4, and feature maps F2, F3, and S4 pass through the ADPN module to obtain feature maps M4, M3, and M2. In the decoder, feature maps M2, M3, and M4 are fed into the detection head to obtain an image with the category and position of light guide plate defects.

[0011] As an improvement of the method for detecting light guide plate defects with multi-scale feature fusion and knowledge distillation of the present invention:

[0012] The operation of each said hierarchical structure is:

[0013] Input feature F l-1 (l = 1, 2, 3, 4) After two convolutional processes, it is divided into two halves along the channel dimension as feature and feature Feature Output features through the feed-forward network Features Output features through the MSEIE module Then the features And the features After the concatenation operation and then through a convolution process, output the feature F l .

[0014] As a further improvement of the light guide plate defect detection method for multi-scale feature fusion and knowledge distillation of the present invention:

[0015] The operation of the MSEIE module is as follows:

[0016] The said features Simultaneously pass through a parallel Conv 3×3 Branch and four EIE branches, and then obtain the output features after the concatenation operation and the convolution operation

[0017] In the said EIE branch, the input features Sequentially pass through adaptive average pooling, convolution, depth convolution, upsampling and EIE Block.

[0018] As a further improvement of the light guide plate defect detection method for multi-scale feature fusion and knowledge distillation of the present invention:

[0019] The operation of the EIE Block is as follows: The input features Calculate the difference with the features After the average pooling operation, and then add it to the features After convolution.

[0020] As a further improvement of the light guide plate defect detection method for multi-scale feature fusion and knowledge distillation of the present invention:

[0021] The operation of the ADPN module is as follows:

[0022] The feature maps F2, F3 and S4 pass through the first DASI module to obtain the feature map P3. Subsequently, the feature map P3 is downsampled and upsampled respectively. The downsampled feature map and the feature map F4 pass through the first Fusion module for dimensional feature aggregation to obtain the feature map N4. The upsampled feature map and the feature map F2 pass through the second Fusion module for dimensional feature aggregation to obtain the feature map N2; Then, the feature maps N4, N2 and P3 are used as the input of the second DASI module, and the obtained feature map N3 is directly output as the feature map M3; And the feature maps N4, the downsampled feature map of N3, and the downsampled features of the feature map P3 Figure 1Starting from the third Fusion module, the feature map M4 is obtained, and the feature maps N2, the upsampled feature maps of the feature map N3, and the upsampled features of the feature map P3 Figure 1 Starting from the fourth Fusion module, the feature map M2 is obtained.

[0023] As a further improvement of the light guide plate defect detection method with multi-scale feature fusion and knowledge distillation of the present invention:

[0024] The preprocessing includes grayscale enhancement, edge detection, least squares fitting, perspective transformation, and tilt correction of the light guide plate image, and then it is cropped to a unified size.

[0025] As a further improvement of the light guide plate defect detection method with multi-scale feature fusion and knowledge distillation of the present invention:

[0026] The offline training includes two stages: offline pre-training and offline distillation training.

[0027] As a further improvement of the light guide plate defect detection method with multi-scale feature fusion and knowledge distillation of the present invention:

[0028] The process of the offline pre-training is as follows:

[0029] The ILGPDD dataset is used for offline pre-training. The labeled pictures in the training set and the validation set are input into the model in sequence, the loss function value is calculated, and the model parameters are iteratively optimized by backpropagation. After reaching the preset number of epochs, the training ends; then the test set is input into the model, and the model performance is verified through evaluation metrics, and the multi-scale feature fusion model RT-DETR-LGP that is well pre-trained offline is obtained.

[0030] As a further improvement of the light guide plate defect detection method with multi-scale feature fusion and knowledge distillation of the present invention:

[0031] The process of the offline distillation training is as follows:

[0032] The well-offline pre-trained scale feature fusion model RT-DETR-LGP is used as both the teacher model and the student model at the same time. The offline distillation training is carried out by using the distillation method based on CWKD. The training set and the validation set of the ILGPDD dataset are input into the model in sequence. By calculating the KL divergence, the difference between the channel distributions of the teacher model and the student model is evaluated and added as the distillation loss to the total loss function, and the backpropagation and parameter update are carried out to obtain a multi-scale feature fusion model RT-DETR-LGP with better performance;

[0033] Then the test set is input into the multi-scale feature fusion model RT-DETR-LGP with better performance, and the model performance is verified through evaluation metrics.

[0034] As a further improvement of the light guide plate defect detection method based on multi-scale feature fusion and knowledge distillation of the present invention:

[0035] The construction process of the ILGPDD dataset is as follows:

[0036] Collect the production pictures of the light guide plate on the production line. After preprocessing the collected pictures in the computer, divide the pictures with light guide plate defects into training set, validation set and test set after manually annotating the light guide plate defect types and target boxes;

[0037] The light guide plate defect types include scratches, missing engraving, pressing damage, white stains, black spots, scars, white shadows.

[0038] The beneficial effects of the present invention are mainly reflected in:

[0039] (1) The multi-scale feature fusion model of the present invention effectively solves the challenge of difficult feature extraction of small defects in the light guide plate, reduces the influence of the scale difference between various defects in the image on the effectiveness of modeling: and combines the self-distillation strategy based on CWKD to further improve the detection accuracy and generalization without increasing the model size, achieving the balance of the detection accuracy and detection speed of the light guide plate surface defects;

[0040] (2) The MSEIE module proposed by the present invention combines the CSP idea, lightens the model while deeply extracting defect feature information, improves the sensitivity of the model to edge information, and enhances the model's recognition ability for small and medium targets;

[0041] (3) The ADPN module proposed by the present invention can make each scale of features have detailed context information, and through the diffusion mechanism, make the features with rich context information diffuse to each detection scale, improving the multi-scale detection ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The following further describes in detail the specific embodiments of the present invention with reference to the accompanying drawings.

[0043] Figure 1 It is a schematic structural diagram of the multi-scale feature fusion model of the present invention;

[0044] Figure 2 It is a schematic structural diagram of the backbone network of the multi-scale feature fusion model of the present invention;

[0045] Figure 2 In:

[0046] (a) is the structural diagram of the backbone network, (b) is the structural diagram of the stage in the backbone network, (c) is the structural diagram of the MSEIE module in the stage, and (d) is the schematic structural diagram of the EIE Block in the MSEIE module;

[0047] Figure 3 It is a schematic structural diagram of the ADPN module of the present invention;

[0048] Figure 4 It is a schematic structural diagram of the DASI module;

[0049] Figure 5 It is a schematic structural diagram of the Fusion module;

[0050] Figure 5 Among them:

[0051] (a) is a schematic structural diagram of the Fusion module, and (b) is a schematic structural diagram of RepConv in the Fusion module;

[0052] Figure 6 It is a schematic diagram of the knowledge distillation technology based on channel knowledge extraction;

[0053] Figure 7 It is a relationship diagram of the distillation temperature T versus AP (%) under different loss weights α in the knowledge distillation experiment of the present invention;

[0054] Figure 7 Among them, from left to right are: the relationship with AP, the relationship with AP 50 relationship, the relationship with AP 75 relationship. Specific embodiments

[0055] The present invention will be further described below in conjunction with specific embodiments, but the protection scope of the present invention is not limited thereto:

[0056] Embodiment 1. A method for detecting light guide plate defects by multi-scale feature fusion and knowledge distillation, as Figure 1 shown. First, construct the multi-scale feature fusion (RT-DETR-LGP) model of the present invention, and then use the self-distillation strategy based on CWKD to train the multi-scale feature fusion (RT-DETR-LGP) model. Without increasing the model size, the detection accuracy and generalization are further improved, achieving a balance between the detection accuracy and detection speed of the light guide plate surface defects, so as to obtain a multi-scale feature fusion (RT-DETR-LGP) model for online use for detecting light guide plate defects in the actual production environment. The specific process is as follows:

[0057] 1. Construct a multi-scale feature fusion (RT-DETR-LGP) model

[0058] The multi-scale feature fusion (RT-DETR-LGP) model of the present invention is improved based on the RT-DETR model as the baseline network, and the overall structure is as Figure 1As shown, it includes a backbone network, an encoder, and a decoder. In the backbone network part, a CNN network with 4 hierarchical structures (stages) is used to replace the original ResNet18 network. Each stage is composed of an MSEIE module combined with a CSP (Cross Stage Partial) structure. In the encoder part, an ADPN module is used to replace the original FPN-PAN module.

[0059] 1.1 Backbone Network

[0060] Due to the large number of small targets and low-contrast defects in the light guide plate defects, and the lack of important high-frequency feature information due to the insufficient feature extraction ability of the RT-DETR model. Based on this, in the present invention, an MSEIE module is used in each stage of the backbone network of the multi-scale feature fusion (RT-DETR-LGP) model to extract features of different scales to enhance the defect edge information, and at the same time, the CSP structure idea is combined to lightweight the model.

[0061] The backbone network of the multi-scale feature fusion (RT-DETR-LGP) model is a feature extraction network similar to VGG and ResNet, as Figure 2 (a) shows, it contains a convolution and four sequentially arranged stages l (l = 1, 2, 3, 4), and each stage l adopts the CSP structure, as Figure 2 As shown in b, the CSP structure is an architecture optimization method for improving the computational efficiency and feature utilization rate of neural networks. The feature map is divided into two parts, one part is directly transmitted to the subsequent layer, and the other part is fused with the former after convolution calculation, which can reduce redundant calculations, improve gradient flow, and contribute to optimizing the training effect of deep networks.

[0062] Specifically, the input feature F of stage1 l-1 After being processed by two convolutions (Conv), it will be divided into two halves along the channel dimension into feature and feature which are respectively sent into two parallel paths. One path is a simple feed-forward network (Feed-Forward Network, FFN), and the input is feature and then the output feature is The other path contains an MSEIE module, and the input is feature and then the output feature is The formula for dividing the input feature tensor into two halves along the channel dimension after convolution processing is:

[0063]

[0064] where Fl-1 is the input feature, and Conv represents the convolution operation;

[0065]

[0066] where C l is the number of channels of the feature map.

[0067]

[0068] The final output of stage1 is:

[0069]

[0070] where Concat is the concatenation operation.

[0071] The structure of the MSEIE module is as Figure 2 (c) shown. The input feature simultaneously passes through a parallel Conv 3×3 branch and four EIE branches, and then obtains the output feature after the concatenation operation and the 1×1 convolution operation In the EIE branch, first, adaptive average pooling with different scales is adopted, and then 1×1 convolution is performed to obtain four local region features of different sizes of 3×3, 6×6, 9×9, and 12×12, and multi-level information of the image is obtained:

[0072]

[0073] where represents adaptive average pooling, {k (m)} = {3, 6, 9, 12}, m = 1, …, 4.

[0074] Secondly, depth convolution is used to capture context information at multiple scales, and bilinear interpolation is used for upsampling to align the features extracted at different scales to the same scale:

[0075]

[0076] Furthermore, the feature map containing different local feature information is enhanced with edge information through the Edge Information Enhancement (EIE) Block. The structure of the EIE Block is as Figure 2 (d) shown. The EIE Block smooths the feature through the average pooling operation, extracts low-frequency information, then calculates the difference from the feature , obtains the enhanced edge feature through 1×1 convolution, and combines it with the feature Add them to obtain a feature map with enhanced edge information, increasing the model's sensitivity to edge information:

[0077]

[0078] After connecting the feature L 3×3 output from the Conv l-1 branch and the feature output from the EIE branch, the local and contextual features with enhanced edge information are fused through a 1×1 convolution, representing the mutual relationships between various channels:

[0079]

[0080] where L l-1 represents the output feature of the input feature after passing through the Conv 3×3 branch. The 1×1 convolution is used as a channel fusion mechanism to integrate features with different receptive field sizes and obtain a unified feature representation. In this way, the MSEIE module can capture extensive contextual information, improve the model's perception of multi-scale features, without compromising the integrity of local contextual features.

[0081] 1.2 Aggregate Diffusion Pyramid Network (ADPN)

[0082] Due to the inconsistent sizes of the surface defects of the light guide plate, the model needs to handle problems of different scales during training and recognition. Even for the same type of defect, it may appear in different sizes in different images. These challenges of multi-scale problems require the model to be able to handle defects of different sizes and have strong scale invariance and generalization ability.

[0083] To address the multi-scale challenges in the detection of light guide plate surface defects, the present invention designs an Aggregate Diffusion Pyramid Network (ADPN) module to achieve multi-scale feature fusion in order to capture a more comprehensive array of defect feature information. As Figure 3 shown, the ADPN module mainly consists of two parts: the DASI module and the Fusion module. The DASI module is used to aggregate feature information of different dimensions, and the Fusion module is used to aggregate feature information of the same dimension. After upsampling and downsampling, the aggregated feature map will spread to other dimensions for further feature fusion. This "aggregation + diffusion" feature fusion method enables each scale of features to have detailed contextual information, which helps to detect subtle features in the defect images of the light guide plate, thereby enhancing the model's detection ability.

[0084] The ADPN module receives the feature maps F2 and F3 extracted from the backbone network, as well as the feature map S4 processed by the AIFI module. After the first DASI module aggregates the feature information, a new feature map P3 is obtained. Subsequently, through its unique "diffusion" mechanism, the ADPN module down-samples and up-samples the feature map P3 respectively. The down-sampled feature map and the feature map F4 are used for dimensional feature aggregation through the first Fusion module to obtain the feature map N4, and the up-sampled feature map and the feature map F2 are used for dimensional feature aggregation through the Fusion module to obtain the feature map N2. Then, the feature maps N4, N2, and P3 are used as the input of the second DASI module for further feature aggregation, and a highly mixed feature map N3 is directly output as the feature map M3. The feature maps N4, the down-sampled feature map of N3, and the down-sampled feature of P3 Figure 1 are used for the diffusion and fusion of feature information through the third Fusion module to obtain a new feature map M4. At the same time, the feature maps N2, the up-sampled feature map of N3, and the up-sampled feature of P3 Figure 1 are used to obtain a new feature map M2 through the Fusion module.

[0085] In multiple down-sampling stages of the model, high-dimensional features may lose information about small objects, while low-dimensional features may not provide enough context. DASI can adaptively select appropriate features for fusion according to the size and characteristics of the object, as Figure 4 shown. The DASI module first aligns high-dimensional features with low-dimensional features and the current layer features through operations such as convolution and interpolation, and then divides them into four equal segments in the channel dimension to obtain and where h i , l i , and u i represent the i-th partition features of high-dimensional, low-dimensional, and current layer features respectively. The partitioning method is:

[0086]

[0087] where, represents the value obtained by the activation function applied to u i , represents the selective aggregation result of each partition. If α > 0.5, the model gives priority to fine-grained features, while if α < 0.5, it emphasizes context features. After merging in the channel dimension, the operations Conv(), Convolution, batch normalization (BN), and rectified linear unit (ReLU) are denoted by and δ() respectively, and the final output obtains features

[0088] After the DASI module aggregates features of three different dimensions, high, medium, and low, the output is spread to high-dimensional and low-dimensional spaces through upsampling and downsampling respectively. It is fused with adjacent features of the same dimension through the Fusion module to form a new feature, further extracting deep defect information. The Fusion module is similar to the CSP structure, as Figure 5 (a) shows, which consists of convolution and RepConv, and fuses the output through element-wise addition. RepConv is as Figure 5 (b) shows, which is a reparameterized convolution. During training, it contains multiple branches, including 1×1 convolution, 3×3 convolution, and identity mapping, extracting defect features in different scale directions and enhancing the model's modeling ability. In the inference stage, the multi-branch structure is reparameterized into an equivalent 3×3 convolution RepConv, which contains all the feature information learned by the multi-branch structure during training, and can significantly improve the calculation speed and reduce the latency.

[0089] 1.3. Multi-scale Feature Fusion (RT-DETR-LGP) Model

[0090] Integrating Steps 1.1 - 1.2, the multi-scale feature fusion (RT-DETR-LGP) model of the present invention is constructed. The input light guide plate image first passes through 4 stages of the backbone network to extract three different scale image features of feature map F2, feature map F3, and feature map F4. In the hybrid encoder, feature map F4 passes through the decoupled single-scale feature interaction (AIFI) to obtain feature map S4. Feature maps F2, F3, and S4 pass through ADPN to obtain feature maps M4, M3, and M2. Finally, feature maps M2, M3, and M4 are highly mixed and sent to the detection head for decoding and detection to obtain an image with the light guide plate defect category and location.

[0091] The backbone network of the multi-scale feature fusion (RT-DETR-LGP) model of the present invention is mainly used to extract multi-scale image features in the light guide plate image, which helps to enhance feature fusion in subsequent processing. The efficient hybrid encoder further refines the extracted features, and efficiently fuses feature maps of different dimensions by decoupling single-scale feature interaction (AIFI) and aggregating diffusion pyramid network (ADPN) to enhance the understanding of various object scales and contexts within the image, improving the model's expressive ability. This design enables the encoder to effectively manage the computational complexity and improve the quality of the features fed into the decoder. Finally, by introducing IoU-aware to calculate the loss, more valuable queries are screened out and sent to the detection head together, and bipartite graph matching is performed between the output and the ground truth to determine the position and category of the target.

[0092] The MSEIE module designed in the present invention, combined with the CSP idea, constructs four stages as components of the backbone network. Through multi-branch pooling and depth convolution operations, feature information of different scales is extracted. At the same time, the EIE Block is used to enhance the edge information and perform fusion, which helps to extract the global features of the light guide plate defects, thereby improving the detection accuracy of the model for small and medium-sized defects. In addition, this design also reduces the number of parameters of the model and realizes the lightweight of the model.

[0093] The hybrid encoder in the neck is used to further refine the extracted features. The RT-DETR network adopts the feature fusion method of AIFI+FPN-PAN, and its multi-scale feature fusion ability is weak. When facing light guide plate defects with large scale changes, it is difficult for the model to balance the detection accuracy of defects of different sizes. The present invention proposes the ADPN module to promote multi-scale information fusion. In ADPN, the DASI module focuses on the adaptive selection and fine fusion of high-dimensional and low-dimensional features to enhance the saliency of small defects, combines with the Fusion module to aggregate feature information, reduces the loss of fine-grained information, and then fuses information of different dimensions through the diffusion mechanism to enhance the detection accuracy of the model for multi-scale defects.

[0094] 2. Offline Training of the RT-DETR-LGP Model

[0095] 2.1 Knowledge Distillation

[0096] To further improve the accuracy and generalization of the model without introducing additional computational costs, the offline training of the multi-scale feature fusion (RT-DETR-LGP) model of the present invention uses a knowledge distillation technique based on Channel-Wise Knowledge Distillation (CWKD). Through additional supervised training transferred from the teacher model, the KL divergence between the predictions of the teacher and student models is minimized. The present invention adopts a self-distillation method as Figure 6 shown, that is, the teacher model is restricted to the student itself to guide the student model to learn. Since the model neck fuses the high-frequency and low-frequency features of the input image and has rich semantic information, the output features of the model neck are selected for knowledge transfer and minimization of KL divergence. CWKD softly aligns the activations of the corresponding channels between the teacher and student models and converts them into probability distributions, so that the KL divergence can be used to measure the difference to better utilize the knowledge in each channel.

[0097] Suppose the teacher model and the student model are represented as T and S respectively, and the activation maps from T and S are represented as y T and y S . The CWKD loss can be expressed as:

[0098]

[0099] where φ(·) is used to convert the activation values into probability distributions, as shown in Equation (9):

[0100]

[0101] where c = 1, 2, …, C is the channel index, i is the spatial position index of the channel, is a hyperparameter (temperature). If a larger is used, the probability distribution is smoother, which can make the model focus on a wider spatial region of each channel. Through softmax normalization, the inconsistency in the magnitude between the outputs of different layers or channels inside the model can be eliminated, which helps knowledge distillation. If the number of channels between the teacher and student models does not match, a 1×1 convolutional layer is used to upsample the number of channels of the student model. is used to calculate the KL divergence to evaluate the difference between the channel distributions of the teacher and student models:

[0102]

[0103] The KL divergence is an asymmetric metric. It can be seen from Equation (11) that if is very large, then should be as large as to minimize the KL divergence. Otherwise, if is very small, the KL divergence is minimized receives less attention. Therefore, the student model can better learn the distribution of the foreground significant regions, while learning less about the activations corresponding to the background regions of the teacher model, which helps the knowledge distillation learning of the light guide plate defect detection model under complex backgrounds.

[0104] The total loss function is formulated as:

[0105]

[0106] where \(y\) is the true label value, \(l(\cdot)\) is the original task loss calculated using the prediction and the label, and \(\alpha\) is a hyperparameter used to balance the two losses. In the early stage of training, the soft labels from the teacher model are easier to learn. As the training continues, the performance of the student model will match that of the teacher model, so the hard labels will be more helpful to the student.

[0107] 2.2 Training Dataset

[0108] The dataset used for offline training of the present invention is named Industrial Light Guide Plate DefectDataset (ILGPDD). Table 1 details the quantities of various defect categories existing in the dataset.

[0109] Table 1 Defect Categories and Quantities in Dataset ILGPDD

[0110]

[0111] ILGPDD was specifically constructed for this study, and its sample data all come from industrial sites. A large number of high-quality light guide plate defect pictures were collected by a 16K line scan camera, and the defect annotation information was screened and annotated using annotation software. This dataset contains 10,831 defect images, and each picture was manually annotated with the light guide plate defect type and the target box. The light guide plate defect types include: scratch (Scratch, SCR), missing (Missing, MIS), dark shadow (Dark shadow, DSH), white stain (Whitestain, WST), dark stain (Dark_stain, DST), scar (Scar, SCA), white shadow (White_shadow, WSH), and they were divided into a training set, a validation set, and a test set according to the ratio of 6:2:2.

[0112] 2.3 Offline Training Process

[0113] 2.3.1 Offline Training Hardware Selection and Parameter Settings

[0114] The multi-scale feature fusion (RT-DETR-LGP) model is implemented using the Python language and the PyTorch deep learning framework. For the offline training settings, an NVIDIA GeForce RTX 4090 with 24GB video memory (RAM) is used as the hardware configuration, and Ubuntu 20.04.5 is selected as the operating system. The multi-scale feature fusion (RT-DETR-LGP) model is trained for 100 epochs with a batch size of 4. The input image size is set to 640×640, the learning rate is set to 0.0001, and the AdamW optimizer is used for model tuning.

[0115] During the distillation training process, both the teacher model and the student model use the multi-scale feature fusion (RT-DETR-LGP) model, and the parameters used in distillation training are the same as those in ordinary training. The distillation loss weight α is set to 0.2, and the distillation temperature is set to 4.

[0116] 2.3.2. Offline Pre-training Using the ILGPDD Light Guide Plate Defect Dataset

[0117] When performing offline pre-training on the multi-scale feature fusion (RT-DETR-LGP) model, the data in the ILGPDD dataset is used for training. The labeled images in the training set and the validation set are sequentially input into the model. The input samples are scaled to a size of 640×640, and then the backbone network is used to extract the feature information in the images. After extracting the shallow feature information, the neck encoder performs feature fusion on it to integrate high-resolution information and high-semantic information, and obtains feature maps of three sizes: large, medium, and small. Finally, the highly mixed feature maps are sent to the detection head for decoding to generate the light guide plate defect types and candidate prediction boxes. The candidate prediction boxes are continuously optimized according to the loss function, and parameter updates are performed through backpropagation to train the model to obtain weights closer to the labels. After 100 rounds of training, the offline pre-trained multi-scale feature fusion (RT-DETR-LGP) model is finally obtained.

[0118] Then, the labeled test set images are input into the trained model to generate candidate prediction boxes. The intersection over union (IoU) between the candidate prediction boxes and the labeled boxes is calculated. When the IoU is greater than 0.5, it is a positive sample; when it is less than 0.5, it is a negative sample. The average precision mean (AP) is used as the evaluation metric, and AP0.5 = 97.6 is obtained for use in the next step of offline distillation training.

[0119] 2.3.3. Offline Distillation Training Using the ILGPDD Light Guide Plate Defect Dataset

[0120] When the multi-scale feature fusion (RT-DETR-LGP) model is trained for offline distillation, the sample data in the ILGPDD dataset is used for offline distillation training. The multi-scale feature fusion (RT-DETR-LGP) model pre-trained offline in step 2.2.2 is used as both the teacher model and the student model, receiving the input training samples for self-distillation training. The CWKD-based distillation method soft-aligns the activations between the channels of the neck feature map between the teacher model and the student model and converts them into probability distributions. The three aligned feature maps are selected from the three feature maps of different sizes sent to the detection head by the neck, because the neck feature map has rich semantic and contextual information after feature fusion, which can provide more soft label information for the model. The KL divergence is calculated to evaluate the difference between the channel distributions of the teacher model and the student model, and is added to the total loss function as the distillation loss, and back-propagation and parameter updates are performed to finally obtain a multi-scale feature fusion (RT-DETR-LGP) model with better performance. It can achieve lossless improvement of the model accuracy without increasing the model parameters and complexity.

[0121] During the test, the image data in the test set was used for testing. For the input test image, the multi-scale feature fusion (RT-DETR-LGP) model will also scale the image size, extract the feature information in the image and encode it, and finally send it to the detection head for decoding to detect the target position and category. The average precision (AP) is used as the evaluation index, and AP0.5 = 98.0 is obtained. The multi-scale feature fusion (RT-DETR-LGP) model that has been trained by offline pre-training and offline distillation is used for light guide plate defect detection in the online light guide plate production process.

[0122] 3. Online use

[0123] An image acquisition device is set up at the end of the light guide plate production line to collect the light guide plate image, including a multi-angle light source and a line scan camera, so as to comprehensively and clearly collect high-quality images. The acquired high-quality image will undergo a series of preprocessing operations, including grayscale enhancement, edge detection, least squares fitting, and perspective transformation to perform tilt correction on the light guide plate image. The image is then cropped to a uniform size and sent to the multi-scale feature fusion (RT-DETR-LGP) model that can be used online in step 2 to obtain the output of the light guide plate defect type and the image of the target frame.

[0124] 4. Experiment

[0125] 4.1 Evaluation Metrics

[0126] Common evaluation metrics include Precision (P), Recall (R), Average Precision (AP[class]), and AP coco , which are defined as follows:

[0127]

[0128] AP[class] = ∫P(R)d(R) (13)

[0129]

[0130] where TP represents true positive, FP represents false positive, and FN represents false negative. AP[class] is used to measure the detection performance of the model for each class. AP coco represents the average AP[class] value for all classes, which is used as the core metric to evaluate the overall performance of the model, including AP, AP 50 , AP 75 evaluates the model performance at different IoU thresholds of 0.5:0.05:0.95, 0.5, and 0.75, and AP S , AP M , AP L are used to evaluate the detection ability of the model for small, medium, and large targets, respectively.

[0131] 4.2, Ablation Experiments

[0132] The present invention conducts ablation experiments using the dataset ILGPDD (only using offline pre-training as shown in step 2.3.3), and the experimental results are shown in Table 2. The first row of data in the table represents the experimental results of the RT-DETR baseline network with ResNet18 as the backbone network; the second row of data in the table represents the experimental results of the network after replacing the backbone network of the RT-DETR baseline network with the backbone network constructed in step 1.1; the third row of data in the table represents the experimental results of the network after only replacing the FPN-PAN module in the encoder of the RT-DETR baseline network with the ADPN module; the fourth row of data in the table represents the experimental results of the RT-DETR-LGP model of the present invention.

[0133] The MSEIE module and the ADPN module are sequentially integrated into the base model with ResNet18 as the backbone,

[0134] Table 2. Ablation Experiment Results of the MSEIE Module and the ADPN Module

[0135]

[0136] As shown in Table 2, compared with the baseline network RT-DETR, the AP and AP 50 ,AP 75 The AP S 、AP M ,AP L The ADPN structure used in the design has an impact on AP 50 There is a significant boost, which improves it by 0.8%, but with a slight increase in the number of parameters. The designed MSEIE module helps the RT-DETR-LGP model achieve an average FPS of 176, which means that the RT-DETR-LGP model will perform real-time reasoning in a more lightweight way.

[0137] The backbone network of the RT-DETR-LGP model combines the MSEIE module and CSPNet ideas, achieving significant performance improvements in model inference speed. The MSEIE module benefits from the multi-branch structural design, enabling it to deeply mine feature information at multiple scales, and strengthen edge information through the EIE Block, thereby enabling the model to better focus on defect location information. In addition, the hierarchical residual connection combined with the CSP structure not only effectively solves the problem of gradient disappearance, but also enhances the model's detection ability for small and medium targets. The results show that AP S 、AP M An increase of 0.3% and 3.7% respectively.

[0138] The ADPN module designed in the neck improves the detection accuracy of small, medium and large targets by 0.5%, 2.0% and 2.6% respectively, effectively improving and balancing the detection accuracy of defects of different scales. ADPN aggregates feature information of different dimensions through the DASI module and the Fusion module, fully mines fine-grained features, generates richer and highly fused feature representations, and diffuses them to different dimensions for interaction. This "aggregation + diffusion" feature fusion strategy allows features with rich contextual information to be propagated to each detection scale, significantly improving the model's ability to handle multi-scale targets, especially in the complex light guide plate image background, and can more accurately identify and locate defects of different sizes.

[0139] 4.3 Comparative Experiments on Multi-Scale Feature Fusion

[0140] Considering the scale gap between the surface defects of the light guide plate, the present invention designs an ADPN module to fully fuse the feature information of different dimensions, so as to achieve more accurate positioning and classification of defects. To further prove the effectiveness of the ADPN module in fusing multi-scale features, the present invention compares it with other multi-scale feature fusion methods, such as FPN [1] , BiFPN [2] , HSFPN [3] and MAFPN [4] . The experimental results are shown in Table 3. The first row of the table represents the experimental results of the network after replacing the FPN-PAN module in the encoder of the RT-DETR baseline network with FPN; the second row of the table represents the experimental results of the network after replacing the FPN-PAN module in the encoder of the RT-DETR baseline network with BiFPN; the third row of the table represents the experimental results of the network after replacing the FPN-PAN module in the encoder of the RT-DETR baseline network with HSFPN; the fourth row of the table represents the experimental results of the network after replacing the FPN-PAN module in the encoder of the RT-DETR baseline network with MAFPN; the fifth row of the table represents the experimental results of the network after replacing the FPN-PAN module in the encoder of the RT-DETR baseline network with ADPN

[0141] Table 3. Comparative experiments using different multi-scale feature fusion models

[0142]

[0143] As can be seen from the table, although FPN and HSFPN have fewer parameters, their detection accuracy is low, and the AP 50 is only 96.7% and 96.5%, showing the limitations of these two methods in practical applications. In addition, FPN has poor detection accuracy for small targets, and the AP S is only 62.0%, and BiFPN has poor detection accuracy for medium targets, and the AP M is only 58.5%. And these feature fusion methods have low detection accuracy for large targets and cannot effectively balance the detection of multi-scale targets

[0144] Compared with these feature fusion methods, ADPN shows better performance in the detection of surface defects of the light guide plate. It not only improves the AP index by 1.3% compared with FPN, but also improves the AP 50 and AP 75They increased by 0.9% and 2.3% respectively. In addition, the detection accuracies of ADPN for small, medium, and large targets reached 66.3%, 63.8%, and 74.6% respectively, proving that it has the best effect in multi-scale feature recognition. These results confirm that ADPN can effectively fuse feature information of different dimensions, reduce the loss of fine-grained information, and enhance the robustness of the model to scaled targets.

[0145] 4.4. Verification of the self-distillation strategy

[0146] Considering that the hard labels contain limited information, only providing information of a single category and not containing any information about other categories, which makes the model ignore the similarities and differences between categories during the training process, thus affecting the generalization ability of the model. Therefore, this paper adopts a knowledge distillation method based on CWKD for self-distillation to enhance the generalization performance of the model. To verify the effectiveness of the self-distillation strategy based on CWKD and the rationality of the selection of its hyperparameters, this paper conducts corresponding ablation experiments by adjusting the temperature parameter to optimize the channel probability map, and then evaluate the influence of different hyperparameter settings on the results of distillation learning. The loss weight α is set to 0.1, 0.2, 0.3, and the distillation temperature T ∈ [1, 5]. The experimental results are as Figure 7 shown.

[0147] According to Figure 7 the data in, it can be observed that when is small, the distillation learning accuracy of the model is low and the knowledge refinement effect is poor. As increases, the distillation accuracy of the model gradually improves and the probability distribution of the channel feature information becomes softer. Under the settings of α = 0.2 and , using CWKD for self-distillation has the best effect, and AP, AP 50 , AP 75 reach 70.3%, 98.0, and 83.5% respectively. Compared with the model before distillation, they increase by 1.4%, 0.4%, and 2.0% respectively. This improvement is because the self-distillation strategy based on CWKD provides more soft label information for the student model by converting feature activations into soft probability values and calculating the KL divergence of the feature distributions between the teacher model and the student model channels. These results strongly confirm that the knowledge distillation strategy based on CWKD can provide more learnable information entropy for the model, thus significantly enhancing the generalization ability of the model.

[0148] 4.5. Comparative experiments

[0149] To evaluate the effectiveness and generalizability of the RT-DETR-LGP model and the CWKD self-distillation strategy, a comparative experiment was conducted on the light guide plate defect dataset (ILGPDD) by comparing the RT-DETR-LGP model with mainstream object detection models. The mainstream object detection models include Faster R-CNN [5] , RetinaNet [6] , FSAF [7] , ATSS [8] , DETR [9] , DINO

[10] , AYOLOv3-tiny

[11] and HM-YOLOv5

[12] . The results of the comparative experiment are shown in Table 4

[0150] Table 4. Comparative Experiment Results of RT-DETR-LGP Model and Existing Models

[0151]

[0152]

[0153] In Table 4, RT-DETR-LGP represents the RT-DETR-LGP model constructed in the present invention that only adopts the offline pre-training method described in 2.3.2, and RT-DETR-LGP(+CWKD) represents the RT-DETR-LGP model constructed in the present invention that adopts the offline distillation training method described in 2.3.3

[0154] As can be seen from Table 4, the AP, AP 50 , AP 75 of the RT-DETR-LGP model proposed in this paper reached 68.9%, 97.6%, and 81.5% respectively. Compared with the traditional two-stage object detection model Faster R-CNN, the RT-DETR-LGP model is more efficient, with improvements of 5.1%, 0.6%, and 6.5% in AP, AP 50 , AP 75 respectively. Compared with traditional single-stage object detection models using multi-level feature extraction, such as RetinaNet, FSAF, and ATSS, the AP increased by 9.2%, 8.7%, and 5.3% respectively, and the AP 50 increased by 2.6%, 2.0%, and 0.8% respectively. Compared with object detection models using global feature extraction, such as DETR and DINO, the AP increased by 12.7% and 2.8% respectively, and the AP 50 increased by 2.3% and 0.2% respectively. Compared with AYOLOv3-Tiny, the AP, AP 50 , AP 75They were increased by 2.2%, 1.2%, and 3.7% respectively. These results indicate that by using the edge information enhancement and multi-scale feature fusion strategies, the RT-DETR-LGP model can effectively improve the accuracy of detecting surface defects of light guide plates.

[0155] In terms of the detection accuracy of defects of different sizes, traditional object detection models not only have a large number of parameters but also have low detection accuracy for objects of different scales. Although DETR and DINO have good detection accuracy for large objects, their detection accuracy for small objects is low, and the AP S is only 49.2% and 61.5% respectively. Although AYOLOv3-Tiny has few parameters, its detection accuracy for medium-sized objects is low, and the AP M is only 55.7%. Although HM-YOLOv5 has a fast detection speed, its detection accuracy for large objects is low, and the AP L is only 65.5%. The AP S 、AP M 、AP L of the RT-DETR-LGP model reached 66.3%, 63.8%, and 74.6% respectively, achieving a balance in the detection accuracy of defects of different scales while maintaining an efficient detection speed of 176 FPS.

[0156] In addition, by using the self-distillation strategy based on CWKD, the AP, AP 50 、AP 75 of the model were increased by 1.4%, 0.4%, and 2.0% respectively, and the AP S 、AP M 、AP L were increased by 0.6%, 4.9%, and 2.9% respectively, achieving a lossless accuracy improvement without changing the number of model parameters.

[0157] The references involved in the above text are as follows:

[0158] [1] LIN T Y, DOLLáR P, GIRSHICK R, et al. Feature Pyramid Networks for Object Detection;proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), F21-26 July 2017, 2017[C]. https: / / doi.org / 10.1109 / CVPR.2017.106

[0159] [2]TAN M,PANG R,LE Q V.EfficientDet:Scalable and Efficient Object Detection;proceedings of the 2020 IEEE / CVF Conference on Computer Vision and Pattern Recognition(CVPR),F 13-19 June 2020,2020[C]. https: / / doi.org / 10.1109 / CVPR42600.2020.01079

[0160] [3]CHEN Y,ZHANG C,CHEN B,et al.Accurate leukocyte detection based on deformable-DETR and multi-level feature fusion for aiding diagnosis of blood diseases[J].Computers in Biology and Medicine,2024,170:107917.

[0161] https: / / doi.org / https: / / doi.org / 10.1016 / j.compbiomed.2024.107917

[0162] [4]ZHAO X,CHEN J,LIU M,et al.Multi-scale Attention-Based Feature Pyramid Networks for Object Detection;proceedings of the Image and Graphics,Cham,F 2021 / / ,2021[C].Springer International Publishing.

[0163] [5]REN S,HE K,GIRSHICK R,et al.Faster R-CNN:Towards Real-Time Object Detection with Region Proposal Networks[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2017,39(6):1137-1149. https: / / doi.org / 10.1109 / TPAMI.2016.2577031

[0164] [6] LIN T Y, GOYAL P, GIRSHICK R, et al. Focal Loss for Dense Object Detection;proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), F 22-29 Oct. 2017, 2017[C]. https: / / doi.org / 10.1109 / ICCV.2017.324

[0165] [7] ZHU C, HE Y, SAVVIDES M. Feature Selective Anchor-Free Module for Single-Shot Object Detection;proceedings of the 2019 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), F 15-20 June 2019, 2019[C]. https: / / doi.org / 10.1109 / CVPR.2019.00093

[0166] [8] ZHANG S, CHI C, YAO Y, et al. Bridging the Gap Between Anchor-Based and Anchor-Free Detection via Adaptive Training Sample Selection;proceedings of the 2020 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), F 13-19 June 2020, 2020[C]. https: / / doi.org / 10.1109 / CVPR42600.2020.00978

[0167] [9] CARION N, MASSA F, SYNNAEVE G, et al. End-to-End Object Detection with Transformers;proceedings of the Computer Vision–ECCV 2020, Cham, F 2020 / / , 2020[C]. Springer International Publishing.

[0168]

[10] ZHANG H,LI F,LIU S,et al.DINO:DETR with Improved DeNoising AnchorBoxes for End-to-End Object Detection[J].2022,abs / 2203.03605:

[0169]

[11] YAO J, LI J.AYOLOv3-Tiny: An improved convolutional neural networkarchitecture for real-time defect detection of PAD light guide plates[J]. Computers in Industry, 2022,136:103588. https: / / doi.org / https: / / doi.org / 10.1016 / j.compind.2021.103588

[0170]

[12] LI J, YANG Y. HM-YOLOv5: A fast and accurate network for defect detection of hot-pressed light guide plates [J]. Engineering Applications of Artificial Intelligence, 2023, 117: 105529. https: / / doi.org / https: / / doi.org / 10.1016 / j.engappai.2022.105529

[0171] Finally, it should be noted that the above examples are only some specific embodiments of the present invention. Obviously, the present invention is not limited to the above embodiments, and there are many variations. All variations that can be directly derived or associated with the content disclosed by a person skilled in the art should be considered as the protection scope of the present invention.

Claims

1. A light guide plate defect detection method based on multi-scale feature fusion and knowledge distillation, characterized in that: It includes collecting the light guide plate image, and then inputting the light guide plate image into the offline trained multi-scale feature fusion model RT-DETR-LGP after preprocessing in the host computer; The multi-scale feature fusion model RT-DETR-LGP is improved by using the RT-DETR model as the baseline network, including a backbone network, an encoder and a decoder. The backbone network includes a CNN network with a 4-layer structure, and outputs image features of three different scales, namely, feature map F2, feature map F3 and feature map F4. In the encoder, feature map F4 is obtained by decoupling the single-scale feature interaction to obtain feature map S4, and feature maps F2, feature map F3 and feature map S4 are obtained by the ADPN module to obtain feature maps M4, feature map M3 and feature map M2. In the decoder, feature maps M2, M3 and M4 are sent to the detection head to obtain an image with the defect category and position of the light guide plate.

2. The method for detecting defects in light guide plates by multi-scale feature fusion and knowledge distillation according to claim 1, characterized in that: The operations of each of the hierarchical structures are: Input feature F l-1 (l = 1, 2, 3, 4) After two convolutions, it is divided into features in half along the channel dimension. and Features feature Output features through the feedforward network feature Output features through MSEIE module Then the features and Features After the connection operation and then a convolution operation, the output feature F l .

3. The method for detecting defects in light guide plates by multi-scale feature fusion and knowledge distillation according to claim 2, characterized in that: The operations of the MSEIE module are: The features At the same time, a parallel Conv 3×3 The output features are obtained after the connection operation and convolution operation. In the EIE branch, the input feature It goes through adaptive average pooling, convolution, depthwise convolution, upsampling and EIEBlock in sequence.

4. The method for detecting defects in light guide plates by multi-scale feature fusion and knowledge distillation according to claim 3, characterized in that: The operation of the EIE Block is: input feature After the average pooling operation, the feature The difference is then convolved with the feature Add.

5. The method for detecting defects of light guide plates by multi-scale feature fusion and knowledge distillation according to claim 4, characterized in that: The operations of the ADPN module are: The feature map F2, the feature map F3 and the feature map S4 are passed through the first DASI module to obtain the feature map P3. Subsequently, the feature map P3 is downsampled and upsampled respectively. The downsampled feature map and the feature map F4 are dimensionally aggregated through the first Fusion module to obtain the feature map N4. The upsampled feature map and the feature map F2 are dimensionally aggregated through the second Fusion module to obtain the feature map N2. Then, the feature map N4, the feature map N2 and the feature map P3 are used as the input of the second DASI module to obtain the feature map N3, which is directly output as the feature map M3. The feature map N4, the feature map N3 after downsampling, and the feature map P3 after downsampling are passed through the third Fusion module to obtain the feature map M4. The feature map N2, the feature map N3 after upsampling, and the feature map P3 after upsampling are passed through the fourth Fusion module to obtain the feature map M2.

6. The method for detecting defects in light guide plates by multi-scale feature fusion and knowledge distillation according to claim 5, characterized in that: The preprocessing includes grayscale enhancement, edge detection, least squares fitting, perspective transformation, and tilt correction of the light guide plate image, and then cutting it into a uniform size.

7. The method for detecting defects in light guide plates by multi-scale feature fusion and knowledge distillation according to claim 6, characterized in that: The offline training includes two stages: offline pre-training and offline distillation training.

8. The method for detecting defects in light guide plates by multi-scale feature fusion and knowledge distillation according to claim 7, characterized in that: The process of offline pre-training is as follows: The ILGPDD dataset is used for offline pre-training. The labeled images in the training set and validation set are input into the model in turn. The loss function value is calculated and the model parameters are iteratively optimized through backpropagation. The training ends after the preset number of epochs. Then the test set is input into the model, and the model performance is verified by evaluation indicators to obtain the offline pre-trained multi-scale feature fusion model RT-DETR-LGP.

9. The method for detecting defects in light guide plates by multi-scale feature fusion and knowledge distillation according to claim 8, characterized in that: The process of offline distillation training is as follows: The offline pre-trained scale feature fusion model RT-DETR-LGP is used as both the teacher model and the student model, and the CWKD-based distillation method is used for offline distillation training. The training set and the validation set of the ILGPDD dataset are sequentially input into the model, and the difference between the channel distributions of the teacher model and the student model is evaluated by calculating the KL divergence, and added to the total loss function as the distillation loss, and back propagation and parameter update are performed to obtain a multi-scale feature fusion model RT-DETR-LGP with better performance; The test set is then input into the multi-scale feature fusion model RT-DETR-LGP with better performance, and the model performance is verified by evaluation indicators.

10. The method for detecting defects of light guide plates by multi-scale feature fusion and knowledge distillation according to claim 9, characterized in that: The construction process of the ILGPDD dataset is as follows: Collect the production pictures of light guide plates on the production line, pre-process the collected pictures in the computer, and divide the pictures with light guide plate defects into training set, verification set and test set after manually marking the light guide plate defect type and target frame; The defect types of the light guide plate include scratches, missing marks, bruises, white stains, black spots, scars, and white shadows.

Citation Information

Cited By

  • Method for reducing gradient deviation in quantitative training process of multi-task network

    CN121303210A

  • Automobile part surface flaw and defect detection method and system based on feature heterogeneity measurement

    CN122550607A