Mini-led wafer surface defect detection method and device based on improved detr model

By improving the DETR model for Mini-LED wafer surface defect detection, and utilizing wavelet convolutional residual network, spatial-channel convolutional scale feature interaction module, and content-guided dynamic shuffling fusion module, the method solves the problems of feature extraction difficulty and insufficient accuracy in Mini-LED wafer surface defect detection, and achieves efficient and accurate Mini-LED wafer surface detection.

CN120747037BActive Publication Date: 2025-11-28HUAQIAO UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511149664.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-18
Publication Date
2025-11-28
Estimated Expiration
2045-08-18

AI Technical Summary

Technical Problem

Traditional manual visual inspection methods are inefficient and highly subjective, making it difficult to meet the needs of large-scale production. Furthermore, traditional image processing algorithms have difficulty extracting features and have poor generalization ability when detecting defects on the surface of Mini-LED wafers. The original DETR model has insufficient accuracy and slow convergence speed when dealing with tiny defects.

Method used

A method for detecting surface defects on Mini-LED wafers based on an improved DETR model is constructed, including a wavelet convolutional residual network, a spatial-channel convolutional scale feature interaction module, and a content-guided dynamic shuffling fusion module. Through bidirectional feature fusion and an end-to-end RT-DETR detection head, the method achieves the detection of surface defects on Mini-LED wafers.

Benefits of technology

It offers highly efficient surface inspection for Mini-LED wafers, making it particularly suitable for surface inspection of Mini-LED wafers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747037B_ABST
    Figure CN120747037B_ABST
Patent Text Reader

Abstract

The application discloses a Mini-LED wafer surface defect detection method and device based on an improved DETR model, and relates to the field of image processing.The method comprises the following steps: obtaining a wafer image to be detected and inputting the wafer image into a trained Mini-LED wafer surface defect detection model, first passing through a wavelet convolution residual network to obtain a first feature map, a second feature map and a third feature map, inputting the third feature map into a space-channel convolution scale feature interaction module for feature enhancement after channel dimension transformation through a first convolution layer, and obtaining an enhanced feature map; performing bidirectional feature fusion through an upsampling layer, a content-guided dynamic shuffling fusion module and a downsampling layer in combination with the second feature map and the third feature map, obtaining an eighth feature map, a tenth feature map and a twelfth feature map, and splicing the eighth feature map, the tenth feature map and the twelfth feature map in a channel dimension through a first splicing layer to obtain spliced features and input the spliced features into an RT-DETR detection head to obtain the category corresponding to each defect in the wafer image.The application solves the problems of insufficient detection precision and slow convergence speed.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of image processing, in particular to a Mini-LED wafer surface defect detection method and device based on an improved DETR model. BACKGROUND

[0002] Mini-LED, as a new generation of display technology, has great potential in the high-end display field due to its high brightness, high contrast, low power consumption and other advantages. However, the size of Mini-LED chips is extremely small, and various surface defects such as scratches, contamination, cracks, and foreign matter can easily occur during wafer manufacturing. These tiny defects not only affect product yield, but also cause display screen problems such as dead pixels, uneven brightness, and other issues, which seriously hinder the industrialization process of Mini-LED.

[0003] Traditional manual visual inspection methods are inefficient and subjective, making it difficult to meet large-scale production needs. Although automatic optical inspection systems based on machine vision have improved, traditional image processing algorithms have limitations such as difficulty in feature extraction and poor generalization when facing complex and diverse defect types on Mini-LED wafer surfaces. In recent years, the development of deep learning technology has brought new breakthroughs in defect detection. DETR (Detection Transformer) is an end-to-end object detection model that uses self-attention mechanisms to model global features and performs well in object detection. However, the original DETR model has problems such as insufficient small target detection accuracy and slow convergence when dealing with tiny defects on Mini-LED wafers, and needs targeted improvements to meet the special needs of Mini-LED wafer surface defect detection. SUMMARY

[0004] The present application aims to address the above-mentioned technical problems by providing a Mini-LED wafer surface defect detection method and device based on an improved DETR model.

[0005] In a first aspect, the present application provides a Mini-LED wafer surface defect detection method based on an improved DETR model, comprising the following steps:

[0006] An improved DETR model-based Mini-LED wafer surface defect detection model is constructed and trained to obtain a trained Mini-LED wafer surface defect detection model. The Mini-LED wafer surface defect detection model includes a wavelet convolution residual network, a first convolution layer, a spatial-channel convolution scale feature interaction module, a second convolution layer, an up-sampling layer, a content-guided dynamic shuffle fusion module, a down-sampling layer, a first splicing layer, and an RT-DETR detection head.

[0007] The wafer image to be detected is acquired and input into the trained Mini-LED wafer surface defect detection model. First, a wavelet convolution residual network is used to obtain first, second and third feature maps of different scales. The third feature map is input into a spatial-channel convolution scale feature interaction module after channel dimension transformation by a first convolution layer for feature enhancement to obtain an enhanced feature map. The enhanced feature map is subjected to channel dimension transformation by a second convolution layer to obtain a fourth feature map. The fourth feature map, the first feature map and the third feature map are subjected to bidirectional feature fusion by an upsampling layer, a content-guided dynamic shuffling fusion module and a downsampling layer to obtain eighth, tenth and twelfth feature maps. The eighth, tenth and twelfth feature maps are spliced in the channel dimension by a first splicing layer to obtain spliced features. The spliced features are input into an RT-DETR detection head to obtain the category corresponding to each defect in the wafer image.

[0008] Preferably, the fourth feature map, the first feature map and the third feature map are subjected to bidirectional feature fusion by an upsampling layer, a content-guided dynamic shuffling fusion module and a downsampling layer to obtain eighth, tenth and twelfth feature maps, which specifically include:

[0009] The fourth feature map is subjected to upsampling to obtain a fifth feature map. The fifth feature map and the second feature map are input into a content-guided dynamic shuffling fusion module for feature fusion to generate a sixth feature map. The sixth feature map is subjected to upsampling to obtain a seventh feature map. The seventh feature map and the first feature map are input into the content-guided dynamic shuffling fusion module for feature fusion to generate an eighth feature map. The eighth feature map is subjected to downsampling to obtain a ninth feature map. The seventh feature map and the ninth feature map are input into the content-guided dynamic shuffling fusion module for feature fusion to generate a tenth feature map. The tenth feature map is subjected to downsampling to obtain an eleventh feature map. The eleventh feature map and the fifth feature map are input into the content-guided dynamic shuffling fusion module for feature fusion to generate a twelfth feature map.

[0010] As preferred, the wavelet convolution residual network comprises, in sequence, a scale transformation layer, a first wavelet convolution residual layer, a second wavelet convolution residual layer, a third wavelet convolution residual layer and a fourth wavelet convolution residual layer; the second wavelet convolution residual layer, the third wavelet convolution residual layer and the fourth wavelet convolution residual layer respectively output a first feature map, a second feature map and a third feature map of different scales; the first wavelet convolution residual layer, the second wavelet convolution residual layer, the third wavelet convolution residual layer and the fourth wavelet convolution residual layer all adopt a wavelet convolution residual structure, the wavelet convolution residual structure comprises a third convolution layer, a first batch normalization layer, a first ReLU activation function layer, a wavelet convolution layer, a second batch normalization layer and a second ReLU activation function layer, the first batch normalization layer and the second batch normalization layer both adopt a batch normalization operation, the first ReLU activation function layer and the second ReLU activation function layer apply an activation function, and the third convolution layer adopts a convolution operation with a convolution kernel size of 3x3; the input feature map of the wavelet convolution residual structure is first subjected to the third convolution layer, then sequentially passes through the first batch normalization layer and the first ReLU activation function layer, then enters the wavelet convolution layer, and then passes through the second batch normalization layer to obtain an intermediate feature map; the input feature map of the wavelet convolution residual structure is connected in residual with the intermediate feature map, then passes through the second ReLU activation function layer to obtain the output feature map of the wavelet convolution residual structure, as shown in the following formula:

[0011] ;

[0012] ;

[0013] wherein x represents the input feature map of the wavelet convolution residual structure, represents the intermediate feature map, and y represents the output feature map of the wavelet convolution residual structure; represents a batch normalization operation, represents an activation function, represents a function corresponding to the wavelet convolution layer with a convolution kernel size of 3x3, represents a convolution operation with a convolution kernel size of 3x3.

[0014] ​As preferred, in the space-channel convolution scale feature interaction module, the input feature map of the space-channel convolution scale feature interaction module is first subjected to three fourth convolution layers to obtain a query feature map, a key feature map and a value feature map respectively; the query feature map and the key feature map are input into a space-channel convolution layer to obtain a processed key feature map and a processed value feature map; the space-channel convolution layer comprises a first depth separable convolution layer, a third batch normalization layer, a third ReLU activation function layer, a fifth convolution layer, a first Sigmoid function layer, a first average pooling layer, a sixth convolution layer and a second Sigmoid function layer, the fourth convolution layer, the fifth convolution layer and the sixth convolution layer all adopt a convolution operation with a convolution kernel size of 1x1, the third batch normalization layer adopts a batch normalization operation, and the first Sigmoid function layer and the second Sigmoid function layer adopt a Sigmoid function; the query feature map or the key feature map is sequentially subjected to the first depth separable convolution layer, the third batch normalization layer, the third ReLU activation function layer, the fifth convolution layer and the first Sigmoid function layer to obtain a corresponding spatial weight, the spatial weight is multiplied with the query feature map or the key feature map element by element to obtain a corresponding spatial weighted feature map; the spatial weighted feature map is sequentially subjected to the first average pooling layer, the sixth convolution layer and the second Sigmoid function layer to obtain a corresponding channel weight, the channel weight is multiplied with the spatial weighted feature map element by element to obtain the processed key feature map and the processed value feature map; the processed key feature map and the processed value feature map are added, and then multiplied with the value feature map element by element after being subjected to a second depth separable convolution layer to obtain a first attention feature, and finally subjected to a third depth separable convolution layer and a dropout layer to obtain a second attention feature; the second attention feature is subjected to residual connection and layer normalization operation with the input feature map of the space-channel convolution scale feature interaction module to obtain a third attention feature; the third attention feature is subjected to a feedforward network and then subjected to residual connection and layer normalization operation with the third attention feature to obtain an output feature map of the space-channel convolution scale feature interaction module; the first depth separable convolution layer, the second depth separable convolution layer and the third depth separable convolution layer all adopt a depth separable convolution operation with a convolution kernel size of 3x3, and the first average pooling layer adopts a global average pooling operation; as shown in the following formula:

[0015] ;

[0016] ;

[0017] ;

[0018] ;

[0019] ;

[0020] ;

[0021] ;

[0022] ;

[0023] ;

[0024] wherein, represents an input feature map of the spatial-channel convolution scale feature interaction module, represents an output feature map of the spatial-channel convolution scale feature interaction module, and respectively represent a query feature map and a key feature map, and represent a processed key feature map and a processed value feature map, represents a convolution operation with a kernel size of 1x1, represents an element-wise multiplication operation, and respectively represent spatial weighted feature maps corresponding to the query feature map and the key feature map, represents a depth separable convolution operation with a kernel size of 3x3; represents a batch normalization operation, represents an activation function, represents a layer normalization operation, represents a Sigmoid activation function, represents a function corresponding to the dropout layer, represents a function corresponding to the feedforward network, represents a global average pooling operation; , and respectively represent a first attention feature, a second attention feature and a third attention feature.

[0025] As a preference, the two input feature maps of the content-guided dynamic shuffle fusion module are respectively a first input feature map with a size of and a second input feature map with a size of The second input feature map, the first input feature map is first adjusted in channel by a seventh convolutional layer to obtain an adjusted first input feature map; the adjusted first input feature map and the second input feature map are sequentially subjected to a second splicing layer, a second average pooling layer, an eighth convolutional layer and a Hardsigmoid activation function layer to obtain an attention weight; the attention weight is divided into a first attention weight and a second attention weight, the first attention weight and the second attention weight are multiplied with the adjusted first input feature map and the second input feature map respectively to obtain a first multiplication result and a second multiplication result; the first multiplication result is added with the second input feature map to obtain a first addition result, and the second multiplication result is added with the adjusted first input feature map to obtain a second addition result; the first addition result and the second addition result are subjected to a third splicing layer and a ninth convolutional layer to obtain a cross feature, the cross feature is divided into a first cross feature of one fourth channel and a second cross feature of three fourths channel; the first cross feature is subjected to a channel shuffling operation after a grouped convolutional layer to obtain a shuffled cross feature, the shuffled cross feature is subjected to a fourth splicing layer with the second cross feature to obtain a spliced cross feature; the spliced cross feature is sequentially subjected to a tenth convolutional layer and an eleventh convolutional layer, and then is subjected to residual connection with the spliced cross feature to obtain an output feature of the content-guided dynamic shuffling fusion module; wherein the seventh convolutional layer, the eighth convolutional layer, the ninth convolutional layer, the tenth convolutional layer and the eleventh convolutional layer all adopt a convolution operation with a convolution kernel size of 1x1, the second splicing layer, the third splicing layer and the fourth splicing layer all adopt a splicing operation, and the second average pooling layer adopts a global average pooling operation, as shown in the following formula:

[0026] ;

[0027] ;

[0028] ;

[0029] ;

[0030] ;

[0031] ;

[0032] ;

[0033] ;

[0034] ;

[0035] ;

[0036] wherein, and respectively represent a first input feature map and a second input feature map of the content-guided dynamic shuffle fusion module, represent an adjusted first input feature map; represent a concatenation operation, represent a split operation, represent an element-wise multiplication operation, represent a channel shuffle operation, represent a function corresponding to a group convolution layer, represent a Hard sigmoid activation function, represent attention weights, respectively represent a first attention weight and a second attention weight, and respectively represent a first addition result and a second addition result, represent cross features, respectively represent a first cross feature and a second cross feature, represent shuffled cross features, represent concatenated cross features, represent an output feature map of the content-guided dynamic shuffle fusion module.

[0037] Preferably, the up-sampling layer adopts a nearest neighbor difference value method for up-sampling, and the down-sampling layer adopts a convolution operation with a convolution kernel size of 3x3.

[0038] In a second aspect, the present application provides a Mini-LED wafer surface defect detection device based on an improved DETR model, comprising:

[0039] The model construction module is configured to construct and train a Mini-LED wafer surface defect detection model based on the improved DETR model, to obtain a trained Mini-LED wafer surface defect detection model; the Mini-LED wafer surface defect detection model comprises a wavelet convolution residual network, a first convolution layer, a space-channel convolution scale feature interaction module, a second convolution layer, an up-sampling layer, a content-guided dynamic shuffle fusion module, a down-sampling layer, a first concatenation layer, and an RT-DETR detection head.

[0040] The detection module is configured to acquire a wafer image to be detected and input into the trained Mini-LED wafer surface defect detection model, and first, second and third feature maps of different scales are obtained through a wavelet convolution residual network (WCRN); the third feature map is input into a space-channel convolution scale feature interaction module for feature enhancement after channel dimension transformation through a first convolution layer, and an enhanced feature map is obtained; the enhanced feature map is subjected to channel dimension transformation through a second convolution layer, and a fourth feature map is obtained; the fourth feature map, the first feature map and the third feature map are subjected to bidirectional feature fusion through an up-sampling layer, a content-guided dynamic shuffling fusion module and a down-sampling layer, and an eighth feature map, a tenth feature map and a twelfth feature map are obtained; the eighth feature map, the tenth feature map and the twelfth feature map are spliced in the channel dimension through a first splicing layer, and a spliced feature is obtained; and the spliced feature is input into an RT-DETR detection head, and a category corresponding to each defect in the wafer image is obtained.

[0041] In a third aspect, the present application provides an electronic device, including one or more processors; a storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any of the implementation manners of the first aspect.

[0042] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the method described in any of the implementation manners of the first aspect.

[0043] In a fifth aspect, the present application provides a computer program product, which includes a computer program, and the computer program is executed by a processor to implement the method described in any of the implementation manners of the first aspect.

[0044] Compared with the prior art, the present application has the following beneficial effects:

[0045] (1) The Mini-LED wafer surface defect detection method based on the improved DETR model effectively reduces the parameter quantity and the calculation complexity through the wavelet convolution residual network (WCRN), and at the same time, the receptive field is expanded to capture more rich context information, and the lightweight design is realized.

[0046] (2) The Mini-LED wafer surface defect detection method based on the improved DETR model provided in the application designs a space-channel convolution scale feature interaction (SFIM) module, which significantly improves the recognition ability for small targets and tiny defects by obtaining global information, and a content-guided dynamic reshuffle fusion module (CDRIM) adopts an intelligent dynamic fusion strategy to adaptively integrate different levels of features according to the content, effectively preventing feature information loss. The entire architecture adopts a bidirectional feature fusion path (top-down and bottom-up), fully utilizes shallow detail features and deep semantic features, finally outputs three enhanced feature maps of different scales, and cooperates with an end-to-end RT-DETR detection head, so that the model realizes high efficiency while maintaining high detection accuracy, and is particularly suitable for industrial defect detection and other application scenarios that have strict requirements on real-time performance and accuracy. BRIEF DESCRIPTION OF DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0048] Figure 1 The flowchart of the Mini-LED wafer surface defect detection method based on the improved DETR model of the embodiments of the present application is shown.

[0049] Figure 2 The structure diagram of the Mini-LED wafer surface defect detection model of the Mini-LED wafer surface defect detection method based on the improved DETR model of the embodiments of the present application is shown.

[0050] Figure 3 The structure diagram of the wavelet convolution residual network of the Mini-LED wafer surface defect detection method based on the improved DETR model of the embodiments of the present application is shown.

[0051] Figure 4 The structure diagram of the space-channel convolution scale feature interaction module of the Mini-LED wafer surface defect detection method based on the improved DETR model of the embodiments of the present application is shown.

[0052] Figure 5 The structure diagram of the content-guided dynamic reshuffle fusion module of the Mini-LED wafer surface defect detection method based on the improved DETR model of the embodiments of the present application is shown.

[0053] Figure 6A comparison chart of the detection effect of the Mini-LED wafer surface defect detection method based on the improved DETR model of the embodiment of the present application and the label result;

[0054] Figure 7 A schematic diagram of the Mini-LED wafer surface defect detection device based on the improved DETR model of the embodiment of the present application;

[0055] Figure 8 The hardware structure schematic diagram of the electronic device provided by the embodiment of the present application. DETAILED DESCRIPTION

[0056] In order to make the purpose, technical scheme and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0057] Figure 1 A Mini-LED wafer surface defect detection method based on an improved DETR model is shown, which includes the following steps:

[0058] S1, constructing a Mini-LED wafer surface defect detection model based on an improved DETR model and training to obtain a trained Mini-LED wafer surface defect detection model; the Mini-LED wafer surface defect detection model includes a wavelet convolution residual network, a first convolution layer, a space-channel convolution scale feature interaction module, a second convolution layer, an up-sampling layer, a content-guided dynamic shuffling fusion module, a down-sampling layer, a first splicing layer and an RT-DETR detection head.

[0059] In a specific embodiment, the up-sampling layer uses a nearest neighbor difference value method for up-sampling, and the down-sampling layer uses a convolution operation with a convolution kernel size of 3x3.

[0060] Specifically, the embodiment of the present application constructs a Mini-LED wafer surface defect detection model based on an improved DETR model, which includes a wavelet convolution residual network, a first convolution layer, a space-channel convolution scale feature interaction module, a second convolution layer, an upsampling layer, a content-guided dynamic shuffling fusion module, a downsampling layer, a first splicing layer, and an RT-DETR detection head. The wavelet convolution residual network serves as the backbone network and outputs three feature maps of different scales: a first feature map W1 of 80x80x128, a second feature map W2 of 40x40x256, and a third feature map W3 of 20x20x512. The upsampling layer, the content-guided dynamic shuffling fusion module, and the downsampling layer are repeated multiple times to realize content guidance, dynamic shuffling, and fusion. In the embodiment of the present application, the upsampling layer realizes upsampling through nearest neighbor interpolation, and the downsampling layer realizes downsampling through a convolution operation with a step size of 2 and a convolution kernel size of 3x3.

[0061] In the training process of the Mini-LED wafer surface defect detection model, Mini-LED wafer surface defect images are first collected, and the defect regions in the Mini-LED wafer surface defect images are accurately labeled. The labeling uses the YOLO format to ensure compatibility with the subsequent detection framework. The completed dataset is divided into a training set and a validation set in a ratio of 9:1 to ensure the effectiveness and generalization ability of the model training. The training set is used to train the Mini-LED wafer surface defect detection model, the size of the input image is set to 640x640x3, the initial learning rate is set to 0.0001, AdamW is used as the optimizer, and the training period is 300 rounds. After training, the performance of the model is evaluated using the validation set to verify its effectiveness in the Mini-LED wafer surface defect detection task.

[0062] S2, the wafer image to be detected is input into the trained Mini-LED wafer surface defect detection model, first passes through the wavelet convolution residual network to obtain first, second, and third feature maps of different scales; the third feature map is input into the space-channel convolution scale feature interaction module after channel dimension transformation by the first convolution layer for feature enhancement to obtain an enhanced feature map, and the enhanced feature map is subjected to channel dimension transformation by the second convolution layer to obtain a fourth feature map; the fourth feature map, the first feature map, and the third feature map are subjected to bidirectional feature fusion by the upsampling layer, the content-guided dynamic shuffling fusion module, and the downsampling layer to obtain eighth, tenth, and twelfth feature maps, the eighth, tenth, and twelfth feature maps are spliced in the channel dimension by the first splicing layer to obtain spliced features, and the spliced features are input into the RT-DETR detection head to obtain the category corresponding to each defect in the wafer image.

[0063] In specific embodiments, the fourth feature map, the first feature map and the third feature map are bidirectional feature fused through an upsampling layer, a content-guided dynamic shuffle fusion module and a downsampling layer to obtain an eighth feature map, a tenth feature map and a twelfth feature map, specifically including:

[0064] The fourth feature map passes through the upsampling layer to obtain a fifth feature map; the fifth feature map and the second feature map are input into the content-guided dynamic shuffle fusion module for feature fusion to generate a sixth feature map; the sixth feature map passes through the upsampling layer to obtain a seventh feature map; the seventh feature map and the first feature map are input into the content-guided dynamic shuffle fusion module for feature fusion to generate an eighth feature map; the eighth feature map passes through the downsampling layer to obtain a ninth feature map; the seventh feature map and the ninth feature map are input into the content-guided dynamic shuffle fusion module for feature fusion to generate a tenth feature map; the tenth feature map passes through the downsampling layer to obtain an eleventh feature map, and the eleventh feature map and the fifth feature map are input into the content-guided dynamic shuffle fusion module for feature fusion to generate a twelfth feature map.

[0065] Specifically, the trained Mini-LED wafer surface defect detection model is deployed, and in the inference stage, referring to Figure 2 , the feature map W3 is transformed to 20x20x1024 in channel dimension through the first convolution layer with a convolution kernel size of 1x1, and then input into the spatial-channel convolution scale feature interaction module for feature enhancement to obtain an enhanced feature map. Subsequently, the enhanced feature map passes through the second convolution layer with a convolution kernel size of 1x1 to generate the fourth feature map W4, and then passes through the upsampling layer to obtain the fifth feature map W5. The fifth feature map W5 and the second feature map W2 are input into the content-guided dynamic shuffle fusion module for feature fusion to generate the sixth feature map W6. The sixth feature map W6 passes through the upsampling layer to obtain the seventh feature map W7. The first feature map W1 and the seventh feature map W7 are input into the content-guided dynamic shuffle fusion module for feature fusion to generate the eighth feature map W8. The eighth feature map W8 passes through the downsampling layer to obtain the ninth feature map W9, and the seventh feature map W7 and the ninth feature map W9 are input into the content-guided dynamic shuffle fusion module for feature fusion to generate the tenth feature map W10. The tenth feature map W10 passes through the downsampling layer to obtain the eleventh feature map W11, and the fifth feature map W5 and the eleventh feature map W11 are input into the content-guided dynamic shuffle fusion module for feature fusion to generate the twelfth feature map W12. Finally, the eighth feature map W8, the tenth feature map W10 and the twelfth feature map W12 are spliced in the channel dimension through the first splicing layer, and then input into the RT-DETR detection head for target detection.

[0066] In specific embodiments, the wavelet convolution residual network comprises, connected in sequence, a scale transformation layer, a first wavelet convolution residual layer, a second wavelet convolution residual layer, a third wavelet convolution residual layer, and a fourth wavelet convolution residual layer; the second wavelet convolution residual layer, the third wavelet convolution residual layer, and the fourth wavelet convolution residual layer respectively output first feature maps, second feature maps, and third feature maps of different scales; the first wavelet convolution residual layer, the second wavelet convolution residual layer, the third wavelet convolution residual layer, and the fourth wavelet convolution residual layer all adopt a wavelet convolution residual structure, the wavelet convolution residual structure comprises a third convolution layer, a first batch normalization layer, a first ReLU activation function layer, a wavelet convolution layer, a second batch normalization layer, and a second ReLU activation function layer, the first batch normalization layer and the second batch normalization layer both adopt a batch normalization operation, the first ReLU activation function layer and the second ReLU activation function layer apply an activation function, and the third convolution layer adopts a convolution operation with a convolution kernel size of 3x3; the input feature maps of the wavelet convolution residual structure are first subjected to the third convolution layer, then sequentially pass through the first batch normalization layer and the first ReLU activation function layer, then enter the wavelet convolution layer, and then pass through the second batch normalization layer to obtain intermediate feature maps; the input feature maps of the wavelet convolution residual structure are connected in residual with the intermediate feature maps, then pass through the second ReLU activation function layer to obtain the output feature maps of the wavelet convolution residual structure, as shown in the following formula:

[0067] ;

[0068] ;

[0069] wherein x represents the input feature maps of the wavelet convolution residual structure, represents the intermediate feature maps, and y represents the output feature maps of the wavelet convolution residual structure; represents a batch normalization operation, represents an activation function, represents a function corresponding to the wavelet convolution layer with a convolution kernel size of 3x3, represents a convolution operation with a convolution kernel size of 3x3.

[0070] Specifically, reference is made to Figure 3 ​The wavelet convolutional residual network (WCRN) comprises five stages connected in sequence. Stage zero is a scale transformation layer, which is characterized by transforming the size of the input feature map of the wavelet convolutional residual network from 640x640x3 to 160x160x64, and then entering the wavelet convolutional residual layer stage: stage one uses a first wavelet convolutional residual layer to output a 160x160x64 feature map; stage two uses a wavelet convolutional residual layer to output a first 80x80x128 feature map W1, stage three uses a wavelet convolutional residual layer to output a second 40x40x256 feature map W2; stage four uses a wavelet convolutional residual layer to output a third 20x20x512 feature map W3.

[0071] Further, the first wavelet convolutional residual layer, the second wavelet convolutional residual layer, the third wavelet convolutional residual layer and the fourth wavelet convolutional residual layer are all wavelet convolutional residual structures. The input feature map x of the wavelet convolutional residual structure is first subjected to a convolution operation with a convolution kernel size of 3x3 by a third convolution layer, then is sequentially processed by a first batch normalization layer and a first ReLU activation function layer, then enters a wavelet convolution layer with a convolution kernel size of 3x3, and then is subjected to a second batch normalization layer to obtain an intermediate feature map x1, the intermediate feature map x1 is connected in residual with the input feature map x of the wavelet convolutional residual structure to obtain a residual connected feature map, and finally the residual connected feature map is subjected to a second ReLU activation function layer to obtain the output feature map of the wavelet convolutional residual structure .

[0072] In specific embodiments, in the space-channel convolution scale feature interaction module, the input feature map of the space-channel convolution scale feature interaction module is first subjected to three fourth convolution layers to obtain a query feature map, a key feature map and a value feature map respectively; the query feature map and the key feature map are input into a space-channel convolution layer to obtain a processed key feature map and a processed value feature map; the space-channel convolution layer comprises a first depth separable convolution layer, a third batch normalization layer, a third ReLU activation function layer, a fifth convolution layer, a first Sigmoid function layer, a first average pooling layer, a sixth convolution layer and a second Sigmoid function layer, the fourth convolution layer, the fifth convolution layer and the sixth convolution layer all adopt a convolution operation with a convolution kernel size of 1x1, the third batch normalization layer adopts a batch normalization operation, and the first Sigmoid function layer and the second Sigmoid function layer adopt a Sigmoid function; the query feature map or the key feature map is sequentially subjected to the first depth separable convolution layer, the third batch normalization layer, the third ReLU activation function layer, the fifth convolution layer and the first Sigmoid function layer to obtain a corresponding spatial weight, the spatial weight is multiplied with the query feature map or the key feature map element by element to obtain a corresponding spatial weighted feature map; the spatial weighted feature map is sequentially subjected to the first average pooling layer, the sixth convolution layer and the second Sigmoid function layer to obtain a corresponding channel weight, the channel weight is multiplied with the spatial weighted feature map element by element to obtain the processed key feature map and the processed value feature map; the processed key feature map and the processed value feature map are added, and then subjected to a second depth separable convolution layer to be multiplied with the value feature map element by element to obtain a first attention feature, and finally subjected to a third depth separable convolution layer and a dropout layer to obtain a second attention feature; the second attention feature is subjected to residual connection and layer normalization operation with the input feature map of the space-channel convolution scale feature interaction module to obtain a third attention feature; the third attention feature is subjected to a feedforward network and then subjected to residual connection and layer normalization operation with the third attention feature to obtain an output feature map of the space-channel convolution scale feature interaction module; the first depth separable convolution layer, the second depth separable convolution layer and the third depth separable convolution layer all adopt a depth separable convolution operation with a convolution kernel size of 3x3, and the first average pooling layer adopts a global average pooling operation; as shown in the following formula:

[0073] ;

[0074] ;

[0075] ;

[0076] ;

[0077] ;

[0078] ;

[0079] ;

[0080] ;

[0081] ;

[0082] wherein, represents an input feature map of the spatial-channel convolution scale feature interaction module, represents an output feature map of the spatial-channel convolution scale feature interaction module, and respectively represent a query feature map and a key feature map, and represent a processed key feature map and a processed value feature map, represents a convolution operation with a kernel size of 1x1, represents an element-wise multiplication operation, and respectively represent spatial weighted feature maps corresponding to the query feature map and the key feature map, represents a depth separable convolution operation with a kernel size of 3x3; represents a batch normalization operation, represents an activation function, represents a layer normalization operation, represents a Sigmoid activation function, represents a function corresponding to the dropout layer, represents a function corresponding to the feedforward network, represents a global average pooling operation; , and respectively represent a first attention feature, a second attention feature and a third attention feature.

[0083] Specifically, reference is made to Figure 4The input feature map of the spatial-channel convolution scale feature interaction module (SFIM) is decomposed into three feature maps of the same dimension, i.e., a query feature map Q, a key feature map K and a value feature map V, through three fourth convolution layers with a convolution kernel size of 1x1. The query feature map Q and the key feature map K are processed as follows: in the spatial operation, first, a first depth separable convolution layer with a convolution kernel size of 3x3 is used to capture local spatial features, and then a third batch normalization layer and a third ReLU activation function layer are performed. Then, a fifth convolution layer with a convolution kernel size of 1x1 is used to compress all channels into a single-channel spatial attention map, and a first Sigmoid activation function layer is performed to obtain a spatial weight. This spatial weight is multiplied element by element with the original input to realize the weighting of different spatial positions, and a corresponding spatial weighted feature map is obtained. Next, in the channel operation, first, a first average pooling layer is used to perform a global average pooling operation on the spatial weighted feature map. The global average pooling operation is to average all spatial positions in each channel to compress the HxW spatial information of each channel into a single value of 1x1, and obtain the global statistical information of the channel. Then, a sixth convolution layer with a convolution kernel size of 1x1 is used to learn the mutual relationship between channels, and the number of channels is kept unchanged. After a second Sigmoid activation function layer, a channel weight is obtained. The channel weight is multiplied element by element with the spatial weighted feature map to realize the importance weighting of different channels, and a processed query feature map and a processed key feature map are obtained. The processed query feature map and the processed key feature map are added, and then multiplied element by element with the value feature map V after a second depth separable convolution layer with a convolution kernel size of 3x3. Finally, a second attention feature is obtained through a third depth separable convolution layer with a convolution kernel size of 3x3 and a dropout layer. Subsequently, through residual connection, layer normalization (LayerNorm) operation and feedforward network, the output feature map of the spatial-channel convolution scale feature interaction module is finally obtained.

[0084] In specific embodiments, the two input feature maps of the content-guided dynamic shuffle fusion module are a first input feature map with a size of and a second input feature map with a size of The second input feature map, the first input feature map is first subjected to a seventh convolutional layer for channel adjustment to obtain an adjusted first input feature map; the adjusted first input feature map and the second input feature map are sequentially subjected to a second splicing layer, a second average pooling layer, an eighth convolutional layer and a Hardsigmoid activation function layer to obtain an attention weight; the attention weight is divided into a first attention weight and a second attention weight, the first attention weight and the second attention weight are multiplied with the adjusted first input feature map and the second input feature map respectively to obtain a first multiplication result and a second multiplication result; the first multiplication result is added with the second input feature map to obtain a first addition result, and the second multiplication result is added with the adjusted first input feature map to obtain a second addition result; the first addition result and the second addition result are subjected to a third splicing layer and a ninth convolutional layer to obtain a cross feature, the cross feature is divided into a first cross feature of one fourth channel and a second cross feature of three fourths channel; the first cross feature is subjected to a channel shuffling operation after a grouped convolutional layer to obtain a shuffled cross feature, and the shuffled cross feature and the second cross feature are subjected to a fourth splicing layer to obtain a spliced cross feature; the spliced cross feature is sequentially subjected to a tenth convolutional layer and an eleventh convolutional layer, and then is connected with the spliced cross feature in a residual manner to obtain an output feature of the content-guided dynamic shuffling fusion module; wherein the seventh convolutional layer, the eighth convolutional layer, the ninth convolutional layer, the tenth convolutional layer and the eleventh convolutional layer all adopt a convolution operation with a convolution kernel size of 1x1, the second splicing layer, the third splicing layer and the fourth splicing layer all adopt a splicing operation, and the second average pooling layer adopts a global average pooling operation, as shown in the following formula:

[0085] ;

[0086] ;

[0087] ;

[0088] ;

[0089] ;

[0090] ;

[0091] ;

[0092] ;

[0093] ;

[0094] ;

[0095] wherein, and respectively represent a first input feature map and a second input feature map of the content-driven dynamic reshuffling integration module, represents an adjusted first input feature map; represents a concatenation operation, represents a split operation, represents an element-wise multiplication operation, represents a channel shuffle operation, represents a function corresponding to a group convolution layer, represents a Hard sigmoid activation function, represents an attention weight, respectively represent a first attention weight and a second attention weight, and respectively represent a first addition result and a second addition result, represents a cross feature, respectively represent a first cross feature and a second cross feature, represents a shuffled cross feature, represents a concatenated cross feature, represents an output feature map of the content-driven dynamic reshuffling integration module.

[0096] Specifically, referring to Figure 5 , a content-driven dynamic reshuffling integration module (CDRIM) receives two input feature maps, a first input feature map and a second input feature map . The size of the first input feature map is The first input feature map is first adjusted in channel by a seventh convolution layer with a convolution kernel size of 1×1, so that the number of channels is equal to the size of the second input feature map The size of the adjusted first input feature map is matched with the size of the second input feature map. The adjusted first input feature map and the second input feature map are spliced in the channel dimension through a second splicing layer to form a merged feature map. The merged feature map is globally averaged through a second average pooling layer, and then generates attention weights through an eighth convolutional layer with a convolution kernel size of 1x1 and a Hard sigmoid activation function layer. The attention weights are divided into two parts, i.e., a first attention weight and a second attention weight, which are multiplied by the corresponding adjusted first input feature map and second input feature map respectively to obtain a first multiplication result and a second multiplication result. The adjusted first input feature map is added to the second multiplication result, and the second input feature map is added to the first multiplication result. The two results are spliced in the channel dimension through a third splicing layer, and the spliced feature is first adjusted in the channel number through a ninth convolutional layer with a convolution kernel size of 1x1 to obtain cross features. Then the cross features are divided into two parts: a first cross feature of one-fourth of the channel and a second cross feature of three-fourths of the channel. The first cross feature of one-fourth of the channel enters a grouped convolutional layer for processing, and the remaining second cross feature of three-fourths of the channel remains unchanged. The grouped convolutional layer uses a convolution kernel with a size of 1 1, the number of groups is equal to the number of input channels, and the grouped convolutional layer divides the input channels into several groups, each group performs convolution independently to reduce the number of parameters and the amount of calculation. The part after the grouped convolutional layer performs a channel shuffle operation: reshapes the features into a new shape, performs dimension permutation, reshapes again, and obtains channel shuffled features. The purpose of the channel shuffle operation is to exchange information between features of different groups. Then the channel shuffled features and the second cross feature without channel shuffle are spliced in the channel dimension to obtain spliced cross features. Finally, the spliced cross features pass through a tenth convolutional layer and an eleventh convolutional layer with a convolution kernel size of 1x1 in series, and are connected in residual connection with the spliced cross features to obtain the output features of the content-guided dynamic shuffle fusion module.

[0097] The content-guided dynamic shuffle fusion module, the upsampling layer and the downsampling layer in the embodiments of the present application can be multiple or one but repeatedly used multiple times, Figure 2 The multiple content-guided dynamic shuffle fusion modules, the upsampling layers and the downsampling layers in the embodiments of the present application are only one kind of connection structure of the dynamic shuffle fusion modules, the upsampling layers and the downsampling layers in the embodiments of the present application.

[0098] The embodiments of the present application are described below through specific experiments.

[0099] The experimental results are shown in Table 1, and the detection effects are as follows Figure 6As shown, when only using the wavelet convolution residual network (WRCN), the model parameter quantity is significantly reduced from 19.88M to 13.92M, and the computational complexity is reduced from 57.0 GFLOPs to 42.5 GFLOPs, which is due to the inherent lower complexity of wavelet convolution, making the model more suitable for industrial deployment scenarios that highly value lightweight architecture. When the spatial-channel convolution scale feature interaction (SFIM) module is implemented alone, the mAP50 is improved by 0.77% compared with the baseline, and the AP50 of detecting "foreign matter" and "lump" defects shows a large increase, which is attributed to the effective coordination of SFIM between the spatial domain and the channel domain, significantly improving the detection accuracy of defects with large morphological changes. After introducing the content-guided dynamic reshuffle fusion (CDRIM) module, the model shows significant improvement in mAP50 performance, which is good at integrating multi-source features from different layers or branches, and reallocating multi-scale information through the content-guided mechanism, thereby greatly improving the overall detection capability.

[0100] The importance of each module is further verified through ablation experiments, and the results are shown in Table 1, which is the ablation experiment results of the Mini-DETR model, wherein mAP50 is the average precision mean, Foreign is the average precision of "foreign matter" defects, Glue-res is the average precision of "residual glue" defects, Incomp is the average precision of "incomplete" defects, Lump is the average precision of "lump" defects, Dirt is the average precision of "dirt" defects, and Scratch is the average precision of "scratch" defects.

[0101] The results show that: when the wavelet convolution residual network (WRCN) is removed, the mAP50 of the model decreases from 90.90% to 90.39%, and the performance decreases most significantly in the "foreign matter" and "incomplete" defect categories, highlighting the important role of wavelet convolution in providing an extended receptive field for Mini-DETR; removing the SFIM module causes the mAP50 to drop to 90.66%, especially in the "dirt" defect category, indicating that spatial feature interaction through the SFIM module is more effective than multi-head self-attention interaction alone; removing the CDRIM module results in the largest performance degradation, with the mAP50 dropping to 90.02%, which indirectly reflects the excellent ability of the CDRIM module to facilitate effective feature exchange between the WRCN and SFIM modules. When all three modules are integrated, the model achieves an mAP50 of 90.90%, which is 2.07% higher than the baseline, and more importantly, this enhanced performance only uses 68.76% of the baseline parameter quantity and 62.81% of the computational complexity, which provides an excellent precondition for subsequent edge device deployment in industrial environments.

[0102] Table 1

[0103]

[0104] Further with reference Figure 7 , as an implementation of the method shown in the above figures, the present application provides an embodiment of a Mini-LED wafer surface defect detection device based on an improved DETR model, which corresponds to the method embodiment shown in Figure 1 , and the device can be specifically applied to various electronic devices.

[0105] The embodiment of the present application provides a Mini-LED wafer surface defect detection device based on an improved DETR model, which comprises:

[0106] The model construction module 1 is configured to construct and train a Mini-LED wafer surface defect detection model based on an improved DETR model to obtain a trained Mini-LED wafer surface defect detection model; the Mini-LED wafer surface defect detection model comprises a wavelet convolution residual network, a first convolution layer, a space-channel convolution scale feature interaction module, a second convolution layer, an up-sampling layer, a content-guided dynamic shuffling fusion module, a down-sampling layer, a first splicing layer and an RT-DETR detection head;

[0107] The detection module 2 is configured to obtain a wafer image to be detected and input into the trained Mini-LED wafer surface defect detection model, and first, second and third feature maps of different scales are obtained through the wavelet convolution residual network; the third feature map is input into the space-channel convolution scale feature interaction module after channel dimension transformation through the first convolution layer for feature enhancement to obtain an enhanced feature map, and the enhanced feature map is input into the second convolution layer for channel dimension transformation to obtain a fourth feature map; the fourth feature map, the first feature map and the third feature map are fused through the up-sampling layer, the content-guided dynamic shuffling fusion module and the down-sampling layer to obtain eighth, tenth and twelfth feature maps, and the eighth, tenth and twelfth feature maps are spliced in the channel dimension through the first splicing layer to obtain spliced features, and the spliced features are input into the RT-DETR detection head to obtain a category corresponding to each defect in the wafer image.

[0108] Figure 8 The hardware structure schematic diagram of the electronic device provided by the embodiment of the present application is shown in FIG. 8. Figure 8 As shown in FIG. 8, the electronic device of the embodiment comprises a processor 801 and a memory 802; the memory 802 is used to store computer execution instructions; the processor 801 is used to execute the computer execution instructions stored in the memory to realize each step executed by the electronic device in the above-mentioned embodiment. For details, please refer to the related description in the foregoing method embodiment.

[0109] Optionally, the memory 802 can be independent or integrated with the processor 801.

[0110] When the memory 802 is independent, the electronic device further includes a bus 803 for connecting the memory 802 and the processor 801.

[0111] The embodiment of the application further provides a computer storage medium, and the computer storage medium stores computer execution instructions. When the processor 801 executes the computer execution instructions, the method described above is realized.

[0112] The embodiment of the application further provides a computer program product, and the computer program product includes a computer program. When the computer program is executed by the processor 801, the method described above is realized.

[0113] In the embodiments of the application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the modules is only a logical function division. In actual implementation, another division mode can be used. For example, a plurality of modules can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the modules shown or discussed can be indirect coupling or communication connection through some interfaces, devices or modules, and can be electrical, mechanical or other forms.

[0114] The modules described as separate components can or can not be physically separate, and the components shown as modules can or can not be physical units, i.e. can be located in one place, or can be distributed on a plurality of network units. Some or all of the modules can be selected according to actual needs to implement the embodiment of the application.

[0115] In addition, the functional modules in each embodiment of the application can be integrated in one processing unit, or each module can be physically present alone, or two or more modules can be integrated in one unit. The unit formed by the above modules can be realized in the form of hardware, or in the form of hardware plus software function unit.

[0116] The integrated modules realized in the form of software function modules can be stored in a computer readable storage medium. The software function modules described above are stored in a storage medium, and include a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or the processor 801 to execute part of the steps of the method of each embodiment of the application.

[0117] It should be appreciated that the processor 801 described above can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), or the like. The general-purpose processor can be a microprocessor or the processor 801 can also be any conventional processor 801, and the like. The steps of the methods disclosed in the present application can be directly embodied as the execution of the processor 801 in hardware, or be executed by a combination of hardware and software modules in the processor 801.

[0118] The memory 802 can include a high-speed RAM memory, and can also include a non-volatile storage NVM, such as at least one disk memory, and can also be a U disk, a mobile hard disk, a read-only memory, a magnetic disk or an optical disk, and the like.

[0119] The bus 803 can be an industry standard architecture (ISA), a peripheral component interconnect (PCI) bus, or an extended industry standard architecture (EISA) bus, or the like. The bus 803 can be divided into an address bus, a data bus, a control bus, and the like. For the sake of convenience, the bus 803 in the drawings of the present application does not limit to only one bus 803 or one type of bus 803.

[0120] The storage medium described above can be realized by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk, and the like. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0121] An example storage medium is coupled to the processor 801 such that the processor 801 can read information from, and can write information to, the storage medium. Of course, the storage medium can be a part of the processor 801. Consistent with the teachings provided herein, the processor 801 and the storage medium can be located in a single ASIC or a plurality of ASICs. Alternatively, the processor 801 and the storage medium can be located in different devices or components of a computing device or electronic system.

[0122] Those skilled in the art can understand that all or part of the steps of the above-mentioned method embodiments can be completed by relevant hardware of program instructions. The foregoing program can be stored in a computer readable storage medium. When the program is executed, the steps of the above-mentioned method embodiments are executed; and the foregoing storage medium includes: ROM, RAM, magnetic disk or optical disk and various media that can store program codes.

[0123] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for detecting surface defects on Mini-LED wafers based on an improved DETR model, characterized in that, Includes the following steps: A Mini-LED wafer surface defect detection model based on an improved DETR model is constructed and trained to obtain a trained Mini-LED wafer surface defect detection model. The Mini-LED wafer surface defect detection model includes a wavelet convolutional residual network, a first convolutional layer, a spatial-channel convolutional scale feature interaction module, a second convolutional layer, an upsampling layer, a content-guided dynamic shuffling fusion module, a downsampling layer, a first stitching layer, and an RT-DETR detection head. The wafer image to be detected is acquired and input into the trained Mini-LED wafer surface defect detection model. First, it passes through the wavelet convolutional residual network to obtain first, second, and third feature maps at different scales. The third feature map undergoes channel dimension transformation through the first convolutional layer and is then input into the spatial-channel convolutional scale feature interaction module for feature enhancement, resulting in an enhanced feature map. The enhanced feature map undergoes channel dimension transformation through the second convolutional layer to obtain a fourth feature map. The fourth, first, and third feature maps undergo bidirectional feature fusion through an upsampling layer, a content-guided dynamic shuffling fusion module, and a downsampling layer to obtain an eighth, tenth, and twelfth feature map. The eighth, tenth, and twelfth feature maps are then stitched together in the channel dimension through the first stitching layer to obtain stitched features. These stitched features are input into the RT-DETR detection head to obtain the category corresponding to each defect in the wafer image.

2. The method for detecting surface defects on Mini-LED wafers based on the improved DETR model according to claim 1, characterized in that, The fourth, first, and third feature maps are fused bidirectionally through an upsampling layer, a content-guided dynamic shuffling fusion module, and a downsampling layer to obtain the eighth, tenth, and twelfth feature maps, specifically including: The fourth feature map is upsampled to obtain the fifth feature map; the fifth and second feature maps are input into the content-guided dynamic shuffling fusion module for feature fusion to generate the sixth feature map; the sixth feature map is upsampled to obtain the seventh feature map; the seventh and first feature maps are input into the content-guided dynamic shuffling fusion module for feature fusion to generate the eighth feature map; the eighth feature map is downsampled to obtain the ninth feature map; the seventh and ninth feature maps are input into the content-guided dynamic shuffling fusion module for feature fusion to generate the tenth feature map; the tenth feature map is downsampled to obtain the eleventh feature map; the eleventh and fifth feature maps are input into the content-guided dynamic shuffling fusion module for feature fusion to generate the twelfth feature map.

3. The method for detecting surface defects on Mini-LED wafers based on the improved DETR model according to claim 1, characterized in that, The wavelet convolutional residual network comprises a scale transformation layer, a first wavelet convolutional residual layer, a second wavelet convolutional residual layer, a third wavelet convolutional residual layer, and a fourth wavelet convolutional residual layer connected in sequence. The second, third, and fourth wavelet convolutional residual layers output first, second, and third feature maps at different scales, respectively. All three layers employ a wavelet convolutional residual structure, which includes a third convolutional layer, a first batch normalization layer, a first ReLU activation function layer, a wavelet convolutional layer, a second batch normalization layer, and a second ReLU activation function layer. Both the first and second batch normalization layers employ batch normalization operations. The first and second ReLU activation function layers are applied with... The activation function is used in the third convolutional layer, which employs a 3×3 kernel. The input feature map of the wavelet convolutional residual structure first passes through the third convolutional layer, then sequentially through the first batch of normalization layers and the first ReLU activation function layer, before entering the wavelet convolutional layer. It then passes through the second batch of normalization layers to obtain an intermediate feature map. The input feature map of the wavelet convolutional residual structure is then residually concatenated with the intermediate feature map, and finally passed through the second ReLU activation function layer to obtain the output feature map of the wavelet convolutional residual structure, as shown in the following equation: ; ; Where x represents the input feature map of the wavelet convolution residual structure. y represents the intermediate feature map, and y represents the output feature map of the wavelet convolution residual structure. This indicates a batch normalization operation. express Activation function This represents the function corresponding to a wavelet convolutional layer with a kernel size of 3×3. This indicates a convolution operation with a kernel size of 3×3.

4. The method for detecting surface defects on Mini-LED wafers based on the improved DETR model according to claim 1, characterized in that, In the spatial-channel convolutional scale feature interaction module, the input feature map first passes through three fourth convolutional layers to obtain a query feature map, a key feature map, and a value feature map, respectively. The query feature map and the key feature map are then input into the spatial-channel convolutional layer to obtain the processed key feature map and the processed value feature map, respectively. The spatial-channel convolutional layer includes a first depthwise separable convolutional layer, a third batch normalization layer, a third ReLU activation function layer, a fifth convolutional layer, a first Sigmoid function layer, a first average pooling layer, a sixth convolutional layer, and a second Sigmoid function layer. The moid function layer, the fourth, fifth and sixth convolutional layers all use convolution operations with a kernel size of 1×1, the third batch normalization layer uses batch normalization, and the first and second Sigmoid function layers use the Sigmoid function. The query feature map or key feature map first passes through the first depthwise separable convolutional layer, the third batch normalization layer, the third ReLU activation function layer, the fifth convolutional layer and the first Sigmoid function layer in sequence to obtain the corresponding spatial weights. The spatial weights are then multiplied element-wise with the query feature map or key feature map. The corresponding spatially weighted feature map is obtained; the spatially weighted feature map is sequentially passed through the first average pooling layer, the sixth convolutional layer, and the second sigmoid function layer to obtain the corresponding channel weights. The channel weights are multiplied element-wise with the spatially weighted feature map to obtain the processed key feature map and the processed value feature map; the processed key feature map and the processed value feature map are added together, and then passed through the second depthwise separable convolutional layer and multiplied element-wise with the value feature map to obtain the first attention feature. Finally, the second attention feature is obtained by passing through the third depthwise separable convolutional layer and the abandonment layer; the second... The attention feature is subjected to residual connection and layer normalization operations with the input feature map of the spatial-channel convolutional scale feature interaction module to obtain the third attention feature; the third attention feature is then subjected to residual connection and layer normalization operations with the input feature map of the spatial-channel convolutional scale feature interaction module after passing through the feedforward network to obtain the output feature map of the spatial-channel convolutional scale feature interaction module; the first depthwise separable convolutional layer, the second depthwise separable convolutional layer, and the third depthwise separable convolutional layer all adopt depthwise separable convolution operation with a kernel size of 3×3, and the first average pooling layer adopts global average pooling operation; as shown in the following formula: ; ; ; ; ; ; ; ; ; in, This represents the input feature map of the spatial-channel convolutional scale feature interaction module. This represents the output feature map of the spatial-channel convolutional scale feature interaction module. and These represent the query feature map and the key feature map, respectively. and This represents the processed key feature map and the processed value feature map. This represents a convolution operation with a kernel size of 1×1. This indicates an element-wise multiplication operation. and These represent the spatially weighted feature maps corresponding to the query feature map and the key feature map, respectively. This indicates a depthwise separable convolution operation with a kernel size of 3×3; This indicates a batch normalization operation. express Activation function Presentation layer normalization operation, This represents the Sigmoid activation function. This represents the function corresponding to the abstention layer. This represents the function corresponding to the feedforward network. This indicates a global average pooling operation; , and These represent the first attention feature, the second attention feature, and the third attention feature, respectively.

5. The method for detecting surface defects on Mini-LED wafers based on the improved DETR model according to claim 1, characterized in that, The content-guided dynamic shuffling fusion module has two input feature maps of size [missing information]. The first input feature map and its size are The second input feature map is obtained by first passing the first input feature map through the seventh convolutional layer for channel adjustment. The adjusted first input feature map and the second input feature map are then passed sequentially through the second concatenation layer, the second average pooling layer, the eighth convolutional layer, and the Hard sigmoid activation function layer to obtain attention weights. The attention weights are divided into a first attention weight and a second attention weight. The first and second attention weights are multiplied by the adjusted first and second input feature maps, respectively, to obtain a first multiplication result and a second multiplication result. The first multiplication result is added to the second input feature map to obtain a first addition result. The second multiplication result is added to the adjusted first input feature map to obtain a second addition result. The first and second addition results are passed through a third concatenation layer and a ninth convolutional layer to obtain cross features. The cross features are divided into a first cross feature with one-quarter channels and a second cross feature with three-quarter channels. The first cross features are passed through a grouped convolutional layer and then undergo channel shuffling to obtain shuffled cross features. The shuffled cross features and the second cross features are passed through a fourth concatenation layer to obtain concatenated cross features. The concatenated cross features are sequentially passed through the tenth and eleventh convolutional layers, and then residually connected with the concatenated cross features to obtain the output features of the content-guided dynamic shuffling fusion module. Specifically, the seventh, eighth, ninth, tenth, and eleventh convolutional layers all employ 1×1 convolution operations, the second, third, and fourth concatenated layers all employ concatenation operations, and the second average pooling layer employs global average pooling, as shown in the following formula: ; ; ; ; ; ; ; ; ; ; in, and These represent the first and second input feature maps of the content-guided dynamic shuffling fusion module, respectively. This represents the adjusted first input feature map; This indicates a splicing operation. This indicates a splitting operation. This indicates an element-wise multiplication operation. This indicates a channel mixing operation. This represents the function corresponding to the grouped convolutional layer. This represents the Hard sigmoid activation function. Indicates attention weights, These represent the first attention weight and the second attention weight, respectively. and These represent the first and second addition results, respectively. Indicates cross features, These represent the first cross feature and the second cross feature, respectively. Indicates the cross-washing characteristics. Indicates splicing and intersection features. This represents the output feature map of the dynamic shuffling and fusion module guided by the content.

6. The method for detecting surface defects on Mini-LED wafers based on the improved DETR model according to claim 1, characterized in that, The upsampling layer uses the nearest neighbor difference method for upsampling, and the downsampling layer uses a convolution operation with a kernel size of 3×3.

7. A Mini-LED wafer surface defect detection device based on an improved DETR model, characterized in that, include: The model building module is configured to build and train a Mini-LED wafer surface defect detection model based on an improved DETR model, resulting in a trained Mini-LED wafer surface defect detection model. The Mini-LED wafer surface defect detection model includes a wavelet convolutional residual network, a first convolutional layer, a spatial-channel convolutional scale feature interaction module, a second convolutional layer, an upsampling layer, a content-guided dynamic shuffling fusion module, a downsampling layer, a first stitching layer, and an RT-DETR detection head. The detection module is configured to acquire the wafer image to be detected and input it into the trained Mini-LED wafer surface defect detection model. First, the image passes through the wavelet convolutional residual network to obtain first, second, and third feature maps at different scales. The third feature map undergoes channel dimension transformation through the first convolutional layer and is then input into the spatial-channel convolutional scale feature interaction module for feature enhancement, resulting in an enhanced feature map. The enhanced feature map undergoes channel dimension transformation through the second convolutional layer to obtain a fourth feature map. The fourth, first, and third feature maps undergo bidirectional feature fusion through an upsampling layer, a content-guided dynamic shuffling fusion module, and a downsampling layer to obtain an eighth, tenth, and twelfth feature map. The eighth, tenth, and twelfth feature maps are then stitched together in the channel dimension through the first stitching layer to obtain stitched features. These stitched features are input into the RT-DETR detection head to obtain the category corresponding to each defect in the wafer image.

8. An electronic device, comprising: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Display screen defect detection method and system based on cascade multilayer feature fusion network

    CN118154603A

  • Multi-scale defect detection method based on deep learning network

    CN119722580A