A Deep Learning-Based Screen Mura Detection Method and Device

Through the deep learning screen Mura detection method, the EIOU loss function and attention structure weighted features are used, combined with multi-layer detection heads, the detection speed and accuracy of Mura defects in the display screen is improved, and the problems of poor consistency and high error detection rate in traditional methods are solved.

CN116958046BActive Publication Date: 2025-07-25NEW VISION MICROELECTRONICS INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310679958.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-08
Publication Date
2025-07-25
Estimated Expiration
2043-06-08

AI Technical Summary

Technical Problem

Traditional human eye detection methods are difficult to consistently detect Mura defects on display screens, and the detection speed is slow and the error detection rate is high.

Method used

The screen Mura detection method based on deep learning is adopted, and image processing is performed through the pre-trained screen Mura detection model, using the EIOU loss function and attention structure weighted features, combined with the multi-layer detection head to improve detection performance.

Benefits of technology

It improves the speed and accuracy of Mura defect detection, and solves the problems of slow detection speed and high error detection rate in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116958046B_ABST
    Figure CN116958046B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for screen Mura detection based on deep learning. The method includes: obtaining a display panel image; inputting the display panel image into the backbone network of a pre-trained screen Mura detection model to obtain feature maps of multiple different scales; the pre-trained screen Mura detection model is obtained by training an initial screen Mura detection model using an object loss function including an EIOU loss function and a training set; inputting the feature maps of multiple different scales into the attention network of the pre-trained screen Mura detection model to obtain attention-weighted feature maps; inputting the attention-weighted feature maps into the neck network of the pre-trained screen Mura detection model to obtain fused feature maps of multiple different scales; and inputting the fused feature maps of multiple different scales into the multi-layer detection heads of the pre-trained screen Mura detection model to obtain the types and location information of Mura defects included in the display panel image. The present invention can improve the speed and accuracy of detecting Mura defects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of industrial defect detection, and particularly relates to a method and device for detecting screen Mura based on deep learning. Background Art

[0002] With the development and progress of science and technology, display screens have been widely present in people's lives. However, due to process problems during industrial production, various defects exist in the manufacturing process of display screens, such as point defects, line defects, Mura, etc. (The term "Mura" comes from a Japanese word, originally meaning spot, and now specifically refers to the defects on the display screen. When the display screen shows a constant brightness, the unevenness of the display area), which affect the luminous uniformity of the display screen, the yield rate of the display screen, the screen life, etc.

[0003] Mura defects are a kind of defective defects that seriously affect the picture quality of the display screen, mainly manifested as uneven brightness display within the screen area. There are many types and various shapes of Mura defects, and they have low contrast, non-fixed positions, and unclear edge contours. Compared with other optical defects of the display screen, Mura defects are more difficult to detect. The traditional human eye detection method mainly relies on the experience and subjective feelings of the detector to detect Mura defects and evaluate their grades, which results in inconsistent determination results for the same Mura defect by different detectors. Summary of the Invention

[0004] In order to solve the above problems existing in the related technologies, the present invention provides a method and device for detecting screen Mura based on deep learning. The technical problems to be solved by the present invention are realized through the following technical solutions:

[0005] The present invention provides a method for detecting screen Mura based on deep learning, including:

[0006] Obtaining a display screen panel image to be detected;

[0007] Inputting the display screen panel image into the backbone network of a pre-trained screen Mura detection model to obtain feature maps of multiple different scales; the pre-trained screen Mura detection model is obtained by training an initial screen Mura detection model using an object loss function containing an EIOU loss function and a training set;

[0008] Inputting the feature maps of multiple different scales into the attention network of the pre-trained screen Mura detection model to obtain attention-weighted feature maps;

[0009] Inputting the attention-weighted feature maps into the neck network of the pre-trained screen Mura detection model to obtain fusion feature maps of multiple different scales;

[0010] Input the fusion feature maps of the multiple different scales into the multi-layer detection heads of the pre-trained screen Mura detection model to obtain the types and location information of Mura defects included in the display panel image.

[0011] In some embodiments, the fusion feature maps of the multiple different scales include: a first-scale fusion feature map, a second-scale fusion feature map, a third-scale fusion feature map, and a fourth-scale fusion feature map; the attention-weighted feature maps include: a first weighted feature map, a second weighted feature map, a third weighted feature map, and a fourth weighted feature map; the step of inputting the attention-weighted feature maps into the neck network of the pre-trained screen Mura detection model to obtain fusion feature maps of multiple different scales includes:

[0012] Perform residual, convolution, upsampling, and fusion processing on the first weighted feature map, the second weighted feature map, the third weighted feature map, and the fourth weighted feature map to obtain a first convolutional feature map, a second convolutional feature map, a third convolutional feature map, and a first fusion feature map;

[0013] Perform residual processing on the first fusion feature map to obtain the first-scale fusion feature map;

[0014] Determine the second-scale fusion feature map according to the first-scale fusion feature map and the third convolutional feature map;

[0015] Determine the third-scale fusion feature map according to the second-scale fusion feature map and the second convolutional feature map;

[0016] Determine the fourth-scale fusion feature map according to the third-scale fusion feature map and the first convolutional feature map.

[0017] In some embodiments, the step of performing residual processing on the first fusion feature map to obtain the first-scale fusion feature map includes:

[0018] After performing convolution processing on the first fusion feature map, obtain a first sub-convolutional feature map;

[0019] After performing convolution processing on the first sub-convolutional feature map, obtain a second sub-convolutional feature map;

[0020] Perform splicing processing on the second sub-convolutional feature map and the first fusion feature map to obtain the first-scale fusion feature map.

[0021] In some embodiments, the step of determining the second-scale fusion feature map according to the first-scale fusion feature map and the third convolutional feature map includes:

[0022] After subjecting the first-scale fusion feature map to convolution processing, it is fused with the third convolution feature map to obtain a second fusion feature map;

[0023] Subject the second fusion feature map to the residual processing to obtain the second-scale fusion feature map.

[0024] In some embodiments, the process of performing residual, convolution, upsampling, and fusion processing on the first weighted feature map, the second weighted feature map, the third weighted feature map, and the fourth weighted feature map to obtain the first convolution feature map, the second convolution feature map, the third convolution feature map, and the first fusion feature map includes:

[0025] Perform residual and convolution processing on the first weighted feature map in sequence to obtain the first convolution feature map;

[0026] After performing upsampling processing on the first convolution feature map, it is fused with the second weighted feature map to obtain a first sub-fusion feature map;

[0027] Perform residual and convolution processing on the first sub-fusion feature map in sequence to obtain the second convolution feature map;

[0028] After performing upsampling processing on the second convolution feature map, it is fused with the third weighted feature map to obtain a second sub-fusion feature map;

[0029] Perform residual and convolution processing on the second sub-fusion feature map in sequence to obtain the third convolution feature map;

[0030] After performing upsampling processing on the third convolution feature map, it is fused with the fourth weighted feature map to obtain the first fusion feature map.

[0031] In some embodiments, the multi-layer detection head includes: multiple groups of convolutional layers corresponding one-to-one to the fusion features of the multiple different scales Figure 1 The process of inputting the fusion feature maps of the multiple different scales into the multi-layer detection head of the pre-trained screen Mura detection model to obtain the types and location information of Mura defects included in the display panel image includes:

[0032] Input the fusion feature maps of the multiple different scales into the corresponding multiple groups of convolutional layers to correspondingly obtain multiple groups of detection results;

[0033] Perform non-maximum suppression processing on the multiple groups of detection results to obtain a group of final detection results; the group of final detection results includes: the types and location information of Mura defects included in the display panel image.

[0034] In some embodiments, the target loss function includes the sum of the weights of a classification loss function, the EIOU loss function, and a confidence loss function; the classification loss function corresponds to a first weight coefficient, the EIOU loss function corresponds to a second weight coefficient, and the confidence loss function corresponds to a third weight coefficient.

[0035] In some embodiments, the training set includes multiple sample images containing various Mura defects. The multiple sample images belong to different grayscale images, and each original sample image is provided with a position label and a category label. Among them, at least one sample image with a Mura defect located at the edge of the screen is obtained by mirror-flipping the original sample image with a Mura defect located at the edge of the screen and then splicing the mirror-flipped image with the original sample image.

[0036] In some embodiments, before inputting the display screen panel image into the backbone network of the pre-trained screen Mura detection model to obtain feature maps of multiple different scales, the method further includes:

[0037] Obtain multiple original sample images; the multiple original sample images contain various Mura defects, and the multiple original sample images belong to different grayscale images, and each original sample image is provided with a position label and a category label;

[0038] Perform data augmentation processing on the multiple original sample images to obtain the training set;

[0039] Use the training set, the target loss function, and a preset learning rate to train the initial screen Mura detection model until the loss value calculated using the target loss function meets a preset condition, and then stop training to obtain the pre-trained screen Mura detection model.

[0040] The present invention also provides a defect detection device, which is deployed with a pre-trained screen Mura detection model. The defect detection device is located on a production line and is used to obtain in real time an image of a display panel to be detected. The image of the display panel is input into the backbone network of the pre-trained screen Mura detection model to obtain feature maps of multiple different scales. The feature maps of multiple different scales are input into the attention network of the pre-trained screen Mura detection model to obtain attention-weighted feature maps. The attention-weighted feature maps are input into the neck network of the pre-trained screen Mura detection model to obtain fused feature maps of multiple different scales. The fused feature maps of multiple different scales are input into the multi-layer detection heads of the pre-trained screen Mura detection model to obtain the types and location information of Mura defects included in the image of the display panel. The pre-trained screen Mura detection model is obtained by training an initial screen Mura detection model using an objective loss function including an EIOU loss function and a training set.

[0041] The present invention has the following beneficial technical effects:

[0042] Through the proposed deep learning-based display screen Mura defect detection method, the present invention regards the Mura defect detection task as a regression problem based on the global image. By adding an attention structure to weight the backbone features, it can better extract the target information of Mura defects. By setting the multi-layer detection heads to increase the low-scale feature information, it can improve the detection performance for tiny Mura defects. Moreover, by using the objective loss including the EIOU loss function to train the detection model, it can add width and height constraints to the prediction boxes, thereby improving the sensitivity of the detection model to small targets. Finally, it improves the speed and accuracy of detecting Mura defects and solves the problems of slow speed and high false detection rate in detecting Mura defects by traditional methods.

[0043] The following will further elaborate on the present invention in conjunction with the drawings and embodiments. Description of the Drawings

[0044] Figure 1 It is a flowchart of a deep learning-based screen Mura detection method provided by an embodiment of the present invention;

[0045] Figure 2 It is a schematic structural diagram of an exemplary attention network provided by an embodiment of the present invention;

[0046] Figure 3 It is a schematic structural diagram of an exemplary neck network provided by an embodiment of the present invention. Detailed Embodiments

[0047] The present invention will be further described in detail below in conjunction with specific embodiments, but the implementation manners of the present invention are not limited thereto.

[0048] In the description of the present invention, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality" means two or more, unless otherwise specifically defined.

[0049] In the description of this specification, the description with reference to terms such as "an embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.

[0050] Although the present invention has been described herein in connection with various embodiments, however, in the process of implementing the claimed invention, those skilled in the art can understand and achieve other variations of the disclosed embodiments by viewing the accompanying drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "one" does not exclude a plurality. A single processor or other unit can implement several functions recited in the claims. Certain measures are recited in mutually different dependent claims, but this does not mean that these measures cannot be combined to produce good results.

[0051] Figure 1 is a flowchart of a screen Mura detection method based on deep learning provided by an embodiment of the present invention. As Figure 1 shown, the method includes the following steps:

[0052] S101. Obtain a display screen panel image to be detected.

[0053] Here, the display screen panel image to be detected can be one or more, and there is no limitation thereto.

[0054] S102. Input the display panel image into the backbone network of the pre-trained screen Mura detection model to obtain multiple feature maps of different scales. The pre-trained screen Mura detection model is obtained by training the initial screen Mura detection model using an object loss function containing the EIOU loss function and a training set.

[0055] Here, the backbone network can be the CSPDarknet53 structure, which sequentially includes 11 network layers. Among them, the first layer is a 3×3 convolutional layer (Conv) with a stride of 1; the second, fourth, sixth, eighth, and tenth layers are all 3×3 convolutional layers with a stride of 2; the third layer is a CSP structure, which is sequentially composed of a convolutional network block, a concatenation layer, and a 1×1 convolutional layer with a stride of 1. The convolutional network block consists of two parallel branches. One branch is a 1×1 convolutional layer with a stride of 1, and the other branch sequentially has a 1×1 convolutional layer with a stride of 1 and a residual block, and the residual block includes a 1×1 convolutional layer with a stride of 1 and a 3×3 convolutional layer with a stride of 1; the fifth layer is a CSP structure, which is sequentially composed of a convolutional network block, a concatenation layer, and a 1×1 convolutional layer with a stride of 1. The convolutional network block consists of two parallel branches. One branch is a 1×1 convolutional layer with a stride of 1, and the other branch sequentially has a 1×1 convolutional layer with a stride of 1 and two residual blocks (each residual block includes a 1×1 convolutional layer with a stride of 1, a 3×3 convolutional layer with a stride of 1, and a residual layer); the seventh and ninth layers are both CSP structures, which are sequentially composed of a convolutional network block, a concatenation layer, and a 1×1 convolutional layer with a stride of 1. The convolutional network block consists of two parallel branches. One branch is a 1×1 convolutional layer with a stride of 1, and the other branch sequentially has a 1×1 convolutional layer with a stride of 1 and eight residual blocks (each residual block includes a 1×1 convolutional layer with a stride of 1 and a 3×3 convolutional layer with a stride of 1); the eleventh layer is a CSP structure, which is sequentially composed of a convolutional network block, a concatenation layer, and a 1×1 convolutional layer with a stride of 1. The convolutional network block consists of two parallel branches. One branch is a 1×1 convolutional layer with a stride of 1, and the other branch sequentially has a 1×1 convolutional layer with a stride of 1 and four residual blocks (each residual block includes a 1×1 convolutional layer with a stride of 1 and a 3×3 convolutional layer with a stride of 1). Exemplarily, the 11 network layers sequentially included in CSPDarknet53, as well as the sizes of the feature maps output by each layer and the number of channels of the feature maps output by each network layer, are specifically shown in Table 1:

[0056]

[0057]

[0058] Table 1

[0059] It should be noted that the input and output sizes of each network layer shown in Table 1 are only exemplary, and the actual input and output sizes are determined according to actual needs.

[0060] Here, the obtained feature maps of multiple different scales are respectively the first scale feature map output by the 5th layer of the backbone network, the second scale feature map output by the 7th layer, the third scale feature map output by the 9th layer, and the fourth scale feature map output by the 9th layer.

[0061] S103. Input multiple feature maps of different scales into the attention network of the pre-trained screen Mura detection model to obtain attention-weighted feature maps.

[0062] For example, the structure of the attention network is as follows Figure 2 As shown. Figure 2 As shown, the attention network includes: channel attention model, two point multiplication operation layers ( Figure 2 Adopted in Representation) and spatial attention model; channel attention model includes: maximum pooling layer, average pooling layer, two multilayer perceptron (MLP) layers, a summation layer ( Figure 2 Adopted in Representation) and a sigmoid activation layer; the spatial attention model includes: a concatenation layer, a 1×1 convolution layer and a sigmoid activation layer.

[0063] Here, each of the feature maps of multiple different scales will pass through an attention network for attention weighting. For example, taking the feature map of the first scale as an example, when the feature map of the first scale enters the attention network, it is first processed by the channel attention model. Specifically, the feature information after the maximum pooling in the channel dimension and the average pooling in the channel dimension of the feature map of the first scale respectively passes through the MLP and then is summed. After the summation result passes through the sigmoid activation, a weight vector in the channel dimension is obtained; then, the feature information after the maximum pooling in the spatial dimension and the average pooling in the spatial dimension of the feature map of the first scale is processed by the spatial attention model. Specifically, the feature information after the maximum pooling and the average pooling of the feature map of the first scale is concatenated and then input into a 1×1 convolutional layer. After the result of the convolutional processing passes through the sigmoid activation, a weight vector in the spatial dimension is obtained; finally, both the weight vector in the channel dimension and the weight vector in the spatial dimension are multiplied element-wise to the feature map of the first scale. In this way, the attention weighting of the feature map of the first scale is completed, and the first weighted feature map corresponding to the feature map of the first scale is obtained; for the feature maps of the second scale, the third scale, and the fourth scale, the weighted feature maps of each are obtained using the same principle, that is, the second weighted feature map corresponding to the feature map of the second scale, the third weighted feature map corresponding to the feature map of the third scale, and the fourth weighted feature map corresponding to the feature map of the fourth scale are obtained.

[0064] Here, through attention weighting, the weight of important feature information can be increased, which is beneficial for the model to accurately output the position and classification information of the target.

[0065] S104. Input the attention-weighted feature map into the neck network of the pre-trained screen Mura detection model to obtain fused feature maps of multiple different scales.

[0066] Here, as Figure 3 shown, the neck network includes: three residual convolutional blocks C1, three upsampling networks S, six fusion layers R, three convolutional layers J, and four residual blocks C2; each residual convolutional block includes a residual block C2 and a convolutional layer J; each residual block C2 includes: two convolutional layers J and a concatenation layer P.

[0067] Here, the processing process of the neck network for the feature maps weighted by attention (i.e., the first weighted feature map, the second weighted feature map, the third weighted feature map, and the fourth weighted feature map) is as follows: A residual convolution block C1 is used to perform residual and convolution processing on the first weighted feature map input at the Input_1 end in sequence to obtain the first convolutional feature map. Then, an upsampling network S is used to perform upsampling processing on the first convolutional feature map, and then a fusion layer R is used to fuse the upsampled feature map with the second weighted feature map input at the Input_2 end to obtain the first sub-fused feature map; Then, a residual convolution block C1 is used to perform residual and convolution processing on the first sub-fused feature map in sequence to obtain the second convolutional feature map; Then, an upsampling network S is used to perform upsampling processing on the second convolutional feature map, and then a fusion layer R is used to fuse the upsampled feature map with the third weighted feature map input at the Input_3 end to obtain the second sub-fused feature map; Then, a residual convolution block C1 is used to perform residual and convolution processing on the second sub-fused feature map in sequence to obtain the third convolutional feature map; Then, an upsampling network S is used to perform upsampling processing on the third convolutional feature map, and then a fusion layer R is used to fuse the upsampled feature map with the fourth weighted feature map input at the Input_4 end to obtain the first fused feature map; Then, the first fused feature map is input into a residual block C2. After the first fused feature map passes through the convolution processing of the first convolutional layer J in the residual block C2, the first sub-convolutional feature map is obtained; Then, after the first sub-convolutional feature map passes through the convolution processing of the second convolutional layer J in the residual block C2, the second sub-convolutional feature map is obtained; Then, after the second sub-convolutional feature map and the first fused feature map pass through the concatenation processing of the concatenation layer P in the residual block C2, the first-scale fused feature map is obtained, and the first-scale fused feature map is output by Output_1. After obtaining the first-scale fused feature map, a convolutional layer J is used to perform convolution processing on the first-scale fused feature map. Then, a fusion layer R is used to fuse the convolution processing result with the above-mentioned third convolutional feature map to obtain the second fused feature map; Then, a residual block C2 is used to perform residual processing on the second fused feature map to obtain the second-scale fused feature map, and the second-scale fused feature map is output by Output_2. After obtaining the second-scale fused feature map, a convolutional layer J is used to perform convolution processing on the second-scale fused feature map. Then, a fusion layer R is used to fuse the convolution processing result with the above-mentioned second convolutional feature map to obtain the third fused feature map; Then, a residual block C2 is used to perform residual processing on the third fused feature map to obtain the third-scale fused feature map, and the third-scale fused feature map is output by Output_3.After obtaining the third-scale fused feature map, a convolutional layer J is used to perform convolutional processing on the third-scale fused feature map. Then, a fusion layer R is used to fuse the convolutional processing result with the above-mentioned first convolutional feature map to obtain a fourth fused feature map. Then, a residual block C2 is used to perform residual processing on the fourth fused feature map to obtain a fourth-scale fused feature map, which is output by Output_4.

[0068] Here, through the processing of the neck network, feature fusion can be performed on feature information of four different scales, so that the output feature information contains both the position information of the shallow features and the semantic information of the deep features, which is beneficial to the model accurately outputting the position and classification information of the target. Moreover, the Res Block (residual block) used also prevents the network from experiencing gradient explosion due to the relatively deep network structure.

[0069] S105: Input the fused feature maps of multiple different scales into the multi-layer detection heads of a pre-trained screen Mura detection model to obtain the types and position information of Mura defects contained in the display panel image.

[0070] Here, the multi-layer detection heads include: multiple groups of convolutional layers corresponding one by one to the fused features of multiple different scales Figure 1 The fused feature maps of multiple different scales can be input into multiple groups of corresponding convolutional layers to correspondingly obtain multiple groups of detection results. Non-maximum suppression processing is performed on the multiple groups of detection results to obtain a group of final detection results. A group of final detection results includes: the types and position information of Mura defects contained in the display panel image. Exemplarily, the multi-layer detection heads can include 4 groups of convolutional layers, and each group of convolutional layers can include at least one convolutional layer.

[0071] Here, setting the multi-layer detection heads can fully obtain the shallow information of the input image, which is beneficial to improving the detection performance of the model for tiny Mura defects.

[0072] In some embodiments, before the above S102, the method further includes:

[0073] S201: Obtain multiple original sample images; the multiple original sample images contain various Mura defects, and the multiple original sample images belong to different grayscale images; each original sample image is provided with a position label and a category label.

[0074] Exemplarily, multiple original sample images may include images with 9 different Mura defects such as speckle Mura, edge Mura, banded Mura, and block Mura. Moreover, these original sample images belong to 10 different grayscale images such as 256-level, 240-level, 224-level, and 192-level. Each original sample image is marked with a position label (a label box for identifying the Mura defect) and a class label of the Mura defect. Training the model with these images can improve the robustness of the model and enable the network model to detect multiple Mura defects under different grayscale levels.

[0075] S202. Perform data augmentation on multiple original sample images to obtain a training set.

[0076] To avoid overfitting problems and enhance the generalization ability of the model, data augmentation can be performed on the original sample images, and the images after feature enhancement are used as the training set for subsequent model training. Specifically, each original sample image can be rotated, moved horizontally and vertically respectively to increase the diversity of the samples. Additionally, for at least one original sample image with Mura defects located at the screen edge, each original sample image can be mirror-flipped, and the image obtained after mirror-flipping is spliced with the original sample image to obtain an image after feature enhancement. In this way, the edge data features are increased, enabling the model to better distinguish Mura defects located at the screen edge when training the model with the image after feature enhancement.

[0077] S203. Use the training set, the target loss function, and the preset learning rate to train the initial screen Mura detection model until the loss value calculated using the target loss function meets the preset conditions and then stop training to obtain a pre-trained screen Mura detection model.

[0078] Here, the target loss function includes: the sum of the weights of the classification loss function L cls , the EIOU loss function L box and the confidence loss function L obj ; the classification loss function L cls corresponds to the first weight coefficient α, the EIOU loss function L box corresponds to the second weight coefficient β, and the confidence loss function L obj corresponds to the third weight coefficient λ.

[0079] Exemplarily, the formula of the target loss function is as follows: L total =αL cls +βL box +λL obj . L cls Adopt binary cross-entropy loss, L clsUsed to represent the deviation between the predicted Mura defect classification and the corresponding label classification, N is the total number of sample images input in each batch during training of the model, a i is the true class label of the i-th sample image among the N sample images (a i has a value of 0 or 1), x 1 i is the confidence corresponding to the class to which the i-th sample image predicted by the screen Mura detection model belongs (x 1 i has a value that is a decimal between 0 and 1). L box is used to identify the error between the predicted bounding box and the label bounding box, A i is the label bounding box of the i-th sample image among the N sample images, B i is the location of the Mura defect in the i-th sample image predicted by the screen Mura detection model (i.e., the predicted bounding box), b i p is the coordinate of the center point of the label bounding box of the i-th sample image, b i gt the coordinate of the center point of the predicted bounding box of the i-th sample image, d(b i p , b i gt ) is the Euclidean distance between b i p and b i gt , w i p is the width of the label bounding box of the i-th sample image, w i gt is the width of the predicted bounding box of the i-th sample image, d(w i p , w i gt ) is the difference between w i p and w i gt , h i p is the height of the label bounding box of the i-th sample image, h i gt is the height of the predicted bounding box of the i-th sample image, d(h i p , h i gt ) is the difference between h i p and h i gt , wi c is the width of the minimum bounding rectangle of the label box and the predicted box of the i-th sample image, h i c is the height of the minimum bounding rectangle of the label box and the predicted box of the i-th sample image. Compared with the traditional IOU loss, L box introduces the length and width information of the label box and the predicted box, solves the problem that the loss function remains unchanged due to the proportional change of the aspect ratio, which affects the detection performance of small targets, and improves the detection performance of the model for small targets. L obj =(1 - L gr ) + L gr ×L cls , L obj is used to indicate whether the target to be detected is included in the predicted box. L gr adopts binary cross-entropy loss and is used to indicate whether there is a detection target in the predicted box. x 2 i is the category to which the i-th sample image predicted by the screen Mura detection model belongs. The value of x 2 i is 0 or 1. Among them, 0 means that the predicted box corresponding to the i-th sample image predicted by the screen Mura detection model does not contain any Mura defects, and 1 means that the predicted box corresponding to the i-th sample image predicted by the screen Mura detection model contains a certain type of Mura defect.

[0080] Exemplarily, α = 0.5, β = 0.05, λ = 1.0.

[0081] Here, the set α, β, and λ hyperparameters, learning rate, and optimizer can be used, applied to the initial screen Mura detection model, and the initial screen Mura detection model is iteratively trained using the training set. After training, when the loss value calculated using the target loss function in a certain iteration is less than the preset value, or when the loss values calculated using the target loss function in several consecutive iterations no longer decrease, it means that the model converges. At this time, a pre-trained screen Mura detection model can be obtained. Here, the training of the initial screen Mura detection model can adopt existing training methods.

[0082] The present invention also provides a defect detection device. The defect detection device is deployed with a pre-trained screen Mura detection model. Moreover, the defect detection device is located on a production line and is used to obtain in real time an image of a display panel to be detected. The image of the display panel is input into the backbone network of the pre-trained screen Mura detection model to obtain feature maps of multiple different scales. The feature maps of multiple different scales are input into the attention network of the pre-trained screen Mura detection model to obtain an attention-weighted feature map. The attention-weighted feature map is input into the neck network of the pre-trained screen Mura detection model to obtain fused feature maps of multiple different scales. The fused feature maps of multiple different scales are input into the multi-layer detection heads of the pre-trained screen Mura detection model to obtain the types and location information of Mura defects included in the image of the display panel. The pre-trained screen Mura detection model is obtained by training an initial screen Mura detection model using an object loss function including an EIOU loss function and a training set. For example, the defect detection device can be a camera on a production line for detecting screen Mura defects. The camera captures an image of the display screen and inputs it into the pre-trained screen Mura detection model. The model outputs the location and classification results of Mura defects, thereby realizing the detection of defects.

[0083] Through the proposed deep learning-based display screen Mura defect detection method, the present invention regards the Mura defect detection task as a regression problem based on the global image. By adding an attention structure to weight the backbone features, it is possible to better extract Mura defect target information. By setting the multi-layer detection heads to increase the low-scale feature information, the detection performance for tiny Mura defects can be improved. Moreover, by training the detection model using an object loss function including an EIOU loss function, width and height constraints can be added to the prediction box, thereby improving the sensitivity of the detection model to small targets. Ultimately, the speed and accuracy of detecting Mura defects are improved, and the problems of slow speed and high false detection rate in detecting Mura defects using traditional methods are solved.

[0084] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, which should all be regarded as belonging to the protection scope of the present invention.

Claims

1. A method for detecting screen Mura based on deep learning, characterized in that, Including: Obtain a display panel image to be detected; Input the display panel image into the backbone network of a pre-trained screen Mura detection model to obtain multiple feature maps of different scales; The pre-trained screen Mura detection model is obtained by training an initial screen Mura detection model using an object loss function including an EIOU loss function and a training set; Input the multiple feature maps of different scales into the attention network of the pre-trained screen Mura detection model to obtain attention-weighted feature maps; The attention-weighted feature maps include: a first weighted feature map, a second weighted feature map, a third weighted feature map, and a fourth weighted feature map; Input the attention-weighted feature maps into the neck network of the pre-trained screen Mura detection model to obtain multiple feature maps of different scales for fusion, and the multiple feature maps of different scales for fusion include: a first-scale fusion feature map, a second-scale fusion feature map, a third-scale fusion feature map, and a fourth-scale fusion feature map. This step specifically includes: Perform residual, convolution, upsampling, and fusion processing on the first weighted feature map, the second weighted feature map, the third weighted feature map, and the fourth weighted feature map to obtain a first convolutional feature map, a second convolutional feature map, a third convolutional feature map, and a first fusion feature map; Perform residual processing on the first fusion feature map to obtain the first-scale fusion feature map; Determine the second-scale fusion feature map based on the first-scale fusion feature map and the third convolutional feature map; Determine the third-scale fusion feature map based on the second-scale fusion feature map and the second convolutional feature map; Determine the fourth-scale fusion feature map based on the third-scale fusion feature map and the first convolutional feature map; Input the multiple feature maps of different scales for fusion into the multi-layer detection head of the pre-trained screen Mura detection model to obtain the types and location information of Mura defects included in the display panel image. The multi-layer detection head includes: multiple groups of convolutional layers corresponding one-to-one to the multiple feature maps of different scales for fusion. This step specifically includes: Input the multiple feature maps of different scales for fusion into the corresponding multiple groups of convolutional layers to correspondingly obtain multiple groups of detection results; Perform non-maximum suppression processing on the multiple groups of detection results to obtain a set of final detection results; the set of final detection results includes: the types and location information of Mura defects included in the display panel image.

2. The method for detecting screen Mura based on deep learning according to claim 1, wherein The performing residual processing on the first fusion feature map to obtain the first-scale fusion feature map includes: After performing convolution processing on the first fusion feature map, obtain a first sub-convolutional feature map; After performing convolution processing on the first sub-convolutional feature map, obtain a second sub-convolutional feature map; Perform splicing processing on the second sub-convolutional feature map and the first fusion feature map to obtain the first-scale fusion feature map.

3. The method for detecting screen Mura based on deep learning according to claim 1, characterized in that The determining the second-scale fusion feature map based on the first-scale fusion feature map and the third convolutional feature map includes: The first-scale fusion feature map is convolved and then fused with the third convolutional feature map to obtain a second fusion feature map; The second fusion feature map is subjected to the residual processing to obtain the second-scale fusion feature map.

4. The method for detecting screen Mura based on deep learning according to claim 1, wherein The performing residual, convolution, upsampling, and fusion processing on the first weighted feature map, the second weighted feature map, the third weighted feature map, and the fourth weighted feature map to obtain a first convolutional feature map, a second convolutional feature map, a third convolutional feature map, and a first fusion feature map includes: Performing residual and convolution processing on the first weighted feature map in sequence to obtain the first convolutional feature map; The first convolutional feature map is upsampled and then fused with the second weighted feature map to obtain a first sub-fusion feature map; Performing residual and convolution processing on the first sub-fusion feature map in sequence to obtain the second convolutional feature map; The second convolutional feature map is upsampled and then fused with the third weighted feature map to obtain a second sub-fusion feature map; Performing residual and convolution processing on the second sub-fusion feature map in sequence to obtain the third convolutional feature map; The third convolutional feature map is upsampled and then fused with the fourth weighted feature map to obtain the first fusion feature map.

5. The method for detecting screen Mura based on deep learning according to claim 1, wherein The target loss function includes the sum of the weights of a classification loss function, the EIOU loss function, and a confidence loss function; the classification loss function corresponds to a first weight coefficient, the EIOU loss function corresponds to a second weight coefficient, and the confidence loss function corresponds to a third weight coefficient.

6. The method for detecting screen Mura based on deep learning according to claim 1 or 5, characterized in that The training set includes multiple sample images containing various Mura defects, the multiple sample images belong to different grayscale images, and each original sample image is provided with a position label and a category label. Among them, at least one sample image with a Mura defect located at the screen edge is obtained by mirror-flipping the original sample image with a Mura defect located at the screen edge and then splicing the mirror-flipped image with the original sample image.

7. The method for detecting screen Mura based on deep learning according to claim 1, wherein Before inputting the display panel image into the backbone network of the pre-trained screen Mura detection model to obtain feature maps of multiple different scales, the method further includes: Obtaining multiple original sample images; the multiple original sample images contain various Mura defects, and the multiple original sample images belong to different grayscale images, and each original sample image is provided with a position label and a category label; Performing data augmentation processing on the multiple original sample images to obtain the training set; Using the training set, the target loss function, and a preset learning rate to train the initial screen Mura detection model until the loss value calculated using the target loss function meets a preset condition, and then stopping the training to obtain the pre-trained screen Mura detection model.

8. A defect detection device, characterized in that, A pre-trained screen Mura detection model is deployed, and the defect detection device is located on a production line and is used to execute the steps of any one of the above claims 1 to 7.

Citation Information

Patent Citations

  • Display screen defect detection method, training method, device, equipment and medium

    CN114677377A

  • Semantic segmentation-based image composite defect detection method and system

    CN114820579A