Liquid crystal screen defect detection method based on deformable convolution and feature fusion

By introducing DCNv2 variable convolution, GhostConv and EMA attention mechanisms, as well as IGF modules in the YOLOv8n model, the model's shortcomings in LCD screen defect detection are solved, and higher detection accuracy and speed are achieved.

CN120014413APending Publication Date: 2025-05-16WUHAN INST OF TECH

Patent Information

Application Number
CN202510157692.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

After the existing YOLOv8 model is pre-trained on natural data sets, it is difficult to effectively detect defects in the LCD screen field, and the model is difficult to adaptively capture defect characteristics on the LCD screen, especially the small defect detection effect is poor.

Method used

By replacing the normal convolution in the backbone network of the YOLOv8n model with DCNv2 variable convolution, introducing GhostConv and EMA attention mechanisms in the feature extraction module, and using the IGF module in the feature fusion part, an improved YOLOv8n model is built to improve the accuracy of LCD screen defect detection.

Benefits of technology

It realizes more accurate detection of defects of small LCD screens, improves detection accuracy and speed, average average accuracy (mAP) reaches 96.5%, and detection speed reaches 142.9 FPS, significantly improving the actual detection effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014413A_ABST
    Figure CN120014413A_ABST
Patent Text Reader

Abstract

The invention discloses a liquid crystal screen defect detection method based on deformable convolution and feature fusion, and belongs to the technical field of screen detection. The method comprises the following steps: S1, preprocessing a sample; s2, constructing an improved YOLOv8n model on the basis of the YOLOv8n model, wherein the improved YOLOv8n model has the following adjustments: (1) replacing CBS standard convolution of a third layer, a fifth layer and a seventh layer of a backbone network Backbone in the YOLOv8n model with DCNv2 variable convolution; (2) a C2f module in the YOLOv8n model is replaced with a CEG module; (3) a Concat module in the YOLOv8n model is replaced with an IGF module; s3, training and testing the improved YOLOv8n model by adopting the preprocessed sample to obtain an optimal model; and S4, acquiring a liquid crystal screen image and detecting through the optimal model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of screen detection, and in particular relates to a liquid crystal screen defect detection method based on deformable convolution and feature fusion. Background Art

[0002] Small liquid crystal display screens are important components in the electronic product manufacturing industry. In the process of automated production, due to equipment process failures and human operation, defects such as breakage, bright spots, bright lines and leakage may occur. These defects seriously affect the quality of finished products, and thus become a weak link in the quality control of downstream product performance and service life. The early detection method was mainly manual detection, which usually involves naked eye observation by workers, and has the disadvantages of low efficiency, strong subjectivity and high labor intensity. Therefore, it is an urgent problem for the industrial manufacturing industry to realize automated detection of defects in LCD production lines, improve detection accuracy and speed up detection efficiency. With the development of science and technology, deep learning has been widely used in various fields such as pattern recognition, natural language processing, and computer vision. Compared with traditional image processing technology, target detection technology based on deep learning can automatically capture effective features in a large data set, analyze the language information of the image, and greatly reduce human intervention while ensuring detection accuracy, which has broad application prospects. In recent years, deep learning combined with target detection technology has been changing with each passing day, and many advanced target detection algorithms have emerged in the field of computer vision, which have achieved remarkable results in defect detection tasks. At present, the field of target detection can be roughly divided into two categories. One is the one-stage target detection algorithm represented by the YOLO series and SSD algorithms, which have the advantages of fast detection speed and strong real-time performance. The other is the detection algorithm represented by R-CNN, which has the characteristics of good feature extraction effect and high detection accuracy. Considering the accuracy and real-time performance in industrial detection, we use the one-stage detection algorithm YOLO series. The YOLO algorithm is an excellent one-stage target detection algorithm proposed by Ultralytics. After being developed by different researchers, multiple versions have been produced. Considering the industrial applicability, we use the eighth-generation version YOLOv8 as the basic model for improvement. Although YOLOv8 performs well on large natural data sets such as COCO as an advanced target detection model, there are still the following problems for small LCD screen detection: 1. The YOLOv8 model is pre-trained on natural data sets, and the original model cannot detect defects related to the LCD field well. 2. The types of surface defects of LCD screens are diverse and the scales are different. It is difficult for the original model to adaptively capture defect features. 3. As the number of network layers increases, small defects in the LCD screen tend to disappear gradually during the feature extraction process, resulting in poor performance of the original model in detecting small defects. Summary of the invention

[0003] In order to solve the above problems, an embodiment of the present invention provides a method for detecting LCD screen defects based on deformable convolution and feature fusion. The technical solution is as follows: The embodiment of the present invention provides a method for detecting defects in a liquid crystal screen based on deformable convolution and feature fusion, the method comprising: S1: Preprocess the sample.

[0004] S2: The improved YOLOv8n model is constructed based on the YOLOv8n model. The improved YOLOv8n model has the following adjustments: (1) The CBS standard convolutions in the 3rd, 5th and 7th layers of the backbone network in the YOLOv8n model are replaced with DCNv2 variable convolutions. (2) The C2f module in the YOLOv8n model is replaced with the CEG module. (3) The Concat module in the YOLOv8n model is replaced with the IGF module.

[0005] S3: Use the preprocessed samples to train and test the improved YOLOv8n model to obtain the optimal model.

[0006] S4: Acquire the LCD screen image and detect it through the optimal model.

[0007] The technical solution provided by the embodiment of the present invention has the following beneficial effects: the mean average precision (mAP) of the method of the present invention reaches 96.5%, which is 12.9% higher than the benchmark model, and the accuracy and recall rate reach 95.4% and 93.7% respectively, which are 2.9% and 13.3% higher than the benchmark model respectively, and the detection speed of the algorithm reaches 142.9FPS. The test results show that compared with the benchmark model and other mainstream detection models, the model of the present invention has better detection accuracy and faster detection speed, and can better meet the actual detection needs of small LCD screen defect detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 is a flow chart of a method for detecting defects in a liquid crystal screen based on deformable convolution and feature fusion provided by an embodiment of the present invention; Figure 2 is a detailed flow chart of step S1; Figure 3 This is the structural diagram of the improved YOLOv8n model; Figure 4 This is a schematic diagram of the DCNv2 variable convolution processing process; Figure 5 This is a flow chart of the DCNv2 variable convolution processing process; Figure 6 It is a flow chart of the CEG module processing process; Figure 7This is a schematic diagram of the Ghostconv convolution process; Figure 8 This is a schematic diagram of the processing of the EMA attention mechanism; Fig. 9 This is a flowchart of the EMA attention mechanism processing process; Fig.10 It is a flow chart of the IGF module processing; Fig.11 It is a schematic diagram of the processing of the IGF module; Fig.12 This is a schematic diagram of the structure of the CBS standard convolution; Fig.13 It is a schematic diagram of the structure of the CEG module; Fig.14 It is a schematic diagram of the structure of the GhostBottleneck module; Fig.15 It is a schematic diagram of the structure of the SPPF module; Fig.16 It is a schematic diagram of the structure of the detection head; Fig.17 is a detailed flow chart of step S3; Fig.18 It is the appearance picture of some defective samples; Fig.19 It is an actual detection effect diagram of the method of the present invention; Fig. 20 This is a comparison chart of the detection effect of the method of this patent and different mainstream models; Fig.21 is the confusion matrix plot of the standard model; Fig. 22 It is the confusion matrix diagram of the patented method; Fig.23 It is a characteristic thermal perception map of different models in the comparative experiment. DETAILED DESCRIPTION

[0009] In order to make the objectives, technical solutions and advantages of the present invention more clear, the present invention will be further described in detail below with reference to the accompanying drawings.

[0010] Example 1 See also Figure 1 and 3 Embodiment 1 provides a method for detecting defects in a liquid crystal screen based on deformable convolution and feature fusion, the method comprising: S1: Preprocess the sample.

[0011] S2: An improved YOLOv8n model is constructed based on the YOLOv8n model. The improved YOLOv8n model has the following adjustments: (1) The CBS standard convolutions in the 3rd, 5th and 7th layers of the backbone network Backbone in the YOLOv8n model are replaced with DCNv2 variable convolutions. DCNv2 variable convolutions enable the model to better capture targets with different aspect ratios on the LCD screen in the backbone feature extraction part. The offset parameter in the deformable convolution is used to control the size of the receptive field in the model detection process, so as to avoid the model being unable to handle the features of irregularly shaped defects and enhance the model's ability to extract and express complex morphological defect features. Specifically, the idea of ​​DCNv2 variable convolution is that convolution should not be limited to a simple rectangle, but the optimal convolution kernel structure may exist at different stages, different feature maps, and even different pixels. This convolution adds an offset and a modulation factor m to the sampling position of the ordinary convolution, so that the sampling position of the convolution kernel is variable, thereby producing a larger receptive field. First, perform a regular convolution operation on the input feature map to obtain a set of offset prediction results, where the offset and the input feature map size are the same, and the number of channels of the offset feature map is 2N. This is because for a two-dimensional image, we need to calculate the offset in both the horizontal and vertical directions. 2 represents the two values ​​of the offset (X, Y), and N represents the number of pixels in the convolution kernel. For example, when the convolution kernel size is 3*3, N is 9. For a deformable convolution, variable does not mean that the physical shape of the convolution kernel is variable, but that the sampling position of the convolution kernel is variable. The flexibility of feature extraction is enhanced by dynamically adjusting the sampling position and weighting the sampling results, so that the convolution can adapt to the shape and scale changes of the target in space.

[0012] (2) The C2f module in the YOLOv8n model is replaced by the CEG module. The CEG module introduces the GhostConv module and the EMA attention mechanism. The GhostConv module reduces the amount of calculation by generating "Ghost" features, thereby significantly improving the efficiency of feature extraction. In the feature extraction process, GhostConv reduces the computational overhead by simplifying the convolution calculation and optimizing the feature generation, while maintaining a high feature expression capability. In addition, GhostConv improves the real-time performance of the model, making it more efficient in real-time defect detection tasks. The EMA attention mechanism reduces the noise and instability that may occur during the training process by weighted smoothing of the feature map, thereby improving the model's sensitivity to key features. This mechanism uses historical feature information during the optimization process, allowing the model to better capture long-term dependencies, avoiding the impact of instantaneous fluctuations during training on the final feature extraction, and maintaining the detailed expression of the feature map. Especially when dealing with complex backgrounds and high-noise environments, it can effectively reduce information loss and overfitting, and improve the accuracy and robustness of target detection. In summary, the CEG module is based on C2f and uses GhostConv to achieve feature reuse, avoiding the problem of excessive computational overhead caused by too many redundant features. At the same time, a dynamic weight selection mechanism is added to the place where the gradient flow is richest during the feature extraction process to ensure that the model training process can focus more on the interaction between different scales and enhance the overall performance of the model. Specifically, Ghostconv is used for feature extraction operations. The core idea of ​​this convolution is that the feature map generated by the original convolution calculation often has high redundancy. This is because the feature maps extracted between different levels are similar. Therefore, Ghostconv divides the traditional convolution into two steps: 1. Main convolution operation. 2. Linear transformation generates Ghost features. In the main convolution operation, a part of the core feature map is generated by a small number of conventional convolutions. This step is inevitable. Then, some simple linear operations are used to generate the Ghost feature map for the core feature map. This operation can obtain the remaining feature map through extremely low-cost simple linear operations. Finally, the core feature map and the generated Ghost feature map are fused to obtain the final output result, which effectively reduces the problem of excessive computation and parameters caused by feature map redundancy. The EMA mechanism is introduced to dynamically and smoothly update the model parameters through exponential sliding average, effectively improving the stability and robustness of feature representation. At each parameter update, EMA weights and fuses the current and historical parameters with a smoothing coefficient, which reduces the instantaneous fluctuations in training and enables the model to more robustly capture long-term feature dependencies. By integrating historical information, EMA improves the model's noise resistance and generalization performance in complex scenarios, ensuring the recognition accuracy and consistency of target detection in diverse defect scenarios.Combining the ideas of GhostConv and EMA, a dynamic weight selection mechanism is added to the place with the richest gradient flow in the entire CEG module. The input features will be spliced ​​after passing through GhostBottleneck twice, and then convolution is used to readjust the features and channels. Here, the model splices together feature maps from multiple paths. These feature maps may contain different spatial and semantic information, and reflect the details extracted by different convolution paths. The gradient information at this fusion point is extremely rich and diverse, which reflects the model's sensitivity to loss in different feature dimensions.

[0013] (3) The Concat module in the YOLOv8n model is replaced by the IGF module. The IGF module fuses features of different scales and re-weights them through stepwise convolution, depthwise separable convolution and weighted operations in different directions, effectively improving the adaptability and recognition of multi-layer features. At the same time, IGF dynamically assigns weights to features through the Sigmoid activation function, suppresses the interference of invalid features, and enhances the expressiveness of significant features. This structure can better retain important detail information and global semantics, reduce redundant features and information loss that may be introduced by direct Concat, and thus improve the detection accuracy and robustness of the model. Constructing the IGF module to replace the single Concat operation effectively captures multi-scale contextual information by focusing on the interdependence between distant pixels, thereby enhancing the central features of the image. The IGF module first average pools the input to reduce sensitivity, smooth features, and reduce local noise. Subsequently, features are extracted through a 1×1 convolution layer to maintain channel consistency. Using horizontal and vertical convolution operations, the IGF module achieves deep fusion and feature interaction between different levels, ensuring the optimal use of key features while reducing redundant information. This design significantly enhances the network's capabilities in feature extraction and information fusion, thereby improving the overall performance and accuracy of the defect detection model. Among them, the CEG module includes the first CBS standard convolution, the Splite module, the first GhostBottleneck module, the second GhostBottleneck module, the first Concat module, the EMA attention mechanism, and the second CBS standard convolution. The first CBS standard convolution, the Splite module, the first GhostBottleneck module, the second GhostBottleneck module, the first Concat module, the EMA attention mechanism, and the second CBS standard convolution are connected in sequence, and the first GhostBottleneck module residual is connected to the first Concat module. The first Concat module splices the other output of the Splite module, the residual output of the first GhostBottleneck module, and the output of the second GhostBottleneck module.Among them, the IGF module includes the second Concat module, the AvgPool module, the first conv2d1*1 module, the DWConv1*n module, the DWConvn*1 module, the second conv2d1*1 module, the sigmoid activation function, the third Concat module, the multi module and the Add module. The second Concat module, the AvgPool module, the first conv2d1*1 module, the DWConv1*n module, the DWConvn*1 module, the second conv2d1*1 module and the sigmoid activation function are connected in sequence. The multi module is used to multiply the weight feature map output by the sigmoid activation function by the two original feature maps element by element to obtain the enhanced feature map. The Add module is used to perform alternating residual connections between the enhanced feature map and the two original feature maps. The third Concat module is used to splice the two feature maps after the residual connection again in the channel dimension.

[0014] S3: Use the preprocessed samples to train and test the improved YOLOv8n model to obtain the optimal model.

[0015] S4: Get the LCD screen image and test it through the optimal model. Before getting the LCD screen image, the camera can be calibrated to ensure that the industrial camera can capture valid sample images with high resolution and low noise, so as to avoid the impact of data quality on the test results.

[0016] The present invention replaces the ordinary convolution in the Backbone part of the YOLOv8 model with a deformable convolution, thereby realizing the learning of multi-scale features and features with different aspect ratios of small LCD screens. By adding an offset offset and a modulation factor m at the sampling position, the sampling position of the convolution kernel is made variable, thereby generating a larger receptive field to adapt to irregular shape defects. The improved model can more accurately capture line defects and leakage defects in small LCD screens. The present invention enhances the utilization of features in the model in feature extraction operations by constructing a CEG module and integrating the ideas of GhostConv convolution and EMA attention, avoiding the introduction of too many redundant features. At the same time, the back-propagated gradient information can be used through a dynamic weight selection mechanism to automatically learn and select which feature channels or spatial positions are more important in a specific context, thereby making the gradient utilization of the model more efficient. Improve the detection accuracy of point defects and scratch defects. The present invention avoids the problem of poor multi-scale information fusion of the model caused by simple Concat operation by constructing an IGF module. By re-weighting the features of different scales after fusion, the adaptability and recognition of multi-layer features are effectively improved, so that deep fusion and feature interaction are achieved between different levels, ensuring the optimal use of key features while reducing redundant information. In summary, the present invention improves the accuracy and efficiency of defect detection of small LCD screens, can better capture the details and irregular defects on the LCD screen, improves the detection ability of small target defects and defects with large scale changes, and further enhances the model's utilization of features through global perception and multi-scale feature fusion methods in the feature fusion part, thereby improving the model's detection and generalization capabilities for defects.

[0017] Example 2 See also Figure 1-17 Embodiment 2 provides a method for detecting defects in a liquid crystal screen based on deformable convolution and feature fusion, the method comprising: S1: Preprocess the sample. Preprocessing includes: S101: Manually clean the collected LCD screen samples (specifically high-quality images) to remove blurred images, images with insufficient contrast, images containing impurities, etc. Try to ensure uniform defect standards.

[0018] S102: Divide the cleaned samples into a training set, a validation set, and a test set. Specifically, the samples may be divided into a training set, a validation set, and a test set in a ratio of 7:2:1.

[0019] S103: Perform data enhancement on the training set. The data enhancement methods include rotation, translation, scaling and / or cropping, etc., and the enhancement methods that change the quality of the original image are avoided as much as possible.

[0020] S104: Label the sample to obtain a TXT file in YOLO format. The TXT file contains five normalized parameters, namely, category, ratio of the center X coordinate of the box to the image width, ratio of the center Y coordinate of the box to the image height, ratio of the width of the box and the height of the box to the image; the categories include Dot defect, Linedefect, Scratch and Leakage.

[0021] S2: An improved YOLOv8n model is constructed based on the YOLOv8n model. The improved YOLOv8n model has the following adjustments: (1) The original YOLOv8 model was trained and the experimental results were observed. It was found that ordinary convolution could not capture defects with large scale changes. Therefore, the CBS standard convolutions of the 3rd, 5th and 7th layers of the backbone network Backbone in the YOLOv8n model were replaced with DCNv2 variable convolutions. This patent only replaces part of the CBS standard convolutions of the YOLOv8n model. After a large number of experiments by the applicant, it was found that in the scenario and new model of this patent, only the 3rd, 5th and 7th layers need to be adjusted, and the processing effect is the best.

[0022] See also Figure 6 , the working process of DCNv2 variable convolution in the embodiment of the present invention is: S211: Calculate the offset of the sampling position based on the additional convolutional layer and modulation factor m; where the input feature Figure X After adding the convolutional layer, we get: ;in, and are the weights and biases of the additional convolutional layer, The dimension is ; Input features Figure X The dimension is ;in, is the number of channels of the input feature map, H and W are the width and height of the feature map respectively, and the size of the convolution kernel is K*K. ;in, and They are the weights and biases of the convolution layer for calculating the modulation factor. Each sampling point of the convolution kernel has a modulation factor, and the dimension of m is .

[0023] S212: By offset Calculate the eigenvalues ​​of the sampling points by bilinear interpolation .

[0024] S213: Use the modulation factor m to adjust the contribution of each sampling point: .

[0025] S214: Obtain output feature map through convolution operation: ;in, is the weight of the convolution kernel.

[0026] (2) Replace the C2f module in the YOLOv8n model with the CEG module. The GhostConv module can improve the model's utilization of features and avoid excessive redundant features that lead to poor feature extraction. The EMA attention mechanism can improve the stability and robustness of feature representation and ensure the recognition accuracy and consistency of target detection in a variety of defect scenarios.

[0027] (3) The Concat module in the YOLOv8n model is replaced by the IGF module. The IGF module effectively captures multi-scale contextual information by focusing on the interdependence between long-distance pixels, achieves deep fusion and feature interaction, and ensures the optimal use of key features.

[0028] See also Figure 3 , the structure of the adjusted improved YOLOv8n model includes the input terminal Input, the backbone network Backbone, the neck network neck and the detection head head. The backbone network Backbone includes: the 0th and 1st layers are CBS standard convolutions. The 2nd, 4th, 6th and 8th layers are CEG modules for feature extraction. The 3rd, 5th and 7th layers are DCNv2 variable convolutions for 2x downsampling. The 9th layer is the SPPF module for maximum pooling to aggregate LCD screen defect features of various scales. The detection head head includes a small detection head, a medium detection head and a large detection head. The neck network neck includes: the 10th and 13th layers are Upsumple modules for upsampling. The 12th, 15th, 18th and 21st layers are CEG modules for feature extraction. The 16th and 19th layers are CBS standard convolutions. The 11th layer is the IGF module for fusing the feature maps of the 10th and 6th layers. The 14th layer is an IGF module, which is used to fuse the feature maps of the 13th and 4th layers. The 17th layer is an IGF module, which is used to fuse the feature maps of the 16th and 12th layers. The 20th layer is an IGF module, which is used to fuse the feature maps of the 19th and 9th layers. The 15th, 18th, and 21st layers are output to the small detection head, the medium detection head, and the large detection head, respectively.

[0029] Among them, see Fig.13, the CEG module includes the first CBS standard convolution, the Splite module, the first GhostBottleneck module, the second GhostBottleneck module, the first Concat module, the EMA attention mechanism and the second CBS standard convolution. The first CBS standard convolution, the Splite module, the first GhostBottleneck module, the second GhostBottleneck module, the first Concat module, the EMA attention mechanism and the second CBS standard convolution are connected in sequence, and the first GhostBottleneck module is residually connected to the first Concat module. The first Concat module splices the other output of the Splite module, the residual output of the first GhostBottleneck module and the output of the second GhostBottleneck module. Among them, the first GhostBottleneck module and the second GhostBottleneck module include the first GhostConv unit, the second GhostConv unit and the Add unit connected in sequence, and the input is also residually connected to the Add unit.

[0030] See also Figure 6 , the working process of the CEG module is: S221: Generate an initial feature map through the first CBS standard convolution.

[0031] S222: The initial feature map is grouped through the Splite module, one part is output to the first GhostBottleneck module, and the other part is directly output to the first Concat module.

[0032] S223: Processed by the first Ghost Bottleneck module and output to the second Ghost Bottleneck module, and the residual is connected to the first Concat module.

[0033] S224: Processing is performed through the second GhostBottleneck module to obtain a feature map.

[0034] S225: Concatenate the initial feature map, the residual output of the first GhostBottleneck module, and the output of the second GhostBottleneck module through the first Concat module.

[0035] S226: Dynamic weight selection via EMA attention mechanism.

[0036] S227: Rescale the features and channels through the second CBS standard convolution.

[0037] See also Fig. 9, the working process of the EMA attention mechanism is: S401: Input features Figure X Divided into G groups, the dimension of each sub-feature map is (C / G, H, W); among them, the input feature Figure X The dimensions are (C,H,W).

[0038] S402: Each sub-feature map after grouping is processed in a different way. Some sub-feature maps are subjected to average pooling in the horizontal and vertical directions. The grouped G features are , after horizontal pooling and vertical pooling, we can get: , the number of channels is (C / G)*1*W; , the number of channels is (C / G)*1*H; the pooled features are reapplied to the original feature map through Concat and 1*1 convolution operations and then Reweight to obtain the reweighted feature map: ;in, Represents the weight obtained after sigmoid. Represents element-wise multiplication.

[0039] S403: performing group normalization on the re-weighted feature maps.

[0040] S404: Perform spatial attention calculation on the normalized feature map, and redistribute and normalize the weight of each spatial position through Avgpool and Softmax activation to obtain the weight of the spatial position.

[0041] S405: Perform channel attention calculation on another part of the feature maps in the initial grouping, and obtain the importance weight of each channel through Avgpool and Softmax symmetric operations.

[0042] S406: Combine the importance weight of the channel and the weight of the spatial position to obtain the final output.

[0043] Among them, see Fig.11The IGF module includes a second Concat module, an AvgPool module, a first conv2d1*1 module, a DWConv1*n module, a DWConvn*1 module, a second conv2d1*1 module, a sigmoid activation function, a third Concat module, a multi module, and an Add module. The second Concat module, the AvgPool module, the first conv2d1*1 module, the DWConv1*n module, the DWConvn*1 module, the second conv2d1*1 module, and the sigmoid activation function are connected in sequence. The multi module is used to multiply the weight feature map output by the sigmoid activation function by the two original feature maps element by element to obtain an enhanced feature map. The Add module is used to perform an alternating residual connection between the enhanced feature map and the two original feature maps. The third Concat module is used to splice the two feature maps after the residual connection again in the channel dimension.

[0044] See also Fig.10 , the working process of the IGF module is: S231: The two feature maps are concatenated in the channel dimension through the second Concat module to obtain a new feature map: , ;in, and Represent two inputs of the same dimension, Indicates stacking two feature maps along the channel dimension.

[0045] S232: The new feature map is pooled through the AvgPool module Perform average pooling, and the pooling window size is , then the result is a global eigenvector , represents the average value of each channel: .

[0046] S233: Perform a 1*1 convolution operation through the first conv2d1*1 module to recombine the features in the channel dimension; wherein the weight of the convolution kernel is , then the result after convolution is: ;in, .

[0047] S234: Perform two depth convolutions in different directions through the DWConv1*n module and the DWConvn*1 module to extract independent features of each channel, and obtain the feature map after two depth convolutions: ; Among them, the weight of the depth convolution is ,but .

[0048] S235: Perform a 1*1 convolution operation through the second conv2d1*1 module to reintegrate the channel features. The result after convolution is ;in, , the dimension is ; Definition and Similar, for general description.

[0049] S236: The sigmoid activation function is used to limit the value of each element to the range of [0,1] to obtain the weight feature map: .

[0050] S237: weight feature map through multi module The enhanced weights are obtained by element-wise multiplication with the respective original feature maps: , .

[0051] S238: Add the enhanced weights through the Add module The information interaction between different scales is realized by alternating residuals with the original feature map to obtain the feature map: , .

[0052] S239: The two feature maps after the residual connection are concatenated again in the channel dimension through the third Concat module to obtain the final output feature map ;in, .

[0053] S3: Use the preprocessed samples to train and test the improved YOLOv8n model to obtain the optimal model. Fig.17 , the specific process is: S301: The improved YOLOv8n model is used to train the training set. Three detection heads of different scales are used to extract features of input information of different scales through two independent branches of classification and regression, and finally the category and location information of the target are obtained.

[0054] S302: After obtaining the category and location information of the defect sample, the detection head performs loss calculation with the label data.

[0055] S303: Calculate the gradient through CIOU loss function and back propagation to update the parameters. S304: Test the improved YOLOv8n model on the test set to determine the hyperparameters with the best effect.

[0056] In the above process, an epochs value is set to iterate and optimize the network multiple times.

[0057] S4: Acquire the LCD screen image and detect it through the optimal model.

[0058] Example 3 Embodiment 3 provides a method for detecting defects in a liquid crystal screen based on deformable convolution and feature fusion, the method comprising: S1 Sample acquisition environment setup: Perform camera calibration based on the characteristics of the LCD screen to be tested, and select the optimal parameters of the camera under the imaging environment. If there are backlight or reflection problems in the environment, you can provide appropriate external light sources for fill light to ensure that the details of the photographed object are clearly displayed. When fill light, you need to adjust the angle and intensity of the light source to avoid highlight areas or reflection problems caused by excessive lighting. You can also use a soft light cover or diffuser to make the light soft and uniform, and reduce shadows and noise interference caused by strong light. In addition, if necessary, you can adjust the camera's exposure parameters and white balance so that the color and brightness of the image meet the detection requirements, thereby improving the quality and stability of the collected image. Fig.18 Some of the acquired defect images are shown.

[0059] S2 Sample data cleaning and division: Manually clean the collected LCD sample data. Since there may be noise, outliers or samples that do not meet the detection standards in the data, they need to be screened one by one to remove images that are blurred, lack contrast or contain impurities. After cleaning, the data is divided into training sets, validation sets and test sets according to task requirements to ensure that the data is evenly distributed and fully covered. The division ratio should be reasonably set according to the number of samples and the requirements of the detection target. For example, a common division ratio is 70% training set, 20% validation set and 10% test set to support accurate training and performance evaluation of the model.

[0060] S3 performs data enhancement on the training set: In order to enhance the generalization ability of the model, data enhancement is only performed on the divided training set. Common data enhancement methods include rotation, translation, scaling, cropping, etc. By randomly rotating the angle, slightly translating or scaling, the model can adapt to LCD samples of different angles, positions and sizes; the data-enhanced training set can effectively simulate a variety of actual scenes and improve the robustness and generalization ability of the model to complex environments. However, it should be noted that when enhancing the LCD screen, the enhancement method that changes the original display characteristics of the LCD screen should be avoided as much as possible, such as changing the real color displayed by the LCD screen, or adding defect types that cannot appear in the original image. Use the LabelImg tool for each sample image, select the YOLO type for the annotation format, and mark the area and type of defects in each image. During the annotation process, try to ensure that the box selection area fits the defect area to avoid introducing too many other normal area features that affect the training results. After annotation, a TXT file in YOLO format will be generated, which contains five normalized parameters, namely the category, the ratio of the center X coordinate of the box to the image width, the ratio of the center Y coordinate of the box to the image height, and the ratio of the box width and the box height to the image.

[0061] S4 embeds DCNv2 convolution in the YOLOv8 model for training: before training, the default parameter 640 of img_size in YOLOv8 is used as the input image size, and then the ordinary convolution in the original YOLOv8 is replaced with the DCNv2 convolution structure, specifically replacing the convolution operations with 3, 5 and 7 layers. The replacement scheme of the present invention is the optimal result obtained by multiple experiments on our dataset. The number of replacements can be appropriately increased or decreased for other different datasets. The core of deformable convolution lies in: 1. How to calculate the offset (offset) and modulation factor m of the sampling position through additional convolution layers; 2. How to use the calculated offset to adjust the sampling position of the convolution kernel to flexibly adapt to the deformation of the input features, thereby generating a new feature map. DCNv2 introduces a modulation factor m based on DCNv1, and further enhances the control of each sampling position by accurately processing complex image transformations and nonlinear features. Assuming the input feature Figure X The dimension is ,.in, is the number of channels of the input feature map, H and W are the width and height of the feature map respectively, and the size of the convolution kernel is K*K. First, the offset Generated by an additional convolutional layer, the input feature map is obtained after the additional convolutional layer: .in and are the weight and bias of the additional convolutional layer, which are two learnable parameters. The dimension is The subsequent modulation factor m is also calculated by adding a convolutional layer: . and They are the weights and biases of the convolution layer that calculate the modulation factor. Each convolution kernel sampling point has a modulation factor, so the dimension of m is , the obtained modulation factor needs to be constrained between [0,1] by the sigmoid activation function in order to control the contribution of each sampling point. The final calculation process of the modulation factor is: After obtaining the offset and modulation factor m, assume that the center position of the feature map convolution kernel is , the sampling position of the traditional convolution kernel is: .in is the offset of the corresponding center point. For example, for a 3*3 convolution kernel, It may be (-1, -1), (0, -1), (1, -1), ..., (1, 1), etc. In DCNv2, the new sampling position is determined by the offset However, since the sampling position may not be an integer, it is necessary to calculate the eigenvalues ​​of the sampling points through bilinear interpolation. ,After obtaining the eigenvalue of each adjusted sampling point, the modulation factor m is used to adjust the contribution of each sampling point: Among them, m has been normalized by the sigmoid activation function. When m=0, it means that the position does not contribute to the output feature, and when m=1, it means that the position fully participates in the calculation. The final output feature map Y is at position The convolution operation is performed on the adjusted sampling position to obtain: ,.in, is the weight of the convolution kernel, is the eigenvalue after modulation.

[0062] S5 introduces the GhostConv convolution module: In order to improve the model's ability to capture features without increasing the amount of computation, we use Ghostconv in the original C2f module to perform feature extraction operations. The core idea of ​​this convolution is that the feature maps generated by the original convolution calculation often have high redundancy. This is because the feature maps extracted between different levels are similar. Therefore, Ghostconv divides the traditional convolution into two steps: 1. Main convolution operation. 2. Linear transformation generates Ghost features. In the main convolution operation, a part of the core feature map is generated through a small number of conventional convolutions. This step is inevitable. Then, some simple linear operations are performed on the core feature map to generate the Ghost feature map. This operation can obtain the remaining feature map through extremely low-cost simple linear operations. Finally, the core feature map and the generated Ghost feature map are fused to obtain the final output result, which effectively reduces the problem of excessive computation and parameters caused by feature map redundancy. The principle of GhostConv convolution is as follows Figure 7 shown.

[0063] S6 introduces a dynamic weight selection mechanism: the input features will be spliced ​​after two GhostBottleneck operations, and then convolution will be used to readjust the features and channels. Our design adds a dynamic weight selection mechanism between the spatial and channel adjustments, where the model splices together feature maps from multiple paths. These feature maps may contain different spatial and semantic information, and reflect the details extracted by different convolution paths. The gradient information at this fusion point is extremely rich and diverse, which reflects the model's sensitivity to losses in different feature dimensions. In this context, the dynamic weight selection mechanism can use the back-propagated gradient information to automatically learn and select which feature channels or spatial positions are more important in a specific context, thereby making the model's gradient utilization more efficient. The dynamic weight selection mechanism EMA structure is as follows: Figure 8 As shown. The specific principle is as follows: First, we group the input feature map. For input X with dimensions (C, H, W), we first divide the feature map into G groups, and the dimensions of each sub-feature map are (C / G, H, W). The purpose of grouping is to reduce the amount of calculation and enable the model to process the features in a more fine-grained manner, so as to better capture local and global information. After grouping, each feature map will go through different processing paths. Some feature maps will go through average pooling in the horizontal and vertical directions. Assume that the G group features after grouping are , after horizontal pooling and vertical pooling, we can get: , the number of channels is (C / G)*1*W; , the number of channels is (C / G)*1*H. The pooled features are reapplied to the original diagnostic image through a Concat and 1*1 convolution operation and then Reweight: .in, Represents the weight obtained after sigmoid. Represents element-by-element multiplication, giving different weights to each feature point. Then, the re-weighted feature map is subjected to Group Norm (group normalization), which can effectively improve the training stability and convergence speed of the model. At the same time, the EMA module fuses the attention of two different dimensions, CA and SE, and performs spatial attention calculation on the feature map after Group Norm. Through Avgpool and Softmax activation, the weight of each spatial position is redistributed and normalized to adjust the importance of each position. For the other part of the feature map in the initial grouping, channel attention calculation is performed, and the importance weight of each channel is obtained through Avgpool and Softmax symmetric operations. Finally, the weights of the two are combined to obtain the final output Y. In general, the EMA attention mechanism combines multiple attention calculation paths, including center weight calculation, channel attention, and spatial attention. In each path, the features are subjected to multiple pooling, convolution, activation, and fusion operations to gradually capture the importance information of different scales and dimensions. Finally, by fusing this information, the EMA attention mechanism can enhance the feature expression ability of the model, thereby improving the overall performance of the model.

[0064] S7 builds the IGF feature extraction module and embeds it into the neck network: We can see the Concat operation in many network structures, especially in the process of multi-scale feature extraction. By splicing features from different layers or different paths in the channel dimension, the concat operation provides the model with an intuitive and simple way to combine features. However, although the concat operation has the advantages of simplicity and no additional computation, its ability in feature fusion is limited, and there are some problems that cannot be ignored. First, the concat operation only connects the features in the channel dimension, and essentially does not further model the correlation or significance between these features. Features from different sources may carry information of different scales or types, but simply splicing cannot effectively guide the model to distinguish or emphasize the importance of these features. Especially in multi-scale feature fusion, due to the significant differences in the expressive power of features, the spliced ​​feature map usually has an unbalanced information distribution, which cannot fully utilize the advantages of each part of the features. Therefore, to address this problem, we propose a new feature fusion structure, which we will name IGF (Information Guide Fusion). This structure is as follows Fig.11 For information at different levels, we regard it as two inputs, and the two inputs with the same dimensions are recorded as , First, these two feature maps are concatenated in the channel dimension to obtain a new feature map: , ,.in, Indicates that the two input feature maps are stacked according to the channel dimension, and the concatenated feature map After average pooling operation. Assume that the pooling window size is , then the result is a global eigenvector , represents the average value of each channel: Next, the pooled feature map After a series of convolution operations. First, a 1*1 convolution operation is performed to recombine features in the channel dimension. Assume that the weight of the convolution kernel is , then the result after convolution is: ,in, , which represents the feature representation after the feature channel is reduced in dimension. Then, two depthwise convolutions (DWConv) are performed in different directions. Depthwise convolution is a convolution operation performed independently on each channel, which helps to extract independent features of each channel. Assume that the weight of depthwise convolution is ,but ,here It is the feature after two depth convolutions. Feature map After another 1*1 convolution, the purpose is to reintegrate the channel features. The result after convolution is recorded as , the dimension is : , and then through the sigmoid activation function, the value of each element is limited to the range of [0,1] to obtain the weight feature map: In the feature fusion part, we use a combination of feature weighting and residual connection to weight the features using the generated weights, and then perform residual connection with the original input to enhance the expressiveness of the features. First, the generated weight feature map Multiply each element by its original feature map to enhance important features and suppress minor features: , , and obtain their respective enhanced weights Then, the information interaction between different scales is realized by alternating residuals, which enhances the model's ability to capture contextual features: , Finally, the two feature maps after the residual connection are concatenated again in the channel dimension to obtain the final output feature map: ,.in, .

[0065] S8 uses the constructed YOLO-DEI model for training: the improved YOLO-DEI is used to train the data set obtained in S3. Specifically, during the training process, the loss value is calculated by calculating the difference between the output result of the input training data and the expected output result, and then the back propagation algorithm is used to perform gradient backpropagation to update the parameters. This process sets an epochs value to iteratively optimize the network multiple times, so that it gradually learns the sample features and finally obtains a detection model that meets the requirements of downstream tasks. The network structure of the improved YOLOv8n model is shown in the figure. Figure 3 As shown, it includes the input, backbone network Backbone, neck network neck and detection head head; the model first crops the original RGB image to a size of 640*640 and then inputs it into the Backbone part of YOLO-DEI through the Input. This part is used to extract the location and semantic information of the defect image. The output feature map is then sent to the neck part for multi-scale feature fusion. The feature map information output by different levels of the Backbone is fused and interacted. During the fusion process, the IGF module is used to ensure that different semantic information can be better used. At the same time, irrelevant information is suppressed, and more weight is given to important information. Finally, three different detection heads are set to adapt to the three types of targets in the feature map: large, medium and small. The small detection head is used to detect the shallower feature maps to prevent the small target features from gradually disappearing during the extraction process as the network deepens. The other two detection heads are used to detect the middle and deep feature maps. The feature maps have a larger size and can better contain the semantic information of the entire image. Figure 3 Backbone contains a total of 10 layers. Layer 0 and layer 1 are mainly used to expand the three channels of RGB and reduce the feature map to extract deep features. Due to two times of 2-fold downsampling, the size of the feature map after these two layers is reduced to 1 / 4 of the original. The second layer is the CEG module we built. This module is mainly used for feature extraction operations without changing the feature map size and the number of channels. The third layer is the DCNv2 module we use. This module also performs downsampling operations. The feature map extracted by CEG is downsampled by 2 times and reduced by 1 / 2. The size of the feature map here has become 1 / 8 of the input size. The subsequent layers 4-8 are consistent with the 2nd and 3rd layers, which are used to repeatedly reduce the feature map and extract important features in the map. After the first 8 layers, the size of the feature map becomes 1 / 32 of the original. The feature map output by this layer has rich semantic information, but the feature map is small. The 9th layer is the SPPF pooling layer, which gathers LCD screen defect features of various scales here through the maximum pooling operation. The calculation formula for the feature map size is , ,.in, and Represents the width and height of the output feature map, H and W represent the width and height of the input feature map, K represents the size of the convolution kernel, P represents the padding of the feature map, and S represents the step size of the convolution. The output feature map from the 9th layer SPPF will be input to the neck part for multi-scale feature fusion operation. First, the feature map of the 9th layer is upsampled. The purpose of this operation is to prepare for feature fusion and ensure that the two fused features have the same dimension. The feature size of the 10th layer after 2 times upsampling will become 40*40*256. The feature map is fused with the feature map of the 6th layer to obtain the feature map of the 11th layer. The subsequent upsampling fusion operation is similar. The feature map is enlarged by the nearest neighbor interpolation method, and then fused with the shallow feature to obtain a new feature map. Specifically, the output feature map of the 13th layer is fused with the feature map of the 4th layer. After upsampling and fusion, the size of the feature map is restored to 80*80 and sent to the small detection head for detection. At the same time, the features of this layer will continue to be downsampled to obtain a 40*40 feature map and fuse it with the feature map of the 12th layer. After that, it will be sent to the medium detection head for detection after feature extraction by the CEG module. The 40*40 feature map will be downsampled again to obtain a 20*20 feature map and fuse it with the feature map of the 9th layer, and finally sent to the large detection head for detection. The fusion part of the entire model adopts the PAN bottom-up and top-down combination method. By fusing the features between different layers multiple times, semantic information of different scales is fully gathered on the output feature map. In the detection head part, the original structure of YOLOv8 is continued to be used, and three detection heads of different scales are adopted. The input information of different scales is extracted through two independent branches of classification and regression to finally obtain the category and location information of the target. After obtaining the category and location information of the defect sample, the Head calculates the loss with the label data, and then calculates the gradient through the CIOU loss function and back propagation to update the parameter optimization model, making the classification and location of defects more accurate. Finally, the model is verified on the verification set to obtain the improved model indicators.

[0066] S9 tests the model on the test set to determine the hyperparameters with the best effect: we export the trained model as a file with the suffix .pth, apply it to the test set for prediction, and obtain the optimal parameters based on the test results. The optimal experimental parameters in the present invention are: the image input size is 640*640, the training round is 300 epochs, the optimizer we use is Adam, the momentum parameter is 0.937, the initial learning rate is 0.01, the batch size is 24, and the workers are 8. The specific model evaluation indicators use accuracy, recall, average precision, parameter quantity, calculation amount and other indicators. The introduction of each indicator is as follows: Accuracy: , where TP represents the number of defects detected correctly and FP represents the number of defects detected incorrectly. A higher precision indicates that the model has a lower false positive rate. Recall: , FN refers to the number of targets that have not been detected, and recall refers to the proportion of all real targets that have been correctly detected, indicating whether the model has fully detected the target. A high recall rate indicates that the model has a low missed detection rate. Average precision: , AP stands for average accuracy, k stands for the number of sample categories, and 50 stands for the average accuracy of each type of target when the IoU (intersection over union) threshold is 0.5. mAP50 is an important indicator for measuring the accuracy of target detection, which represents the mean of the average accuracy on different categories. Parameters: The parameter refers to the total number of all trainable parameters in the model, which usually indicates the complexity of the model. The larger the parameter, the more complex the model, and the storage requirements and computing resource requirements also increase accordingly. A smaller parameter helps improve the efficiency of the model on devices with limited resources. Computational Amount: The computational amount is usually measured in FLOPs (floating point operations per second), which indicates the computing resources required by the model during forward reasoning. The larger the computational amount, the more computing resources and time the model requires during reasoning. For real-time target detection applications, models with smaller computational amounts are usually more suitable. Simulation experiment: To verify the feasibility of the technical solution of the present invention, we conducted a simulation experiment. The hardware environment of the experiment was a single NVDIA RTX 3090 (24GB) GPU, an Inter(R) Xeon(R) Platinum 8362 processor, the operating system was Ubuntu18.04, the programming language was Python3.8.1, the deep learning framework was Pytorch2.0, the CUDA version was 11.8, and the specific hyperparameters were set as follows: the image input size was 640*640, the training round was 300 epochs, the optimizer we used was Adam, the momentum parameter was 0.937, the initial learning rate was 0.01, the batch size was 24, and the workers were 8. Fig.19 The recognition results of the present invention are shown in the figure. It can be seen from the figure that the network can accurately identify various types of defects on the surface of the LCD screen in the figure and accurately locate them. Ablation experiment: To further verify the effectiveness of the combination of the modules of the present invention, we use ablation experiments to prove the advantages of using these three innovations in combination. The experimental results are shown in Table 1. In this experiment, YOLOv8n is used as the benchmark model. The hyperparameters and software and hardware environments in the ablation experiment are exactly the same. '√' represents that this module is added to this group of experiments.

[0067] Table 1 YOLO-DEI ablation experiment DCNv2 CEG IGF Precision / % Recall / % mAP50 / % Params / m FLOPs / G 92.7 82.7 85.5 3.0 8.1 √ 96.3 82.6 86.4 3.1 7.6 √ 95.7 83.9 88.1 2.8 7.7 √ 93.8 93.2 95.1 4.0 10.2 √ √ 94.1 93.8 89.5 2.9 7.2 √ √ 94.9 88.1 95.3 4.0 9.6 √ √ 95.1 93.4 95.9 3.8 9.7 √ √ √ 95.4 93.7 96.5 3.7 9.1

[0068] Through the experimental data in the table, we can find that the accuracy of the YOLOv8 benchmark model can only reach 85.5%, which is far from meeting the requirement of high accuracy for industrial production detection. After using the DCNv2 structure in the backbone of the model, the structure uses deformable convolution kernels, which is better than ordinary convolution in capturing features, so the accuracy is slightly improved and the model calculation complexity is reduced. After the improved CEG module is applied to the network, more feature maps are generated by simple linear operations, so that the model obtains more features, and because linear transformation does not introduce too many additional parameters and is simple to calculate, it is better than the original benchmark model in terms of accuracy, number of parameters, and complexity. After adding the IGF module proposed in this paper to the model, the feature fusion part is no longer a simple splicing operation, but a selective fusion of features of different values. Smaller weights are given to unimportant features and larger weights are given to important features, so that the features sent to the detection head contain more position information and semantic information than before. Therefore, the accuracy of the model is greatly improved after the introduction of the IGF module. Although the computational complexity and the number of parameters have increased, the accuracy has increased by 11.2%. At the same time, we can see from the table that the accuracy of different modules combined has an upward trend. Finally, the YOLO-DEI module proposed in this paper through a large number of experiments has improved by 12.9% in accuracy.

[0069] Comparative experiment: To prove the advancedness of the improved model of the present invention, we conducted comparative experiments with several mainstream target detection algorithms, such as SSD, FasterRCNN, Retinanet, Fcos and other versions of YOLO, under the same data set and with the same parameters. The specific experimental results are shown in Table 2.

[0070] Table 2 Comparative test Model Precision / % Recall / % Map50 / % Params / M FLOP / G FPS Faster RCNN-r50 77.2 76.1 62.5 137.1 370.2 22.7 SSD 75.3 75.5 59.1 26.3 62.7 156 YOLOv3 82.3 76.6 73.4 61.9 156.6 43 RetinaNet-r50 79.7 81.3 78.7 37.9 170.1 47.5 FCOS 88.3 79.8 86.8 32.2 161.9 52.6 YOLOV3-EfficientNet 76.6 74.5 73.3 7.2 9.5 71 YOLOv5s 86.2 78.9 84.9 7.23 17.16 74 YOLOX-s 85.5 80.7 83.9 8.97 26.93 65.5 YOLOv7 87.2 84.1 85.1 37.6 106.5 58.4 YOLOv7-tiny 85.9 79.4 82.7 6.3 13.9 108.4 YOLOv8 (baseline) 92.7 82.7 85.5 3.0 8.1 156 Ours 95.4 93.7 96.5 3.7 9.1 142.9

[0071] By observing Table 2, we can find that the two-stage network FasterRcnn-r50 has much higher accuracy than other one-stage networks of the same period, but because the network needs to generate candidate target boxes before classification and regression operations can be performed, the model speed is slow and cannot achieve real-time detection. In addition, due to the need to train two networks, the number of parameters and computational complexity of the two-stage network are much larger than other one-stage networks. As a standard one-stage algorithm, the SSD algorithm has stronger real-time performance and faster speed than the two-stage algorithm, but it is not as effective as the two-stage algorithm for small target detection and its accuracy is slightly lower than FasterRcnn. As one-stage algorithms, YOLOv3, Retinanet-r50, and fcos have different improvements in speed and accuracy, but because the network has been improved for different detection targets and many additional parameters have been introduced, the complexity and number of parameters of the model have been greatly improved. Compared with the same type of one-stage algorithm SSD, the number of parameters has increased by 1.5-3 times, and the computational complexity has increased by about 2 times. Compared with other networks, subsequent networks such as YOLOv3-efficientnet and YOLOX have solved the speed and complexity problems of previous models while maintaining high accuracy. Compared with other algorithms, the algorithm proposed in this paper far exceeds other models in mAP50 and mAP50-95. Although the detection speed is slightly lower than the baseline model, it can also meet the real-time requirements of LCD screen defect detection.

[0072] In summary, the patented algorithm significantly improves the detection accuracy while meeting the real-time detection requirements. The model parameter quantity and complexity are only slightly increased compared to the benchmark model, and it has high application value in the field of LCD screen defect detection. Fig. 20 As shown, through observation we can find that the two-stage network FasterRcnn is more accurate in locating the detected defects, but point defects are missed because the target small model cannot capture the features well. For other one-stage networks, the defect types can be basically detected, but due to the large variations in the shapes and scales of different types of defects, especially the large variations in the length and width of line defects, there may even be bright lines across the entire screen. Therefore, from the results, it can be seen that the model used in this patent for comparative testing is difficult to completely and accurately locate the overall position of the line defects, and because the background color of the display panel changes greatly during the detection process, there may be overlapping recognition frames in the baseline model YOLOv8. The new model proposed in the present invention solves these problems very well. Fig.21 and 22The confusion matrix diagrams before and after the improvement are shown. It can be seen from the experiments that for the two types of point defects and line defects that are difficult to detect by other models, the model proposed in this paper makes targeted modifications to the algorithm and uses deformable convolution and the IGF module proposed in this paper to solve the problem of large scale variation of line defects and inaccurate identification of other types of defects. In summary, the method proposed in the present invention is superior to other methods mentioned in the comparative experiments in the article in terms of LCD screen defect detection.

[0073] Heat map experiment: In order to study the effect of feature extraction within the model, in this section we conduct heat map exploration on the output part of different models. According to the weighted sum of gradient information, we can get the importance score of each pixel to the target category to produce a heat map. The specific results are as follows: Fig.23 As shown. In the figure, we can see that the two-stage detection network FasterRcnn is more sensitive to the other three larger defects, but it is difficult to identify small defects. The other one-stage networks can identify the four defects mentioned in this article, but due to the differences in the feature extraction parts, the position recognition performance of different defects is also different, especially for the two types of samples, line defects and scratches. The position of the defect cannot be accurately identified, resulting in a large deviation between the regression frame and the real frame. Based on the new model proposed in this invention, we can clearly see that for different types of defects in the data set, the model can accurately identify them, and can also adapt to different scales of defect features. For the two types of samples, line defects and scratches, the model proposed in this article can accurately identify the shape and size of the entire defect, so as to make the regression of the target frame more accurate and improve the detection accuracy.

[0074] Generality Experiments: To verify the generalization ability of the proposed YOLO-DEI model, we used alternative defect datasets for testing. Specifically, we used the NEU-DET steel dataset provided by Northeastern University, which contains both minor corrosion spots and larger block defects. This dataset includes 1,800 samples covering six defect categories - cracks, inclusions, spots, surface pits, rolling shrinkage marks, and scratches - providing a solid foundation for evaluating the generalization performance of the model. The experimental results are shown in Table 3, which show that the YOLO-DEI model continues to surpass the original YOLOv8 model in both mAP50% and mAP50–95%. It is worth noting that the YOLO-DEI model maintains a high detection speed with an FPS of 122 and an image processing time of approximately 7.5 milliseconds, thereby improving accuracy without reducing efficiency.

[0075] Table 3 Generality experiment Model Map50 / % Map50-95 / % Params / M FLOP / G FPS YOLOv8 76.2 42.1 3.0 8.1 122 YOLO-DEI 77.3 44.8 3.7 9.1 129

[0076] In summary, the present invention aims at the phenomenon of missed detection, wrong detection and inaccurate positioning in the defect detection link of small LCD screens, and proposes a YOLO-DEI LCD screen defect detection model. Through a large number of comparative experiments and ablation experiments, it is proved that the model can effectively solve the problem of inaccurate positioning due to large scale changes of line defects and the problem of easy missed detection due to small point defects, and the algorithm is almost consistent with the original model in terms of parameter quantity and calculation amount, while maintaining the accuracy improvement without increasing additional calculation cost. This paper improves the feature extraction part, and uses DCNv2 deformable convolution to replace the original ordinary convolution, so that the feature extraction part can more accurately extract the location of the defect and adapt to the shape of the defect. At the same time, ghost convolution is combined with C2f to enhance the reusability of features without increasing additional costs. In the feature fusion part, this paper proposes an IGF structure to further fuse important information between different features to improve the overall performance of the model, providing an effective solution to the above-mentioned problems. A large number of ablation experiments and comparative experiments have proved that the improved model can effectively solve the problems of missed detection, wrong detection, inaccurate positioning, etc. in the detection of small LCD screens. The improved module of the present invention is universal and can also show good performance on other similar data sets. Compared with the current mainstream algorithms, the model proposed by the present invention has excellent performance in both detection accuracy and detection speed.

Claims

1. A method for detecting defects in a liquid crystal display based on deformable convolution and feature fusion, characterized in that: The method comprises: S1: Pre-process the sample; S2: An improved YOLOv8n model is constructed based on the YOLOv8n model, and the improved YOLOv8n model has the following adjustments: (1) Replace the CBS standard convolutions in the 3rd, 5th, and 7th layers of the backbone network in the YOLOv8n model with DCNv2 variable convolutions; (2) Replace the C2f module in the YOLOv8n model with the CEG module; (3) Replace the Concat module in the YOLOv8n model with the IGF module; Among them, the CEG module includes a first CBS standard convolution, a Splite module, a first GhostBottleneck module, a second GhostBottleneck module, a first Concat module, an EMA attention mechanism and a second CBS standard convolution; the first CBS standard convolution, the Splite module, the first GhostBottleneck module, the second GhostBottleneck module, the first Concat module, the EMA attention mechanism and the second CBS standard convolution are connected in sequence, the first GhostBottleneck module residual connection first Concat module, the first Concat module another output of the Splite module, the residual output of the first GhostBottleneck module and the output of the second GhostBottleneck module are spliced; Among them, the IGF module includes a second Concat module, an AvgPool module, a first conv2d1*1 module, a DWConv1*n module, a DWConvn*1 module, a second conv2d1*1 module, a sigmoid activation function, a third Concat module, a multi module and an Add module. The second Concat module, the AvgPool module, the first conv2d1*1 module, the DWConv1*n module, the DWConvn*1 module, the second conv2d1*1 module and the sigmoid activation function are connected in sequence. The multi module is used to multiply the weight feature map output by the sigmoid activation function by the two original feature maps element by element to obtain an enhanced feature map; the Add module is used to perform an alternating residual connection between the enhanced feature map and the two original feature maps; the third Concat module is used to splice the two feature maps after the residual connection again in the channel dimension; S3: Use the preprocessed samples to train and test the improved YOLOv8n model to obtain the optimal model; S4: Acquire the LCD screen image and detect it through the optimal model.

2. The method according to claim 1, characterized in that The sample pretreatment includes: S101: Manually clean the collected LCD screen samples to remove blurred images, images with insufficient contrast, and images containing impurities; S102: Divide the cleaned samples into a training set, a validation set, and a test set; S103: Performing data enhancement on the training set, wherein the data enhancement method includes rotation, translation, scaling and / or cropping; S104: Label the sample to obtain a TXT file in YOLO format; The TXT file contains five normalized parameters, namely, category, the ratio of the center X coordinate of the box to the image width, the ratio of the center Y coordinate of the box to the image height, the ratio of the width of the box and the height of the box to the image; the categories include Dot defect, Line defect, Scratch and Leakage.

3. The method according to claim 2, characterized in that The improved YOLOv8n model includes an input terminal Input, a backbone network Backbone, a neck network neck and a detection head head; The backbone network Backbone includes: Layers 0 and 1 are both CBS standard convolutions; Layers 2, 4, 6, and 8 are CEG modules, used for feature extraction; The 3rd, 5th, and 7th layers are all DCNv2 variable convolutions, which are used for 2x downsampling; The 9th layer is the SPPF module, which is used to perform maximum pooling to aggregate LCD screen defect features of various scales; The detection head includes small detection head, medium detection head and large detection head; The neck network includes: The 10th and 13th layers are Upsumple modules, which are used for upsampling; Layers 12, 15, 18, and 21 are all CEG modules, used for feature extraction; The 16th and 19th layers are both CBS standard convolutions; The 11th layer is the IGF module, which is used to fuse the feature maps of the 10th and 6th layers; The 14th layer is the IGF module, which is used to fuse the feature maps of the 13th and 4th layers; The 17th layer is the IGF module, which is used to fuse the feature maps of the 16th and 12th layers; The 20th layer is the IGF module, which is used to fuse the feature maps of the 19th and 9th layers; The 15th, 18th and 21st layers are output to the small, medium and large detection heads respectively.

4. The method according to claim 1, characterized in that The working process of the DCNv2 variable convolution is: S211: Calculate the offset of the sampling position based on the additional convolutional layer and modulation factor m; Among them, the input feature map X is obtained after the additional convolution layer: ;in, and are the weights and biases of the additional convolutional layer, The dimension is ; The dimension of the input feature map X is ;in, is the number of channels of the input feature map, H and W are the width and height of the feature map respectively, and the size of the convolution kernel is K*K; in, ;in, and They are the weights and biases of the convolution layer for calculating the modulation factor. Each sampling point of the convolution kernel has a modulation factor, and the dimension of m is ; S212: By offset Calculate the eigenvalues ​​of the sampling points by bilinear interpolation ; S213: Use the modulation factor m to adjust the contribution of each sampling point: ; S214: Obtain the output feature map through convolution operation: ;in, is the weight of the convolution kernel.

5. The method according to claim 1, characterized in that: The working process of the CEG module is: S221: Generate an initial feature map through the first CBS standard convolution; S222: The initial feature map is grouped by the Splite module, one part is output to the first GhostBottleneck module, and the other part is directly output to the first Concat module; S223: Processed by the first GhostBottleneck module and output to the second GhostBottleneck module, and the residual is connected to the first Concat module; S224: Processing by the second GhostBottleneck module; S225: splicing the initial feature map, the residual output of the first GhostBottleneck module, and the output of the second GhostBottleneck module through the first Concat module; S226: Dynamic weight selection via EMA attention mechanism; S227: re-adjust features and channels through the second CBS standard convolution; The first GhostBottleneck module and the second GhostBottleneck module both include a first GhostConv unit, a second GhostConv unit and an Add unit connected in sequence, and the input is also residually connected to the Add unit.

6. The method according to claim 5, characterized in that The working process of the EMA attention mechanism is as follows: S401: Divide the input feature map X into G groups, and the dimension of each sub-feature map is (C / G, H, W); wherein the dimension of the input feature map X is (C, H, W); S402: Each sub-feature map after grouping is processed in a different way. Some sub-feature maps are subjected to average pooling in the horizontal and vertical directions. The grouped G features are , after horizontal pooling and vertical pooling, we can get: , the number of channels is (C / G)*1*W; , the number of channels is (C / G)*1*H; the pooled features are reapplied to the original feature map through Concat and 1*1 convolution operations and then Reweight to obtain the reweighted feature map: ;in, Represents the weight obtained after sigmoid. Represents element-wise multiplication; S403: performing group normalization on the re-weighted feature maps; S404: Perform spatial attention calculation on the normalized feature map, and redistribute and normalize the weight of each spatial position through Avgpool and Softmax activation to obtain the weight of the spatial position; S405: Perform channel attention calculation on another part of the feature maps in the initial group, and obtain the importance weight of each channel through Avgpool and Softmax symmetric operations; S406: Combine the importance weight of the channel and the weight of the spatial position to obtain the final output.

7. The method according to claim 6, characterized in that The working process of the IGF module is: S231: The two feature maps are concatenated in the channel dimension through the second Concat module to obtain a new feature map: , ;in, and Represent two inputs of the same dimension, Indicates stacking two feature maps according to the channel dimension; S232: The new feature map is pooled through the AvgPool module Perform average pooling, and the pooling window size is , then the result is a global eigenvector , which represents the average value of each channel: ; S233: Perform a 1*1 convolution operation through the first conv2d1*1 module to recombine the features in the channel dimension; wherein the weight of the convolution kernel is , then the result after convolution is: ;in, ; S234: Perform two depth convolutions in different directions through the DWConv1*n module and the DWConvn*1 module to extract independent features of each channel, and obtain the feature map after two depth convolutions: ; Among them, the weight of the depth convolution is ,but ; S235: Perform a 1*1 convolution operation through the second conv2d1*1 module to reintegrate the channel features. The result after convolution is ;in, , the dimension is ; S236: The sigmoid activation function is used to limit the value of each element to the range of [0,1] to obtain the weight feature map: ; S237: weight feature map through multi module Multiply each element by its original feature map to get the enhanced weights: , ; S238: Add the enhanced weights through the Add module The information interaction between different scales is realized by alternating residuals with the original feature map to obtain the feature map: , ; S239: The two feature maps after the residual connection are concatenated again in the channel dimension through the third Concat module to obtain the final output feature map ;in, .

8. The method according to claim 3, characterized in that The improved YOLOv8n model is trained and tested using the preprocessed samples to obtain the optimal model, including: S301: The improved YOLOv8n model is used to train the training set. Three detection heads of different scales are used to extract features of input information of different scales through two independent branches of classification and regression, and finally the category and location information of the target are obtained; S302: After obtaining the category and location information of the defect sample, the detection head performs loss calculation with the label data; S303: Calculate the gradient through CIOU loss function and back propagation to update the parameters; S304: Test the improved YOLOv8n model on the test set to determine the hyperparameters with the best effect; In the above process, an epochs value is set to iterate and optimize the network multiple times.

9. The method according to claim 8, characterized in that The hyper parameters are: image input size is 640*640, epochs is 300, optimizer is Adam, momentum parameter is 0.937, initial learning rate is 0.01, batch size is 24, and workers is 8.

Citation Information

Patent Citations

  • Traffic remote sensing target detection method

    CN119169268A

  • Road defect detection method based on DRR module and SDFM

    CN119418285A

Cited By

  • Locomotive cab personnel leaving detection method based on YOLOV8s and related equipment

    CN121170766A

  • Wafer defect detection method and system based on YOLO-Label comparison

    CN121437520A

  • A wafer defect detection method and system based on YOLO-Label comparison

    CN121437520B