Display screen defect detection method and system based on frequency perception guidance
Through the frequency-aware guided display defect detection method, combined with split heterogeneous feature extraction, feature enhancement enhancement and adaptive feature interaction, the diversity and noise interference problems in the detection of LCD display defects is solved, and the detection accuracy and efficiency are improved.
Patent Information
- Application Number
- CN202510291764.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-07-04
AI Technical Summary
The existing LCD display defect detection technology faces problems such as diverse defect shapes, significant changes in different defect types, fewer defect target pixels and background noise interference, resulting in low detection accuracy and low efficiency.
Using a display screen defect detection method based on frequency perception guidance, through the combination of feature extraction network, cross-fusion network and detection layer, the defect perception features are extracted using split heterogeneous feature extraction module and frequency perception mechanism, combined with feature enhancement module and adaptive feature interaction module, the LM-CIoU loss function is designed for training, and the detection accuracy is improved.
It improves the detection ability of complex small target defects, reduces redundant information, enhances the expression and detection accuracy of defect characteristics, and improves the adaptability and efficiency of the model.
Smart Images

Figure CN120259201A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of defect detection, and particularly relates to a display screen defect detection method and system based on frequency perception guidance. Background Art
[0002] Defect detection of liquid crystal display screens can ensure the quality and reliability of products, and is crucial for safeguarding consumer rights and interests, enhancing the market competitiveness of electronic products, and ensuring production efficiency. Effective defect detection can timely discover problems in the production process, reduce the defective rate, and optimize the production process.
[0003] However, there are many difficulties in the surface defect detection of liquid crystal display screens, including the diversity of defect shapes, different shapes, few pixels contained in defect targets, and significant variations between different defect types. The above detection difficulties make it difficult to accurately and efficiently detect defects in liquid crystal display screens. The research on liquid crystal display screen defect detection has experienced a transformation from traditional manual detection methods to modern deep learning-based automated detection technologies. In traditional detection methods, manual visual inspection is usually relied on, which is inefficient and leads to a decrease in detection accuracy. With the development of machine vision technology, automatic detection methods based on image processing have begun to be applied to the defect detection of liquid crystal display screens. However, traditional image processing methods are prone to losing defect details during the processing, which restricts the accuracy of liquid crystal screen defect detection. In recent years, deep learning technology has demonstrated excellent performance in computer vision tasks, and more and more researchers have applied it to the detection of liquid crystal display screen defects. Deep learning methods can automatically learn image features and perform complex object detection.
[0004] Deep learning technology demonstrates excellent performance in feature extraction and model generalization. However, when implementing defect detection tasks, current deep learning methods also face a series of challenges:
[0005] (1) The defect shapes of liquid crystal display screens are diverse and different, making it difficult to detect multiple defect types simultaneously.
[0006] (2) There are significant variations between different defect types of liquid crystal display screens. The general detection model cannot effectively fuse adjacent layer features, and the fusion process does not select a large number of defect features, introducing redundant information and having a low detection accuracy for multi-scale defect targets.
[0007] (3) The defect targets in liquid crystal display screens contain few pixels, the features are not obvious, and there is background noise interference, resulting in a low detection accuracy of the object detection algorithm.
[0008] (4) The currently used loss function has deficiencies in identifying liquid crystal display screen defects with diverse shapes and cannot effectively capture the features of these defects. Summary of the Invention
[0009] In order to solve at least one deficiency of the prior art, the object of the present invention is to provide a display screen defect detection method and system based on frequency perception guidance to improve the accuracy of display screen defect detection.
[0010] In order to achieve the above object, according to some embodiments, in the first aspect of the present invention, a display screen defect detection method based on frequency perception guidance is provided, including:
[0011] Obtain the image to be detected;
[0012] Input the image to be detected into the trained display screen defect detection model for detection to obtain the display screen defect detection result;
[0013] Wherein, the display screen defect detection model includes a feature extraction network, a cross-fusion network and a detection layer connected in sequence. The feature extraction network includes multiple layers of feature extraction structures. Each layer of feature extraction structure includes a split heterogeneous feature extraction module and a frequency perception mechanism. The input feature of each feature extraction structure is obtained by downsampling the output of the split heterogeneous feature extraction module in the previous layer; the split heterogeneous feature extraction module is used to extract the defect perception feature in the input feature through split convolution, and the frequency perception mechanism is used to extract the high-frequency feature and low-frequency feature in the input feature, and fuse the low-frequency feature with the defect perception feature to obtain the guided perception feature; the cross-fusion network is used to perform cross-fusion on the concatenated feature obtained by concatenating the guided perception feature and the high-frequency feature; the detection layer is used to detect the feature after cross-fusion by the cross-fusion network to obtain the defect detection result.
[0014] Preferably, the feature extraction network includes five layers of feature extraction structures. The first four layers of feature extraction structures include a split heterogeneous feature extraction module and a frequency perception mechanism. The fifth layer of feature extraction structure includes a split heterogeneous feature extraction module, a feature enhancement module and a frequency perception mechanism;
[0015] The feature enhancement module includes a feature enhancement branch and a fine-grained feature perception branch. The feature enhancement branch groups the input feature map according to the number of channels of the input feature map, splices and fuses the grouped feature maps after adaptive max pooling and adaptive average pooling processing, and passes through a convolution module, a depth convolution module, and a convolution module in sequence, generates an attention score through the Sigmoid function, multiplies the attention score by the grouped feature maps, and finally obtains the feature map output by the feature enhancement branch through pointwise convolution processing;
[0016] The fine-grained feature perception branch processes the input feature map through a depth convolution module and a pointwise convolution module in sequence to obtain the feature map output by the fine-grained feature perception branch;
[0017] The feature maps output by the feature enhancement branch and the feature maps output by the fine-grained feature perception branch are added element-wise and fused to obtain the feature maps output by the feature enhancement and enhancement module.
[0018] Preferably, the frequency perception mechanism obtains the low-frequency component, low-high frequency component, high-low frequency component, and high-high frequency component of the input feature map through wavelet transform;
[0019] The low-frequency component is normalized after passing through a depth convolution module with a convolution kernel size of 3x3 and a convolution module with a convolution kernel size of 1x1 in sequence, and is multiplied element-wise with the low-frequency component LL to output the low-frequency feature after establishing a residual connection;
[0020] The low-high frequency component, high-low frequency component, and high-high frequency component are concatenated and fused to obtain a fused feature. The fused feature is processed through a convolution module with a convolution kernel size of 1x1 and a Softmax layer, and is multiplied element-wise with the fused feature after establishing a residual connection to output the high-frequency feature.
[0021] Preferably, the cross-fusion network includes three fusion stage networks with decreasing cross-fusion nodes in sequence; the first layer of the cross-fusion stage is composed of the concatenation and fusion of the guiding perception feature and the high-frequency feature of the frequency perception mechanism, the second layer of the cross-fusion stage is composed of the fusion of adjacent nodes in the first layer of the cross-fusion stage using an adaptive feature interaction module, and the third layer of the cross-fusion stage is composed of the fusion of adjacent nodes in the second layer of the cross-fusion stage using an adaptive feature interaction module;
[0022] Preferably, the adaptive feature interaction module upsamples the input low-level feature and adds it to the input high-level feature element-wise to obtain a rough feature. The rough feature is processed through a convolution module and a Sigmoid activation function to generate a focus weight mask, and the focus weight mask is multiplied by the rough feature to obtain a spatially information-focused feature;
[0023] The spatially information-focused feature is divided along the channel dimension and processed by global max pooling and global average pooling respectively and then concatenated. After scaling and adjustment, it is normalized through the Softmax function and multiplied element-wise with the weight mask obtained by processing the spatially information-focused feature using a convolution module to obtain the output feature of the adaptive feature interaction module.
[0024] Preferably, the loss function in the training process of the display defect detection model uses the shape perception loss function L M-CIoU :
[0025]
[0026] where A max is the area of the largest prediction box, A maxis the minimum predicted bounding box area, B is the ground truth bounding box, α is a positive constant used to adjust the weight of the area ratio, ∩ is the intersection area, σ is the Sigmoid activation function, M is the shape loss penalty coefficient, IoU represents the intersection over union of the ground truth bounding box and the predicted bounding box, b, b gt represent the center points of the predicted bounding box and the ground truth bounding box respectively, ρ represents the Euclidean distance, ρ 2 (b, b gt ) represents the square of the distance between the centers of the ground truth bounding box and the predicted bounding box, c represents the diagonal length of the smallest bounding box enclosing the ground truth bounding box and the predicted bounding box, and v represents the consistency of the aspect ratios of the two bounding boxes.
[0027] In a second aspect of the present invention, there is provided a display defect detection system based on frequency perception guidance, including:
[0028] A data acquisition module configured to acquire an image to be detected;
[0029] A defect detection module configured to input the image to be detected into a trained display defect detection model for detection to obtain a display defect detection result;
[0030] Wherein, the display defect detection model includes a feature extraction network, a cross-fusion network, and a detection layer connected in sequence. The feature extraction network includes multiple layers of feature extraction structures. Each layer of feature extraction structure includes a split heterogeneous feature extraction module and a frequency perception mechanism. The input feature of each feature extraction structure is obtained by downsampling the output of the split heterogeneous feature extraction module in the previous layer; the split heterogeneous feature extraction module is used to extract defect perception features in the input feature through split convolution, and the frequency perception mechanism is used to extract high-frequency features and low-frequency features in the input feature, and fuse the low-frequency features with the defect perception features to obtain guided perception features; the cross-fusion network is used to perform cross-fusion on the concatenated features obtained by concatenating the guided perception features and the high-frequency features; the detection layer is used to detect the features after cross-fusion by the cross-fusion network to obtain a defect detection result.
[0031] In a third aspect of the present invention, there is provided an electronic device, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to complete the steps of the above-mentioned display defect detection method based on frequency perception guidance.
[0032] In a fourth aspect of the present invention, there is provided a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the steps of the above-mentioned display defect detection method based on frequency perception guidance are completed.
[0033] In a fifth aspect of the present invention, there is provided a computer program product comprising computer programs / instructions which, when executed by a processor, implement the steps of the above-mentioned method for detecting display screen defects based on frequency perception guidance.
[0034] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0035] The present invention provides a method and system for detecting display screen defects based on frequency perception guidance, designs a frequency perception mechanism, increases the skip connections between low frequencies and high frequencies, and improves the ability to detect complex small target defects. The low-frequency features extracted by the frequency perception mechanism are spliced and fused with the defect perception features extracted by the split heterogeneous feature extraction module to enhance the processing of defect images with complex textures and structures and retain the detailed information of the images. The high-frequency features of the frequency perception mechanism are input into the cross-fusion network to improve the model's ability to detect complex small target defects. At the same time, while maintaining the extraction of tiny feature information, the frequency information of the features is effectively extracted and utilized.
[0036] By designing a split heterogeneous feature extraction module, the present invention uses split convolution operations to reduce the dimension of data, extracts useful defect perception features while removing redundant information, enhances the ability of defect feature modeling and expression, improves the detection accuracy of the model through multi-level feature fusion, and significantly reduces the number of parameters, thereby improving the detection efficiency on the premise of ensuring the detection accuracy.
[0037] The present invention designs a feature enhancement and strengthening module to enable the network to focus on the local detailed textures in the global features, obtain global perception features, and improve the detection accuracy of display screen defects; the feature enhancement branch enhances the ability of defect feature representation and gives the model control over features of different channels; while reducing the calculation and the number of parameters, the fine-grained feature perception branch maintains the effective extraction ability of the input features and improves the model's ability to identify and classify defects.
[0038] The present invention designs a cross-fusion network to make full use of the low-frequency and high-frequency features of the frequency perception mechanism, effectively extract and utilize the frequency information of the features; designs a three-layer cross-fusion network with decreasing order to fully fuse the defect features extracted by adjacent layers, reduce redundant information, and improve the detection of defects of different scales; uses the designed adaptive feature interaction module to fuse features of different scales, capture the details and features of defect targets at different scales, improve the adaptability of the model at different scales, and generate adaptive defect features.
[0039] The present invention designs an L M-CIoU loss function to sense the shape loss of different defects and improve the detection of defects with various shapes on the display screen.
[0040] Advantages of additional aspects of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not unduly limit the present invention.
[0042] Figure 1 is a flowchart of the method in Embodiment 1 of the present invention;
[0043] Figure 2 is a schematic diagram of the overall structure of the display defect detection model in Embodiment 1 of the present invention;
[0044] Figure 3 is a schematic diagram of the split isomerism feature extraction module;
[0045] Figure 4 is a schematic diagram of the frequency perception mechanism;
[0046] Figure 5 is a schematic diagram of the feature enhancement and strengthening module;
[0047] Figure 6 is a schematic diagram of the cross-fusion network;
[0048] Figure 7 is a schematic diagram of the adaptive feature interaction module. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0049] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0050] Embodiment 1
[0051] Embodiment 1 of the present invention provides a display defect detection method guided by frequency perception, as Figures 1 - 7 shown, including:
[0052] Obtain an image to be detected;
[0053] Input the image to be detected into a trained display defect detection model for detection to obtain a display defect detection result;
[0054] Among them, the display screen defect detection model includes a feature extraction network, a cross-fusion network, and a detection layer connected in sequence. The feature extraction network includes multiple layers of feature extraction structures. Each layer of feature extraction structure includes a split heterogeneous feature extraction module and a frequency perception mechanism. The input feature of each feature extraction structure is obtained by downsampling the output of the split heterogeneous feature extraction module in the previous layer. The split heterogeneous feature extraction module is used to extract defect perception features in the input feature through split convolution. The frequency perception mechanism is used to extract high-frequency features and low-frequency features in the input feature, and fuse the low-frequency features with the defect perception features to obtain guided perception features. The cross-fusion network is used to perform cross-fusion on the concatenated features obtained by concatenating the guided perception features and the high-frequency features. The detection layer is used to detect the features after cross-fusion by the cross-fusion network to obtain the defect detection result.
[0055] Existing deep learning-based display screen defect detection methods, although they can automatically learn image features and perform complex object detection, still have the following problems:
[0056] (1) Existing object detection models cannot accurately detect display screen defects with various shapes.
[0057] (2) Existing detectors have poor detection effects on defects with significant changes between different types. The general detection model cannot effectively fuse features of adjacent layers, and redundant information is introduced during the fusion process as a large number of defect features are not selected.
[0058] (3) Liquid crystal display screens contain a large number of small targets. The defective targets have fewer pixels, the features are not obvious, and there is background noise interference.
[0059] (4) Existing loss functions have difficulties in detecting diverse-shaped defects in display screens.
[0060] Based on this, in this embodiment, a display screen defect detection model is built to achieve efficient and accurate display screen defect detection. First, obtain the display screen defect image, preprocess the image, and divide the data set into a training set, a test set, and a validation set. Subsequently, build each module in the display screen defect detection model. After completing the model building, use the designed shape perception loss function to train the model, and select the model with the optimal weights for deployment to be used for the display screen defect detection task in the industrial field.
[0061] The display screen defect detection model includes a feature extraction network, a cross-fusion network, and a detection layer connected in sequence, as Figure 2As shown in the figure. The feature extraction network includes multiple layers of feature extraction structures, which are five layers in this embodiment. The feature extraction network realizes feature extraction through a frequency perception mechanism and a split heterogeneous feature extraction module. The frequency perception mechanism is used to extract low-frequency features and high-frequency features. The low-frequency features include the details and texture features of defects. The split heterogeneous feature extraction module is used to reduce the dimension of the data, extract useful defect perception features while removing redundant information. Then, the low-frequency features extracted by the frequency perception mechanism are spliced and fused with the defect perception features extracted by the split heterogeneous feature extraction module to obtain the fused guiding perception features, enhancing the processing of defect images with complex textures and structures and retaining the detail information of the images. In addition, at the end of the feature extraction network, that is, in the fifth layer of feature extraction structure, a feature enhancement and strengthening module is designed to make the network focus on the local detail textures in the global features and obtain the global perception features. Then, the obtained guiding perception features are input into the cross-fusion network, and the high-frequency features of the frequency perception mechanism are fused with the guiding perception features to obtain high-frequency aggregation features, improving the model's ability to detect complex small target defects, realizing more efficient feature integration, and further improving the prediction accuracy of the model. The high-frequency aggregation features and the guiding perception features are input into the cross-fusion network, and a three-layer cross-fusion network with decreasing layers is adopted to fully fuse the defect features extracted from adjacent layers, reduce redundant information, improve the detection ability for defects of different scales, use an adaptive feature interaction module to fuse features of different scales, capture the details and features of defect targets at different scales, improve the adaptability of the model at different scales, and generate adaptive defect features; the adaptive defect features are input into the detection layer to obtain the defect detection results.
[0062] Specifically, the feature extraction network consists of five layers of feature extraction structures. The first four layers (D1 layer to D4 layer) include a split heterogeneous feature extraction module and a frequency perception mechanism, and the fifth layer (D5 layer) includes a split heterogeneous feature extraction module, a feature enhancement and strengthening module, and a frequency perception mechanism. The input of each layer is the output of the split heterogeneous feature extraction module of the previous layer, and each layer downsamples the input feature map by a factor of two.
[0063] The split heterogeneous feature extraction module is used to reduce the dimension of the data, extract useful defect perception features while removing redundant information, and obtain defect perception features.
[0064] Such as Figure 3As shown in the figure, the split heterogeneous feature extraction module first passes through a convolution module with a convolution kernel size of 1x1 to reduce the number of channels and the number of model parameters, obtaining feature map A1. Feature map A1 is divided into two branches. The first branch divides feature map A1 into two groups, A and B, along the channel dimension. The features in group A are transformed using a convolution module with a convolution kernel size of 1×3, and the features in group B are transformed using a convolution module with a convolution kernel size of 3×1 to improve the robustness of spatial feature extraction. Secondly, the transformed features are added element-wise to achieve information transfer and fusion within and between channels. The transformed group A features and group B features are concatenated and fused to restore the feature dimension, obtaining feature map A2. Finally, a residual connection is established between feature map A2 and feature map A1 and added element-wise, and then passed through the SiLU activation function to accelerate the model convergence, and the output is feature map A3. The convolution module consists of convolution, batch normalization layer, and SiLU activation function. The second branch passes feature map A1 through a convolution module with a convolution kernel size of 1x1 to reduce the number of channels, and the output is feature map A4. A residual connection is established between feature map A3 output by the first branch and feature map A4 output by the second branch and multiplied element-wise, which helps the gradient to be more effectively backpropagated during training. Finally, a 1x1 convolution layer is used to further mix the feature channels. The design of the split heterogeneous feature extraction module enhances the ability of feature modeling and expression, improves the detection accuracy through multi-level feature fusion, and significantly reduces the number of parameters, thus improving the model training efficiency.
[0065] The specific data processing process of the split heterogeneous feature extraction module can be expressed by the following formula:
[0066] A,B = Split(Conv1(f));
[0067]
[0068] where f is the feature map input to the split heterogeneous feature extraction module, Split is the splitting operation, SiLU is the activation function, Cat is the concatenation operation, is the element-wise addition, is the element-wise multiplication, Conv1 is the convolution module with a convolution kernel size of 1x1, Conv 1×3 is the convolution module with a convolution kernel size of 1×3, Conv 3×1 is the convolution module with a convolution kernel size of 3×1, A and B are the feature maps divided along the channel dimension; Q1 is the feature map obtained by the first branch of the split heterogeneous feature extraction module, and Q2 is the feature map output by the split heterogeneous feature extraction module.
[0069] The frequency perception mechanism is used to extract low-frequency and high-frequency features from the input feature map.
[0070] The feature map input by the frequency perception mechanism is the same as the feature map input by the corresponding split heterogeneous feature extraction module of this layer. As Figure 4 shown, first, the input feature map is subjected to wavelet transform, decomposed into components of different frequencies, and the high-frequency and low-frequency features in the data are extracted to enhance the representation ability of the features. Secondly, the features after wavelet transform are processed through four different paths, and each path corresponds to a specific frequency component: the low-frequency component LL (low frequency - low frequency), the low-high frequency component LH (low frequency - high frequency), the high-low frequency component HL (high frequency - low frequency), and the high-high frequency component HH (high frequency - high frequency). In the LL path, the features pass through a depth convolution module with a convolution kernel size of 3x3 and a convolution module with a convolution kernel size of 1x1 to extract the refined features of the defects, and then are normalized through the Softmax layer, and a residual edge is established with the low-frequency component LL for element-wise multiplication to accelerate the training of the network to extract low-frequency features. Secondly, the low-high frequency component LH, the high-low frequency component HL, and the high-high frequency component HH are concatenated and fused to obtain the fused feature F a , and then processed through a convolution module with a convolution kernel size of 1x1 and the Softmax layer, so that the network can identify and enhance the frequency components containing important texture and detail features, and establish a residual edge with the fused feature F a for element-wise multiplication to accelerate the training of the network to extract high-frequency features.
[0071] Finally, the low-frequency features are fused with the split heterogeneous feature extraction module of the feature extraction network to obtain the guided perception features, and the high-frequency features are fused with the guided perception features to obtain the high-frequency aggregation features, which improve the model's ability to detect complex small target defects. At the same time, while maintaining the extraction of tiny feature information, the frequency information of the features is effectively extracted and utilized.
[0072] The specific process of the frequency perception mechanism can be expressed by the following formula:
[0073] {LL, LH, HL, HH} = WT(x);
[0074]
[0075] where x is the feature map input by the frequency perception mechanism, is element-wise multiplication, Conv1 is a convolution module with a convolution kernel size of 1x1, Conv3 is a convolution module with a convolution kernel size of 3x3, Cat is the concatenation operation, Softmax is the activation function, LL, LH, HL, and HH are the low-low frequency, low-high frequency, high-low frequency, and high-high frequency components respectively, F1 is the low-frequency feature, F2 is the high-frequency feature, and WT() represents wavelet transform.
[0076] As Figure 2 , Figure 5As shown in the figure, the feature enhancement module is set in the fifth-layer feature extraction structure of the feature extraction network, and is used to enhance the output features of the split heterogeneous feature extraction module, so that the network focuses on the local detail textures in the global features and obtains global perception features.
[0077] The input of the feature enhancement module is the output of the split heterogeneous feature extraction module in the fifth-layer feature extraction structure. For the input features, they are first grouped by grouped convolution to obtain the feature map G n (n is the number of channels of the input feature map), and the input features are divided into n groups of channels using the defined learning rate hyperparameter. Then each channel passes through two branches, namely the feature enhancement branch and the fine-grained feature perception branch.
[0078] In the feature enhancement branch, first, the grouped feature map G n passes through adaptive max pooling and adaptive average pooling to achieve per-channel feature aggregation to capture the spatial information between different channels. The feature map G n passes through adaptive max pooling to reduce the dimension of the feature map, reduce the computational amount, and at the same time retain important feature information. Passing through adaptive average pooling retains the feature distribution of the feature map while reducing the dimension, avoiding the loss of key information and improving the defect detection accuracy. Then, the features passing through the adaptive max pooling layer and the adaptive average pooling layer are concatenated and fused to obtain the feature map G1. The fused feature map G1 passes through a 1x1 convolutional layer to reduce the number of channels, and then is passed to the depth convolutional layer and the 1x1 convolutional layer to perform a compression operation on the spatial dimension of the generated feature map. Then, an attention score is generated through the Sigmoid function, and the generated attention score is multiplied by the grouped feature map G n to highlight the key information within each channel group, obtaining the feature map G2. Finally, pointwise convolution is used to perform deep fusion on the features within the channel group to achieve information interaction and feature enhancement between channels, obtaining the feature map G3. This process not only enhances the ability to represent defect features, but also endows the model with the control of different channel features.
[0079] The fine-grained feature perception branch is composed of depthwise separable layers, including a depth convolutional module with a convolution kernel size of 3x3 and a pointwise convolutional module with a convolution kernel size of 1x1. While reducing the calculation and the number of parameters, it maintains the effective extraction ability of the input features, and the output obtains the feature map G4. Depth convolution is responsible for feature extraction within the channel, while pointwise convolution acts as the role of feature fusion and mapping, improving the model's ability to identify and classify defects.
[0080] Finally, the feature map G3 obtained by the feature enhancement branch and the feature map G4 obtained by the fine-grained feature perception branch are added and fused element by element to enhance the model's extraction of defect features and improve the defect detection capability, thus obtaining the feature map G5. This process can be expressed as:
[0081] B1 = GConv(B);
[0082]
[0083] B3 = PConv1(DWConv3(B1));
[0084]
[0085] Among them, B is the feature map input by the feature enhancement module, GConv is the group convolution operation, AMP is the adaptive maximum pooling, AAP is the adaptive average pooling, PConv1 is the point-by-point convolution module with a convolution kernel size of 1x1, and DWConv3 is the depth convolution module with a convolution kernel size of 3x3. is an element-by-element addition operation, is an element-by-element multiplication operation, Cat is a concatenation operation, and S is a Sigmoid activation function.
[0086] The cross-fusion network includes a three-layer fusion stage network with descending cross-fusion nodes. The first cross-fusion stage is composed of the splicing and fusion of the high-frequency features of the guided perception feature and the frequency perception mechanism, the second cross-fusion stage is composed of the fusion of the adjacent nodes of the first cross-fusion stage using the adaptive feature interaction module, and the third cross-fusion stage is composed of the fusion of the adjacent nodes of the second cross-fusion stage using the adaptive feature interaction module. This process combines the three-layer lightweight network structure that descends in sequence with the multi-scale feature fusion technology to fully integrate the multi-scale defect features extracted by the backbone network and realize efficient recognition of multi-scale defects.
[0087] The adaptive feature interaction module adaptively fuses feature maps of different sizes to provide more comprehensive and rich visual context.
[0088] like Figure 7 As shown, the specific steps of the adaptive feature interaction module for adaptive fusion include:
[0089] The first step is spatial information focusing. First, the adaptive feature interaction module receives high-level (High) and low-level (Low) features as inputs. It upsamples the low-level features to twice the feature size of the high-level features, performs an element-wise addition operation on the high-level features and the upsampled low-level features to obtain a rough feature S1. Then, the fused rough feature S1 is passed through a convolution module with a kernel size of 1x1 to adjust the number of channels and fuse the features. A convolution module with a kernel size of 3x3 is used to further extract more detailed defect features, and a Sigmoid activation function is used to generate a focusing weight mask. The generated focusing weight mask is multiplied by the rough feature S1 to adjust the rough feature and obtain the spatial information focused feature S2.
[0090] The second step is feature fusion, which includes two branches. The first branch divides the feature S2 into m groups along the channel dimension to obtain the feature S m , and the grouped feature S m is respectively passed through global max pooling and global average pooling to capture feature information at different scales. The pooled features are concatenated to form a feature that combines multiple scale information. The processed feature is adjusted in feature size through a scaling operation, and then normalized through a Softmax function to obtain the feature S3. The second branch performs a convolution operation using a convolution module with a kernel size of 1×1, and performs refined processing on the similar channel features within each group, thereby generating a weight mask for the attention mechanism. This mask can identify and reflect the mutual relationships and dependencies between the channel features. In this way, the model can pay more attention to those feature channels that are more critical for task completion, thereby improving the quality of feature representation. Then, the weight mask is applied to the refined feature S3, multiplied element-wise with the feature S3, and finally the features of each group are aggregated to obtain the feature S4. This process can be expressed as:
[0091]
[0092] F m = GConv(F1);
[0093]
[0094] where S is the Sigmoid activation function, Conv1 is the convolution module with a kernel size of 1x1, Conv3 is the convolution module with a kernel size of 3x3, UP is the upsampling operation by a factor of two, F H is the high-level feature, F L is the low-level feature, GConv is the grouped convolution, F m is the m groups of features obtained through grouped convolution, is the element-wise addition, Element-wise multiplication, Cat is the concatenation operation, Softmax is the activation function, RS is the scaling operation, GMP is the global max pooling, and GAP is the global average pooling.
[0095] As Figure 6 shown, in the cross-fusion network, the first cross-fusion stage consists of nodes A1, A2, A3, A4, and A5, where each node is formed by splicing and fusing the guiding perception features and the high-frequency features of the frequency perception mechanism. Specifically:
[0096] The A1 node fuses the output feature E1 obtained by splicing and fusing the output feature of the split heterogeneous feature extraction module in the first-layer feature extraction structure (D1) and the low-frequency feature of the frequency perception mechanism with the high-frequency feature of the frequency perception mechanism, and outputs the feature Q1.
[0097] The A2 node fuses the output feature E2 obtained by splicing and fusing the output feature of the split heterogeneous feature extraction module in the second-layer feature extraction structure (D2) and the low-frequency feature of the frequency perception mechanism with the high-frequency feature of the frequency perception mechanism, and outputs the feature Q2.
[0098] The A3 node fuses the output feature E3 obtained by splicing and fusing the output feature of the split heterogeneous feature extraction module in the third-layer feature extraction structure (D3) and the low-frequency feature of the frequency perception mechanism with the high-frequency feature of the frequency perception mechanism, and outputs the feature Q3.
[0099] The A4 node fuses the output feature E4 obtained by splicing and fusing the output feature of the split heterogeneous feature extraction module in the fourth-layer feature extraction structure (D4) and the low-frequency feature of the frequency perception mechanism with the high-frequency feature of the frequency perception mechanism, and outputs the feature Q4.
[0100] The A5 node is the feature obtained by splicing and fusing the output feature of the feature enhancement and strengthening module and the high-frequency feature of the frequency perception mechanism. Specifically, the output feature O of the split heterogeneous feature extraction module in the fifth-layer feature extraction structure (D5) A , input the feature O A into the feature enhancement and strengthening module to obtain the feature O B , and fuse the feature O after passing through the feature enhancement and strengthening module B with the low-frequency feature of the frequency perception mechanism to output the feature E5, and fuse the feature E5 with the high-frequency feature of the frequency perception mechanism to output the feature Q5.
[0101] The second cross-fusion stage consists of nodes C1, C2, C3, and C4, where each node is formed by fusing adjacent nodes using the adaptive feature interaction module. Specifically,
[0102] The C1 node is composed of the feature Q4 and the feature Q5 fused by an adaptive feature interaction module, and outputs the feature X1.
[0103] The C2 node is composed of the feature Q3 and the feature Q4 fused by an adaptive feature interaction module, and outputs the feature X2.
[0104] The C3 node is composed of the feature Q2 and the feature Q3 fused by an adaptive feature interaction module, and outputs the feature X3.
[0105] The C4 node is composed of the feature Q1 and the feature Q2 fused by an adaptive feature interaction module, and outputs the feature X4.
[0106] The third-layer cross-fusion stage is composed of the M1, M2, and M3 nodes, where each node is composed of adjacent nodes fused by an adaptive feature interaction module. Specifically,
[0107] The M1 node is composed of the feature X3 and the feature X4 fused by an adaptive feature interaction module, and outputs the feature Y1.
[0108] The M2 node is composed of the feature X2 and the feature X3 fused by an adaptive feature interaction module, and outputs the feature Y2.
[0109] The M3 node is composed of the feature X1 and the feature X2 fused by an adaptive feature interaction module, and outputs the feature Y3.
[0110] After the display defect detection model is built, the designed shape perception loss function is used for model training, and the model with the optimal weight is selected for deployment to be used for the display defect detection task in the industrial field.
[0111] Shape perception loss L M-CIoU Consists of the CIoU loss function and the shape loss penalty coefficient M, used to perceive the shape loss of different defects and improve the detection of various-shaped defects on the display. The shape perception loss function L M-CIoU Is expressed by the following formula:
[0112]
[0113] Among them, A max Is the area of the largest prediction box, A max Is the area of the smallest prediction box, B is the ground truth box, α is a positive constant used to adjust the weight of the area ratio, ∩ is the intersection area, σ is the Sigmoid activation function, and M is the shape loss penalty coefficient.
[0114]
[0115] Among them, IoU represents the intersection over union of the ground truth box and the prediction box, b, bgt represent the center points of the predicted bounding box and the ground truth bounding box respectively, ρ(x) represents the Euclidean distance, ρ 2 (b,b gt ) represents the square of the distance between the center of the ground truth bounding box and the predicted bounding box, c represents the diagonal length of the smallest bounding box enclosed by the ground truth bounding box and the predicted bounding box, w gt 、h gt represent the width and height of the ground truth bounding box respectively, w and h represent the width and height of the predicted bounding box respectively.
[0116] In this embodiment, by designing a split heterogeneous feature extraction module, the split convolution operation is used to reduce the dimension of the data, extract useful defect perception features while removing redundant information, and enhance the ability of defect feature modeling and expression.
[0117] In this embodiment, a frequency perception mechanism is designed to increase the skip connection between low frequencies and high frequencies, improving the ability to detect complex small target defects. The low-frequency features extracted by the frequency perception mechanism are spliced and fused with the defect perception features extracted by the split heterogeneous feature extraction module to enhance the processing of defect images with complex textures and structures and retain the detailed information of the images. The high-frequency features of the frequency perception mechanism are input into the feature fusion network to improve the model's ability to detect complex small target defects.
[0118] In this embodiment, a feature enhancement and strengthening module is designed, including a feature enhancement branch and a fine-grained feature perception branch, enabling the network to focus on the local detailed textures in the global features, obtaining global perception features, and improving the accuracy of display screen defect detection; the feature enhancement branch enhances the ability to represent defect features and endows the model with control over different channel features; the fine-grained feature perception branch reduces the calculation and the number of parameters while maintaining the effective extraction ability of the input features, and improves the model's ability to identify and classify defects.
[0119] In this embodiment, a cross-fusion network is designed to make full use of the low-frequency and high-frequency features of the frequency perception mechanism, effectively extract and utilize the frequency information of the features; a three-layer cross-fusion network with decreasing order is designed to fully fuse the defect features extracted by adjacent layers, reduce redundant information, and improve the detection of defects at different scales; the designed adaptive feature interaction module is used to fuse features at different scales, capture the details and features of defect targets at different scales, improve the model's adaptability at different scales, and generate adaptive defect features. Design L M-CIoU loss function to sense the shape loss of different defects and improve the detection of various-shaped defects on the display screen.
[0120] Embodiment 2
[0121] This embodiment provides a display screen defect detection system based on frequency perception guidance, including:
[0122] A data acquisition module, configured to acquire an image to be detected;
[0123] A defect detection module, configured to input the image to be detected into a trained display defect detection model for detection to obtain a display defect detection result;
[0124] Wherein, the display defect detection model includes a feature extraction network, a cross-fusion network and a detection layer connected in sequence. The feature extraction network includes multiple layers of feature extraction structures. Each layer of feature extraction structure includes a split heterogeneous feature extraction module and a frequency perception mechanism. The input feature of each feature extraction structure is obtained by downsampling the output of the split heterogeneous feature extraction module in the previous layer; the split heterogeneous feature extraction module is used to extract defect perception features in the input feature through split convolution, and the frequency perception mechanism is used to extract high-frequency features and low-frequency features in the input feature, and fuse the low-frequency features with the defect perception features to obtain guided perception features; the cross-fusion network is used to perform cross-fusion on the concatenated features obtained by concatenating the guided perception features and the high-frequency features; the detection layer is used to detect the features after cross-fusion by the cross-fusion network to obtain a defect detection result.
[0125] It should be noted here that each module in this embodiment corresponds to the steps of the method in Embodiment 1 one by one, and the specific implementation process is the same, so it will not be repeated here.
[0126] Embodiment 3
[0127] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to complete the steps of the method in Embodiment 1.
[0128] The processor may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor may also include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0129] The memory may include one or more computer-readable media, and the computer-readable media may be non-transitory. The memory may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable media in the memory is used to store at least one computer program, and the at least one computer program is used to be executed by the processor to implement a method for detecting display screen defects based on frequency-aware guidance provided in the embodiments of the present disclosure.
[0130] Those skilled in the art can understand that the electronic device provided in this embodiment may include more or fewer components, or combine certain components, or adopt different component arrangements.
[0131] Embodiment 4
[0132] This embodiment provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the steps of the method in Embodiment 1 are completed.
[0133] Embodiment 5
[0134] This embodiment provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of the method in Embodiment 1 are implemented.
[0135] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A display defect detection method based on frequency perception guidance, characterized in that Including: Obtain the image to be detected; Input the image to be detected into the trained display defect detection model for detection to obtain the display defect detection result; Among them, the display defect detection model includes a feature extraction network, a cross-fusion network, and a detection layer connected in sequence. The feature extraction network includes multiple layers of feature extraction structures. Each layer of feature extraction structure includes a split heterogeneous feature extraction module and a frequency perception mechanism. The input feature of each feature extraction structure is obtained by downsampling the output of the split heterogeneous feature extraction module in the previous layer. The split heterogeneous feature extraction module is used to extract defect perception features in the input feature through split convolution. The frequency perception mechanism is used to extract high-frequency features and low-frequency features in the input feature, and fuse the low-frequency feature and the defect perception feature to obtain a guiding perception feature. The cross-fusion network is used to perform cross-fusion on the concatenated feature obtained by concatenating the guiding perception feature and the high-frequency feature. The detection layer is used to detect the feature after cross-fusion by the cross-fusion network to obtain the defect detection result.
2. The method for detecting display screen defects based on frequency perception guidance according to claim 1, wherein The feature extraction network includes five layers of feature extraction structures. The first four layers of feature extraction structures include a split heterogeneous feature extraction module and a frequency perception mechanism. The fifth layer of feature extraction structure includes a split heterogeneous feature extraction module, a feature enhancement and strengthening module, and a frequency perception mechanism; The feature enhancement and strengthening module includes a feature enhancement branch and a fine-grained feature perception branch. The feature enhancement branch groups the input feature map according to the number of channels of the input feature map, splices and fuses the grouped feature maps after adaptive max pooling and adaptive average pooling processing, and passes through a convolution module, a depth convolution module, and a convolution module in sequence. After processing, an attention score is generated through the Sigmoid function, the attention score is multiplied by the grouped feature map, and finally a pointwise convolution is used to obtain the feature map output by the feature enhancement branch; The fine-grained feature perception branch passes the input feature map through a depth convolution module and a pointwise convolution module in sequence to obtain the feature map output by the fine-grained feature perception branch; The feature map output by the feature enhancement branch and the feature map output by the fine-grained feature perception branch are added and fused element by element to obtain the feature map output by the feature enhancement and strengthening module.
3. The method for detecting display screen defects based on frequency perception guidance according to claim 1, characterized in that, The split heterogeneous feature extraction module processes the input feature map using two branches. The first branch divides the input feature map and performs feature transformation on it through a convolution module respectively, then performs element-by-element addition, splicing and fusion, and establishes a residual edge for element-by-element addition in sequence. After activation by an activation function, it multiplies element by element with the feature map output by the second branch processed by the convolution module to establish a residual edge.
4. The method for detecting display screen defects based on frequency perception guidance according to claim 1, characterized in that The frequency perception mechanism obtains the low-frequency component, low-high frequency component, high-low frequency component, and high-high frequency component of the input feature map through wavelet transform; The low-frequency component is normalized after passing through a depth convolution module with a convolution kernel size of 3x3 and a convolution module with a convolution kernel size of 1x1 in sequence, and is multiplied element by element with the low-frequency component to establish a residual edge and then output the low-frequency feature; The low-high frequency components, high-low frequency components, and high-high frequency components are concatenated and fused to obtain a fused feature. The fused feature is processed through a convolutional module with a kernel size of 1x1 and a Softmax layer, and a residual connection is established with the fused feature for element-wise multiplication and then the high-frequency feature is output.
5. The method for detecting display screen defects based on frequency perception guidance according to claim 1, characterized in that The cross-fusion network includes three fusion stage networks with gradually decreasing cross-fusion nodes; the first layer of the cross-fusion stage is formed by concatenating and fusing the high-frequency features of the guided perception feature and the frequency perception mechanism, the second layer of the cross-fusion stage is composed of adjacent nodes in the first layer of the cross-fusion stage fused by an adaptive feature interaction module, and the third layer of the cross-fusion stage is composed of adjacent nodes in the second layer of the cross-fusion stage fused by an adaptive feature interaction module; Preferably, the adaptive feature interaction module upsamples the input low-level feature and adds it to the input high-level feature element-wise to obtain a rough feature. The rough feature is processed through a convolutional module and a Sigmoid activation function to generate a focused weight mask, and the focused weight mask is multiplied by the rough feature to obtain a spatially information focused feature; The spatially information focused feature is divided along the channel dimension and processed by global max pooling and global average pooling respectively and then concatenated. After scaling and adjustment, it is normalized by a Softmax function and multiplied element-wise by the weight mask obtained by processing the spatially information focused feature with a convolutional module to obtain the output feature of the adaptive feature interaction module.
6. The method for detecting display screen defects based on frequency perception guidance according to claim 1, wherein, The loss function in the training process of the display defect detection model adopts the shape perception loss function L M-CIoU : Among them, A max is the area of the largest prediction box, A max is the area of the smallest prediction box, B is the ground truth box, α is a positive constant used to adjust the weight of the area ratio, ∩ is the intersection area, σ is the Sigmoid activation function, M is the shape loss penalty coefficient, IoU represents the intersection over union of the ground truth box and the prediction box, b, b gt represent the center points of the prediction box and the ground truth box respectively, ρ represents the Euclidean distance, ρ 2 (b, b gt ) represents the square of the distance between the centers of the ground truth box and the prediction box, c represents the diagonal length of the smallest bounding box formed by the ground truth box and the prediction box, and v represents the consistency of the aspect ratios of the two boxes.
7. A display defect detection system based on frequency perception guidance, characterized in that It includes: A data acquisition module configured to acquire an image to be detected; A defect detection module configured to input the image to be detected into a trained display defect detection model for detection to obtain a display defect detection result; Among them, the display defect detection model includes a feature extraction network, a cross-fusion network, and a detection layer connected in sequence. The feature extraction network includes multiple layers of feature extraction structures. Each layer of the feature extraction structure includes a split heterogeneous feature extraction module and a frequency perception mechanism. The input feature of each feature extraction structure is obtained by downsampling the output of the split heterogeneous feature extraction module in the previous layer; the split heterogeneous feature extraction module is used to extract defect perception features in the input feature through split convolution, and the frequency perception mechanism is used to extract high-frequency and low-frequency features in the input feature and fuse the low-frequency feature with the defect perception feature to obtain a guided perception feature; the cross-fusion network is used to perform cross-fusion on the concatenated feature obtained by concatenating the guided perception feature and the high-frequency feature; the detection layer is used to detect the feature after cross-fusion by the cross-fusion network to obtain a defect detection result.
8. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to complete the steps of the method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, For storing computer instructions, when the computer instructions are executed by the processor, the steps of the method according to any one of claims 1-6 are completed.
10. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, the steps of the method according to any one of claims 1-6 are implemented.