Glass defect detection method based on improved YOLOv5

By improving the YOLOv5 model, introducing the dimensionally decoupled feature extraction module DDFM and the corner block attention module CoPA, and constructing the DDC-YOLO model, problems such as diversity, complexity, and high transmittance in glass defect detection are solved, and efficient and accurate industrial real-time detection is achieved.

CN120612307AActive Publication Date: 2025-09-09ANHUI LANSHI GLASS TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510713325.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-09
Estimated Expiration
2045-05-30

AI Technical Summary

Technical Problem

Existing glass defect detection methods are unable to meet the needs of high-speed, high-precision industrial real-time detection when faced with problems such as the diversity, complexity, high transmittance, low contrast and difficulty in detecting small targets of glass defects.

Method used

An improved YOLOv5 model is adopted, and the dimensionally decoupled feature extraction module DDFM and the corner block attention module CoPA are introduced to construct a DDC-YOLO model. By reconstructing the network architecture, the perception ability of low-contrast features, multi-scale defects and small targets is enhanced. Combined with the multi-angle feature extraction and fusion mechanism, an efficient detection framework adapted to the characteristics of glass defects is constructed.

Benefits of technology

It achieves high-precision, low-latency glass defect identification on high-speed production lines, meeting the stringent requirements of industrial-grade glass product quality inspection, and the inspection process takes <200ms per piece.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612307A_ABST
    Figure CN120612307A_ABST
Patent Text Reader

Abstract

The invention discloses a glass defect detection method based on improved YOLOv5, and the method comprises the steps: S1, collecting multispectral image data of a glass product production line through a high-dynamic-range industrial camera, constructing a glass defect detection data set through data processing, and dividing the data set; s2, introducing a dimension decoupling feature extraction module DDFM based on a YOLOv5 model, adding a corner block attention module CoPA, and fusing the DDFM and the CoPA to replace an original C3 module to obtain a DDC-YOLO model; and S3, training and verifying the constructed DDC-YOLO model by adopting the data set, and then detecting the defects of the transparent material in real time by adopting the DDC-YOLO model. By reconstructing a YOLOv5 network architecture and combining with the particularity of glass defect detection, the perception ability of the model to low-contrast characteristics, multi-scale defects and tiny targets is enhanced, and meanwhile, the detection precision and real-time requirements are balanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of defect target detection, and in particular to a glass defect detection method based on improved YOLOv5. Background Art

[0002] Glass, as a key industrial material, is widely used in applications such as architecture, automobiles, electronic displays, and home furnishings. Due to its transparency, weather resistance, and aesthetic appeal, glass products are subject to extremely high quality requirements. However, during the glass production process, various defects, such as bubbles, stones, impurities, scratches, cracks, and optical distortion, inevitably appear on or within the glass due to factors such as raw materials, production processes, and the environment. These defects not only affect the appearance and performance of the glass but can also reduce its mechanical strength and safety performance, leading to product failure and even safety accidents. Therefore, efficient and accurate defect detection during the glass production process is crucial.

[0003] Traditional glass defect detection relies primarily on manual visual inspection or simple optical instruments. Manual inspection is typically performed by experienced workers using strong light or a light source at a specific angle. However, this method is inefficient, labor-intensive, and susceptible to subjective factors, making it difficult to meet the high-speed, high-precision requirements of modern production lines. While optical or laser-based instrumentation (such as transmission, scattering, and interferometry) has improved detection efficiency to a certain extent, it still has limitations for detecting tiny defects, complex backgrounds, or high-transmittance glass. Furthermore, the equipment is costly and lacks adaptability.

[0004] In recent years, with the rapid development of computer vision and deep learning technologies, machine learning-based defect detection methods have gradually become a research hotspot. In particular, object detection algorithms have shown great potential in the field of industrial defect detection. Among them, the YOLO (You Only Look Once) family of algorithms has attracted much attention due to its fast speed and high accuracy. YOLOv5, a classic version of the YOLO family, performs well in general object detection tasks, but it still faces many challenges in the specific scenario of glass defect detection:

[0005] 1) Diversity and complexity of glass defects: Glass defects vary greatly in shape, size, and distribution, such as point bubbles, linear scratches, and planar optical deformation. The traditional YOLOv5 design may not be able to adapt to the detection needs of multi-scale defects.

[0006] 2) High light transmittance and low contrast: The high light transmittance of glass results in low contrast between defects and the background. In particular, subtle defects may only appear as slight grayscale changes in imaging, which places higher demands on the algorithm's feature extraction capabilities.

[0007] 3) Small target detection is difficult: Many glass defects (such as microcracks or tiny bubbles) have a very small pixel ratio. The YOLOv5 network structure easily loses the feature information of small targets during the downsampling process, resulting in missed detection or false detection.

[0008] 4) Industrial real-time requirements: Glass production lines usually require high-speed continuous detection. Ordinary detection algorithms may have difficulty meeting real-time requirements while ensuring accuracy.

[0009] To solve the above problems, the present invention proposes a glass defect detection method based on improved YOLOv5. Summary of the Invention

[0010] The purpose of the present invention is to provide a glass defect target detection method based on improved YOLOv5, which can effectively improve the detection accuracy and speed of glass defects.

[0011] In order to solve the above technical problems, the technical solution adopted by the present invention is: the glass defect detection method based on the improved YOLOv5 includes the following steps:

[0012] S1: Use a high dynamic range industrial camera to collect multispectral image data from a glass product production line. After data processing, a glass defect detection dataset is constructed and then divided.

[0013] S2: Based on the YOLOv5 model, the dimensionally decoupled feature extraction module DDFM is introduced, and the corner block attention module CoPA is added. The dimensionally decoupled feature extraction module DDFM and the corner block attention module CoPA are fused to replace the original C3 module to improve the three-stage network architecture and obtain the DDC-YOLO model;

[0014] S3: Use the glass defect dataset to train and verify the constructed DDC-YOLO model, and then use the DDC-YOLO model to detect transparent material defects in real time.

[0015] By adopting the above technical solution, through reconstructing the network architecture of YOLOv5 and combining the particularity of glass defect detection, the model's perception ability of low-contrast features, multi-scale defects and tiny targets is enhanced, while balancing the detection accuracy and real-time requirements. The introduction of the dimensional decoupling feature extraction module DDFM is to gradually decouple the dimensions from three dimensions, two dimensions to one dimension. Each branch gradually shifts the center of gravity of feature extraction to the specified dimension, reducing the information loss caused by the dimensional change, and finally fuses the output feature maps of the three branches as the output, which belongs to the feature extraction module. In view of the high light transmittance, defect diversity and complexity of industrial scenes of glass, the present invention introduces a multi-angle feature extraction and fusion mechanism and a lightweight attention module, and constructs a set of efficient detection frameworks adapted to the characteristics of glass defects. It can achieve high-precision and low-latency defect recognition in high-speed production lines and meet the stringent requirements of industrial-grade glass product quality inspection.

[0016] Preferably, in step S1, image enhancement processing is performed by dynamic threshold segmentation and reflection suppression algorithm to construct a glass defect detection dataset containing bubbles, scratches, and impurity defects; the specific steps are:

[0017] S11 data acquisition: Industrial cameras are deployed on the glass production line, covering key processes such as cutting, polishing, and tempering. They capture 4K or higher high-resolution images and short-duration videos. The data includes common defects (bubbles, scratches, cracks, and impurities) and covers different lighting conditions (strong light, weak light, and reflections), glass thickness (2-12mm), and surface conditions (transparent, frosted, and coated). The cameras also record production line status, such as normal production, equipment restart after maintenance, or material batch changes, to ensure data diversity.

[0018] S12 data preprocessing: Remove blurry or invalid images, and use dynamic threshold segmentation and reflection suppression algorithms to enhance the image of the retained qualified data. First, use polarized light decomposition to eliminate glass surface reflections: calculate the diffuse reflection component and weaken the specular reflection to reduce ambient light interference and improve the contrast of the defect area. Then, dynamically calculate the defect boundary threshold (mean + 0.5 times the standard deviation), and integrate morphological closing operations to enhance defect continuity. Finally, the segmentation result is weightedly fused with the original image (enhancement factor α = 0.3) to enhance the defect contrast and make small defects more prominent. The enhanced data set is then annotated, using rectangular boxes to mark the defect locations, and the defects are divided into four categories: bubbles, scratches, cracks, and impurities. At the same time, a small number of difficult examples (such as small scratches or translucent cracks) are retained to improve model robustness.

[0019] S13 dataset construction: A stratified sampling strategy was used to ensure a balanced distribution of different glass types (flat / curved) and defect sizes (<1mm / 1-5mm / / >5mm). This dataset was then divided into training, validation, and test sets in a 7:2:1 ratio. The test set must include data from all production line equipment models to verify generalization. To facilitate iterative optimization, data subsets were created by production quarter, and basic metadata (such as ambient temperature and humidity) were recorded.

[0020] Preferably, the improvement of the three-stage network architecture to construct the DDC-YOLO model in step S2 specifically includes:

[0021] S21: Introduce the dimension decoupled feature extraction module DDFM into the feature extraction network to perform independent feature modeling on height, width and channel dimensions;

[0022] S22: Add the corner block attention module CoPA to the head network and generate a spatial weight map based on the statistical comparison of the four corner regions;

[0023] S23: The dimensionally decoupled feature extraction module DDFM and the corner block attention module CoPA are integrated to construct a dual-correlation collaborative module DC-Block, which replaces the original C3 module and reconstructs the feature pyramid network.

[0024] Adopting this technical solution, a Dimension-Decomposed Feature Extraction Module (DDFM) is introduced. This module strengthens the cross-dimensional correlation of glass textures by independently decoupling the feature representations of height, width, and channel dimensions. Depthwise separable convolutions are not used to replace traditional convolutions within the model. Instead, the three variants of convolution within the module—depthwise, banded, and pointwise convolutions—are simply components of the DDFM. Unlike the existing MCA module, although it also employs a three-branch architecture, it is an attention mechanism. The DDFM in this technical solution processes feature maps differently, both in terms of method and purpose. MCA directly uses pooling to reduce dimensionality from three dimensions to one dimension, then weights the initial input after calculating weights through convolution. A Corner-referenced Patch Attention (CoPA) module is designed to generate spatial attention weights based on statistical comparisons of four corner regions, precisely enhancing defect feature responses. The YOLO framework is reconstructed by integrating DDFM and CoPA, replacing the original C3 module with a fusion module to achieve decoupled enhancement and dynamic fusion of multi-scale defect features. Through this method, real-time and efficient detection of glass surface defects can be achieved, improving detection accuracy and speed.

[0025] Preferably, the specific steps of implementing dimensional decoupling by the dimension decoupling feature extraction module DDFM in step S21 are:

[0026] S211: The input feature map is rotated along the three dimensions of height, width, and channel to obtain three sets of feature maps.

[0027] S212: performing feature extraction on the three sets of feature maps obtained in step S211 using depthwise convolution, respectively, to extract three two-dimensional correlation features, thereby obtaining three sets of feature maps with two-dimensional correlation features;

[0028] S213: The feature maps obtained in step S213 are then combined in pairs, and their intersection feature dimensions are concatenated accordingly;

[0029] S214: Finally, use band convolution and point convolution to compress the corresponding dimensional features, remove redundant features, and obtain three sets of compressed feature maps;

[0030] S215: The three sets of compressed feature maps are then superimposed and their average is taken and output.

[0031] Preferably, the specific steps of the corner block attention module CoPA module in step S22 generating spatial attention weights are:

[0032] S221 Gridded Feature Analysis: Divide the input feature map into regular grid blocks of n_win×n_win, and calculate the mean μ and standard deviation σ of each grid block;

[0033] S222 Corner reference difference calculation: Select the upper left, upper right, lower left, and lower right corner blocks as reference benchmarks, scan each grid block, and calculate the statistical difference between the current grid block and the grid blocks in the four corners; the formula is:

[0034] Δμ=Σ|μ_current-μ_corner_i|;

[0035] Δσ=Σ|σ_current-σ_corner_i|(i=1~4);

[0036] Among them, μ_current represents the mean value of the current network block scanned when scanning each network block, σ_current is the standard deviation of the current network block; μ_corner_i represents the mean value of the corner block, σ_corner_i represents the standard deviation of the corner block, and i is the corner number; Δμ represents the mean difference between the grid block and the corner block; Δσ represents the standard deviation difference between the grid block and the corner block;

[0037] S223 dynamic weight generation: The difference value is mapped to the weight coefficient through the Sigmoid activation function to obtain the block-level weight W. The formula is:

[0038] W=Sigmoid(Δμ)+Sigmoid(Δσ)-0.5;

[0039] S224 Spatial Attention Enhancement: The block-level weight W is reconstructed into a pixel-level spatial weight map and multiplied point by point with the original feature map, that is, the weight map is weighted onto the input feature map.

[0040] Preferably, in step S23, the dual-correlation collaborative module DC-Block comprises a dual-branch collaborative structure, namely a first branch and a second branch; wherein,

[0041] The first branch performs dimension decoupling feature extraction through the dimension decoupling feature extraction module DDFM, focusing on capturing the cross-dimensional correlation characteristics of the glass texture;

[0042] The second branch generates a spatial attention mask through the corner block attention module CoPA to enhance the characteristic response of edge cracks and small defects;

[0043] The outputs of the first branch and the second branch are cross-fused and then feature reshaped by 1×1 convolution, so that the output dimension is consistent with the input.

[0044] Preferably, the specific steps of step S3 are:

[0045] S31: Use the training set to train the constructed DDC-YOLO model, then use the validation set to verify the constructed DDC-YOLO model, and then use the test set to test its detection ability;

[0046] S32: Deploy the trained DDC-YOLO model to the production line for testing and output the test results in real time.

[0047] Preferably, the specific steps of step S31 are:

[0048] S311 training model: The training set is input into DDC-YOLO, which is improved based on YOLO, for training, which goes through the warm-up phase, main training phase, and fine-tuning phase respectively;

[0049] In the warm-up phase, a linear learning rate is used (the learning rate lr is increased from 1e-5 to 1e-2), some layers of the backbone network are frozen, and only the detection head module is trained. The CIoU loss function is then used to initialize the weights to avoid early gradient instability.

[0050] During the main training phase (epoch 50-450), all network layers are unfrozen and cosine annealing learning rate scheduling is enabled (base_lr = 1e-2, min_lr = 1e-4). A dynamic positive sample allocation strategy is introduced to dynamically adjust the GT allocation threshold based on the match between the predicted box and the true value box.

[0051] During the fine-tuning phase (epoch 450-500), we enabled cross-GPU synchronized batch normalization (SyncBN) and used the exponential moving average (EMA) to update model parameters.

[0052] S312 Validation Model: Configure the loss function and use the validation set to validate the trained DDC-YOLO model to evaluate model performance.

[0053] The S313 test model sets the input image to its original size. If the size is less than an integer multiple of 32, the edges are padded. The improved non-maximum suppression (NMS) algorithm is then used to filter the image using a category-independent IoU threshold (0.7) and a confidence threshold (0.001 initially, adaptively adjusted based on target density). This reduces redundant boxes in target detection and improves detection accuracy.

[0054] Preferably, the loss function configured in step S312 includes three parts of loss, namely positioning loss, classification loss and confidence loss.

[0055] The positioning loss, i.e., the bounding box regression loss, uses the SIoU metric and introduces the angle cost term and the distance cost term to optimize the box regression direction. The formula is:

[0056] L box =1-IoU+(Δ+Ω) / 2;

[0057] Among them, L box is the bounding box regression loss, IoU represents the intersection-over-union ratio between the predicted box and the ground-truth box, Δ represents the corner cost term, and Ω represents the distance cost term;

[0058] The category classification loss is an improved Focal Loss loss function that applies a modulation coefficient of γ to difficult samples. This is an improved version of the cross entropy that reduces the weight of easy-to-classify samples by adjusting the factor to solve the imbalance problem between positive and negative samples. The formula is:

[0059] L cls =-α t *(1-p t ) γ *log(p t );

[0060] Among them, L cls is the category classification loss, αt represents the category balance factor, p t represents the model's predicted probability for the true category, γ is the modulation coefficient; γ = 2.0 is preferred;

[0061] Confidence loss is to increase the dynamic weight factor to alleviate the imbalance problem of positive and negative samples in dense target scenarios; the formula is:

[0062]

[0063] Among them, L obj represents the confidence loss, λ pos represents the dynamic weight of positive samples, λ neg represents the dynamic weight of negative samples, y is the true label (1 for target, 0 for background), The confidence score for the model prediction;

[0064] Total loss function L total For weighted summation, the formula is:

[0065] L total =0.05L box +0.5L cls +L obj .

[0066] Preferably, the specific steps of using the trained model for detection in step S32 are:

[0067] S321 image preprocessing: adjust brightness and contrast to reduce reflection interference; automatically identify glass edges and crop the effective detection area;

[0068] S322 defect detection: It first adopts a two-stage detection, specifically: first quickly scan the entire glass, then carefully inspect the suspicious areas; automatically identify four types of defects: bubbles, cracks, scratches, and impurities;

[0069] S323 result processing: merge adjacent defect frames of the same type, filter out tiny noise (<5 pixels) and normal texture areas; record the defect location, type, and size;

[0070] S324 system linkage: outputs test results to the display device in real time, automatically marks unqualified products and triggers the sorting device.

[0071] The entire process takes less than 200ms per piece, meeting the real-time inspection requirements of the production line. Inspection parameters can be flexibly adjusted according to different glass types (such as flat and curved).

[0072] Compared with the prior art, the present invention has the following characteristics and beneficial effects:

[0073] 1) The algorithm of the present invention designs a dimensionally decoupled feature extraction module (DDFM). By decoupling the height, width, and channel three-dimensional features and extracting them separately, it then integrates the relevant features between each dimension and finally combines the features of each dimension for feature fusion. It gradually transitions from extracting local refined features to integrating global features, thus optimizing the network's feature extraction capability and accelerating the model convergence speed.

[0074] 2) The algorithm in this paper involves the Corner Block Attention (CoPA) module. By using the mean and standard deviation to simulate the normal feature distribution and calculating the inter-block correlation based on this, it locates the distribution of abnormal regions and assigns corresponding weights, thereby enhancing the feature map and improving the model's detection accuracy. This module also implements a parameter-free construction of the attention module, enhancing model performance without increasing the number of parameters. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] Figure 1 Flowchart of the glass defect detection method based on improved YOLOv5 of the present invention;

[0076] Figure 2 This is a model diagram of DDC-YOLO in the glass defect detection method based on improved YOLOv5 of the present invention;

[0077] Figure 3 This is a diagram of the DDFM module in the glass defect detection method based on improved YOLOv5 of the present invention;

[0078] Figure 4 This is a module diagram of the CoPA method for glass defect detection based on improved YOLOv5 of the present invention. DETAILED DESCRIPTION

[0079] To illustrate the technical solution disclosed in the present invention in detail, further description is given below in conjunction with the accompanying drawings.

[0080] Example: Figure 1 As shown, the glass defect detection method based on the improved YOLOv5 includes the following steps:

[0081] S1: Use a high dynamic range industrial camera to collect multispectral image data from a glass product production line. After data processing, a glass defect detection dataset is constructed and then divided.

[0082] In step S1, image enhancement processing is performed by dynamic threshold segmentation and reflection suppression algorithm to construct a glass defect detection dataset containing bubbles, scratches, and impurity defects; the specific steps are:

[0083] S11 data acquisition: Industrial cameras are deployed on the glass production line, covering key processes such as cutting, polishing, and tempering. They capture 4K or higher high-resolution images and short-duration videos. The data includes common defects (bubbles, scratches, cracks, and impurities) and covers different lighting conditions (strong light, weak light, and reflections), glass thickness (2-12mm), and surface conditions (transparent, frosted, and coated). The cameras also record production line status, such as normal production, equipment restart after maintenance, or material batch changes, to ensure data diversity.

[0084] S12 data preprocessing: Remove blurry or invalid images, and use dynamic threshold segmentation and reflection suppression algorithms to enhance the image of the retained qualified data. First, use polarized light decomposition to eliminate glass surface reflections: calculate the diffuse reflection component and weaken the specular reflection to reduce ambient light interference and improve the contrast of the defect area. Then, dynamically calculate the defect boundary threshold (mean + 0.5 times the standard deviation), and integrate morphological closing operations to enhance defect continuity. Finally, the segmentation result is weightedly fused with the original image (enhancement factor α = 0.3) to enhance the defect contrast and make small defects more prominent. The enhanced data set is then annotated, using rectangular boxes to mark the defect locations, and the defects are divided into four categories: bubbles, scratches, cracks, and impurities. At the same time, a small number of difficult examples (such as small scratches or translucent cracks) are retained to improve model robustness.

[0085] S13 Dataset Construction: A stratified sampling strategy was used to ensure a balanced distribution of different glass types (flat / curved) and defect sizes (<1mm / 1-5mm / >5mm). This dataset was then divided into training, validation, and test sets in a 7:2:1 ratio. The test set must include data from all production line equipment models to verify generalization. To facilitate iterative optimization, data subsets were created by production quarter, and basic metadata (such as ambient temperature and humidity) were recorded.

[0086] S2: Based on the YOLOv5 model, the dimensionally decoupled feature extraction module DDFM is introduced, and the corner block attention module CoPA is added. The DDFM and CoPA are fused to replace the original C3 module to improve the three-stage network architecture to obtain the DDC-YOLO model. The introduction of the dimensionally decoupled feature extraction module DDFM gradually decouples dimensions from three dimensions, two dimensions to one dimension. Each branch gradually shifts the center of gravity of feature extraction to the specified dimension, reducing information loss caused by dimensional changes. Finally, the output feature maps of the three branches are fused as the output, which belongs to the feature extraction module.

[0087] This technical solution does not use depthwise separable convolution to replace traditional convolution within the model. Instead, the three variants of convolution within the module, namely depthwise convolution, banded convolution, and pointwise convolution, are only used as components of the DDFM module. Unlike the existing MCA module, although the existing MCA module also uses a three-branch architecture, our DDFM module processes feature maps in a different way and for a different purpose. MCA directly uses pooling to perform dimensionality compression from three dimensions to one dimension, and then weights the initial input after calculating the weights through convolution. It is also an attention mechanism.

[0088] like Figure 2 As shown, the improvement of the three-stage network architecture in step S2 to construct the DDC-YOLO model specifically includes:

[0089] S21: Introducing the dimensionally decoupled feature extraction module (DDFM) into the feature extraction network to independently model the height, width, and channel dimensions. By independently decoupling the feature representation of height, width, and channel dimensions, the cross-dimensional correlation of glass texture is enhanced.

[0090] like Figure 3 As shown, the specific steps of the dimension decoupling feature extraction module DDFM in step S21 to achieve dimensional decoupling are:

[0091] S211 3D axis rotation and reorganization: The input feature map is rotated along the height, width, and channel axes in sequence to obtain three sets of feature maps; specifically:

[0092] Input feature map X_in∈R CxHxW Perform rotation operations along the dimension axes in sequence:

[0093] X_hw=X_in∈R CxHxW ;

[0094] X_hc=Rotate_H(X_in)∈R WxHxC (width and channel dimension permutation);

[0095] X_cw=Rotate_W(X_in)∈R HxCxW (channel and height dimension permutation);

[0096] Among them, X_in represents the input feature map, X_hw represents the input feature map with the original dimension axis maintained; X_hc represents the feature map after X_in is rotated and reorganized along the height axis; X_cw represents the feature map after X_in is rotated and reorganized along the width axis; Rotate_H represents rotation along the height axis, and Rotate_W represents rotation along the width axis;

[0097] ∈R a1xa2xa3Indicates that the channel, height, and width dimensions of the feature map are a1, a2, and a3 respectively;

[0098] S212 Two-dimensional correlation modeling: The three sets of feature maps obtained in step S211 are respectively subjected to feature extraction using deep convolution, and three two-dimensional correlation features are respectively extracted, thereby obtaining three sets of feature maps with two-dimensional correlation features; that is, deep convolution is performed on the three sets of feature maps after rotation while restoring the dimensional order, specifically:

[0099] HW correlation features: F_hw = DWConv3×3(X_hw)∈R CxHxW (capturing height-width correlation);

[0100] HC correlation feature: F_hc = Rotate_H(DWConv3×3(X_hc))∈R CxHxW (capturing height-channel correlation);

[0101] CW correlation feature: F_cw=Rotate_W(DWConv3×3(X_cw))∈R CxHxW (capturing channel-width correlation);

[0102] Among them, DWConv3×3 represents the depth convolution with a convolution kernel size of 3×3; F_hw represents the output feature map that captures the height-width correlation feature, F_hc represents the output feature map that captures the height-channel correlation feature, and F_cw represents the output feature map that captures the channel-width correlation feature;

[0103] S213 Cross Feature Fusion: The feature maps obtained in step S213 are then combined in pairs, and their intersection feature dimensions are concatenated accordingly, specifically:

[0104] H-dimensional concatenation: F_h=Concat(F_hw,F_hc)∈R Cx2HxW ;

[0105] W-dimensional splicing: F_w=Concat(F_hw,F_cw)∈R CxHx2W ;

[0106] C-dimensional concatenation: F_c=Concat(F_hc,F_cw)∈R 2CxHxW ;

[0107] Among them, Concat means concatenating two feature maps in the specified dimension; F_h means the output feature map after concatenating F_hw and F_hc in the height dimension, F_w means the output feature map after concatenating F_hw and F_cw in the width dimension, and F_c means the output feature map after concatenating F_hc and F_cw in the channel dimension.

[0108] S214 feature compression and aggregation: Finally, use banded convolution and point convolution to compress the corresponding dimensional features, remove redundant features, and obtain three sets of compressed feature maps, specifically:

[0109] Out1=Conv 3x1 (F_h)∈R CxHxW ;

[0110] Out2=Conv 1x3 (F_w)∈R CxHxW ;

[0111] Out3=Conv 1x1 (F_c)∈R CxHxW ;

[0112] Among them, Conv 3x1 Represents a banded convolution with a convolution kernel size of 3x1, Conv 1x3 Represents a banded convolution with a convolution kernel size of 1x3, Conv 1x1 Indicates point convolution with a convolution kernel size of 1x1; Out i The output of the feature map after the corresponding convolution processing (i = 1 to 3);

[0113] S215 feature fusion output: The three sets of compressed feature maps are then superimposed and averaged before output. Specifically, the three feature maps obtained are weighted fused with the input feature map through residual connection and output:

[0114] X_out=(W1*Out1⊕W2*Out2⊕W3*Out3⊕W4*X_in);

[0115] Among them, W i represents the learnable weight corresponding to each branch (i=1~4); X_out represents the final output feature map of the module;

[0116] The DDFM module breaks through the spatial constraints of conventional convolution through three-dimensional axis rotation, explicitly modeling the cross-dimensional correlation of height, width, and channel. At the same time, the combination of point convolution and band convolution achieves a balanced optimization of cross-dimensional feature interaction and redundant information filtering.

[0117] S22: Add the corner block attention module CoPA to the head network and generate a spatial weight map based on the statistical comparison of the four corner areas; design the corner block attention module CoPA, such as Figure 4 As shown;

[0118] The specific steps of generating spatial attention weights by the corner block attention module CoPA module in step S22 are as follows:

[0119] S221 Gridded Feature Analysis: Divide the input feature map into regular grid blocks of n_win×n_win, and calculate the mean μ and standard deviation σ of each grid block; the formula is:

[0120] Mean: μ = AvgPool(X_in)∈R 1×n_win×n_win ;

[0121] Standard deviation: σ = StdPool(X_in)∈R 1×n_win×n_win ;

[0122] Among them, X_in represents the input feature map of the module; AvgPool means calculating the mean value in each grid block after the feature map is divided into blocks, and StdPool means calculating the standard deviation in each grid block after the feature map is divided into blocks;

[0123] S222 Corner reference difference calculation: Select the upper left, upper right, lower left, and lower right corner blocks as reference benchmarks, scan each grid block, and calculate the statistical difference between the current grid block and the grid blocks in the four corners; the formula is:

[0124] Δμ=Σ|μ_current-μ_corner_i|;

[0125] Δσ=Σ|σ_current-σ_corner_i|(i=1~4);

[0126] Among them, μ_current represents the mean value of the current network block scanned when scanning each network block, σ_current is the standard deviation of the current network block; μ_corner_i represents the mean value of the corner block, σ_corner_i represents the standard deviation of the corner block, and i is the corner number; Δμ represents the mean difference between the grid block and the corner block; Δσ represents the standard deviation difference between the grid block and the corner block;

[0127] S223 dynamic weight generation: The difference value is mapped to the weight coefficient through the Sigmoid activation function to obtain the block-level weight W. The formula is:

[0128] W=Sigmoid(Δμ)+Sigmoid(Δσ)-0.5;

[0129] S224 Spatial Attention Enhancement: The block-level weight W is reconstructed into a pixel-level spatial weight map and multiplied point by point with the original feature map, that is, the weight map is weighted onto the input feature map; specifically:

[0130] Expand the block-level weight W to a pixel-level weight map W_map∈R through bilinear interpolation 1×H×W , and weight it onto the input feature map, the formula is:

[0131] X_out=X_in⊙W_map∈R C×H×W ;

[0132] Among them, W_map represents the expanded pixel-level weight map; X_out represents the final output feature map of the module after weighting the input feature map;

[0133] The CoPA module quantifies the feature distribution by combining the statistical mean shift and standard deviation fluctuation of the corner reference blocks. By comparing the feature distribution differences, it enhances the sensitivity to abnormal areas without introducing additional parameters.

[0134] S23: Fusion of the dimensionally decoupled feature extraction module DDFM and the corner block attention module CoPA to construct a dual-correlation collaborative module DC-Block, replacing the original C3 module and reconstructing the feature pyramid network;

[0135] In step S23, the dual-correlation collaborative module DC-Block includes a dual-branch collaborative structure, namely a first branch and a second branch; wherein,

[0136] The first branch performs dimension decoupling feature extraction through the dimension decoupling feature extraction module DDFM, focusing on capturing the cross-dimensional correlation characteristics of the glass texture;

[0137] The second branch generates a spatial attention mask through the corner block attention module CoPA to enhance the characteristic response of edge cracks and small defects;

[0138] The outputs of the first and second branches are cross-fused and then reshaped by 1×1 convolution to keep the output dimension consistent with the input.

[0139] S3: Use the glass defect dataset to train and verify the constructed DDC-YOLO model, and then use the DDC-YOLO model to detect transparent material defects in real time;

[0140] The specific steps of step S3 are:

[0141] S31: Use the training set to train the constructed DDC-YOLO model, then use the validation set to verify the constructed DDC-YOLO model, and then use the test set to test its detection ability;

[0142] First, the training environment was configured. During the model training phase, a high-performance computing cluster was used to build a distributed training environment, equipped with a multi-GPU server (such as NVIDIA A6000×3), and mixed precision training (AMP) was implemented based on the PyTorch framework. CUDA acceleration was enabled during training, and a unified random seed (seed=626) was set to ensure experimental reproducibility. Asynchronous loading (num_workers=8) and pin_memory technology were used for data processing to optimize data transfer efficiency between video memory and main memory.

[0143] Then load and enhance the training data: The input data is processed through the online enhancement pipeline, including:

[0144] Geometric transformation: random horizontal flip (p = 0.5), ±15° rotation, 0.8-1.2 scale scaling;

[0145] Color perturbation: HSV color space adjustment (hue ±0.1, saturation ±0.7, lightness ±0.4);

[0146] Mosaic enhancement: Four images are stitched together to generate complex background samples, improving the robustness of small target detection;

[0147] The specific steps of step S31 are:

[0148] S311 training model: The training set is input into DDC-YOLO, which is improved based on YOLO, for training, which goes through the warm-up phase, main training phase, and fine-tuning phase respectively;

[0149] In the warm-up phase, a linear learning rate is used (the learning rate lr is increased from 1e-5 to 1e-2), some layers of the backbone network are frozen, and only the detection head module is trained. The CIoU loss function is then used to initialize the weights to avoid early gradient instability.

[0150] During the main training phase (epoch 50-450), all network layers are unfrozen and cosine annealing learning rate scheduling is enabled (base_lr = 1e-2, min_lr = 1e-4). A dynamic positive sample allocation strategy is introduced to dynamically adjust the GT allocation threshold based on the match between the predicted box and the true value box.

[0151] During the fine-tuning phase (epoch 450-500), we enabled cross-GPU synchronized batch normalization (SyncBN) and used the exponential moving average (EMA) to update model parameters.

[0152] S312 Validation Model: Configure the loss function and use the validation set to validate the trained DDC-YOLO model to evaluate model performance.

[0153] S313 test model: The input image is set to its original size. If it is less than an integer multiple of 32, the edges are padded. Then, the improved non-maximum suppression algorithm NMS (Non-Maximum Suppression) is used to filter the image by setting a category-independent IoU threshold (here set to 0.7) and a confidence threshold (0.001 initial threshold, adaptively adjusted with target density) to reduce redundant frames in target detection and improve detection accuracy.

[0154] The loss function configured in step S312 includes three parts of loss, namely positioning loss, classification loss and confidence loss.

[0155] The positioning loss, i.e., the bounding box regression loss, uses the SIoU metric and introduces the angle cost term and the distance cost term to optimize the box regression direction. The formula is:

[0156] L box =1-IoU+(Δ+Ω) / 2;

[0157] Among them, L box is the bounding box regression loss, IoU represents the intersection-over-union ratio between the predicted box and the ground-truth box, Δ represents the corner cost term, and Ω represents the distance cost term;

[0158] The category classification loss is an improved Focal Loss loss function, which applies a modulation coefficient of γ = 2.0 to difficult samples. This is an improved version of cross entropy. By adjusting the factor, the weight of easy-to-classify samples is reduced to solve the imbalance problem of positive and negative samples. The formula is:

[0159] L cls =-α t *(1-p t ) γ *log(p t )

[0160] Among them, L cls is the category classification loss, α t represents the category balance factor, p t represents the model's predicted probability for the true category, and γ is the modulation coefficient; in this embodiment, γ = 2.0;

[0161] Confidence loss is to increase the dynamic weight factor to alleviate the imbalance problem of positive and negative samples in dense target scenarios; the formula is:

[0162]

[0163] Among them, L obj represents the confidence loss, λ pos represents the dynamic weight of positive samples, λ negrepresents the dynamic weight of negative samples, y is the true label, 1 represents the target, and 0 represents the background; The confidence score for the model prediction;

[0164] Total loss function L total For weighted summation, the formula is:

[0165] L total =0.05L box +0.5L cls +L obj ;

[0166] S32: Deploy the trained DDC-YOLO model to the production line for testing and output the test results in real time.

[0167] The specific steps of using the trained model to perform detection in step S32 are:

[0168] S321 image preprocessing: adjust brightness and contrast to reduce reflection interference; automatically identify glass edges and crop the effective detection area;

[0169] S322 defect detection: It first adopts a two-stage detection, specifically: first quickly scan the entire glass, then carefully inspect the suspicious areas; automatically identify four types of defects: bubbles, cracks, scratches, and impurities;

[0170] S323 result processing: merge adjacent defect frames of the same type, filter out tiny noise points <5 pixels and normal texture areas; record the defect location, type and size;

[0171] S324 system linkage: outputs test results to the display device in real time, automatically marks unqualified products and triggers the sorting device.

[0172] The entire inspection process takes less than 200ms per piece, meeting the real-time inspection requirements of the production line. Inspection parameters can be flexibly adjusted according to different glass types (such as flat and curved).

[0173] To verify the detection performance of the DDC-YOLO model proposed in this paper on the glass dataset, we compared it with various existing YOLO improved models. Table 1 below shows the performance comparison of each method on this glass dataset.

[0174] Table 1 Performance comparison of various methods on this glass dataset

[0175]

[0176] The test results in Table 1 show that the DDC-YOLO model obtained by improving the three-stage network architecture by introducing the dimensionally decoupled feature extraction module DDFM, adding the corner block attention module CoPA, and fusing the DDFM and CoPA to replace the original C3 module has a detection suitability of 48.6% for tiny defects such as scratches, pores, and cracks, an increase of 3.2% compared to the original model. At the same time, it has only 5.1M parameters and 8.8G computational complexity, which greatly speeds up training and reduces the training cost of industrial defect detection.

[0177] The above embodiments are only specific implementations of the present invention and are not intended to limit the present invention. Any modification, replacement or improvement made by those skilled in the art through conventional technical means without departing from the core concept of the present invention shall fall within the scope of protection of the present invention.

Claims

1. A glass defect detection method based on improved YOLOv5, characterized in that: The following steps are involved: S1: Use a high dynamic range industrial camera to collect multispectral image data from a glass product production line. After data processing, a glass defect detection dataset is constructed and then divided. S2: Based on the YOLOv5 model, the dimensionally decoupled feature extraction module DDFM is introduced, and the corner block attention module CoPA is added. The dimensionally decoupled feature extraction module DDFM and the corner block attention module CoPA are fused to replace the original C3 module to improve the three-stage network architecture and obtain the DDC-YOLO model; S3: Use the glass defect dataset to train and verify the constructed DDC-YOLO model, and then use the DDC-YOLO model to detect transparent material defects in real time.

2. The glass defect detection method based on improved YOLOv5 according to claim 1, characterized in that: In step S1, image enhancement processing is performed by dynamic threshold segmentation and reflection suppression algorithm to construct a glass defect detection dataset containing bubbles, scratches, and impurity defects; the specific steps are: S11 data acquisition: Industrial cameras are deployed on the glass production line to cover key processes including cutting, polishing, and tempering. The cameras capture 4K or higher high-resolution images and videos, including defects, across different lighting conditions, glass thickness, and surface conditions. The cameras also record production line status to ensure data diversity. S12 data preprocessing: Remove blurry or invalid images and enhance the retained qualified data using dynamic threshold segmentation and reflection suppression algorithms. The enhanced data set is then annotated, using rectangular boxes to mark defect locations and classifying defects into four categories: bubbles, scratches, cracks, and impurities. S13 dataset construction: A stratified sampling strategy was adopted to ensure a balanced distribution of different glass types and defect sizes. The dataset was then divided into training, validation, and test sets in proportion.

3. The glass defect detection method based on improved YOLOv5 according to claim 2, characterized in that: The improvements to the three-stage network architecture in step S2 to construct the DDC-YOLO model specifically include: S21: Introduce the dimension decoupled feature extraction module DDFM into the feature extraction network to perform independent feature modeling on height, width and channel dimensions; S22: Add the corner block attention module CoPA to the head network and generate a spatial weight map based on the statistical comparison of the four corner regions; S23: The dimensionally decoupled feature extraction module DDFM and the corner block attention module CoPA are integrated to construct a dual-correlation collaborative module DC-Block, which replaces the original C3 module and reconstructs the feature pyramid network.

4. The glass defect detection method based on improved YOLOv5 according to claim 3, characterized in that: The specific steps for the dimension decoupling feature extraction module DDFM in step S21 to implement dimension decoupling are: S211: The input feature map is rotated along the three dimensions of height, width, and channel to obtain three sets of feature maps. S212: performing feature extraction on the three sets of feature maps obtained in step S211 using depthwise convolution, respectively, to extract three two-dimensional correlation features, thereby obtaining three sets of feature maps with two-dimensional correlation features; S213: The feature maps obtained in step S213 are then combined in pairs, and their intersection feature dimensions are concatenated accordingly; S214: Finally, use band convolution and point convolution to compress the corresponding dimensional features, remove redundant features, and obtain three sets of compressed feature maps; S215: The three sets of compressed feature maps are then superimposed and their average is taken and output.

5. The glass defect detection method based on improved YOLOv5 according to claim 4, characterized in that: The specific steps of generating spatial attention weights by the corner block attention module CoPA module in step S22 are as follows: S221 Gridded Feature Analysis: Divide the input feature map into regular grid blocks of n_win×n_win, and calculate the mean μ and standard deviation σ of each grid block; S222 Corner reference difference calculation: Select the upper left, upper right, lower left, and lower right corner blocks as reference benchmarks, scan each grid block, and calculate the statistical difference between the current grid block and the grid blocks in the four corners; the formula is: Δμ=Σ|μ_current-μ_corner_i|; Δσ=Σ|σ_current-σ_corner_i|(i=1~4); Among them, μ_current represents the mean value of the current network block scanned when scanning each network block, σ_current is the standard deviation of the current network block; μ_corner_i represents the mean value of the corner block, σ_corner_i represents the standard deviation of the corner block, and i is the corner number; Δμ represents the mean difference between the grid block and the corner block; Δσ represents the standard deviation difference between the grid block and the corner block; S223 dynamic weight generation: The difference value is mapped to the weight coefficient through the Sigmoid activation function to obtain the block-level weight W. The formula is: W=Sigmoid(Δμ)+Sigmoid(Δσ)-0.5; S224 Spatial Attention Enhancement: The block-level weight W is reconstructed into a pixel-level spatial weight map and multiplied point by point with the original feature map, that is, the weight map is weighted onto the input feature map.

6. The glass defect detection method based on improved YOLOv5 according to claim 5, characterized in that: In step S23, the dual-correlation collaborative module DC-Block includes a dual-branch collaborative structure, namely a first branch and a second branch; wherein, The first branch performs dimension decoupling feature extraction through the dimension decoupling feature extraction module DDFM to capture the cross-dimensional correlation characteristics of the glass texture; The second branch generates a spatial attention mask through the corner block attention module CoPA to enhance the characteristic response of edge cracks and small defects; The outputs of the first branch and the second branch are cross-fused through channels, and then feature reshaping is performed through convolution, so that the output dimension is consistent with the input.

7. The glass defect detection method based on improved YOLOv5 according to claim 5, characterized in that: The specific steps of step S3 are: S31: Use the training set to train the constructed DDC-YOLO model, then use the validation set to verify the constructed DDC-YOLO model, and then use the test set to test its detection ability; S32: Deploy the trained DDC-YOLO model to the production line for testing and output the test results in real time.

8. The glass defect detection method based on improved YOLOv5 according to claim 7, characterized in that: The specific steps of step S31 are: S311 training model: The training set is input into DDC-YOLO, which is improved based on YOLO, for training, which goes through the warm-up phase, main training phase, and fine-tuning phase respectively; In the warm-up phase, a linear learning rate is used to warm up, freezing some layers of the backbone network and training only the detection head module. The CIoU loss function is then used to initialize the weights. In the main training phase, all network layers are unfrozen and cosine annealing learning rate scheduling is enabled (base_lr = 1e-2, min_lr = 1e-4). A dynamic positive sample allocation strategy is introduced to dynamically adjust the GT allocation threshold based on the matching degree between the predicted box and the true value box. During the fine-tuning phase, cross-GPU synchronized batch normalization is enabled, and exponential moving average (EMA) is used to update model parameters. S312 Validation Model: Configure the loss function and use the validation set to validate the trained DDC-YOLO model to evaluate model performance. S313 test model: The input image is set to its original size. If it is less than an integer multiple of 32, the edges are padded. Then, the improved non-maximum suppression algorithm (NMS) is used to set the category-independent IoU threshold and confidence threshold for filtering to reduce redundant boxes in target detection.

9. The glass defect detection method based on improved YOLOv5 according to claim 8, characterized in that: The loss function configured in step S312 includes three parts of loss, namely positioning loss, classification loss and confidence loss. The positioning loss, i.e., the bounding box regression loss, uses the SIoU metric and introduces the angle cost term and the distance cost term to optimize the box regression direction. The formula is: L box =1-IoU+(Δ+Ω) / 2; Among them, L box is the bounding box regression loss, IoU represents the intersection-over-union ratio between the predicted box and the ground-truth box, Δ represents the corner cost term, and Ω represents the distance cost term; The category classification loss is an improved Focal Loss loss function, which applies a modulation coefficient of γ to difficult samples; the formula is: L cls =-a t *(1-p t ) γ *log(p t ); Among them, L cls is the category classification loss, α t represents the category balance factor, p t represents the model's predicted probability of the true category, and γ is the modulation coefficient; The confidence loss is to increase the dynamic weight factor, the formula is: Among them, L obj represents the confidence loss, λ pos represents the dynamic weight of positive samples, λ neg represents the dynamic weight of negative samples, y is the true label, where 1 represents the target and 0 represents the background; The confidence score for the model prediction; Total loss function L total For weighted summation, the formula is: L total =0.05L box +0.5L cls +L obj 。 10. The glass defect detection method based on improved YOLOv5 according to claim 8, characterized in that: The specific steps of using the trained model to perform detection in step S32 are: S321 image preprocessing: adjust brightness and contrast; automatically identify glass edges and crop effective detection areas; S322 defect detection: It first adopts a two-stage detection, specifically: first quickly scan the entire glass, then carefully inspect the suspicious areas; automatically identify four types of defects: bubbles, cracks, scratches, and impurities; S323 result processing: merge adjacent defect frames of the same type, filter out tiny noise points <5 pixels and normal texture areas; Record defect location, type, and size; S324 system linkage: outputs test results to the display device in real time, automatically marks unqualified products and triggers the sorting device.

Citation Information

Patent Citations

  • Water level identification method, device and equipment and readable storage medium

    CN114332870A

  • Surface defect detection method based on improved YOLO V4 algorithm

    CN115861170A

  • PCB defect detection method based on LDLFModel

    CN115908356A

  • Nuclear containment vessel defect detecting and positioning method and system based on deep learning

    CN117670795A

  • Improved Yolov10-oriented target detection fusion method

    CN119723272A