Shaving board surface defect detection method based on deep learning
By improving the YOLOv11 object detection model, combining the efficient multi-scale attention mechanism and weighted feature pyramid network, the problem of particleboard surface defect detection accuracy and efficiency is solved, and accurate detection of multiple categories of defects is achieved, thereby reducing labor costs.
Patent Information
- Application Number
- CN202510211620.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-25
AI Technical Summary
The prior art is difficult to accurately detect defects in various categories of particleboard surfaces, especially when the defects are complex and the area is small, the detection accuracy and efficiency are low, and the labor cost is high.
The particleboard surface defect detection method based on deep learning is adopted, and the bounding box regression performance is improved by building an improved YOLOv11 object detection model, the C3k2_EMA module and BiFPN module are used to enhance feature extraction and fusion capabilities, and the bounding box regression performance is improved in combination with the Shape-NWD loss function.
It significantly improves the accuracy and efficiency of particleboard surface defect detection, and can accurately detect defects in multiple categories, including small and complex defects, greatly reducing labor costs.
Smart Images

Figure CN120147248A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning, and specifically to a method for detecting surface defects of particleboard based on deep learning. Background Art
[0002] Particleboard is an important wood-based composite material, which is widely used in furniture manufacturing, building decoration and other fields. The surface quality of particleboard directly affects the appearance and performance of products. Therefore, detecting its surface defects during the production process is an important link in quality control. Traditional methods for detecting surface defects of particleboard mainly include manual detection methods and machine vision methods based on image processing. Manual detection highly depends on the experience and subjective judgment of inspectors, and is easily affected by factors such as fatigue and distraction, resulting in high rates of missed detection and false detection. In addition, manual detection has low efficiency and cannot meet the needs of large-scale continuous production. Traditional machine vision methods based on image processing use predefined features (such as texture, edge or color) to detect defects. However, these methods are difficult to handle particleboards with complex surface textures and diverse defects, and have poor robustness and generalization ability.
[0003] With the rapid development of artificial intelligence technology, the emergence of deep learning, especially convolutional neural network (CNN), provides a new solution for detecting surface defects of particleboard. Deep learning models have the ability to automatically learn complex features from data, can efficiently extract the subtle features of surface defects of particleboard, and avoid the limitations of traditional methods that rely on manually designed features. By training a large-scale labeled dataset, deep learning models show extremely strong robustness to complex backgrounds, illumination changes, and the diversity of defect morphologies. However, in the face of problems such as a large number of types of surface defects of particleboard, complex defects, and tiny areas of defects, it is also difficult for mainstream object detection models to achieve accurate detection. Therefore, a method for detecting surface defects of particleboard based on deep learning is proposed to improve the detection accuracy and efficiency of surface defects of particleboard. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a method for detecting surface defects of particleboard based on deep learning in view of the above-mentioned prior art. The method for detecting surface defects of particleboard based on deep learning can accurately predict various types of surface defects of particleboard, also improve the detection accuracy of tiny defects and complex defects on the surface of particleboard, and at the same time greatly improve the detection accuracy and efficiency of surface defects of particleboard, and greatly reduce the labor cost.
[0005] To achieve the above technical objectives, the technical solution adopted by the present invention is as follows:
[0006] A method for detecting surface defects of particleboard based on deep learning, comprising:
[0007] Step 1: Set up the particleboard surface image acquisition device, acquire the particleboard surface images, and screen out the particleboard images with surface defects;
[0008] Step 2: Manually annotate the surface defects of the particleboard in the screened particleboard images, construct a particleboard surface defect dataset, and divide it into a training set, a validation set, and a test set according to a certain proportion;
[0009] Step 3: Construct an improved YOLOv11 object detection model;
[0010] Step 4: Input the training set and the validation set into the improved YOLOv11 object detection model for training;
[0011] Step 5: Set evaluation metrics, use the evaluation metrics to comprehensively evaluate and analyze the training results, and judge the pros and cons of the model performance;
[0012] Step 6: Use the test set data to test the trained improved YOLOv11 object detection model.
[0013] As a further improved technical solution of the present invention, the specific content of Step 1 is as follows:
[0014] Step 1.1: Set up the particleboard surface image acquisition device, including a conveying mechanism, an image acquisition module, and a host computer; the conveying mechanism is used to convey the particleboard to the image acquisition module in a certain direction; the image acquisition module is used to acquire the particleboard surface image under the control instruction of the host computer and transmit the image to the host computer;
[0015] Step 1.2: Acquire several particleboard surface images according to the method of Step 1.1, and screen out the particleboard images with surface defects; the defect categories include large particles, dust spots, sand leakage, glue spots, and scratches.
[0016] As a further improved technical solution of the present invention, the parameters of the manual annotation in Step 2 include the category information of the defect and the position information of the defect;
[0017] The specific steps of the manual annotation are as follows:
[0018] First, clarify the annotation specifications, select the annotation tool LabelImg for manual annotation, and the annotation content includes the defect category and the specific position of the defect in the image; load the particleboard images with surface defects into the annotation tool LabelImg, and mark the defect areas in the images one by one to ensure the accuracy of the selected defect areas and the defect category labels;
[0019] After completing the annotation, save the annotation file; organize the annotation data, randomly divide the data according to the ratio of 80% training set, 10% validation set, and 10% test set, and store the images and their corresponding annotation files in separate folders for training and optimization of the subsequent model according to training, validation, and test.
[0020] As a further improved technical solution of the present invention, the improved YOLOv11 object detection model in step 3 is divided into three parts: a backbone network, a neck network, and a head network;
[0021] Among them, the backbone network includes a C3k2_EMA module, and the C3k2_EMA module includes a C2f module, an EMA module, and a C3k2 module; the neck network includes a Concat-BiFPN module, and the Concat-BiFPN module includes a Concat module and a BiFPN module.
[0022] As a further improved technical solution of the present invention, the backbone network includes 5 convolutional layers, 4 C3k2_EMA modules, 1 SPPF module, and 1 C2PSA module; the convolutional layers, convolutional layers, C3k2_EMA modules, convolutional layers, C3k2_EMA modules, convolutional layers, C3k2_EMA modules, convolutional layers, C3k2_EMA modules, SPPF module, and C2PSA module are connected in sequence;
[0023] The neck network includes 4 C3k2 modules, 4 Concat-BiFPN modules, 2 Upsample modules, and 2 convolutional layers; the Upsample module, Concat-BiFPN module, C3k2 module, Upsample module, Concat-BiFPN module, C3k2 module, convolutional layer, Concat-BiFPN module, C3k2 module, convolutional layer, Concat-BiFPN module, and C3k2 module are connected in sequence;
[0024] Among them, the second C3k2_EMA module in the backbone network is connected to the second Concat-BiFPN module in the neck network; the third C3k2_EMA module in the backbone network is connected to the first Concat-BiFPN module in the neck network; the C2PSA module in the backbone network is simultaneously connected to the first Upsample module and the fourth Concat-BiFPN module in the neck network; the first C3k2 module in the neck network is connected to the third Concat-BiFPN module in the neck network; the second C3k2 module, the third C3k2 module, and the fourth C3k2 module in the neck network are all connected to the head network.
[0025] As a further improved technical solution of the present invention, the loss function of the improved YOLOv11 object detection model is the bounding box regression loss function Shape-NWD; the calculation process is as follows:
[0026] Calculate the intersection over union (IoU) by computing the ratio of the intersection area to the union area between the predicted bounding box and the ground truth bounding box.
[0027] Calculate the comprehensive distance metric D:
[0028]
[0029] Calculate the shape penalty NWD shape :
[0030]
[0031] In the formula, weight = 2, w gt represents the width of the ground truth bounding box, h gt represents the height of the ground truth bounding box, w represents the width of the predicted bounding box, h represents the height of the predicted bounding box, (x c , y c ) represents the center coordinates of the predicted bounding box, (x c gt , y c gt ) represents the center coordinates of the ground truth bounding box, h h and w w both represent coefficients, D represents the comprehensive distance metric, and C represents a constant;
[0032] Fuse the IoU, D, and NWD with weights to generate the final Shape-NWD loss value. shape
[0033] As a further improved technical solution of the present invention, the evaluation metrics in step 5 include precision P, recall R, F1 score, and mean average precision (mAP).
[0034] The beneficial effects of the present invention are as follows:
[0035] (1) In the YOLOv11 network, the C3k2_EMA module is adopted. That is, after the C2f feature extraction of the backbone network, an efficient multi-scale attention mechanism (EMA) is added and then incorporated into the C3k2 module, enhancing the ability to extract and fuse features at different scales. At the same time, it can adaptively adjust the weights of features at different scales, ensuring that the network focuses on key defect regions, preserving spatial information, avoiding missed detection of tiny defects, and improving the detection accuracy of tiny defects.
[0036] (2) Replace the original Path Aggregation Network (PANet) in the neck network of the YOLOv11 network with a Weighted Feature Pyramid Network (BiFPN), enabling the model to make full use of feature information at different scales, more effectively integrate feature information from different scales, enhance the network's perception ability of fine-grained objects and complex scenes, and improve the detection accuracy of complex defects.
[0037] (3) Replace the original boundary loss function CIoU with Shape-NWD, which can effectively capture changes in the shape and scale of bounding boxes. For bounding boxes of different shapes and scales, it can calculate the loss more accurately. For small target defects, the changes in the shape and scale of the defects have a greater impact on the detection results. Shape-NWD can better adapt to the characteristics of small targets and improve the accuracy of small target defect detection.
[0038] (4) Compared with existing detection models, the present invention improves the detection accuracy of micro-defects and complex defects on the surface of particleboard. At the same time, for traditional particleboard defect detection methods, this method greatly improves the detection accuracy and efficiency of particleboard surface defects, and significantly reduces the labor cost. Description of the Drawings
[0039] Figure 1 It is a schematic diagram of defect categories in the particleboard surface image.
[0040] Figure 2 It is a structural diagram of the Efficient Multi-Scale Attention Mechanism (EMA) module used in the improved YOLOv11 network provided by the present invention.
[0041] Figure 3 It is a structural diagram of the Weighted Feature Pyramid Network (BiFPN) used in the improved YOLOv11 network provided by the present invention.
[0042] Figure 4 It is a structural diagram of the improved YOLOv11 network provided by the present invention.
[0043] Figure 5 It is a schematic diagram of partial particleboard surface defect detection provided by the present invention. Detailed Embodiments
[0044] The following further explains the detailed embodiments of the present invention with reference to the drawings:
[0045] This embodiment provides a particleboard surface defect detection method based on deep learning, including:
[0046] Step 1: Build a particleboard surface image acquisition device, collect particleboard surface images, and screen out particleboard images with surface defects;
[0047] Step 2: Manually annotate the surface defects of the selected particle board images, construct a data set of particle board surface defects, and divide it into a training set, a validation set, and a test set according to a certain proportion;
[0048] Step 3: Construct an improved YOLOv11 object detection model; improve the YOLOv11 object detection model in view of the problems such as a large number of types of particle board surface defects, complex defects, and tiny areas;
[0049] Step 4: Input the training set and the validation set into the improved YOLOv11 object detection model for training;
[0050] Step 5: Set evaluation metrics, comprehensively evaluate and analyze the training results using the evaluation metrics, judge the pros and cons of the model performance, and select the optimal improved YOLOv11 object detection model;
[0051] Step 6: Use the test set data to test the trained optimal improved YOLOv11 object detection model to obtain the defect positions and defect types of the corresponding images after prediction.
[0052] The specific content of the above-mentioned Step 1 is as follows:
[0053] Step 1.1: Build a particle board surface image acquisition device, including a conveying mechanism, an image acquisition module, and a host computer; the image acquisition module includes an industrial camera, a light source, a photoelectric sensor, a PLC controller, and a dark box.
[0054] The conveying mechanism conveys the particle board to the image acquisition module in a certain direction. When the particle board enters the field of view of the industrial camera in the dark box, the photoelectric sensor transmits a signal to the PLC controller, and the PLC controller uploads the acquisition signal to the host computer. The host computer controls the industrial camera to acquire the surface image of the particle board and transmits the image back; the light source is installed in the dark box to ensure that the acquisition area is under uniform illumination conditions. The dark box reduces the influence of the external environment on the acquired image and at the same time shields the influence of external interfering light sources.
[0055] Step 1.2: Acquire several particle board surface images according to the method in Step 1.1, and select the particle board images with surface defects; the selected particle board surface defect images include five defect types: large particles, dust spots, sand leakage, glue spots, and scratches, as Figure 1 shown.
[0056] The parameters of the manual annotation in the above-mentioned Step 2 include the category information and the position information of the defects.
[0057] The specific steps of the manual annotation are as follows: First, clarify the annotation specifications, select the annotation tool LabelImg for manual annotation, and the annotation content includes the defect category and the specific position of the defect in the image (represented by a rectangular box); load the particleboard images with surface defects into the annotation tool LabelImg, mark the defect areas in the images one by one, and ensure the accuracy of the boxed defect areas and defect category labels; after completing the annotation, save the annotation file; organize the annotation data, randomly divide the data according to the ratio of 80% training set, 10% validation set, and 10% test set, and store the images and the corresponding annotation files in separate folders for training and optimization of the subsequent model respectively.
[0058] The improved YOLOv11 object detection model in step 3 is divided into three parts: the backbone network, the neck network, and the head network; the backbone network is responsible for extracting rich feature information from the input image, the neck network further processes and fuses the extracted features, and the head network is responsible for the final bounding box prediction and class judgment.
[0059] Among them, the backbone network includes the C3k2_EMA module, and the C3k2_EMA module is composed of the C2f module, the EMA module, and the C3k2 module; the neck network includes the Concat-BiFPN module, and the Concat-BiFPN module is composed of the Concat module and the BiFPN module.
[0060] As Figure 4 shown, specifically, the backbone network includes 5 convolutional layers, 4 C3k2_EMA modules, 1 SPPF module, and 1 C2PSA module; the convolutional layer, the convolutional layer, the C3k2_EMA module, the convolutional layer, the C3k2_EMA module, the convolutional layer, the C3k2_EMA module, the convolutional layer, the C3k2_EMA module, the SPPF module, and the C2PSA module are connected in sequence.
[0061] The neck network includes 4 C3k2 modules, 4 Concat-BiFPN modules, 2 Upsample modules, and 2 convolutional layers; the Upsample module, the Concat-BiFPN module, the C3k2 module, the Upsample module, the Concat-BiFPN module, the C3k2 module, the convolutional layer, the Concat-BiFPN module, the C3k2 module, the convolutional layer, the Concat-BiFPN module, and the C3k2 module are connected in sequence.
[0062] Among them, the second C3k2_EMA module in the backbone network is connected to the second Concat-BiFPN module in the neck network; the third C3k2_EMA module in the backbone network is connected to the first Concat-BiFPN module in the neck network; the C2PSA module in the backbone network is simultaneously connected to the first Upsample module and the fourth Concat-BiFPN module in the neck network; the first C3k2 module in the neck network is connected to the third Concat-BiFPN module in the neck network; the second C3k2 module, the third C3k2 module, and the fourth C3k2 module in the neck network are all connected to the head network.
[0063] Regarding the problems of many types of surface defects, complex defects, and tiny areas in particleboard, the improvement points of the YOLOv11 network in step 3 above are specifically as follows:
[0064] I. Improve the C3k2 module in the backbone network of the original YOLOv11 network to form the C3k2_EMA module. The C3k2_EMA module adds an efficient multi-scale attention mechanism (EMA) after C2f feature extraction and then adds it to the C3k2 module. The C3k2 module includes the C2f module, and C3k2_EMA is improved by adding an EMA module after the C2f module, that is, C3k2_EMA = C2f + EMA + C3k2. Its core process is: first, perform feature extraction through C2f, then use the EMA mechanism to enhance multi-scale attention, and finally integrate the results into the C3k2 structure to further optimize feature expression. The efficient multi-scale attention (EMA) uses a cross-space learning method to fuse context information of different scales, enabling the model to generate better pixel-level attention for high-level feature maps and adding short-range and long-range dependencies during the local fusion process to embed accurate position information for better performance.
[0065] The structure of the efficient multi-scale attention mechanism (EMA) module is as Figure 2 shown, Figure 2Among them, Group (channel grouping): Group the input to reduce the computational load or extract specific features. XAvgPool (average pooling in the X direction): Perform average pooling in the X-axis direction to extract features in the horizontal direction. YAvgPool (average pooling in the Y direction): Perform average pooling in the Y-axis direction to extract features in the vertical direction. Contact (connection): Concatenate the results processed by XAvgPool and YAvgPool. Conv(1×1) (1×1 convolution): Use 1×1 convolution for feature transformation or channel mixing to reduce the computational complexity. Conv(1×1) (1×1 convolution) respectively outputs the feature maps to 2 Sigmoids. Sigmoid (Sigmoid activation function): Used to normalize data, restricting its range to between (0, 1) for weight calculation. Re-weight (channel weight adjustment, re-weighting): Re-weight the input features according to the weights calculated by Sigmoid to adjust the importance. Conv(3×3) (3×3 convolution): Use 3×3 convolution to extract local features and increase the receptive field. GroupNorm (group normalization): Normalize each group of feature maps to improve the training stability. XY AvgPool (average pooling in the XY direction): Perform average pooling simultaneously in the XY-axis direction to extract overall features. Softmax (Softmax normalization): Used to calculate the probability distribution, normalizing the values to between (0, 1) and the sum of all values being 1. Matmul (matrix multiplication): Used for feature interaction or weight calculation for information fusion. ⊕ (addition operation): Represents the addition fusion of two feature maps. Sigmoid (Sigmoid activation function): Used to normalize the weighted feature values. Re-weight (channel weight adjustment, re-weighting): Re-weight the features again to finally output the optimized features.
[0066] The function of the efficient multi-scale attention mechanism (EMA) module is as follows: First, the input feature map (H×W×C) is grouped according to the grouping factor (Group) to reduce the computational complexity while retaining sufficient feature expression ability. Then, adaptive average pooling is used to extract global features along the row and column directions respectively. Among them, the row pooling (XAvgPool) generates a feature map of size H×1, and the column pooling (YAvgPool) generates a feature map of size W×1. These global features are further fused through 1×1 convolution (Conv(1x1)) to generate attention weights containing multi-scale spatial information. Subsequently, these weights are used to perform channel-wise weighted adjustment (re-weight) on the grouped features. Among them, the dynamic weight distribution of the features is adjusted through the Sigmoid activation function. At the same time, 3×3 convolution (Conv(3x3)) is used to extract local features. Combining global information, multi-scale attention weights are generated through Softmax normalization and matrix operations. These weights act on the input features to achieve dynamic weighted processing, and finally an enhanced feature map of the original shape (H×W×C) is restored, enabling the model to pay more attention to key feature regions, effectively suppressing redundant information, and significantly improving the feature expression ability and the accuracy of object detection.
[0067] Furthermore, the efficient multi-scale attention mechanism includes: 1. Grouped feature processing module (such as "Groups"): The input feature map is grouped according to the number of channels, and each group of features is processed independently to reduce the computational complexity and enhance the local feature extraction ability; 2. Multi-scale feature pooling module (such as "XAvgPool", "YAvgPool", "XY AvgPool"): It includes a global average pooling layer and one-dimensional pooling layers along the height and width, respectively extracting global features and spatial dimension features at different scales, and capturing the global and local information of the input features at multiple scales; 3. Feature interaction and fusion module (such as "Contact" and "Conv(1x1)"): By concatenating multi-scale features and using a 1x1 convolutional layer for channel compression and fusion, the interaction and unified expression of multi-scale information are realized; 4. Attention weight generation module (such as "Sigmoid" and "Softmax"): Based on the multi-scale fused features, dynamic attention weights are generated through non-linear transformations (such as fully connected layers and activation functions) to guide the network to focus on important feature regions; 5. Dynamic weighting module (such as "Re-weight" and "Matmul"): Using the generated attention weights, element-wise weighting is performed on the channels or spatial dimensions of the input features to achieve dynamic enhancement of multi-scale features.
[0068] Second, in this embodiment, the original Path Aggregation Network (PANet) in the neck network of the original YOLOv11 network is replaced with a Weighted Feature Pyramid Network (BiFPN). The Weighted Feature Pyramid Network is a feature fusion network that enhances image object detection. It mainly improves the quality and detection accuracy of feature maps through a bidirectional feature fusion mechanism, and can perform information transmission between high-level features and low-level features to make full use of the advantages of each layer of features. At the same time, BiFPN adopts a weighted fusion strategy, assigns weights according to the contribution degrees of different features, improves the fusion effect, and enables it to perform outstandingly in complex scenarios;
[0069] As Figure 3 shown, the Weighted Feature Pyramid Network (BiFPN) module first receives a set of multi-scale feature maps as input, and dynamically adjusts the importance of different feature maps by introducing a set of learnable weights. These weights are initialized to all 1 in vector form and are optimized during the training process. After each weight undergoes non-linear processing by the Swish activation function, a normalization operation is performed to ensure that the sum of all weights is 1, so as to avoid numerical instability or bias during the training process. Next, the normalized weights are multiplied element-wise with the corresponding feature maps respectively to highlight the importance of each feature map. The weighted feature maps are stacked into a new high-dimensional tensor and fused by channel-wise addition, thereby obtaining an output feature map that synthesizes multi-scale information.
[0070] Figure 3 is the structure diagram of BiFPN, showing the information flow mode of BiFPN, that is, multi-scale feature maps (rectangles of different colors) enter the BiFPN module and are fused through bidirectional connections inside BiFPN. Figure 3 In, BiFPN (Bidirectional Feature Pyramid Network) is a module for feature fusion, which can efficiently fuse features of different levels and optimize the transmission of information flow in a weighted manner. Output can be connected to Conv to perform convolution operations for extracting local features. Figure 3 In, different color blocks represent feature maps of different scales, coming from different levels of the backbone network. The lower layers represent high-resolution feature maps, while the higher layers represent low-resolution but high-semantic feature maps. The arrows indicate the direction of feature flow, and BiFPN adopts bidirectional connections, enabling high-level and low-level features to influence each other.
[0071] Furthermore, the Weighted Feature Pyramid Network includes: 1. A multi-scale feature extraction module (the multi-scale feature extraction module corresponds to Figure 3The feature pyramid part on the left in the figure. The rectangular blocks of different colors represent feature maps of different scales, and high-semantic and high-resolution features are extracted through multiple layers of convolution and downsampling): Extract multi-scale features from the input feature map, including low-resolution high-semantic features and high-resolution detailed features for subsequent feature fusion and optimization. Different-scale features are gradually extracted through multiple layers of convolution operations and downsampling; 2. Context fusion module (the context fusion module corresponds to Figure 3 the BiFPN (Bidirectional Feature Pyramid Network) part in the figure, that is, the bidirectional connection of the nodes within the dashed box): Through top-down and bottom-up bidirectional feature transmission, fuse multi-scale feature information and capture the association between local and global contexts. Adopt a feature pyramid structure and use the hierarchical relationship of the feature maps to achieve layer-by-layer propagation of information; 3. Dynamic weight generation module (the dynamic weight generation module is not shown in Figure 3 the figure): Generate normalized dynamic weights according to the input feature map to control the contribution of different feature layers to the fusion result; The weight generation is achieved through non-linear transformations such as global pooling, fully connected layers, and activation functions, making the weights have dynamic adaptability; 4. Weighted feature fusion module (the weighted feature fusion module corresponds to the nodes in the BiFPN structure and their connected weighted fusion paths): Based on the dynamic weights, perform element-wise weighting on features of different scales and then superimpose and fuse them to generate a unified multi-scale feature representation. Adopt element-wise operations to ensure that the fusion result retains the effective information of different feature layers at the same time; 5. Output feature pyramid module (the output feature pyramid module corresponds to the feature output part on the right of the BiFPN): Output the fused multi-scale feature map and hierarchical feature pyramid. The fusion result has multi-resolution characteristics and adapts to the perception requirements of targets of different sizes.
[0072] Figure 4For the overall schematic diagram of the improved YOLOv11 network structure, where Conv represents feature extraction using a 3×3 convolutional kernel with a stride of 2 and a padding of 1. The C3k2_EMA module represents a C3k2 structure combined with an efficient multi-scale attention mechanism (EMA) for feature extraction. The SPPF module is spatial pyramid pooling, which extracts multi-scale features through multi-scale pooling operations. The C2PSA module combines a spatial attention mechanism to enhance the feature extraction ability. The Concat-BiFPN module improves the fusion effect of features at different levels through feature concatenation and a bidirectional feature pyramid network (BiFPN). The Upsample module is used for upsampling operations to increase the resolution of the feature map. The depthwise separable convolution (DWConv) module is located in the Detect layer, which belongs to the head network, that is, it performs feature extraction through the convolutional kernel, stride, padding, and number of channels, while reducing the computational amount. The Detect layer is used to output the target detection results. Conv2d represents a two-dimensional convolution, with k as the convolutional kernel, s as the stride, p as the padding, c as the number of channels. BatchNorm2d accelerates training and improves the network stability through batch normalization, while the ReLU activation function introduces non-linearity to enhance the network's expressive ability. MAXPool2d performs max pooling operations to reduce the size of the feature map and extract important features. The Concat operation is used to concatenate multiple feature maps together to enhance the expressive ability of the features.
[0073] In this embodiment, the boundary loss function CIoU in the original YOLOv11 network is also replaced with Shape-NWD; the Shape-NWD is a bounding box regression loss function that combines the Wasserstein distance, shape normalization, and dynamic focusing mechanism. It accurately measures the geometric differences between the predicted box and the ground truth box in terms of position and size through the Wasserstein distance. At the same time, it introduces a shape normalization factor to balance the influence of the aspect ratio on the loss and avoid the adverse effects of abnormal target shapes on optimization. In addition, Shape-NWD integrates a dynamic non-monotonic focusing mechanism, which can dynamically adjust the optimization weights according to the difficulty of the target, improving the detection performance for small targets, targets with abnormal aspect ratios, and complex targets.
[0074] The loss function of the improved YOLOv11 object detection model is the bounding box regression loss function Shape-NWD; the calculation process is as follows:
[0075] Calculate the intersection over union (IoU), by calculating the ratio of the intersection area to the union area of the predicted box and the ground truth box, to obtain the IoU between the two. The specific calculation method uses the existing formula;
[0076] Calculate the comprehensive distance metric D:
[0077]
[0078] Calculate the shape penalty NWD shape :
[0079]
[0080] In the formula, weight = 2, D is a comprehensive distance metric, x c and y c are the center coordinates of the predicted bounding box, x c gt and y c gt are the center point coordinates of the ground truth bounding box, h h and w w are coefficients related to the shape, w and h are the width and height of the predicted bounding box, w gt and h gt are the width and height of the ground truth bounding box. C in [[ ]] is a constant related to the dataset; when calculating, the boundary loss function Shape-NWD comprehensively considers the differences in center point coordinates, width and height differences, and shape-related weighting factors between the predicted bounding box and the ground truth bounding box. (x c -x c gt ) 2 and (y c -y c gt ) 2 measure the offset of the center point. By multiplying h h and w w , the offset can be weighted differently according to the shape factors in the horizontal and vertical directions. Part B calculates the width and height differences, which are also normalized by dividing by weight 2 so that the width and height differences have appropriate weights in the entire distance metric. Then, normalization and emphasis on the differences are performed through exponential calculation . When D is larger (i.e., the difference between the predicted bounding box and the ground truth bounding box is larger), the value of [[ ]] is smaller, the value of [[ ]] is closer to 0, which indicates a lower matching degree between the predicted bounding box and the ground truth bounding box in terms of shape; conversely, when D is smaller, the value of [[ ]] is closer to 1, indicating a higher matching degree;
[0081] Fuse the intersection over union IoU, the comprehensive distance metric D, and the shape penalty NWD shape with weights to generate the final Shape-NWD loss value.
[0082] Furthermore, Shape-NWD includes: a bounding box parsing module that parses the input predicted bounding boxes and ground truth bounding boxes, extracts the center point coordinates, width, height, and the top, bottom, left, and right coordinates of the bounding box (the top, bottom, left, and right coordinates of the bounding box are used for the subsequent IoU calculation, and the IoU calculation is a standard formula), providing basic geometric information for subsequent calculations; an intersection over union (IoU) calculation module that calculates the IoU between the predicted bounding box and the ground truth bounding box (the IoU calculation is a standard formula) through the intersection area and union area of the two, as a basic geometric similarity metric; a shape distance calculation module that combines the differences in the width-to-height ratios of the predicted bounding box and the ground truth bounding box, as well as the offset of the center point coordinates, to calculate the comprehensive shape distance D between the two, comprehensively considering the differences in the center point coordinates, width and height differences, and shape-related weighting factors between the predicted bounding box and the ground truth bounding box. A loss integration module that shape performs weighted fusion on IoU, the comprehensive shape distance D, and the shape penalty term NWD
[0083] Specifically, step 4 is as follows: The training set and validation set divided in step 2 are input into the improved YOLOv11 object detection model after step 3 for training, and successively enter the backbone network for feature extraction, the neck network for feature fusion and processing, and finally enter the head network for final bounding box prediction and class judgment.
[0084] The evaluation metrics in step 5 use the commonly used metrics in object detection to evaluate and analyze the test method, including precision (Precision, P), recall (Recall, R), F1-score (F1-score), and mean average precision (mean Average Precision, mAP). The specific formulas are as follows:
[0085]
[0086]
[0087] In the formula, TP represents the number of correctly classified defective regions, FP represents the number of missed detections of defective regions, and FN represents the number of misclassifications of non-defective regions. N represents the total number of classes of the detected objects, i is a single class number, and AP represents the average accuracy of a single class.
[0088] The advantages of the model are evaluated using the above evaluation metrics, and the optimal improved YOLOv11 object detection model is selected.
[0089] Figure 5 This is a schematic diagram of the detection results of surface defects of particleboard provided by the present invention.
[0090] The protection scope of the present invention includes but is not limited to the above embodiments. The protection scope of the present invention shall be subject to the claims, and any substitutions, deformations, and improvements that are easily conceivable by those skilled in the art to this technology shall fall within the protection scope of the present invention.
Claims
1. A particleboard surface defect detection method based on deep learning, characterized in that: include: Step 1: Build a particleboard surface image acquisition device to collect particleboard surface images and screen out particleboard images with surface defects; Step 2: Manually annotate the particle board surface defects in the screened particle board images, construct a particle board surface defect dataset, and divide it into a training set, a validation set, and a test set in proportion; Step 3: Build an improved YOLOv11 target detection model; Step 4: Input the training set and validation set into the improved YOLOv11 target detection model for training; Step 5: Set evaluation indicators and use them to comprehensively evaluate and analyze the training results to determine the performance of the model. Step 6: Use the test set data to test the trained improved YOLOv11 target detection model.
2. The particleboard surface defect detection method based on deep learning according to claim 1 is characterized in that: The step 1 is specifically as follows: Step 1.1, constructing a particleboard surface image acquisition device, including a transmission mechanism, an image acquisition module and a host computer; the transmission mechanism is used to transmit the particleboard to the image acquisition module in a certain direction; the image acquisition module is used to acquire the particleboard surface image under the control instruction of the host computer, and transmit the image to the host computer; Step 1.2: Collect several particleboard surface images according to the method in step 1.1, and screen out particleboard images with surface defects; defect categories include large particles, dust spots, sand leakage, glue spots and scratches.
3. The particleboard surface defect detection method based on deep learning according to claim 1 is characterized in that: The manually annotated parameters in step 2 include defect category information and defect location information; The specific steps of manual annotation are: First, the labeling specification is clarified, and the labeling tool LabelImg is selected for manual labeling. The labeling content includes the defect category and the specific location of the defect in the image. The particleboard image with surface defects is loaded into the labeling tool LabelImg, and the defect area in the image is marked one by one to ensure the accuracy of the defect area selection and defect category label. After completing the annotation, save the annotation file; Organize the labeled data, randomly divide the data into 80% training set, 10% validation set, and 10% test set, and store the images and corresponding annotation files in separate folders for training, validation, and testing for subsequent model training and optimization.
4. The particleboard surface defect detection method based on deep learning according to claim 1 is characterized in that: The improved YOLOv11 target detection model in step 3 is divided into three parts: a backbone network, a neck network, and a head network; Among them, the backbone network includes a C3k2_EMA module, the C3k2_EMA module includes a C2f module, an EMA module and a C3k2 module; the neck network includes a Concat-BiFPN module, and the Concat-BiFPN module includes a Concat module and a BiFPN module.
5. The particleboard surface defect detection method based on deep learning according to claim 4 is characterized in that: The backbone network includes 5 convolutional layers, 4 C3k2_EMA modules, 1 SPPF module and 1 C2PSA module; convolutional layer, convolutional layer, C3k2_EMA module, convolutional layer, C3k2_EMA module, convolutional layer, C3k2_EMA module, convolutional layer, C3k2_EMA module, SPPF module and C2PSA module are connected in sequence; The neck network includes 4 C3k2 modules, 4 Concat-BiFPN modules, 2 Upsample modules and 2 convolutional layers; the Upsample module, the Concat-BiFPN module, the C3k2 module, the Upsample module, the Concat-BiFPN module, the C3k2 module, the convolutional layer, the Concat-BiFPN module, the C3k2 module, the convolutional layer, the Concat-BiFPN module and the C3k2 module are connected in sequence; Among them, the second C3k2_EMA module in the backbone network is connected to the second Concat-BiFPN module in the neck network; the third C3k2_EMA module in the backbone network is connected to the first Concat-BiFPN module in the neck network; the C2PSA module in the backbone network is simultaneously connected to the first Upsample module and the fourth Concat-BiFPN module in the neck network; the first C3k2 module in the neck network is connected to the third Concat-BiFPN module in the neck network; the second C3k2 module, the third C3k2 module, and the fourth C3k2 module in the neck network are all connected to the head network.
6. The particleboard surface defect detection method based on deep learning according to claim 5 is characterized in that: The loss function of the improved YOLOv11 target detection model is the bounding box regression loss function Shape-NWD; the calculation process is: Calculate the intersection over union (IoU) of the predicted box and the true box by calculating the intersection over union (IoU) of the two boxes. Calculate the comprehensive distance metric D: Calculate shape penalty NWD shape : In the formula, weight=2,w gt Indicates the width of the real box, h gt represents the height of the real box, w represents the width of the predicted box, and h represents the height of the predicted box. (x c ,y c ) represents the center coordinate of the prediction box, (x c gt ,y c gt ) represents the center point coordinates of the real box, h h and w w All represent coefficients, D represents comprehensive distance metric, and C represents a constant; The intersection over union (IoU), comprehensive distance metric D and shape penalty NWD shape Weighted fusion is performed to generate the final Shape-NWD loss value.
7. The particleboard surface defect detection method based on deep learning according to claim 1 is characterized in that: The evaluation indicators in step 5 include precision P, recall rate R, F1 value and mean average precision mAP.
Citation Information
Patent Citations
Wood defect detection method based on improved YOLOX model
CN115222685A
Insulator defect detection method in foggy day scene based on improved YOLOv7 algorithm
CN116843636A
Shaving board quality detection method and system based on image recognition
CN117152161A
Photovoltaic cell panel defect detection method and system based on improved YOLOv9s model
CN118918089A
Metal surface defect detection method
CN119251168A
Cited By
Shaving board grading method based on improved decision tree algorithm
CN120147258A
Oil and gas pipeline defect detection method and system
CN120766011A
Vehicle defect detection method based on fusion frequency adaptive expansion convolution
CN121074005A
Intelligent non-contact structure displacement detection method and system based on computer vision
CN121095167A
Computer vision-based intelligent non-contact structural displacement detection method and system
CN121095167B