A classification method for surface defect detection of prefabricated components based on target detection
Through the improved YOLOv8 deep learning network model, the problems of low efficiency and insufficient accuracy of surface defect detection of prefabricated components in the prior art are solved, and fast and accurate defect detection is achieved, reducing costs and missed detection rates.
Patent Information
- Application Number
- CN202510195359.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-02-21
AI Technical Summary
The existing methods for detecting surface defects of prefabricated components rely on manual measurement, which is inefficient and cost-effective, and is difficult to avoid missed inspection, which affects building quality and logistics costs.
The prefabricated component surface defect detection classification method is adopted based on object detection, and the plane images are enhanced by acquisition and data, the data set is constructed, and the improved YOLOv8 deep learning network model is used to detect, identify and classify surface defects.
It realizes rapid and accurate detection of surface defects of prefabricated components, improves detection efficiency and accuracy, reduces labor costs, and reduces missed detection rates.
Smart Images

Figure CN119672029B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of prefabricated building detection, and in particular to a prefabricated component surface defect detection and classification method based on target detection. Background Art
[0002] As an innovative and efficient construction method, prefabricated buildings have gradually become an important direction for solving urbanization and sustainable development of the construction industry. Prefabricated buildings refer to a construction method in which building components are prefabricated and preassembled in factories and assembled on site. Compared with traditional on-site construction, prefabricated buildings have the advantages of short construction time, controllable engineering quality, and reduced construction noise and dust.
[0003] Among them, the quality control of building prefabricated components is the core of ensuring the quality of prefabricated buildings. Any defect in a prefabricated component may inevitably affect the final building quality, and thus bring immeasurable losses to the entire construction project. The quality acceptance of prefabricated component finished products is the last line of defense to ensure the quality, safety and efficiency of prefabricated building construction. It can be said that the quality of the finished component acceptance determines the final quality of the component.
[0004] At present, the existing methods for detecting surface defects of prefabricated components are limited to manual measurement. Although universal, the labor cost is high, and hundreds of components are inspected every day, and there are at least dozens of surface defects on each component. Due to the subjective judgment, fatigue and inspection habits of inspectors, this is a slow, inefficient, expensive, subjective, inaccurate, and time-insensitive process, which is not suitable for automated production tasks; and the method of random inspection is difficult to avoid the problem of missed inspections, which will lead to returns and scrapping, which not only increases logistics costs, increases material and labor costs, but also further reduces the trust of supply, indirectly causing a lot of losses; in serious cases, any surface defect in a prefabricated component may have an inevitable impact on the final building quality, and then bring immeasurable losses to the entire construction project. How to quickly and accurately complete the surface defect detection of prefabricated components is an important issue in the production of concrete prefabricated components. Summary of the invention
[0005] The purpose of the present invention is to provide a prefabricated component surface defect detection and classification method based on target detection to solve the above-mentioned deficiencies in the prior art.
[0006] To achieve the above object, the present invention adopts the following technical solution:
[0007] A method for detecting and classifying surface defects of prefabricated components based on target detection comprises the following steps:
[0008] S1. Collecting plane images of prefabricated components and performing data enhancement processing and surface defect annotation to construct a plane image dataset;
[0009] S2. Build an improved YOLOv8 deep learning network model: Based on the YOLOv8 deep learning network model, integrate the RepGhost bottleneck structure into the C2f module in the Backbone network, denoted as the C2f_RepGhost module; add the EMA attention mechanism module in the Neck network, and replace the PAFPN network structure with the BiFPN_Concat network structure; optimize the bounding box regression loss function in the Head network;
[0010] S3. Use the improved YOLOv8 deep learning network model to detect the plane image dataset and identify the surface defects in prefabricated components.
[0011] Furthermore, the planar image dataset is constructed by the following steps:
[0012] S11, collecting a plane image of the prefabricated component by taking a photo;
[0013] S12, performing data amplification processing including but not limited to saturation adjustment, flipping, color conversion, and Gaussian blur on the collected planar image;
[0014] S13. Use the LabelImg annotation tool to respectively annotate the surface defects in the plane image after data augmentation processing, and construct a plane image dataset.
[0015] Furthermore, the improved YOLOv8 deep learning network model is recorded as LEF_YOLOv8, and the LEF_YOLOv8 deep learning network model is used to detect the plane image data set, and the surface defects in the prefabricated components are identified by the following steps:
[0016] S31, taking the plane image data set as the input feature map, extracting image features from the input feature map through the Conv layer, C2f_RepGhost module and SPPF module in the Backbone network;
[0017] S32, the feature map extracted by the Backbone network is used as the input of the Neck network, and feature fusion is performed through the EMA attention mechanism module, Upsample layer, BiFPN_Concat module, C2f module, and Conv layer to enhance the feature representation ability of the feature map;
[0018] S33. The feature map obtained by fusing the Neck network is used as the input of the Head network, and the surface defects and their categories in the prefabricated components are identified through the detection layer.
[0019] Furthermore, performing image feature extraction on the input feature map through the Backbone network specifically includes the following steps:
[0020] S311, first use two Conv layers to perform convolution downsampling on the feature map in turn and then input it into the C2f_RepGhost module;
[0021] S312: In the C2f_RepGhost module, the input feature map is firstly The convolution layer performs a convolution process, and then splits the input feature map into two parts in the channel dimension through the Split operation. One part is extracted and fused through a series of RepGhost bottleneck structures containing jump connections, and then spliced with the other part, and then through The convolution layer performs a convolution process, restores the concatenated feature map to its original dimension and outputs it;
[0022] The processing process of the C2f_RepGhost module is expressed by the following formula (1):
[0023] (1);
[0024] In formula (1): Represents the input of the C2f_RepGhost module; Represents the output of the C2f_RepGhost module; Represents the RepGhost bottleneck structure; Indicates the number of RepGhost bottleneck structures; represents the convolutional layer; Indicates that the splicing operation is performed on the channel dimension;
[0025] S313, inputting the feature graph processed by the C2f_RepGhost module into the SPPF module for processing;
[0026] The SPPF module processes the input feature map by the maximum pooling method. The specific processing process of the SPPF module is expressed by the following formula (2):
[0027] (2);
[0028] in:
[0029] (3);
[0030] (4);
[0031] (5);
[0032] (6);
[0033] In formula (2)-(6): Indicates output; Represents the convolution operation; Represents a splicing operation; Represents input; express Output after convolution operation; Represents the maximum pooling operation; Express Output after a maximum pooling operation; Express Output after two max pooling operations; Express Output after three max pooling operations.
[0034] Furthermore, performing image feature fusion on the feature map output by the Backbone network through the Neck network specifically includes the following steps:
[0035] S321. For any input feature map , divided into G sub-feature groups in the cross-channel dimension direction through the EMA attention mechanism module to learn different semantics;
[0036] The input feature map The grouping is recorded as:
[0037] , ;
[0038] S322, through two Branch and a Branch three parallel paths to extract attention weight descriptors of grouped feature maps;
[0039] The two The branch uses a 2D global average pooling operation to encode channel information in two spatial directions: one The branch uses a one-dimensional horizontal global average pooling operation, retaining the width dimension and compressing the height to 1. The branch uses a one-dimensional vertical global average pooling operation, retaining the height dimension and compressing the width to 1;
[0040] Said Branch through Convolutional processing to capture multi-scale feature representation;
[0041] S323, two The output of the branches is concatenated and After the convolution processing, the images are segmented according to height and width, and the height and width features are combined to perform group normalization.
[0042] S324, normalizing the grouping output and The output of the branch is subjected to global average pooling, shape transformation and softmax function application respectively;
[0043] S325, calculate the weights based on the output results of S324 and apply them to restore the shape of the grouped feature map to the original size output;
[0044] S326, using the output of the EMA attention mechanism module as the input feature map of the BiFPN_Concat module, and using different feature fusion paths to achieve the fusion of position information and semantic information;
[0045] The BiFPN_Concat module includes several layers of repeatedly stacked fusion blocks, and performs feature fusion through weighted feature fusion mechanism, convolution, and downsampling. The specific processing process of the BiFPN_Concat module is expressed by the following formula (7):
[0046] (7);
[0047] in:
[0048] (8);
[0049] In formulas (7)-(8): Indicates the number of layers of the fusion block; , , , , represents the weight parameter; represents the downsampling operation; Represents a parameter used to prevent numerical instability due to too small a weight parameter; represents the convolutional layer; Indicates The input features of the layer; Indicates The intermediate output features of the layer; express , as well as Output features after weighted fusion.
[0050] Furthermore, the bounding box regression loss function is improved based on the WIoU v3 loss function, denoted as BWIoU; the BWIoU loss function is expressed by the following formula (9):
[0051] (9);
[0052] in:
[0053] (10);
[0054] (11);
[0055] (12);
[0056] (13);
[0057] In formulas (9)-(13): represents the BWIoU loss function; represents intersection and union ratio; Represents the prediction box; represents the real frame; Represents the area of the intersection between the predicted box and the true box, Represents the area of the union of the predicted box and the true box; Represents the normalized distance between the center points of the predicted bounding box and the true bounding box; It represents the exponential function with the natural constant e as the base; and Represents the real box and prediction box The center point coordinates of and Represents the width and height of the minimum bounding box formed by the real box and the predicted box; represents the non-monotonic focusing factor; Represents the outlier value that measures the quality of the anchor box; and represents a hyperparameter; represents the constructed monotone focusing factor coefficient; represents the running average of momentum; is the set threshold, ; and Represents a hyperparameter, with a value range of .
[0058] It can be seen from the above technical solutions that the present invention has the following advantages compared with the prior art:
[0059] (1) The present invention introduces the C2f_RepGhost module in the Backbone network to more effectively extract the multi-scale feature information of the surface of the prefabricated component, thereby improving the performance of the target detection algorithm.
[0060] (2) The present invention avoids dimensionality reduction by introducing the EMA attention mechanism module in the Backbone network. It achieves comprehensive information retention and computational efficiency by reconstructing a part of the channels and evenly distributing spatial semantics among sub-features. It not only globally encodes information to adjust channel weights, but also captures pixel-level relationships through cross-dimensional interactions.
[0061] (3) The present invention replaces the PAFPN network structure with the BiFPN_Concat network structure in the Neck network, utilizes different feature fusion paths, realizes the fusion of position information and semantic information, and assigns higher weights to important parts, so that the network model focuses more on learning key feature information, thereby improving the learning ability of the network.
[0062] (4) The present invention optimizes the bounding box regression loss function of the network model and introduces hyperparameters based on the WIoU v3 loss function. , , by balancing the relative loss weights, reducing the problem of over-penalizing low-quality examples and under-penalizing high-quality examples; this can be achieved by adjusting the parameters , The re-weighting scheme can achieve different levels of box regression accuracy and make the detection model better adapt to different application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 It is a schematic flow chart of the method of the present invention;
[0064] Figure 2 It is a schematic diagram of the network structure of the improved YOLOV8 deep learning network model of the present invention;
[0065] Figure 3 It is a network structure diagram of the EMA attention mechanism module of the present invention;
[0066] Figure 4 It is a network structure diagram of the BiFPN_Concat module of the present invention;
[0067] Figure 5 This is a visual inspection and recognition effect diagram in an embodiment of the present invention;
[0068] Figure 6 This is a diagram of the visual inspection and recognition effect in an embodiment of the present invention. DETAILED DESCRIPTION
[0069] A preferred embodiment of the present invention is described in detail below with reference to the accompanying drawings.
[0070] like Figure 1 The method for detecting and classifying surface defects of prefabricated components based on target detection includes the following steps:
[0071] S1. Collecting plane images of prefabricated components and performing data enhancement processing and surface defect annotation to construct a plane image dataset;
[0072] This is achieved through the following steps:
[0073] S11, collecting a plane image of the prefabricated component by taking a photo;
[0074] S12, performing data amplification processing including but not limited to saturation adjustment, flipping, color conversion, and Gaussian blur on the collected planar image;
[0075] Data augmentation processing can effectively improve the generalization ability of the model: Saturation can improve the generalization ability of the model, reduce the risk of overfitting, and better reflect the diversity and complexity of the real world; flipping can allow the network model to receive component surface defect samples from more angles, helping the model learn more different perspectives; color transformation creates new image samples by changing the color, brightness, contrast and other attributes of the image. This transformation can make the model more robust to lighting conditions, color changes, etc.; Gaussian Blur can blur the details in the image, thereby reducing the detailed features of the image, allowing the model to focus more on the overall structure of the image rather than local details;
[0076] S13. Use the LabelImg annotation tool to respectively annotate the surface defects in the plane image after data augmentation processing, and construct a plane image dataset.
[0077] The plane image dataset described in this preferred embodiment is a self-constructed private dataset, named as a defect dataset (Detection Dataset, DD). The DD dataset is mainly used for the detection of surface defects of prefabricated components and the training and evaluation of network models.
[0078] S2. Build an improved YOLOv8 deep learning network model: Based on the YOLOv8 deep learning network model, integrate the RepGhost bottleneck structure into the C2f module in the Backbone network, denoted as the C2f_RepGhost module; add the EMA attention mechanism module in the Neck network, and replace the PAFPN network structure with the BiFPN_Concat network structure; optimize the bounding box regression loss function in the Head network; the improved YOLOv8 deep learning network model is denoted as LEF_YOLOv8.
[0079] S3. Using the improved YOLOv8 deep learning network model to detect the plane image data set, and identify the surface defects in the prefabricated components; specifically, the following steps are included:
[0080] S31. Take the plane image dataset as the input feature map, and extract image features from the input feature map through the Conv layer, C2f_RepGhost module and SPPF module in the Backbone network:
[0081] S311, first use two Conv layers to perform convolution downsampling on the feature map in turn and then input it into the C2f_RepGhost module;
[0082] S312: In the C2f_RepGhost module, the input feature map is firstly The convolution layer performs a convolution process, and then splits the input feature map into two parts in the channel dimension through the Split operation. One part is extracted and fused through a series of RepGhost bottleneck structures containing jump connections, and then spliced with the other part, and then through The convolution layer performs a convolution process, restores the concatenated feature map to its original dimension and outputs it;
[0083] The C2f_RepGhost module described in this preferred embodiment can more effectively extract multi-scale feature information of the surface of the prefabricated component, thereby improving the performance of the target detection algorithm; the processing process of the C2f_RepGhost module is expressed by the following formula (1):
[0084] (1);
[0085] In formula (1): Represents the input of the C2f_RepGhost module; Represents the output of the C2f_RepGhost module; Represents the RepGhost bottleneck structure; Indicates the number of RepGhost bottleneck structures; represents the convolutional layer; Indicates that the splicing operation is performed on the channel dimension;
[0086] S313, inputting the feature graph processed by the C2f_RepGhost module into the SPPF module for processing;
[0087] The SPPF module processes the input feature map by the maximum pooling method. The specific processing process of the SPPF module is expressed by the following formula (2):
[0088] (2);
[0089] in:
[0090] (3);
[0091] (4);
[0092] (5);
[0093] (6);
[0094] In formula (2)-(6): Indicates output; Represents the convolution operation; Represents a splicing operation; Represents input; express Output after convolution operation; Represents the maximum pooling operation; Express Output after a maximum pooling operation; Express Output after two max pooling operations; Express Output after three max pooling operations;
[0095] S32. The feature map extracted by the Backbone network is used as the input of the Neck network. Feature fusion is performed through the EMA attention mechanism module, Upsample layer, BiFPN_Concat module, C2f module, and Conv layer to enhance the feature representation capability of the feature map.
[0096] The EMA attention mechanism described in this preferred embodiment avoids dimensionality reduction, achieves comprehensive information retention and computational efficiency by reconstructing a portion of the channels and evenly distributing spatial semantics among sub-features; it not only globally encodes information to adjust channel weights, but also captures pixel-level relationships through cross-dimensional interactions. Figure 3 As shown, performing image feature fusion on the feature map output by the Backbone network through the Neck network specifically includes the following steps:
[0097] S321. For any input feature map , divided into G sub-feature groups in the cross-channel dimension direction through the EMA attention mechanism module to learn different semantics;
[0098] The input feature map The grouping is recorded as:
[0099] , ;
[0100] S322, through two Branch and a Branch three parallel paths to extract attention weight descriptors of grouped feature maps;
[0101] The two The branch uses a 2D global average pooling operation to encode channel information in two spatial directions: one The branch uses a one-dimensional horizontal global average pooling operation, retaining the width dimension and compressing the height to 1. The branch uses a one-dimensional vertical global average pooling operation, retaining the height dimension and compressing the width to 1;
[0102] Said Branch through Convolutional processing to capture multi-scale feature representation;
[0103] S323, two The output of the branches is concatenated and After the convolution processing, the images are segmented according to height and width, and the height and width features are combined to perform group normalization.
[0104] S324, normalizing the grouping output and The output of the branch is subjected to global average pooling, shape transformation and softmax function application respectively;
[0105] S325, calculate the weights based on the output results of S324 and apply them to restore the shape of the grouped feature map to the original size output;
[0106] S326, using the output of the EMA attention mechanism module as the input feature map of the BiFPN_Concat module, and using different feature fusion paths to achieve the fusion of position information and semantic information;
[0107] like Figure 4As shown, the BiFPN_Concat module includes several layers of repeatedly stacked fusion blocks, and performs feature fusion through weighted feature fusion mechanism, convolution, and downsampling. The specific processing process of the BiFPN_Concat module is expressed by the following formula (7):
[0108] (7);
[0109] in:
[0110] (8);
[0111] In formulas (7)-(8): Indicates the number of layers of the fusion block; , , , , represents the weight parameter; represents the downsampling operation; Represents a parameter used to prevent numerical instability due to too small a weight parameter; represents the convolutional layer; Indicates The input features of the layer; Indicates The intermediate output features of the layer; express , as well as Output features after weighted fusion;
[0112] S33. The feature map obtained by fusing the Neck network is used as the input of the Head network, and the surface defects and their categories in the prefabricated components are identified through the detection layer.
[0113] The present invention optimizes the bounding box regression loss function of the network model and introduces hyperparameters based on the WIoU v3 loss function. , , denoted as BWIoU (Balanced wise IoU); the BWIoU loss function is expressed by the following formula (9):
[0114] (9);
[0115] in:
[0116] (10);
[0117] (11);
[0118] (12);
[0119] (13);
[0120] In formulas (9)-(13): represents the BWIoU loss function; represents intersection and union ratio; Represents the prediction box; represents the real frame; Represents the area of the intersection between the predicted box and the true box, Represents the area of the union of the predicted box and the true box; Represents the normalized distance between the center points of the predicted bounding box and the true bounding box; It represents the exponential function with the natural constant e as the base; and Represent the real box and prediction box The center point coordinates of and Represents the width and height of the minimum bounding box formed by the real box and the predicted box; represents the non-monotonic focusing factor; Represents the outlier value that measures the quality of the anchor box; and Represents adjustable hyperparameters for fitting different models; represents the constructed monotone focusing factor coefficient; represents the running average of momentum; is the set threshold, ; and Represents the hyperparameter. Based on a large number of experimental results, the optimal hyperparameter range is .
[0121] Specifically, in the WIoU v3 loss function, the quality of the anchor box is is inversely proportional to the value of Defines the assigned anchor box value, thus affecting the weight of the high-quality anchor box in the total loss function. Large values mean that the anchor box quality is low, resulting in smaller gradient gains and reduced harmful gradients.
[0122] In order to make the positioning and size of the target bounding box more balanced, enhance the generalization ability of the model in different target scenarios, and effectively deploy it in different detection environments, this preferred embodiment introduces hyperparameters based on the WIoU v3 loss function. , , by adjusting the hyperparameters , To balance the weight of relative loss, reduce the problem of over-punishing low-quality examples and insufficient punishment for high-quality examples; and by adjusting the parameters , The re-weighting scheme can achieve different levels of box regression accuracy and make the detection model better adapt to different application scenarios.
[0123] Comparative experimental evaluation:
[0124] This preferred embodiment uses precision (P), recall (R), mean average precision (mAP), FLOPS, PARAMS and FPS as evaluation indicators to conduct an effect comparison experiment on the LEF_YOLOv8 deep learning network model. Among them, precision P is expressed as the proportion of the model's predicted real category samples that are actually true; recall R represents the proportion of all actual real category samples that the model successfully predicts as true; mAP stands for average precision, and mAP@0.5 means that the IOU threshold is set to 0.5 when calculating mAP, that is, the detection frame is correct only when the IOU between the detection frame and the real target is greater than 0.5 (this is a common threshold setting for object detection tasks); Params refers to the number of parameters in the model, which can evaluate the complexity and scale of the model; FLOPs is an indicator of computer hardware performance and algorithm complexity. The specific experimental results are shown in Table 1 below:
[0125] Table 1 Comparative experimental results
[0126]
[0127] According to the experimental results, it can be clearly shown that the LEF_YOLOv8 deep learning network model proposed in the present invention improves the accuracy of target detection while reducing the complexity of a certain model. Compared with the mainstream target detection algorithm and the lightweight target detection algorithm in the YOLO series, the LEF_YOLOv8 has significantly reduced parameters and FLOPs indicators, and achieved an average precision mAP of 89.98% for the detection of surface defects of prefabricated components. On the DD data set, compared with YOLOv8n, which has the least model parameters, the smallest computational complexity, and the highest average precision, it reduces the model parameters and computational complexity by 22.58% and 13.58%, respectively, and improves the average precision by 6.44%. In summary, the deep learning network model algorithm proposed in this paper is superior to other algorithms and performs better in completing the task of detecting surface defects of prefabricated components. The LEF_YOLOv8 model has achieved significant accuracy and precision in detecting surface defects of prefabricated components while maintaining a lightweight structure.
[0128] For specific visual inspection and recognition effects, see Figure 5 and Figure 6 .
[0129] The above-described embodiments are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary technicians in this field should all fall within the protection scope determined by the claims of the present invention.
Claims
1. A method for detecting and classifying surface defects of prefabricated components based on target detection, characterized in that: The following steps are involved: S1. Collecting plane images of prefabricated components and performing data enhancement processing and surface defect annotation to construct a plane image dataset; S2. Build an improved YOLOv8 deep learning network model: Based on the YOLOv8 deep learning network model, integrate the RepGhost bottleneck structure into the C2f module in the Backbone network, denoted as the C2f_RepGhost module; add the EMA attention mechanism module in the Neck network, and replace the PAFPN network structure with the BiFPN_Concat network structure; optimize the bounding box regression loss function in the Head network; S3, using the improved YOLOv8 deep learning network model to detect the plane image dataset and identify the surface defects in the prefabricated components; The improved YOLOv8 deep learning network model is recorded as LEF_YOLOv8. The LEF_YOLOv8 deep learning network model is used to detect the plane image data set, and the surface defects in the prefabricated components are identified by the following steps: S31, using the planar image data set as the input feature map, extracting image features from the input feature map through the Conv layer, C2f_RepGhost module and SPPF module in the Backbone network; S32, the feature map extracted by the Backbone network is used as the input of the Neck network, and feature fusion is performed through the EMA attention mechanism module, Upsample layer, BiFPN_Concat module, C2f module, and Conv layer to enhance the feature representation ability of the feature map; S33, using the feature map obtained by fusion of the Neck network as the input of the Head network, and identifying the surface defects and their categories in the prefabricated components through the detection layer; The image feature extraction of the input feature map through the Backbone network specifically includes the following steps: S311, first use two Conv layers to perform convolution downsampling on the feature map in turn and then input it into the C2f_RepGhost module; S312, in the C2f_RepGhost module, the input feature map is first convolved once through a 1×1 convolution layer, and then the input feature map is evenly split into two parts in the channel dimension through a Split operation, one part is spliced with the other part after feature extraction and fusion through a series of RepGhost bottleneck structures containing jump connections, and then convolved once through a 1×1 convolution layer, and the spliced feature map is restored to the original dimension and output; The processing process of the C2f_RepGhost module is expressed by the following formula (1): (1); In formula (1): Represents the input of the C2f_RepGhost module; Represents the output of the C2f_RepGhost module; Represents the RepGhost bottleneck structure; Indicates the number of RepGhost bottleneck structures; represents the convolutional layer; Indicates that the splicing operation is performed on the channel dimension; S313, inputting the feature graph processed by the C2f_RepGhost module into the SPPF module for processing; The SPPF module processes the input feature map by the maximum pooling method. The specific processing process of the SPPF module is expressed by the following formula (2): (2); in: (3); (4); (5); (6); In formula (2)-(6): Indicates output; Represents the convolution operation; Represents a splicing operation; Represents input; express Output after convolution operation; Represents the maximum pooling operation; Express Output after a maximum pooling operation; Express Output after two max pooling operations; Express Output after three max pooling operations.
2. According to claim 1, a method for detecting and classifying surface defects of prefabricated components based on target detection is characterized in that: The planar image dataset is constructed by the following steps: S11, collecting a plane image of the prefabricated component by taking a photo; S12, performing data amplification processing including but not limited to saturation adjustment, flipping, color conversion, and Gaussian blur on the collected planar image; S13. Use the LabelImg annotation tool to respectively annotate the surface defects in the plane image after data augmentation processing, and construct a plane image dataset.
3. According to the target detection-based prefabricated component surface defect detection and classification method of claim 1, it is characterized in that: The image feature fusion of the feature map output by the Backbone network through the Neck network specifically includes the following steps: S321. For any input feature map , divided into G sub-feature groups in the cross-channel dimension direction through the EMA attention mechanism module to learn different semantics; The input feature map The grouping is recorded as: , ; S322, through two Branch and a Branch three parallel paths to extract attention weight descriptors of grouped feature maps; The two The branch uses a 2D global average pooling operation to encode channel information in two spatial directions: one The branch uses a one-dimensional horizontal global average pooling operation, retaining the width dimension and compressing the height to 1. The branch uses a one-dimensional vertical global average pooling operation, retaining the height dimension and compressing the width to 1; Said Branch through Convolutional processing to capture multi-scale feature representation; S323, two The output of the branches is concatenated and After the convolution processing, the images are segmented according to height and width, and the height and width features are combined to perform group normalization. S324, normalizing the grouping output and The output of the branch is subjected to global average pooling, shape transformation and softmax function application respectively; S325, calculate the weights based on the output results of S324 and apply them to restore the shape of the grouped feature map to the original size output; S326, using the output of the EMA attention mechanism module as the input feature map of the BiFPN_Concat module, and using different feature fusion paths to achieve the fusion of position information and semantic information; The BiFPN_Concat module includes several layers of repeatedly stacked fusion blocks, and performs feature fusion through weighted feature fusion mechanism, convolution, and downsampling. The specific processing process of the BiFPN_Concat module is expressed by the following formula (7): (7); in: (8); In formulas (7)-(8): Indicates the number of layers of the fusion block; , , , , represents the weight parameter; represents the downsampling operation; Represents a parameter used to prevent numerical instability due to too small a weight parameter; represents the convolutional layer; Indicates The input features of the layer; Indicates The intermediate output features of the layer; express , as well as Output features after weighted fusion.
4. According to the target detection-based prefabricated component surface defect detection and classification method of claim 1, it is characterized in that: The bounding box regression loss function is improved based on the WIoUv3 loss function and is denoted as BWIoU. The BWIoU loss function is expressed by the following formula (9): (9); in: (10); (11); (12); (13); In formulas (9)-(13): represents the BWIoU loss function; represents intersection and union ratio; Represents the prediction box; represents the real frame; Represents the area of the intersection between the predicted box and the true box, Represents the area of the union of the predicted box and the true box; Represents the normalized distance between the center points of the predicted bounding box and the true bounding box; It represents the exponential function with the natural constant e as the base; and Represent the real box and prediction box The center point coordinates of and Represents the width and height of the minimum bounding box formed by the real box and the predicted box; represents the non-monotonic focusing factor; Represents the outlier value that measures the quality of the anchor box; and represents a hyperparameter; represents the constructed monotone focusing factor coefficient; represents the running average of momentum; is the set threshold, ; and Represents a hyperparameter, with a value range of .
Citation Information
Patent Citations
Photovoltaic station defect identification method, system and device based on target detection algorithm
CN118506220A
Drainage pipeline defect detection method based on improved YOLOv8s
CN119169378A