Reinforcing steel bar intersection point detection method based on lightweight improved YOLOv8
By introducing the ShuffleNetV2 backbone network and dual convolution optimization residual module into the YOLOv8 model, the real-time performance problem of YOLOv8 on mobile terminals is solved, and efficient and accurate recognition of steel bar intersection detection is achieved, which is suitable for complex construction scenarios.
Patent Information
- Application Number
- CN202510630452.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-10-10
AI Technical Summary
The existing YOLOv8 model is difficult to deploy efficiently on mobile or embedded devices, mainly because its backbone network CSPDarkNet has a complex structure and a large number of parameters, and the C2f module has parameter redundancy, resulting in insufficient real-time performance and difficulty meeting the needs of steel bar intersection recognition in complex construction scenarios.
ShuffleNetV2 is used to replace the backbone network structure of YOLOv8, and the C2f module is replaced by dual convolution. The residual module is optimized and the collaborative processing of parallel 3×3 group convolution and 1×1 point convolution is combined to reduce the computational complexity and parameter amount, and improve the lightweight degree of the model.
While maintaining high detection accuracy, it significantly reduces the scale of network parameters and the amount of floating-point operations, improves operational efficiency, and adapts the model to the deployment requirements of mobile or embedded devices, while ensuring real-time performance and accuracy in complex scenarios.
Smart Images

Figure CN120765952A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of building construction equipment, and in particular to a steel bar intersection detection method based on a lightweight improved YOLOv8. Background Art
[0002] In the fields of building construction and bridge engineering, accurate identification of rebar intersections is critical for ensuring structural quality and safety. Existing methods often rely on manual inspection or traditional image processing techniques, which are inefficient and susceptible to environmental interference. With the development of deep learning technology, automated methods based on object detection models have gradually become mainstream solutions. YOLOv8, the most advanced single-stage object detection algorithm in the YOLO family, leverages its efficient end-to-end detection framework to address complex scenarios such as densely distributed and morphologically similar rebar intersections, meeting the demand for real-time detection on construction sites. However, while YOLOv8 offers improved automation, its backbone network, CSPDarkNet, has a complex structure and large number of parameters, making it difficult to efficiently deploy on mobile or embedded devices, resulting in limited real-time performance. Furthermore, the residual design of the C2f module in YOLOv8 introduces parameter redundancy, further limiting the model's applicability in resource-constrained environments. Therefore, a lightweight and improved method that optimizes the characteristics of rebar intersections is urgently needed to achieve efficient and accurate identification in complex construction scenarios. Summary of the Invention
[0003] The purpose of the present invention is to provide a steel bar intersection detection method based on lightweight improved YOLOv8 to solve the problems raised in the above background technology.
[0004] To achieve the above object, the present invention provides the following technical solutions:
[0005] A steel bar intersection recognition method based on a lightweight improved YOLOv8, including:
[0006] Step 1: Collect steel bar image data to create a data set, and divide it into a training set and a test set;
[0007] Step 2: Image preprocessing and annotation;
[0008] Step 3: Reconstruct the YOLOv8 backbone network: use ShuffleNetV2 instead of CSPDarkNet;
[0009] Step 4: Optimize the residual module: Design a dual convolution to replace the C2f module, and enhance the fusion of small target features through the collaborative processing of parallel 3×3 group convolution and 1×1 point convolution;
[0010] Step 5: The marked training set is trained using the improved YOLOv8 model to obtain a trained steel bar intersection recognition YOLOv8 model;
[0011] Step 6: Input the test set into the trained steel bar intersection YOLOv8 model to perform steel bar intersection detection test.
[0012] Furthermore, step 1 includes: constructing multiple scenes to collect steel bar mesh images, so that the data set is closer to the actual application scenario.
[0013] Furthermore, the step 2 includes: using Labelimg software to label the collected steel bar images.
[0014] Furthermore, in step 3, ShuffleNetV2 is used to replace the backbone network structure of YOLOv8. The basic unit network structure of ShuffleNetV2 includes a basic unit structure and a downsampling basic unit. In the basic unit structure, ShuffleNetV2 introduces a channel splitting operation to divide the input feature map into two left and right branches.
[0015] ShuffleNetV2 controls the number of input channels and output channels to be equal, thereby achieving the minimum memory access cost.
[0016] Furthermore, in step 4, dual convolution is used to improve the residual block structure. In the dual convolution, the N convolution kernels are divided into G groups, and each group performs calculations on all channels of the input feature map. The standard convolution calculation amount FL std The formula is:
[0017]
[0018] Where D0 represents the width × height of the feature map, K represents the width × height of the convolution kernel, M represents the number of input channels, and N represents the number of output channels, that is, the number of convolution kernels;
[0019] The convolution calculation formula using dual convolution is:
[0020]
[0021]
[0022] Among them, G represents the number of groups of output channels, FL Dual_3x3 Indicates the computational cost of 3×3 convolution, FL Dual_1x1 Indicates the computational cost of 1×1 convolution, FL Dual Indicates the total amount of calculation;
[0023] Comparing the computational cost of the dual convolution and standard convolution layers, we can derive from Equations 4 and 7 that
[0024] Formula 8, the reduction rate Rd of the calculation amount is:
[0025]
[0026] Furthermore, in step 5, the AdamW optimizer is used to train the model and save the optimal model result.
[0027] Furthermore, the step 6 includes:
[0028] Step 6.1: Input the test set into the optimal model saved in step 5 for testing;
[0029] Step 6.2: Calculate the model performance indicators: accuracy (P), recall (R), mean average precision (mAP), number of parameters, computational complexity (GFLOPs), and model size. The specific formulas are as follows;
[0030]
[0031] Where TP is a true positive, i.e., the model predicts a positive instance, and it is actually a positive instance; FP is a false positive, i.e., the model predicts a positive instance, but it is actually a negative instance.
[0032]
[0033] Where FN is the false negative, that is, the model predicts a negative instance, but it is actually a positive instance;
[0034]
[0035] Where n is the number of target detection categories, AP i is the AP value of each classification, AP stands for average precision, and P stands for accuracy;
[0036] Step 6.3: Verify the performance indicators of the test set. If the accuracy of the test set is close to that of the training set, it means that the model meets the generalization requirements, and finally a lightweight steel bar intersection detection model based on the improved YOLOv8 is obtained.
[0037] Compared with the prior art, the beneficial effects of the present application are: the ShuffleNetV2 network structure is introduced, and the YOLOv8 backbone network structure is replaced by ShuffleNetV2, which effectively reduces the network parameter scale and floating point operation amount on the premise of ensuring the multi-scale feature fusion capability; the DualConv module is introduced to improve the model lightweight of the residual block structure of YOLOv8, and the C2f function is replaced by C2f_Dual function, and the improved residual block structure performs more outstanding in calculation efficiency and model lightweight. Through the two improvement methods, YOLOv8 improves the running efficiency while maintaining high detection accuracy, providing a feasible lightweight scheme for the extraction of steel bar intersection image information. BRIEF DESCRIPTION OF DRAWINGS
[0038] Figure 1 The flowchart of the present application is shown in the figure.
[0039] Figure 2 The ShuffleNetV2 basic unit structure diagram in the present application is shown in the figure, where A is the basic unit structure, and B is the down-sampling basic unit.
[0040] Figure 3 The DualConv convolution filter design diagram in the present application is shown in the figure.
[0041] Figure 4 The improved YOLOv8 framework diagram in the present application is shown in the figure. DETAILED DESCRIPTION
[0042] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0043] Please refer to Figures 1-4 A steel bar intersection recognition method based on lightweight improved YOLOv8, comprising:
[0044] Step 1, collect steel bar image data to make a data set, and distribute it into a training set and a test set, including: building a multi-scenario experimental environment to improve the reliability and scene adaptability of the data set. By building indoor and outdoor environments, various lighting conditions that the steel bar grid may encounter in different application scenarios are simulated. Steel bar grid images are collected under different lighting intensities in these sites, making the data set more close to the actual application scenario.
[0045] Step 2: Use Labelimg to preprocess and annotate the rebar image dataset. This includes annotating the collected rebar images using Labelimg software. The intersection areas in each image are framed and classified according to whether the rebar intersections are bundled. Rebar intersections in the image require two types of annotation: unbundled intersections and bundled intersections. The annotation results are stored in XML format and managed using the VOC format. Once the sample set is annotated, it is automatically saved as an XML tag file. Each XML file contains important parameters such as the image file name, annotation category, coordinate information of the framed area, and image size.
[0046] Step 3: Reconstruct the YOLOv8 backbone network: Use ShuffleNetV2 to replace CSPDarkNet, reduce computational redundancy through channel segmentation and dual-branch structure, combine depthwise separable convolution with channel shuffling operations, and enhance multi-scale feature extraction capabilities while reducing memory usage, thereby improving the positioning accuracy of dense intersections.
[0047] Among them, such as Figure 2 As shown, this step replaces the YOLOv8 backbone network structure with ShuffleNetV2. Its basic unit network structure includes both a basic unit structure and a downsampling basic unit. In the basic unit structure, ShuffleNetV2 introduces a channel split operation, splitting the input feature map into two branches, left and right. This operation is combined with the convolutional layer of the right branch via the identity mapping of the left branch, avoiding redundant multi-channel convolution operations and reducing memory access. The right branch consists of three convolutional layers. The two 1×1 convolutional layers do not use grouped convolution to avoid memory access performance bottlenecks caused by excessive convolution grouping. The 3×3 convolutional layer uses depthwise separable convolution. In the downsampling basic unit, the channel split operation is removed, and depthwise separable convolution and pointwise convolution layers are added to the left branch. In addition, downsampling operations with a stride of 2 are applied to both branches. The outputs of the two branches are combined through a merge operation, doubling the number of channels to accommodate the halved feature map size, and finally enter the channel shuffling module.
[0048] ShuffleNetV2 keeps the number of input channels and output channels equal to achieve the minimum memory access cost (MAC). Assuming the size of the input feature map is H·W·C1, its convolution kernel shape is (C1, C2, 1, 1), and the H and W of the output feature map remain unchanged, the computational cost of a 1x1 convolution can be expressed as B:
[0049] B=H·W·C1·C2 (Formula 1)
[0050] Therefore, the MAC of 1x1 convolution can be expressed as Equation 2.
[0051] MAC=H·W·C1+H·W·C2+C1·C2=H·W·(C1+C2)+C1·C (Formula 2)
[0052] In Equation 2, the terms H·W·C1, H·W·C2, and C1·C represent the memory access costs of the input feature map, output feature map, and weight parameters, respectively. Equation 2 and the mean inequality allow us to derive Equation 3.
[0053]
[0054] From formula 3, we can see that when C1=C2, (C1+C2) 2 =4·C1·C2, at this time (C1+C2) 2 The lower limit is taken, so MAC can achieve the minimum value only when the number of input and output channels is equal.
[0055] Step 4: Optimize the residual module: Figure 3 As shown in the figure, a dual convolution (DualConv) is designed to replace the C2f module. Through the coordinated processing of parallel 3×3 group convolution and 1×1 point convolution, it strengthens the fusion of small target features, utilizes the sparse nature of grouping to reduce the number of parameters, and avoids the loss of detailed information about rebar intersections. This solution significantly reduces the model's computational complexity and energy consumption, achieving high-precision, real-time detection of rebar intersections in complex construction scenarios, meeting the deployment requirements of edge devices.
[0056] Among them, DualConv is used to improve the residual block structure. In DualConv, N convolution kernels are divided into G group structures, and each group performs calculations on all channels of the input feature map. Among them, M / G channels are processed in parallel by 3×3 and 1×1 convolution kernels, while the remaining (MM / G) channels only apply 1×1 convolution. The output features of the two groups of paths are superimposed element by element in the channel dimension, and the diagonal sparse characteristics of the channel block of the convolution kernel are enhanced through the grouping mechanism, so that the convolution kernel with strong correlation can be efficiently learned through a more structured parameter distribution. The standard convolution calculation amount FL std The formula is:
[0057]
[0058] Where D0 represents the width × height of the feature map, K represents the width × height of the convolution kernel, M represents the number of input channels, and N represents the number of output channels, that is, the number of convolution kernels.
[0059] The convolution calculation formula using DualConv is:
[0060]
[0061] Among them, G represents the number of groups of output channels, FL Dual_3x3Indicates the computational cost of 3×3 convolution, FL Dual_1x1 Indicates the computational cost of 1×1 convolution, FL Dual Indicates the total amount of computation.
[0062] Comparing the computational complexity of the dual convolution layer and the standard convolution layer, we can derive (Equation 8) from (Equation 4) and (Equation 7), and the reduction rate Rd of the computational complexity is:
[0063]
[0064] Step 5: Use the improved YOLOv8 model to train the labeled training set to obtain a trained YOLOv8 model for steel bar intersection recognition. The AdamW optimizer is used for model training and the optimal model result is saved.
[0065] Step 6: Input the rebar image test set into the trained rebar intersection YOLOv8 model to perform a rebar intersection detection test, including:
[0066] Step 6.1: Input the test set into the optimal model saved in step 5 for testing;
[0067] Step 6.2: Calculate the model performance indicators: accuracy (P), recall (R), mean average precision (mAP), number of parameters, computational complexity (GFLOPs), and model size. The specific formulas are as follows;
[0068]
[0069] Where TP is a true positive, i.e., the model predicts a positive instance, and it is actually a positive instance; FP is a false positive, i.e., the model predicts a positive instance, but it is actually a negative instance.
[0070]
[0071] In the formula, FN (False Negative) is a false negative, that is, the model predicts a negative instance, but it is actually a positive instance.
[0072] In the formula, FN (False Negative) is a false negative, that is, the model predicts a negative instance, but it is actually a positive instance.
[0073]
[0074] Where n is the number of target detection categories, AP i is the AP value of each category, AP represents the average precision, and P represents the accuracy.
[0075]
[0076] Step 6.3: Verify the performance indicators of the test set. If the accuracy of the test set is close to that of the training set, it means that the model meets the generalization requirements, and finally a lightweight steel bar intersection detection model based on the improved YOLOv8 is obtained.
[0077] In this embodiment, in order to verify the effect of the improved model, the test set after the constructed data set is divided into 8:1:1 is input into the lightweight steel bar intersection detection model based on the improved YOLOv8 and other models for testing. The evaluation results are shown in Table 1.
[0078] Table 1 Comparative experimental results
[0079]
[0080]
[0081] Comparing the data in Table 1, the proposed rebar intersection recognition model demonstrates remarkable lightweight performance: its mAP@50 reaches 0.9774, which is slightly lower than YOLOv8's 0.9888, but outperforms mainstream models such as YOLOv7 and YOLOv6, demonstrating its competitive detection accuracy. Furthermore, the model's GFLOPs and model size are 69.3M and 40.8MB, respectively, demonstrating significant lightweight advantages over other models. This improvement significantly reduces resource utilization and energy consumption by optimizing the network structure and computational redundancy, making the model more suitable for deployment on mobile or embedded devices while ensuring the real-time and accuracy of rebar intersection detection in complex scenarios.
[0082] It should be noted that the experimental environment is based on the Ubuntu 22.04-LST operating system, uses the A5000 graphics card for model training, and deploys the model to the RK3588 development board.
[0083] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A steel bar intersection recognition method based on lightweight improved YOLOv8, characterized in that: include: Step 1: Collect steel bar image data to create a data set, and divide it into a training set and a test set; Step 2: image preprocessing and annotation; Step 3: Reconstruct the YOLOv8 backbone network: use ShuffleNetV2 instead of CSPDarkNet; Step 4: Optimize the residual module: Design a dual convolution to replace the C2f module, and enhance the fusion of small target features through the collaborative processing of parallel 3×3 group convolution and 1×1 point convolution; Step 5: The marked training set is trained using the improved YOLOv8 model to obtain a trained steel bar intersection recognition YOLOv8 model; Step 6: Input the test set into the trained steel bar intersection YOLOv8 model to perform steel bar intersection detection test.
2. A steel bar intersection recognition method based on lightweight improved YOLOv8 according to claim 1, characterized in that, The step 1 includes: constructing multiple scenes to collect steel bar mesh images, so that the data set is closer to the actual application scene.
3. A steel bar intersection recognition method based on lightweight improved YOLOv8 according to claim 1, characterized in that, The step 2 includes: using Labelimg software to mark the collected steel bar images.
4. A steel bar intersection recognition method based on lightweight improved YOLOv8 according to claim 1, characterized in that, In step 3, ShuffleNetV2 is used to replace the backbone network structure of YOLOv8. The basic unit network structure of ShuffleNetV2 includes a basic unit structure and a downsampling basic unit. In the basic unit structure, ShuffleNetV2 divides the input feature map into two branches, left and right, by introducing a channel splitting operation. ShuffleNetV2 controls the number of input channels and output channels to be equal, thereby achieving the minimum memory access cost.
5. A steel bar intersection recognition method based on lightweight improved YOLOv8 according to claim 1, characterized in that, In step 4, dual convolution is used to improve the residual block structure. In the dual convolution, the N convolution kernels are divided into G groups. Each group performs calculations on all channels of the input feature map. The standard convolution calculation amount FL std The formula is: Where D0 represents the width × height of the feature map, K represents the width × height of the convolution kernel, M represents the number of input channels, and N represents the number of output channels, that is, the number of convolution kernels; The convolution calculation formula using dual convolution is: Among them, G represents the number of groups of output channels, FL Dual_3x3 Indicates the computational cost of 3×3 convolution, FL Dual_1x1 Indicates the computational cost of 1×1 convolution, FL Dual Indicates the total amount of calculation; Comparing the computational complexity of the dual convolution layer and the standard convolution layer, we can derive Equation 8 from Equation 4 and Equation 7. The computational reduction rate Rd is:
6. A steel bar intersection recognition method based on lightweight improved YOLOv8 according to claim 1, characterized in that, In step 5, the AdamW optimizer is used to train the model and save the optimal model result.
7. A steel bar intersection recognition method based on lightweight improved YOLOv8 according to claim 1, characterized in that, The step 6 includes: Step 6.1: Input the test set into the optimal model saved in step 5 for testing; Step 6.2: Calculate the model performance indicators: accuracy (P), recall (R), mean average precision (mAP), number of parameters, computational complexity (GFLOPs), and model size. The specific formulas are as follows; Where TP is a true positive, i.e., the model predicts a positive instance, and it is actually a positive instance; FP is a false positive, i.e., the model predicts a positive instance, but it is actually a negative instance. Where FN is the false negative, that is, the model predicts a negative instance, but it is actually a positive instance; Where n is the number of target detection categories, AP i is the AP value of each classification, AP stands for average precision, and P stands for accuracy; Step 6.3: Verify the performance indicators of the test set. If the accuracy of the test set is close to that of the training set, it means that the model meets the generalization requirements, and finally a lightweight steel bar intersection detection model based on the improved YOLOv8 is obtained.