Steel bar binding point detection method based on improved YOLOv8
By improving the local feature enhancement module, bidirectional multi-scale fusion network and task decoupling detection module of YOLOv8, the problem of insufficient accuracy of traditional YOLOv8 in steel bar binding point detection is solved, and efficient and accurate steel bar binding point detection is achieved.
Patent Information
- Application Number
- CN202510874639.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-06-27
AI Technical Summary
Traditional YOLOv8 has problems in the detection of rebar binding points, such as insufficient local detail modeling capabilities, inefficient multi-scale feature fusion, and high task coupling of the detection head, resulting in poor detection accuracy.
YOLOv8 is improved by adopting local feature enhancement module, bidirectional multi-scale fusion network and task decoupling detection module. The steel bar intersection features are enhanced through local spatial attention calculation, depthwise separable convolution and channel attention weighted operations. Multi-scale features are fused with dynamic weights, and independent branches are used to handle classification and regression tasks.
It improves the recall rate of small target detection, reduces the false detection rate, improves the detection accuracy and bounding box positioning accuracy of cross-scale targets, reduces the time consumption of feature fusion calculation, and meets the needs of real-time detection.
Smart Images

Figure CN120726296A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of steel bar node detection under a single layer of steel bars, and specifically to a steel bar binding point detection method based on improved YOLOv8. Background Art
[0002] In the construction industry, quality inspection of rebar binding points is a critical step in concrete structure construction, directly impacting the safety and durability of buildings. Traditional inspection methods rely on manual visual inspection, which suffers from low efficiency, high rates of missed inspections, and strong subjectivity.
[0003] In recent years, deep learning-based object detection technologies (such as the YOLO series of algorithms) have been introduced into this field due to their high efficiency. YOLOv8, as the latest version, has performed well in general object detection tasks. However, in the specific scenario of rebar binding points, YOLOv8 has the following limitations: Inadequate local detail modeling capabilities: Rebar intersections are typically small objects (approximately 5-10 pixels in diameter) with textures similar to the background (such as wooden formwork and concrete debris). YOLOv8's native C2f module's feature extraction method struggles to effectively distinguish these subtle differences, resulting in a high rate of missed detection of small objects. Inefficient multi-scale feature fusion: The traditional unidirectional feature pyramid network (PAN) only fuses features through a top-down path. In scenarios with densely arranged steel bars, the accuracy of cross-scale object detection decreases significantly. High detection head task coupling: Classification and regression tasks share a symmetrical detection head, which causes the feature expressions of the two tasks to interfere with each other, increasing bounding box positioning errors. In summary, these limitations of traditional YOLOv8 lead to low detection accuracy of steel bar binding points. Summary of the Invention
[0004] In view of the above-mentioned defects or deficiencies in the prior art, this application aims to provide a steel bar binding point detection method based on improved YOLOv8 to improve the detection accuracy of steel bar binding points; the detection method includes the following steps: Acquire a construction site image, and pre-process the construction site image to obtain an image to be detected; The image to be detected is input into the improved target detection model, and the lashing point detection is performed through the following collaborative improvement structure: The local feature enhancement module sequentially performs local spatial attention calculation, depthwise separable convolution, and channel attention weighting operations to enhance the detailed feature expression of steel bar intersections; A bidirectional multi-scale fusion network dynamically fuses the top-down first feature path and the bottom-up second feature path, and embeds lightweight residual units at the fusion nodes; the first feature path corresponds to high-level semantic features, and the second feature path corresponds to low-level detail features; The task decoupling detection module uses a classification branch and a regression branch to process classification tasks and regression tasks respectively. The classification branch outputs category confidence through continuous 1×1 convolution, and the regression branch outputs coordinate positioning through depthwise separable convolution and probability distribution prediction layer; The output data of the target detection model is decoded to generate a detection result including the location and category of the binding point.
[0005] According to the technical solution provided in the embodiment of the present application, the preprocessing of the construction site image includes the following steps: Random rotation, scale scaling, mosaic enhancement and HSV color space perturbation are sequentially performed on the construction site image.
[0006] According to the technical solution provided in the embodiment of the present application, the local spatial attention calculation includes the following steps: The image to be detected is divided into multiple local windows, and the attention weights of the query matrix, key matrix, and value matrix are calculated in each local window. The calculation formula is: ; Among them, Q is the query matrix, K is the key matrix, V is the value matrix, d k The key vector dimension is ,and the window partition adopts the sliding window strategy, and the ,window size is a configurable parameter ranging from 5×5 to 7×7.
[0007] According to the technical solution provided in the embodiment of the present application, the regression branch outputs coordinate positioning through depthwise separable convolution and probability distribution prediction layer, including the following steps: The first feature map output by the bidirectional multi-scale fusion network is input into three layers of depth-wise separable convolutional layers connected in series to obtain a second feature map; the convolution kernel size of each layer is 3×3, the stride is 1, and batch normalization and SiLU activation functions are inserted; Modeling the bounding box coordinates in the second feature map as a discrete probability distribution through the probability distribution prediction layer, wherein the bounding box coordinates include the horizontal and vertical coordinates of the center point and the width and height of the target binding point; The discrete probability distribution is converted into continuous space coordinates through integration operation to obtain coordinate positioning.
[0008] According to the technical solution provided in an embodiment of the present application, modeling the bounding box coordinates in the second feature map as a discrete probability distribution through the probability distribution prediction layer includes the following steps: Input the second feature map into the fully connected layer to generate a vector of dimension 4×reg_max, where 4 represents the horizontal and vertical coordinates of the center point and the four coordinate parameters of width and height, and reg_max is the preset number of discrete distribution interval segmentation points, ranging from 8 to 16; A Softmax operation is performed on the reg_max-dimensional vector of each of the coordinate parameters to obtain a discrete probability distribution.
[0009] According to the technical solution provided in the embodiment of the present application, the lightweight residual unit includes two 1×1 convolutional layers connected in series, with a batch normalization layer and a SiLU activation function inserted in the middle, and a residual connection is established between the input and output. The DropPath random depth drop mechanism is set on the residual path, and the drop probability is 0.1 to 0.3.
[0010] According to the technical solution provided in the embodiment of the present application, the classification branch sequentially includes two layers of 1×1 convolutional layers, each layer is followed by batch normalization and SiLU activation function, and the number of output channels is equal to the total number of categories.
[0011] According to the technical solution provided in the embodiment of the present application, the formula for dynamic weight fusion is: ,in, 、 is a trainable scalar parameter, , F1 is the high-level semantic feature, and F2 is the low-level detail feature.
[0012] According to the technical solution provided in the embodiment of the present application, the training loss function of the task decoupling detection module includes: the classification loss corresponding to the classification branch adopts binary cross entropy loss, and the regression loss corresponding to the regression branch adopts distribution focus loss.
[0013] According to the technical solution provided in the embodiment of the present application, the improved target detection model includes the following lightweight deployment: Channel pruning: remove redundant channels in the feature map with weight sparsity higher than 0.8; Parameter quantization: Convert 32-bit floating-point weights to 8-bit fixed-point numbers using symmetric uniform quantization; Operator fusion: Combines convolutional layers, batch normalization layers, and activation functions into a single computational node.
[0014] In summary, this application proposes a method for detecting steel bar binding points based on improved YOLOv8, which includes the following steps: acquiring and preprocessing construction site images, inputting the images to be detected into the improved target detection model, and sequentially performing local spatial attention calculation, depthwise separable convolution, and channel attention weighted operations to enhance the detailed feature expression of steel bar intersections; dynamically weighting the top-down first feature path and the bottom-up second feature path, and embedding lightweight residual units at the fusion nodes; using classification branches and regression branches to process classification tasks and regression tasks respectively, the classification branch outputs category confidence, and the regression branch outputs coordinate positioning through depthwise separable convolution and probability distribution prediction layer; decoding the output data of the target detection model to generate detection results including the binding point location and category.
[0015] Compared with existing technologies, this application demonstrates the following advantages: This method utilizes a local feature enhancement module through a cascade of "sliding window attention → depthwise separable convolution → channel weighting" to specifically enhance the spatial correlation and channel saliency of rebar intersections. This improves the recall rate and reduces the false detection rate for small object detection in densely rebar-dense areas (intersection spacing <20 pixels). Furthermore, the bidirectional multi-scale fusion network (BiFPN) incorporates a dynamic weight fusion mechanism, balancing the feature contributions of top-down and bottom-up paths through learnable parameters. This improves the average accuracy of cross-scale object detection (e.g., clear lashing points in the foreground and blurred nodes in the distance) and reduces the computational time required for feature fusion. Furthermore, the task-decoupled detection module utilizes independent branches for classification and regression. The classification branch focuses on class confidence through 1×1 convolution, while the regression branch achieves sub-pixel localization through probability distribution prediction. The bounding box localization error is reduced from 6.8 pixels in the original YOLOv8 to 3.2 pixels, and the correlation coefficient between classification confidence and localization accuracy is reduced. The application of lightweight residual units and depth-wise separable convolutions also compresses the model parameters and improves the inference speed on edge devices to meet real-time detection needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 This is a flowchart of the steps of the steel bar binding point detection method based on improved YOLOv8 provided in an embodiment of the present application. DETAILED DESCRIPTION
[0017] The present application will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the relevant invention and are not intended to limit the invention. It should also be noted that, for ease of description, only portions relevant to the invention are shown in the accompanying drawings.
[0018] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0019] Example 1 As mentioned in the background technology, in order to solve the problems in the prior art, this application proposes a steel bar binding point detection method based on improved YOLOv8, such as Figure 1 As shown, the following steps are included: S1. Acquire a construction site image and pre-process the construction site image to obtain an image to be detected; Specifically, we use the OpenCV library to read construction site images and perform the following image processing: brightness normalization: using histogram equalization (CLAHE algorithm) to balance lighting variations; size normalization: resizing the images to a uniform 640×640 pixels and padding them to a 1:1 aspect ratio. This eliminates interference from ambient lighting and shooting angle on the model input, ensures consistent input tensor size, and improves the model's robustness to low-light and backlit scenes. Preprocessing takes less than 10ms (for 1080P images).
[0020] S2. Input the image to be detected into the improved target detection model, and perform lashing point detection through the following collaborative improvement structure: S2-1, local feature enhancement module, sequentially performs local spatial attention calculation, depthwise separable convolution, and channel attention weighting operations to enhance the detailed feature expression of steel bar intersections; Furthermore, a direction-sensitive convolution kernel group is introduced into the local feature enhancement module, and the depth-wise separable convolution is used as a direction-sensitive depth-wise separable convolution. The direction-sensitive depth-wise separable convolution uses a multi-directional convolution kernel group, including constrained convolution in four directions: horizontal, vertical, and ±45°. The feature responses of each direction are fused through dynamic weights. The output feature map of the direction-wise depth-wise separable convolution is adjusted for the number of channels through a 1×1 convolution and then input into the channel attention module to ensure the compatibility of the channel weighted operation. This further optimizes the "cross-shaped" geometric features specifically for the intersection of steel bars to improve detection accuracy. Specific implementation methods include: In the depthwise separable convolution stage, a multi-directional convolution kernel group (horizontal, vertical, ±45° four directions) is used to generate a multi-directional feature map; a 3×3 learnable convolution kernel is configured in each direction, and the formula is expressed as: ⊙
[0021] in, Represents the learnable convolution kernel (size 3×3) corresponding to direction θ. θ takes four directions: 0° (horizontal), 45°, 90° (vertical), and 135°. Each direction is initialized independently. The direction-constrained convolution operation is adaptively adjusted through training. dir Indicates the direction constrained convolution operation, based on the conventional convolution, through the mask Force the convolution kernel to only respond to features in a specific direction. For example, a 0° convolution kernel only retains the gradient response in the horizontal direction. represents the direction mask matrix (3×3), which is used to suppress non-main direction responses, X represents the input feature map, the enhanced features from the output of the local spatial attention module; ⊙ represents the element-by-element multiplication (Hadamard product) for applying the direction mask.
[0022] Description of the generation process of multi-directional feature maps: Design the initial convolution kernel for each direction θ, for example, the initial weight of the 0° direction is biased towards horizontal edge detection (similar to the Sobel horizontal kernel). Perform convolution calculations on the input feature map X in four directions respectively, and after each convolution, perform convolution with the corresponding Multiply them together to get the feature maps of four directions. The weight of each direction is calculated by the following formula , add the feature maps of the four directions according to the weights to obtain a multi-directional feature map; Dynamic weight allocation: The weight coefficients of the convolution kernels in each direction are obtained through online learning. The calculation formula is: ; Among them, GAP is a global average pooling operation to achieve adaptive direction weight adjustment; Represents the standard convolution result without applying the direction mask, which is used to evaluate the activation strength of the convolution kernel in that direction; exp represents the exponential function, which is used to amplify the weight difference in the significant direction. The denominator: Softmax normalizes the weights of the four directions to ensure .
[0023] The multi-directional feature map is channel-concatenated with the local attention feature (i.e., enhanced feature), and feature reorganization is achieved through 1×1 convolution.
[0024] It should be noted that the direction-sensitive depthwise separable convolution is an improvement to the original depthwise separable convolution in YOLOv8. By introducing direction constraints and dynamic weight fusion, it improves the geometric feature extraction capability without significantly increasing the order of parameters. According to actual measurements, this design improves the distinction between intersections and linear steel bars by 37.2%, and reduces the error rate of misidentifying ordinary intersections as binding points to 1.8%.
[0025] S2-2, a bidirectional multi-scale fusion network, dynamically weights the top-down first feature path and the bottom-up second feature path, and embeds a lightweight residual unit at the fusion node; wherein the first feature path corresponds to high-level semantic features, and the second feature path corresponds to low-level detail features; S2-3, task decoupling detection module, uses classification branch and regression branch to process classification task and regression task respectively, wherein the classification branch outputs category confidence through continuous 1×1 convolution, and the regression branch outputs coordinate positioning through depthwise separable convolution and probability distribution prediction layer; In a preferred embodiment, the local spatial attention calculation comprises the following steps: The image to be detected is divided into multiple local windows, and the attention weights of the query matrix, key matrix, and value matrix are calculated in each local window. The calculation formula is: ; Among them, Q is the query matrix, K is the key matrix, V is the value matrix, d k The key vector dimension is ,and the window partition adopts the sliding window strategy, and the ,window size is a configurable parameter ranging from 5×5 to 7×7.
[0026] Specifically, the input image to be detected (such as 80×80×256) is divided into 7×7 local windows (each window is 11×11 pixels), and a sliding window strategy (step size 7 pixels) is adopted. Edge overlap is allowed. Within each window, the query matrix Q, key matrix K, and value matrix V are generated, and the weighted feature map (that is, enhanced feature) is output, retaining the window edge information.
[0027] Specifically, the local feature enhancement module: After the input image extracts the initial features through Backbone, it enters the C2f_iRMB_Cascaded module: Local spatial attention: The feature map is divided into 7×7 local windows, and the pixel association weights in the window (i.e., the attention weights of the query matrix, key matrix, and value matrix) are calculated using the above formula to output enhanced features; Depthwise separable convolution: A 3×3 convolution kernel is used to extract spatial features channel by channel; Channel attention weighting: The SE module dynamically scales the channel weights to suppress noisy channels. Through the "local-channel" dual attention mechanism, the texture details of the steel bar intersections are focused on to improve the detection recall rate of small targets (<20×20 pixels). Bidirectional multi-scale fusion network: Construct a BiFPN structure, including: Top-down path: Upsampling high-level features (P5) and fusing them with middle-level features (P4); Bottom-up path: Downsampling low-level features (P3) and fusing them with middle-level features (P4); Furthermore, the formula for dynamic weight fusion is: ,in, 、 is a trainable scalar parameter, , F1 is the high-level semantic feature, and F2 is the low-level detail feature.
[0028] Specifically, the bidirectional information flow achieves cross-scale feature complementarity, and the dynamic weight optimizes the fusion contribution to improve the mAP of multi-scale object detection.
[0029] Specifically, the task-decoupled detection module uses two layers of 1×1 convolution in the classification branch (channel count: 256 → 128 → number of categories) to output the category confidence of each anchor point. The regression branch uses three layers of depthwise separable convolution (channel count: 256 → 256 → 256 → 4 × reg_max) to output the coordinate probability distribution. These independent branches avoid task interference, and the probability distribution prediction achieves sub-pixel positioning, reducing positioning error.
[0030] S3. Decode the output data of the target detection model to generate a detection result including the location and category of the binding point.
[0031] Specifically, the decoding process involves integrating the 4×reg_max vector output by the regression branch and performing non-maximum suppression. Non-maximum suppression sets an IoU threshold of 0.6 and removes overlapping boxes. The detection boxes are then mapped to the original image, generating an analysis report with a heatmap. The binding point location is the coordinates of the binding point, and the binding point category can be: qualified, unbound, or partially bound.
[0032] Furthermore, the detection frame coordinates are mapped to the original image resolution, and the following strategy is adopted for overlay display: qualified lashing points are marked with green rectangular frames (confidence ≥ 0.6); suspected defective points are marked with yellow rectangular frames (0.3 ≤ confidence < 0.6); missed lashing points are marked with red flashing frames (confidence < 0.3), and a PDF report containing position deviation statistics is generated.
[0033] In a preferred embodiment, the preprocessing of the construction site image comprises the following steps: Random rotation, scale scaling, mosaic enhancement and HSV color space perturbation are sequentially performed on the construction site image.
[0034] Specifically, during the model training phase, the following enhancement operations are performed on the original image (based on the Albumentations library): Random rotation: Angle range of ±45°, simulating different shooting angles and enhancing the model's ability to identify inclined lashing points; Scale scaling: The scaling ratio is 0.5-1.5, covering close-up shots and panoramic scenes; Mosaic enhancement: randomly select four images and stitch them together to improve generalization in scenes with densely populated small objects; HSV color space perturbation: Hue shift ±0.1 to simulate different light color temperatures; Saturation is scaled by 0.5-1.5 to enhance adaptability to corroded steel bars; Brightness scaling 0.5-1.5 to simulate strong light / shadow environments.
[0035] The data enhancement of this preprocessing method reduces the generalization error of the model in unlabeled scenes.
[0036] In a preferred embodiment, the regression branch outputs coordinate positioning through depthwise separable convolution and probability distribution prediction layer, including the following steps: The first feature map output by the bidirectional multi-scale fusion network is input into three layers of depth-wise separable convolutional layers connected in series to obtain a second feature map; the convolution kernel size of each layer is 3×3, the stride is 1, and batch normalization and SiLU activation functions are inserted; Modeling the bounding box coordinates in the second feature map as a discrete probability distribution through the probability distribution prediction layer, wherein the bounding box coordinates include the horizontal and vertical coordinates of the center point and the width and height of the target binding point; The discrete probability distribution is converted into continuous space coordinates through integration operation to obtain coordinate positioning.
[0037] Specifically, the depthwise separable convolution layer: the input first feature map (such as 20×20×512) is processed by three layers of convolution: each layer contains Depthwise Conv (3×3 kernel, number of groups = number of input channels) + Pointwise Conv (1×1 kernel, number of channels remains 512); BN+SiLU activation is inserted to prevent gradient disappearance; the output second feature map size is 20×20×512.
[0038] Specifically, probability distribution prediction: the fully connected layer maps 512 channels to 4×16 dimensions (reg_max=16), and each coordinate parameter (center point horizontal coordinate, vertical coordinate, width, height) corresponds to 16 probability values; after Softmax normalization, through integration Continuous coordinates are obtained. The coordinate prediction resolution is increased by 4 times, and the standard deviation of positioning error is reduced to 1.2 pixels.
[0039] In a preferred embodiment, modeling the bounding box coordinates in the second feature map as a discrete probability distribution by the probability distribution prediction layer comprises the following steps: Input the second feature map into the fully connected layer to generate a vector of dimension 4×reg_max, where 4 represents the horizontal and vertical coordinates of the center point and the four coordinate parameters of width and height, and reg_max is the preset number of discrete distribution interval segmentation points, ranging from 8 to 16; A Softmax operation is performed on the reg_max-dimensional vector of each of the coordinate parameters to obtain a discrete probability distribution.
[0040] Specifically, the input second feature map is flattened to 20×20×512 = 204,800 dimensions, then reduced to 4×16 = 64 dimensions via a fully connected layer, resulting in a weight matrix size of 204,800×64. Softmax calculation: Each 16-dimensional vector of coordinate parameters is independently normalized to ensure that the distribution sums to 1. Discrete interval partitioning: The coordinate range (0-640 pixels) is evenly divided into 16 intervals (40 pixel intervals), and the probability value indicates the probability of the target center falling within each interval. End-to-end training is supported, and the backpropagation of the DFL loss function is stable, accelerating convergence.
[0041] In a preferred embodiment, the lightweight residual unit includes two 1×1 convolutional layers connected in series, with a batch normalization layer and a SiLU activation function inserted in the middle, and a residual connection is established between the input and output. A DropPath random depth drop mechanism is set on the residual path, and the drop probability is 0.1 to 0.3.
[0042] Specifically, the lightweight residual unit architecture consists of the following: backbone path: 1×1 Conv (256→256) → BN → SiLU → 1×1 Conv (256→256); residual path: direct input connection, superimposed DropPath (with a dropout probability of 0.2); output: backbone path output × 0.8 + input × 0.2 (training phase). The 1×1 convolution reduces computational complexity, the DropPath prevents overfitting, and the residual connection mitigates vanishing gradients, resulting in fewer module parameters and faster inference.
[0043] In a preferred embodiment, the classification branch comprises two layers of 1×1 convolutional layers in sequence, each layer is followed by batch normalization and SiLU activation function, and the number of output channels is equal to the total number of categories.
[0044] Specifically, the first layer: 1×1 Conv (256→128) + BN + SiLU; the second layer: 1×1 Conv (128→number of categories) + Sigmoid activation; output processing: each anchor point outputs a category-dimensional vector, representing the independent confidence of each category (such as "tied" 0.92, "untied" 0.08).
[0045] In a preferred embodiment, the training loss function of the task decoupling detection module includes: the classification loss corresponding to the classification branch adopts binary cross entropy loss, and the regression loss corresponding to the regression branch adopts distribution focus loss.
[0046] In a preferred embodiment, the improved target detection model includes the following lightweight deployment: Channel pruning: remove redundant channels in the feature map with weight sparsity higher than 0.8; Parameter quantization: Convert 32-bit floating-point weights to 8-bit fixed-point numbers using symmetric uniform quantization; Operator fusion: Combines convolutional layers, batch normalization layers, and activation functions into a single computational node.
[0047] Specifically, channel pruning involves calculating the L1 norm of each channel and removing channels with sparsity greater than 0.8 (e.g., 512 → 384). Parameter quantization involves linearly mapping FP32 weights to the INT8 range (-127 to 127) with a scaling factor of s = 127 / max(|W|). During inference, dequantization restores the floating-point values. Operator fusion combines Conv+BN+SiLU into a single operation to reduce memory accesses. This lightweight deployment reduces model size and memory usage on edge devices.
[0048] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. The above is only the preferred implementation method of this application. It should be pointed out that due to the limitations of textual expression, there are objectively infinite specific structures. For ordinary technicians in this technical field, without departing from the principles of the present invention, they can also make several improvements, modifications or changes, and can also combine the above technical features in an appropriate manner; these improvements, modifications, changes or combinations, or the direct application of the inventive concept and technical solution to other occasions without improvement, should be regarded as the scope of protection of this application.
Claims
1. A steel bar binding point detection method based on improved YOLOv8, characterized in that: The following steps are involved: Acquire a construction site image, and pre-process the construction site image to obtain an image to be detected; The image to be detected is input into the improved target detection model, and the lashing point detection is performed through the following collaborative improvement structure: The local feature enhancement module sequentially performs local spatial attention calculation, depthwise separable convolution, and channel attention weighting operations to enhance the detailed feature expression of steel bar intersections; A bidirectional multi-scale fusion network dynamically fuses the top-down first feature path and the bottom-up second feature path, and embeds lightweight residual units at the fusion nodes; the first feature path corresponds to high-level semantic features, and the second feature path corresponds to low-level detail features; The task decoupling detection module uses a classification branch and a regression branch to process the classification task and regression task respectively. The classification branch outputs the category confidence through continuous 1×1 convolution, and the regression branch outputs the coordinate location through depthwise separable convolution and probability distribution prediction layer. The output data of the target detection model is decoded to generate a detection result including the location and category of the binding point.
2. The steel bar binding point detection method based on improved YOLOv8 according to claim 1 is characterized in that: The preprocessing of the construction site image comprises the following steps: Random rotation, scale scaling, mosaic enhancement and HSV color space perturbation are sequentially performed on the construction site image.
3. The steel bar binding point detection method based on improved YOLOv8 according to claim 1 is characterized in that: The local spatial attention calculation includes the following steps: The image to be detected is divided into multiple local windows, and the attention weights of the query matrix, key matrix, and value matrix are calculated in each local window. The calculation formula is: ; Among them, Q is the query matrix, K is the key matrix, V is the value matrix, d k The key vector dimension is ,and the window partition adopts the sliding window strategy, and the ,window size is a configurable parameter ranging from 5×5 to 7×7.
4. The steel bar binding point detection method based on improved YOLOv8 according to claim 1 is characterized in that: The regression branch outputs coordinate positioning through depth-wise separable convolution and probability distribution prediction layer, including the following steps: The first feature map output by the bidirectional multi-scale fusion network is input into three layers of depth-wise separable convolutional layers connected in series to obtain a second feature map; the convolution kernel size of each layer is 3×3, the stride is 1, and batch normalization and SiLU activation functions are inserted; Modeling the bounding box coordinates in the second feature map as a discrete probability distribution through the probability distribution prediction layer, wherein the bounding box coordinates include the horizontal and vertical coordinates of the center point and the width and height of the target binding point; The discrete probability distribution is converted into continuous space coordinates through integration operation to obtain coordinate positioning.
5. The method for detecting steel bar binding points based on improved YOLOv8 according to claim 4 is characterized in that: Modeling the bounding box coordinates in the second feature map as a discrete probability distribution through the probability distribution prediction layer includes the following steps: Input the second feature map into the fully connected layer to generate a vector of dimension 4×reg_max, where 4 represents the horizontal and vertical coordinates of the center point and the four coordinate parameters of width and height, and reg_max is the preset number of discrete distribution interval segmentation points, ranging from 8 to 16; A Softmax operation is performed on the reg_max-dimensional vector of each of the coordinate parameters to obtain a discrete probability distribution.
6. The steel bar binding point detection method based on improved YOLOv8 according to claim 1 is characterized in that: The lightweight residual unit includes two 1×1 convolutional layers connected in series, with a batch normalization layer and SiLU activation function inserted in the middle, and a residual connection is established between the input and output. The DropPath random depth drop mechanism is set on the residual path with a drop probability of 0.1 to 0.
3.
7. The method for detecting steel bar binding points based on improved YOLOv8 according to claim 1, characterized in that: The classification branch sequentially contains two layers of 1×1 convolutional layers, each followed by batch normalization and SiLU activation function, and the number of output channels is equal to the total number of categories.
8. The method for detecting steel bar binding points based on improved YOLOv8 according to claim 1, characterized in that: The formula for dynamic weight fusion is: ,in, 、 is a trainable scalar parameter, , F1 is the high-level semantic feature, and F2 is the low-level detail feature.
9. The method for detecting steel bar binding points based on improved YOLOv8 according to claim 1, characterized in that: The training loss function of the task decoupling detection module includes: the classification loss corresponding to the classification branch adopts binary cross entropy loss, and the regression loss corresponding to the regression branch adopts distribution focus loss.
10. The method for detecting steel bar binding points based on improved YOLOv8 according to claim 1, characterized in that: The improved target detection model includes the following lightweight deployment: Channel pruning: remove redundant channels in the feature map with weight sparsity higher than 0.8; Parameter quantization: Convert 32-bit floating-point weights to 8-bit fixed-point numbers using symmetric uniform quantization; Operator fusion: Combines convolutional layers, batch normalization layers, and activation functions into a single computational node.
Citation Information
Patent Citations
Hidden forbidden article detection method based on lightweight millimeter wave radar
CN118823311A
Fine-grained phytoplankton microscopic image enhancement classification method and model building method thereof
CN119360376A
Small target detection method under view angle of unmanned aerial vehicle based on self-attention mechanism
CN119992393A
Surface defect detection method and system based on optimized YOLOv8 model
CN120107163A
Road crack detection method, medium and product
US20250174019A1
Cited By
Lightweight AI-based distribution line unmanned aerial vehicle edge end real-time visual identification and target detection method and system
CN121459227A
Ground penetrating radar image disease detection method and system and electronic equipment
CN121686253A
A ground penetrating radar image disease detection method, system and electronic device
CN121686253B
Steel bar binding point detection method based on improved CF-DETR model
CN122156907A
Rebar binding point detection method based on improved CF-DETR model
CN122156907B