A method for detecting the distance of steel bars based on a depth estimation model
By using drones and depth estimation models for steel bar spacing detection during construction, the problem of manual measurement in the prior art is solved, and fast and accurate steel bar spacing detection is achieved, reducing construction costs and improving safety.
Patent Information
- Application Number
- CN202210433623.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-24
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-04-24
AI Technical Summary
During the construction process, the existing manual measurement of steel bar spacing methods are time-consuming and labor-intensive and have low accuracy, resulting in insufficient shear bearing capacity of steel bars and increasing construction costs.
The steel bar distance detection method based on the depth estimation model is used to obtain the steel bar top shot image using a drone, and a binary template is generated through AdaBins depth estimation and FCM binarization algorithm. Real-time detection is performed in combination with the improved YOLO X model to calculate whether the steel bar spacing meets the standards.
It realizes rapid and accurate detection of steel bar spacing, reduces construction costs, improves construction safety, and improves inspection accuracy and real-time performance.
Smart Images

Figure CN115222788B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of construction safety, and particularly to a steel bar distance detection method based on a depth estimation model. Background Art
[0002] During the construction process, the standard binding of steel bars is the basis for safe construction. However, problems such as the size, spacing, linear type, cover thickness of precast beam steel bars, and template positioning are prone to large errors. Among them, the spacing error of steel bars will lead to insufficient shear bearing capacity of steel bars and various safety problems such as brittle failure. Therefore, it is necessary to measure the spacing of steel bars during the binding process to prevent large errors. However, the existing measurement methods mainly rely on manual measurement methods, which are time-consuming and laborious and have low accuracy. If it is found that the spacing of steel bars is too large during the acceptance, either adjustment or reinforcement treatment needs to be carried out, which greatly increases the construction cost. Summary of the Invention
[0003] In view of this, the purpose of the present invention is to provide a steel bar distance detection method based on a depth estimation model, a FCM (Fuzzy C-means Cluster) algorithm for binarizing a depth map, and an improved YOLO X object detection algorithm to establish a real-time detection system for the spacing of bound steel bars during the construction process.
[0004] To achieve the above purpose, the present invention adopts the following technical solutions: A steel bar distance detection method based on a depth estimation model, including the AdaBins depth estimation method using monocular depth estimation, and the FCM algorithm for binarizing a depth map, including the following steps:
[0005] Step S1: Use a drone to take a top-down photo of the steel bars bound on the ground to obtain a target image, and obtain the corresponding depth map according to the monocular depth estimation algorithm AdaBins;
[0006] Step S2: Perform a simple normalization adjustment on the depth map to obtain the corresponding grayscale map, and use the FCM algorithm to perform binary processing on the grayscale map to obtain the corresponding binary template;
[0007] Step S3: Use the obtained binary template to perform a Hadamard product with the original image to retain only the pixel points corresponding to the topmost steel bars in the original image;
[0008] Step S4: Send the preprocessed image containing only the topmost steel bars into the improved YOLO X model to obtain the center point coordinates of each steel bar, calculate whether the spacing between adjacent center point coordinates is within the standard spacing range, and if not, frame the corresponding steel bar position.
[0009] In a preferred embodiment: In step S1, the acquisition of the depth map includes the following steps;
[0010] Step S11: Build an AdaBins depth estimation network model. Specifically, use a U-Net architecture to extract deep semantic features of the input image, and use a ViT model to further process the semantic features to obtain the Bin widths of the input image and the probability map R of each pixel in the input image. The Bin widths is an N-dimensional vector, and the dimension of the probability map R is H×W×N, where H and W are the length and width of the input image. The depth information of each pixel in the input image can be calculated from the Bin widths and R, and the calculation formula is as follows:
[0011]
[0012] where d(i,j) is the depth value of the image coordinate (i,j), and b k is the value of the k-th dimension of the Bin widths, and P k (i,j) is the value of the k-th dimension of the probability map R at the coordinate (i,j);
[0013] Step S12: Use the training weights based on the NYU dataset provided by the official to load the AdaBins network model to achieve the depth prediction of the steel bars.
[0014] In a preferred embodiment: In step S2, normalize the depth map output by AdaBins, and the normalization calculation formula is as follows:
[0015]
[0016] where d(i,j) is the depth value of the image coordinate (i,j), a min is the minimum depth value in the depth map, a max is the maximum depth value in the depth map, x ij is the normalization result of the depth map. Multiply x ij by 255 to obtain a grayscale image with a grayscale value of 0-255.
[0017] In a preferred embodiment: Perform binarization processing on the grayscale image converted from the depth map, and use the FCM algorithm as the implementation method of local binarization to obtain a binarization template.
[0018] In a preferred embodiment: In step S3, the binarization template separates the top-layer steel bars from other layers of steel bars. After performing a Hadamard product of the binarization template and the original image, the RGB values of all pixels except the top-layer steel bars are set to zero.
[0019] In a preferred embodiment: In step S4, improving the training of the YOLO X model and calculating the spacing between adjacent steel bars includes the following steps;
[0020] Step S41: First, construct a dataset containing multi-layer or single-layer steel bars. Then, augment the dataset through methods such as affine transformation, and enhance the collected dataset through image enhancement means such as CLAHE local histogram equalization to generate a sufficiently large dataset with high picture quality. Moreover, the generated dataset must be manually annotated with the location of the target and its central point coordinates.
[0021] Step S42: Build a neural network model required by the YOLOX framework. Specifically, use darknet53 as the backbone network, add the Focus structure and the SPP structure, and adopt the decoupled head idea to be responsible for predicting class information and bounding box information respectively. In addition, use the SimOTA algorithm to automatically match positive samples for each annotation box.
[0022] Step S43: Calculate the distance between adjacent steel bars using the central point coordinates of the steel bars detected by the improved YOLO X model. The distance needs to be mapped to the same range, and the mapping relationship is as follows:
[0023]
[0024] where d is the distance between the central points of adjacent steel bars before mapping, w i and h i represent the width and height of the i-th steel bar respectively, and d' is the mapped distance.
[0025] In step S2,
[0026] Step A1: Set the sliding window size to 25, the number of clusters to 2, and initialize the relationship matrix μ ij ;
[0027] Step A2: Update the cluster center C, and the update formula is as follows:
[0028]
[0029] Step A3: Update the relationship matrix μ ij , and the update formula is as follows:
[0030]
[0031] Step A4: Repeat steps A2 and A3 until the relationship matrix μ ij changes very little, and output the values of the two cluster centers;
[0032] Step A5: The binarization threshold of the current pixel point (i, j) in the grayscale image is the average value of the two cluster centers. Repeat the above steps for each pixel point to obtain the binarization thresholds of all pixel points, and then the local binarization operation can be performed.
[0033] In a preferred embodiment: In step S4, the improvement method of the improved YOLO X includes the following steps:
[0034] Step C1: Add a CBAM attention module after the three outputs of the YOLO X backbone network to enhance the feature extraction ability and detection accuracy of the subsequent FPN Neck;
[0035] Step C2: Extract features from the low-level feature map used to output the 80×80 prediction head in the FPN Neck through 7×7 grouped convolution, normalize the features through layer normalization, and then splice them with the high-level features to improve the detection accuracy of the 40×40 prediction head and the 20×20 prediction head;
[0036] Step C3: Adaptive random mask mechanism, the specific operation is as follows: Obtain the annotation box information of each training image, calculate the area of the annotation box, set 10% of the pixels in the annotation box to 0 or any value, that is, mask 10% of the pixel values, and then send them into the network for training, so as to enhance the network's recognition ability of occluded objects in this way;
[0037] Step C4: This method combines the MAE pre-trained model with YOLO X. Utilize the ability of MAE to obtain the latent features of an image from a small amount of image information. Send the original image into the encoder obtained by MAE pre-training to obtain the latent features of the image, and then add the features to the output of the last layer of the YOLO X backbone network, so that the object detection network simultaneously has the deep semantic information of the binary template-generated image and the latent features of the original image.
[0038] Compared with the prior art, the present invention has the following beneficial effects:
[0039] During the construction process, a drone is equipped to estimate the depth of the top-view of steel bars using the AdaBins depth estimation model. After obtaining the depth map, the corresponding binary template is obtained through normalization and FCM binarization. After performing the Hadamard product of the binary template and the original image, pixels other than those of the top-layer steel bars are filtered out. The YOLO X model is used to detect the steel bar targets in the processed steel bar corrosion image to obtain the midpoint coordinates of the top-layer steel bars. Finally, it is calculated whether the center point coordinate spacing of adjacent steel bars meets the standard. Since the binarization template will have a certain corrosion effect on the steel bars to be detected and reduce the detection accuracy, the present invention proposes an improved YOLO X algorithm. By adding an attention module and improving the multi-scale fusion mechanism, its detection accuracy is improved. A random mask mechanism is added and MAE is used to help the detection network obtain the latent features of the image. The prediction head with the smallest receptive field is removed to improve the detection speed of the network, making the present invention have good real-time performance and accuracy, and being able to accurately calculate the spacing of the top-layer steel bars on the premise of solving the interference between multiple layers of steel bars. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 is a schematic diagram of the working process of the preferred embodiment of the present invention;
[0041] Figure 2 is the depth estimation result of the AdaBins for the top-view of steel bars in the preferred embodiment of the present invention;
[0042] Figure 3 is the effect of the binary template obtained by using the FCM algorithm in the preferred embodiment of the present invention;
[0043] Figure 4 is a schematic diagram of the attention module in the improved YOLO X of the preferred embodiment of the present invention;
[0044] Figure 5 is a schematic diagram of the improvement of multi-scale fusion in the improved YOLO X of the preferred embodiment of the present invention;
[0045] Figure 6 is a schematic diagram of the auxiliary MAE in the improved YOLOX of the preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] The present invention will be further described below with reference to the drawings and embodiments.
[0047] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0048] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should also be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0049] A steel bar distance detection method based on a depth estimation model, referring to Figures 1 to 6 , including the AdaBins depth estimation method using monocular depth estimation and the FCM algorithm for binarizing the depth map, which includes the following steps:
[0050] Step S1: Use a drone to take a top-down photo of the steel bars tied on the ground to obtain a target image, and obtain the corresponding depth map according to the monocular depth estimation algorithm AdaBins, as Figure 2 shown;
[0051] Step S2: Perform simple normalization adjustment on the depth map to obtain the corresponding grayscale map, and use the FCM algorithm to perform binary processing on the grayscale map to obtain the corresponding binary template;
[0052] Step S3: Use the obtained binary template to perform a Hadamard product with the original image to only retain the pixel points corresponding to the topmost steel bars in the original image;
[0053] Step S4: Send the preprocessed image containing only the topmost steel bars into the improved YOLO X model to obtain the center point coordinates of each steel bar, and calculate whether the distance between adjacent center point coordinates is within the standard distance range. If not, frame the corresponding steel bar position.
[0054] In step S1, the acquisition of the depth map includes the following steps;
[0055] Step S11: Build an AdaBins depth estimation network model. Specifically, use the U-Net architecture to extract the deep semantic features of the input image, and use the ViT model to further process the semantic features to obtain the Bin widths of the input image and the probability map R of each pixel of the input image. Bin widths is an N-dimensional vector, and the dimension of the probability map R is H×W×N, where H and W are the length and width of the input image; the depth information of each pixel of the input image can be calculated from Bin widths and R, and the calculation formula is as follows:
[0056]
[0057] where d(i,j) is the depth value of the image coordinate (i,j), b kis the value of the k-th dimension of Binwidths, P k (i,j) is the value of the k-th dimension of the probability map R at the coordinate (i,j);
[0058] Step S12: Load the AdaBins network model using the training weights based on the NYU dataset provided by the official to achieve the depth prediction of steel bars.
[0059] In step S2, normalize the depth map output by AdaBins, and the normalization calculation formula is as follows:
[0060]
[0061] where d(i,j) is the depth value of the image coordinate (i,j), a min is the minimum depth value in the depth map, a max is the maximum depth value in the depth map, x ij is the normalization result of the depth map. Multiply x ij by 255 to obtain a grayscale image with grayscale values from 0 to 255.
[0062] Perform binarization on the grayscale image converted from the depth map, and use the FCM algorithm as the implementation method for local binarization to obtain a binarization template.
[0063] In step S2, the binarization process includes the following steps;
[0064] Step A1: Set the sliding window size to 25, the number of clusters to 2, and initialize the relationship matrix μ ij ;
[0065] Step A2: Update the cluster center C, and the update formula is as follows:
[0066]
[0067] Step A3: Update the relationship matrix μ ij , and the update formula is as follows:
[0068]
[0069] Step A4: Repeat steps A2 and A3 until the relationship matrix μ ij changes very little, and output the values of the two cluster centers;
[0070] Step A5: The binarization threshold of the current pixel point (i,j) of the grayscale image is the average of the two cluster centers. Repeat the above steps for each pixel point to obtain the binarization thresholds of all pixel points, and then the local binarization operation can be performed. The binarization template obtained using the FCM algorithm is as shown in the appendix Figure 3 as shown.
[0071] In step S3, the binarization template divides the top - layer steel bars from the steel bars of other layers. After performing the Hadamard product of the binarization template and the original image, the RGB values of all pixels except those of the top - layer steel bars are set to zero.
[0072] In step S4, improving the training of the YOLO X model and calculating the spacing between adjacent steel bars includes the following steps;
[0073] Step S41: First, construct a data set containing multiple or single - layer steel bars. Then, expand the data set through methods such as affine transformation, and enhance the collected data set through image enhancement means such as CLAHE local histogram equalization to generate a sufficiently large and high - quality data set. And the generated data set must be manually annotated with the location of the target and its center - point coordinates;
[0074] Step S42: Build the neural network model required for the YOLO X framework. Specifically, use darknet53 as the backbone network, add the Focus structure and the SPP structure, and adopt the decoupled head idea to be responsible for predicting class information and bounding box information respectively. In addition, use the SimOTA algorithm to automatically match positive samples for each annotated box;
[0075] Step S43: Calculate the spacing between adjacent steel bars using the center - point coordinates of the steel bars detected by the improved YOLO X model. The spacing needs to be mapped to the same range, and the mapping relationship is as follows:
[0076]
[0077] where d is the distance between the centers of adjacent steel bars before mapping, w i and h i respectively represent the width and height of the i - th steel bar, and d' is the mapped spacing.
[0078] In step S4, the improvement method of the improved YOLO X includes the following steps:
[0079] Step C1: Add a CBAM attention module after the three outputs of the YOLO X backbone network to enhance the feature extraction ability and detection accuracy of the subsequent FPN Neck;
[0080] Step C2: Extract features from the low-level feature map used to output the 80×80 prediction head in the FPN Neck through 7×7 grouped convolution, normalize the features through layer normalization, and then concatenate them with the high-level features to improve the detection accuracy of the 40×40 prediction head and the 20×20 prediction head; Since steel bars are objects with extreme aspect ratios, the prediction head is required to have a relatively large receptive field. However, because the receptive field of the 80×80 prediction head is the smallest, the YOLO X model obtained after training is used in this invention to perform object detection on 50 steel bar images. It is found that the confidence levels of the 6400 prediction boxes output by the 80×80 prediction head are all lower than the threshold and are thus filtered out. The 40×40 and 20×20 prediction heads both contribute to the finally predicted prediction boxes. Therefore, in this paper, the YOLO Head structure with a size of 80×80 in YOLO X is discarded to achieve the lightweight of the network. And extract features from the low-level feature map used to output the 80×80 prediction head in the FPN Neck through 7×7 grouped convolution, normalize the features through layer normalization, and then concatenate them with the high-level features to achieve the improvement of multi-scale fusion. The improvement schematic diagram is as shown in Appendix Figure 5 as follows.
[0081] Step C3: Adaptive random mask mechanism, the specific operation is as follows: Obtain the annotation box information of each training image, calculate the area of the annotation box, set 10% of the pixels in the annotation box to 0 or any value, that is, mask 10% of the pixel values, and then send them into the network for training. In this way, the network's ability to recognize occluded objects is enhanced; In addition, in this invention, the original input image is sent into the encoder obtained by MAE pre-training to obtain the latent features of the image, and then this feature is added to the output of the last layer of the YOLOX backbone network, so that the object detection network has both the deep semantic information of the binary template-generated image and the latent features of the original image, strengthening the detection ability of the network. The improvement schematic diagram is as shown in Appendix Figure 6 as follows.
[0082] Step C4: This method combines the MAE pre-trained model with YOLO X, and utilizes the ability of MAE to obtain the latent features of the image from a small amount of image information. The original image is sent into the encoder obtained by MAE pre-training to obtain the latent features of the image, and then this feature is added to the output of the last layer of the YOLO X backbone network, so that the object detection network has both the deep semantic information of the binary template-generated image and the latent features of the original image.
[0083] During the construction process, a drone is equipped to use the AdaBins depth estimation model to estimate the depth of the top-view of the steel bars. After obtaining the depth map, the corresponding binary template is obtained through normalization and FCM binarization. After performing the Hadamard product of this binary template with the original image, other pixels except the top-layer steel bars are filtered out. The YOLO X model is used to detect the steel bar targets in the processed steel bar corrosion image to obtain the middle point coordinates of the top-layer steel bars. Finally, it is calculated whether the center point coordinate spacing of adjacent steel bars meets the standard. Since the binarization template will have a certain corrosion effect on the steel bars to be detected and reduce the detection accuracy, the present invention proposes an improved YOLO X algorithm. By adding an attention module and improving the multi-scale fusion mechanism, its detection accuracy is improved. A random mask mechanism is added and MAE is used to help the detection network obtain the latent features of the image. The prediction head with the smallest receptive field is removed to improve the detection speed of the network, making the present invention have good real-time performance and accuracy, and being able to accurately calculate the spacing of the top-layer steel bars on the premise of solving the interference between multiple layers of steel bars.
Claims
1. A steel bar distance detection method based on a depth estimation model, characterized in that: The AdaBins depth estimation method using monocular depth estimation and the FCM algorithm for binarizing the depth map include the following steps: Step S1: Use a drone to take a downward photo of the steel bars tied to the ground to obtain a target image, and obtain the corresponding depth map according to the monocular depth estimation algorithm AdaBins; Step S2: Simply normalize the depth map to obtain the corresponding grayscale map, and use the FCM algorithm to perform binary processing on the grayscale map to obtain the corresponding binary template; Step S3: Use the obtained binary template to perform a Hadamard product with the original image to only retain the pixel points corresponding to the top-layer steel bars in the original image; Step S4: Send the preprocessed image containing only the top-layer steel bars into the improved YOLO X model to obtain the center point coordinates of each steel bar, and calculate whether the distance between adjacent center point coordinates is within the standard distance range. If not, frame the corresponding steel bar position; In step S4, the training of the improved YOLO X model and the calculation of the distance between adjacent steel bars include the following steps; Step S43: Calculate the distance between adjacent steel bars using the center point coordinates of the steel bars detected by the improved YOLO X model. The distance needs to be mapped to the same range, and the mapping relationship is as follows: Among them, d is the distance between the centers of adjacent steel bars before mapping, w i and h i respectively represent the width and height of the i-th steel bar, and d′ is the spacing after mapping; In step S4, the improvement method of the improved YOLO X includes the following steps: Step C1: Add a CBAM attention module after the three outputs of the YOLO X backbone network to enhance the feature extraction ability and detection accuracy of the subsequent FPN Neck; Step C2: Extract features from the low-level feature map used to output the 80×80 prediction head in the FPN Neck through 7×7 grouped convolutions, normalize the features through layer normalization, and then splice them with the high-level features to improve the detection accuracy of the 40×40 prediction head and the 20×20 prediction head; Step C3: Adaptive random mask mechanism, the specific operation is as follows: Obtain the annotation box information of each training image, calculate the area of the annotation box, set 10% of the pixels in the annotation box to 0 or any value according to 10% of the annotation box area, that is, mask 10% of the pixel values, and then send them into the network for training. In this way, the network's ability to recognize occluded objects is enhanced; Step C4: This method combines the MAE pre-trained model with YOLO X. Utilize the ability of MAE to obtain the latent features of an image from a small amount of image information. Send the original image into the encoder obtained by MAE pre-training to obtain the latent features of the image, and then add the features to the output of the last layer of the YOLO X backbone network, so that the object detection network simultaneously has the deep semantic information of the binary template-generated image and the latent features of the original image.
2. The steel bar distance detection method based on a depth estimation model according to claim 1, wherein: In step S1, the acquisition of the depth map includes the following steps; Step S11: Build an AdaBins depth estimation network model. Specifically, use the U-Net architecture to extract the deep semantic features of the input image, and use the ViT model to further process the semantic features to obtain the Bin widths of the input image and the probability map R of each pixel in the input image. The Bin widths is an N-dimensional vector, and the dimension of the probability map R is H×W×N, where H and W are the length and width of the input image; the depth information of each pixel in the input image is calculated from the Bin widths and R, and the calculation formula is as follows: (1) where d(i,j) is the depth value of the image coordinate (i,j), and b k is the value of the k-th dimension of Bin widths, and P k (i,j) is the value of the k-th dimension of the probability map R at the coordinate (i,j); Step S12: Load the AdaBins network model with the training weights based on the NYU dataset provided by the official to achieve the depth prediction of steel bars.
3. A steel bar distance detection method based on a depth estimation model according to claim 1, characterized in that: In step S2, normalize the depth map output by AdaBins, and the normalization calculation formula is as follows: where d(i, j) is the depth value of the image coordinate (i, j), d min is the minimum depth value in the depth map, d max is the maximum depth value in the depth map, x ij is the normalization result of the depth map. Multiply x ij by 255 to obtain a grayscale image with a grayscale value of 0 - 255.
4. The steel bar distance detection method based on a depth estimation model according to claim 3, characterized in that: Perform binarization processing on the grayscale image converted from the depth map, and use the FCM algorithm as the implementation method of local binarization to obtain a binarization template.
5. The steel bar distance detection method based on a depth estimation model according to claim 1, wherein: In step S3, the binarization template divides the top-layer steel bars from other layers of steel bars. After multiplying the binarization template with the original image by Hadamard product, the RGB values of all pixels except the top-layer steel bars are set to zero.
6. The steel bar distance detection method based on a depth estimation model according to claim 1, characterized in that: Before step S43, improving the training of the YOLO X model and calculating the spacing between adjacent steel bars further includes the following steps Step S41: First, construct a dataset containing multiple layers or a single layer of steel bars, then amplify the dataset through the affine transformation method, and enhance the collected dataset through the means of CLAHE local histogram equalization image enhancement to generate a sufficiently large and high-quality dataset; and the generated dataset must be manually annotated with the location of the target and its center point coordinates. Step S42: Build the neural network model required by the YOLO X framework. Specifically, use darknet53 as the backbone network, add the Focus structure and the SPP structure, and adopt the decoupled head idea to be responsible for predicting class information and bounding box information respectively; in addition, use the SimOTA algorithm to automatically match positive samples for each annotated box.
7. A steel bar distance detection method based on a depth estimation model according to claim 1, characterized in that: In step S2, Step A1: Set the sliding window size to 25, the number of clusters to 2, and initialize the relationship matrix μ ij ; Step A2: Update the cluster center C, and the update formula is as follows: Step A3: Update the relationship matrix μ ij , and the update formula is as follows: Step A4: Repeat steps A2 and A3 iteratively until the relationship matrix μ ij changes very little, and output the values of the two cluster centers; Step A5: The binarization threshold of the current pixel point (i,j) in the grayscale image is the average value of the two cluster centers. Repeat the above steps for each pixel point to obtain the binarization threshold of all pixel points, and then the local binarization operation can be performed.
Citation Information
Patent Citations
Target detection method, apparatus and computer device based on RGBD image
WO2021249351A1