Precast beam reinforcement cage quality detection method and system based on depth vision

Through the quality detection method of prefabricated beam reinforcement frame based on depth vision, the depth map and RGB map combined with ResNet-FPN and Mask R-CNN models are used to realize the precise segmentation and quality detection of steel bar examples, solving the problem of insufficient detection efficiency and accuracy in the existing technology, and meeting the needs of high-quality and large-scale production in bridge construction.

CN120182175AActive Publication Date: 2025-06-20CCCC HIGHWAY BRIDGES NATIONAL ENGINEERING RESEARCH CENTRE CO LTD +3

Patent Information

Application Number
CN202510108983.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-06-20
Estimated Expiration
2045-01-23

AI Technical Summary

Technical Problem

The existing prefabricated beam reinforced frame quality inspection technology has serious shortcomings in applicability and efficiency, and cannot meet the requirements of high-quality and large-scale production of prefabricated beams in bridge construction.

Method used

Using a detection method based on depth vision, the depth map and RGB map of the steel bar frame are obtained through a binocular structured light camera, and ResNet-FPN is used as the backbone network of Mask R-CNN, combining edge detection branches and feature fusion modules to achieve accurate segmentation and quality detection of steel bar examples.

Benefits of technology

It improves the accuracy and efficiency of the quality inspection of steel bars, can more accurately outline the contour of the steel bars, avoid segmentation errors caused by the lack of edge features, and meets the fast and efficient inspection requirements in the production process of prefabricated beams.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182175A_ABST
    Figure CN120182175A_ABST
Patent Text Reader

Abstract

The invention discloses a precast beam reinforcement cage quality detection method and system based on depth vision. The method comprises the following steps: acquiring a depth map and an RGB map of a reinforcement cage based on a binocular structured light camera; the method comprises the following steps: collecting steel reinforcement framework picture data, preprocessing original data, and completing data labeling; segmenting the labeled data set into a training set, a verification set and a test set; the ResNet-FPN is used as a backbone network of the Mask R-CNN and is used for extracting multi-scale features of the image; an edge detection model is constructed based on ResNet-FPN; taking the rest part of the edge detection model except the backbone network as edge detection branches and integrating the edge detection branches into the original Mask R-CNN model; introducing a feature fusion module into the mask branch, wherein the module combines shape and position features in the edge detection branch; and according to a reinforcement instance segmentation result, calculating the actual diameter of the reinforcement and the distance between adjacent reinforcements, and performing evaluation according to a design standard. The problems that the method is only suitable for a single-layer reinforcing mesh and the calculation efficiency is low are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of bridge engineering, in particular to the quality inspection technology of precast beam steel skeletons. Specifically, it is a method and system for quality inspection of precast beam steel skeletons based on depth vision. Background Art

[0002] In modern bridge construction, precast structures exhibit many advantages, and precast concrete beams have been widely used. The quality of the steel bar work in precast beams is crucial for the quality of the entire bridge project, directly affecting key performance indicators such as the structural strength, safety, and durability of the bridge.

[0003] In terms of industry standards, the current standards have clear and definite regulations on important parameters in the steel bar work of precast beams, such as steel bar spacing, steel skeleton size, and cover thickness. These regulations aim to ensure the quality and performance of precast beams and their safety and reliability during use.

[0004] In the actual production process of precast beams, according to the construction process, before concrete pouring, the steel bar work must be reported for inspection. And the quality inspection of the steel skeleton is a key part of the steel bar work inspection. At present, in this field, in most cases, the quality inspection of the steel skeleton mainly relies on manual means.

[0005] However, the manual inspection method has obvious disadvantages. First, in terms of production tasks, precast beam factories usually undertake a large number of precast beam production tasks and often need to inspect the steel skeletons of thousands of beams. The efficiency of manual operation is extremely low, seriously restricting the production progress and unable to meet the high-efficiency requirements of large-scale precast beam production. Second, in terms of the accuracy and objectivity of inspection, manual inspection is inevitably affected by the subjective factors of inspectors. The inspection results of different inspectors may vary greatly, and due to the limitations of manual measurement, its inspection accuracy is also difficult to guarantee, which brings great uncertainty to the quality control of precast beams.

[0006] With the development of technology, computer vision technology brings new possibilities to solve the above problems. Some research and patents attempt to apply this technology to the quality inspection of steel skeletons. However, the existing inspection schemes based on computer vision technology have many deficiencies. For example, many published patent applications mainly focus on the quality inspection of single-layer steel bar meshes, while the actual precast beams often have multi-layer steel skeletons, and this method cannot meet the actual inspection needs. In addition, some patents use point cloud data for steel bar spacing detection. Although theoretically the required data can be obtained, in actual applications, due to the complexity of point cloud data processing, the detection speed is affected, and the overall detection efficiency is greatly reduced, unable to meet the fast and efficient inspection requirements in the precast beam production process. Summary of the Invention

[0007] In summary, the existing quality inspection technology for precast beam steel skeletons has serious deficiencies in terms of applicability and efficiency, and cannot meet the requirements of high-quality and large-scale production of precast beams in bridge construction. There is an urgent need for an innovative inspection method and system to solve the problems existing in the existing technology and improve the quality and efficiency of quality inspection for precast beam steel skeletons.

[0008] In view of the above defects or improvement requirements of the existing technology, as the first aspect of the present invention, the present invention provides a quality inspection method for precast beam steel skeletons based on depth vision, including:

[0009] S1. Obtain the depth map and RGB map of the steel skeleton based on a binocular structured light camera;

[0010] S2. Collect the picture data of the steel skeleton, perform preprocessing operations such as cropping and denoising on the original data, and then complete data annotation; and divide the annotated data set into a training set, a validation set, and a test set;

[0011] S3. Model construction, the specific steps are as follows:

[0012] Use ResNet-FPN as the backbone network of Mask R-CNN to extract multi-scale features of the image;

[0013] Build an edge detection model based on ResNet-FPN;

[0014] Use the remaining part of the edge detection model except the backbone network as the edge detection branch, and directly integrate it into the original MaskR-CNN model to avoid repeated feature extraction calculations;

[0015] Introduce a feature fusion module in the mask branch, which combines the shape and position features in the edge detection branch;

[0016] S4. Quality inspection of the steel skeleton, the specific steps are as follows:

[0017] By improving the mask branch of the Mask R-CNN model, combining the shape and position features provided by the edge detection branch, and using the feature fusion module to perform mask prediction to generate the mask of each steel bar instance, separating the steel bars from the background and other interfering elements;

[0018] According to the segmentation result of the steel bar instance, calculate the actual diameter of the steel bar and the distance between adjacent steel bars, and evaluate according to the design standard.

[0019] Further, the depth map described in S1 is divided into three layers: upper-layer steel bars, lower-layer steel bars, and the bottom ground or bench; when photographing the steel bar skeleton from top to bottom, the plane of the top-layer steel bar mesh is closest to the camera, the plane of the lower-layer steel bars is the second closest, and the plane of the ground or bench is the farthest.

[0020] Further, the depth map can be used to eliminate the interference between the upper and lower layers of steel bars during the calculation of the steel bar spacing by extracting the depth range of the lower-layer steel bars and removing the pixel points of the lower-layer steel bars, as shown in the following formula:

[0021] S = {(x, y)|d min ≤ depth(x, y) ≤ d max , (x, y) ∈ Depth}

[0022]

[0023] In the formula, (x, y) is the pixel coordinate of the depth map Depth, depth(x, y) is the depth value at this coordinate, Mask is a mask with the same size as the RGB map, rgb(x, y) is the pixel value at this coordinate, and d min and d max respectively represent the upper and lower limits of the depth range where the lower-layer steel bars are located.

[0024] Further, the output of the ResNet-FPN backbone network described in S3 is the feature maps of P2 - P5 layers, where P2 has the highest resolution and P5 has the lowest resolution but contains higher-level semantic information.

[0025] Further, after the feature maps of P2 - P5 layers undergo corresponding deconvolution operations, the size of the feature maps is restored to the original input size, the four feature maps are merged in the second dimension, and then after a 1×1 convolution operation, the number of channels is converted to 1 and input into the sigmoid function to output the edge detection result.

[0026] Further, a feature fusion module is introduced in the mask branch described in S3, and the specific formula is expressed as:

[0027] F = f(F b ) + F m

[0028] where F is the output fused feature, F b is the input mask feature, F m is the mask branch feature, and f represents the 1×1 convolution and ReLU activation operations.

[0029] Further, the method for detecting the distance between adjacent steel bars in S4 is as follows:

[0030] After extracting all the top-layer steel bar instances, the Zhang-Suen image thinning algorithm is used to extract the centerlines of the steel bar instances. Then, one of the steel bar instances is selected, and the intersection points of the centerline of this instance and other centerlines are calculated to obtain the intersection pixel coordinates (x, y).

[0031] On the depth map obtained by the binocular structured light camera, the depth value Z of this point is obtained. Substituting it into the following formula, the three-dimensional physical coordinates (X, Y, Z) of this point can be obtained;

[0032] Calculate the distance between two points in three-dimensional space, and the steel bar spacing information can be obtained:

[0033]

[0034] where, f x and f y are the focal lengths of the camera in the x and y directions respectively, c x and c y are the principal point coordinates of the camera respectively. The focal length and the principal point can be obtained from the internal parameter matrix of the camera, and u and v are the pixel coordinates of the intersection point.

[0035] Furthermore, the actual diameter detection method of the steel bar in S4 is as follows:

[0036] In the calculation of the steel bar diameter, the eight-direction search method is used. Select several points on the centerline of each steel bar instance. Each point searches for the number of pixel points belonging to the steel bar mask in the directions of 0°, 45°, 90°, 135°, 180°, 235°, 270°, and 315° respectively until the edge of the steel bar. Calculate the steel bar edge point corresponding to the minimum value according to the following formula;

[0037] Convert the edge to three-dimensional physical coordinates to obtain the steel bar diameter at this point. Take the average value among the non-zero steel bar diameters to complete the diameter detection:

[0038] d1 = n1 + n5 + 1

[0039] d2 = 2 1 / 2 (n2 + n6 + 1)

[0040] d3 = n3 + n7 + 1

[0041] d4 = 2 1 / 2 (n4 + n8 + 1)

[0042] d = min(d1, d2, d3, d4)

[0043] Among them, n1 - n8 are the number of pixels in 8 directions, d1 represents the pixel width in the 0 - degree and 180 - degree directions, d2 represents the pixel width in the 45 - degree and 235 - degree directions, d3 represents the pixel width in the 90 - degree and 270 - degree directions, d4 represents the pixel width in the 135 - degree and 315 - degree directions, and d represents the final pixel width.

[0044] According to the second aspect of the present invention, there is provided a quality inspection system for precast beam steel bar skeletons based on depth vision, including:

[0045] An image acquisition unit for simultaneously obtaining a depth map and an RGB map of the steel bar skeleton based on a binocular structured light camera;

[0046] A data processing unit for collecting steel bar skeleton image data, performing pre - processing operations such as cropping and denoising on the original data and then completing data annotation; and splitting the annotated data set into a training set, a validation set, and a test set;

[0047] A model construction unit for completing the following steps:

[0048] Using ResNet - FPN as the backbone network of Mask R - CNN to extract multi - scale features of the image;

[0049] Constructing an edge detection model based on ResNet - FPN;

[0050] Taking the remaining part of the edge detection model except the backbone network as the edge detection branch and directly integrating it into the original MaskR - CNN model to avoid repeated feature extraction calculations;

[0051] Introducing a feature fusion module in the mask branch, which combines the shape and position features in the edge detection branch to assist in the accurate prediction of the mask;

[0052] A steel bar skeleton quality inspection unit for completing the following steps:

[0053] By improving the mask branch of the Mask R - CNN model, combining the shape and position features provided by the edge detection branch, and using the feature fusion module to perform mask prediction to generate a mask for each steel bar instance, separating the steel bars from the background and other interfering elements to achieve instance segmentation of the steel bars;

[0054] According to the steel bar instance segmentation result, calculating the actual diameter of the steel bars and the distance between adjacent steel bars, and evaluating according to the design standard.

[0055] As a third aspect of the present invention, there is also provided a computer-readable storage medium having stored thereon a computer program, which when executed by a processor, implements the method for detecting the quality of precast beam steel bar skeletons based on depth vision as claimed in claims 1-8.

[0056] Generally speaking, compared with the prior art by the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:

[0057] 1. In the method for detecting the quality of precast beam steel bar skeletons based on depth vision of the present invention, by using the feature maps of different levels output by ResNet-FPN, through deconvolution, feature fusion and convolution operations, combined with the information of the edge detection branch, the contour of the steel bars can be outlined more precisely during the steel bar instance segmentation. Especially for the edge detail parts, it effectively overcomes the problem of edge feature loss caused by traditional convolution and pooling operations, and improves the segmentation accuracy to a new level, providing a reliable image segmentation basis for the precise detection of steel bar quality.

[0058] 2. In the method for detecting the quality of precast beam steel bar skeletons based on depth vision of the present invention, a unique feature fusion module is introduced into the mask branch of the improved Mask R-CNN model. This feature fusion module can fuse the shape and position features obtained by the edge detection branch with the features of the mask branch, realizing the efficient integration of information. In this way, the steel bar instance segmentation process is optimized, so that when performing steel bar instance segmentation, the model can make full use of the detailed information brought by edge detection, while ensuring the overall calculation efficiency, greatly improving the accuracy of instance segmentation. Whether for complete steel bar instances or incomplete lower-layer steel bar instances that may appear in the actual scene, their contours can be outlined more accurately, avoiding segmentation errors caused by the lack of edge features, and finally achieving the remarkable technical effects of improving the instance segmentation accuracy, avoiding repeated calculations, and enhancing the model calculation and segmentation efficiency, providing a more reliable and efficient tool for the quality detection of precast beam steel bar skeletons. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 is a flowchart of the method for detecting the quality of precast beam steel bar skeletons based on depth vision in a preferred embodiment;

[0060] Figure 2 is a flowchart of the quality detection of the steel bar skeleton in a preferred embodiment;

[0061] Figure 3 is a schematic diagram of the steel bar skeleton model in a preferred embodiment;

[0062] Figure 4 is a schematic diagram of removing the lower-layer steel bars in a preferred embodiment;

[0063] Figure 5Schematic diagram of the depth value distribution range of the preferred implementation;

[0064] Figure 6 Schematic diagram of the steel bar instance segmentation dataset of the preferred implementation;

[0065] Figure 7 Schematic diagram of the edge detection model based on ResNet-FPN of the preferred implementation;

[0066] Figure 8 Schematic diagram of the edge detection result of the preferred implementation;

[0067] Figure 9 Schematic diagram of the improved Mask R-CNN model of the preferred implementation;

[0068] Figure 10 Schematic diagram of the comparison of steel bar instance segmentation of the preferred implementation;

[0069] Figure 11 Schematic diagram of the inspection result of the steel bar skeleton quality of the preferred implementation;

[0070] Figure 12 Schematic diagram of the precast beam steel bar skeleton quality inspection system based on depth vision of the preferred implementation. Specific implementation mode

[0071] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0072] Embodiment 1

[0073] Please refer to Figure 1 , this Embodiment 1 provides a method for inspecting the quality of a precast beam steel bar skeleton based on depth vision, including:

[0074] (1) Image acquisition

[0075] Based on a binocular structured light camera, obtain the depth map and RGB map of the steel bar skeleton. The specific method is as follows:

[0076] At the precast beam production site, set up a binocular structured light camera. The position and angle of the camera are carefully adjusted to ensure that the steel bar skeleton of the precast beam can be completely and clearly photographed from top to bottom. Through the binocular structured light camera, obtain the depth map and RGB map of the steel bar skeleton at the same time.

[0077] Please refer to Figure 3 and Figure 4, the depth map shows an obvious three-layer structure, namely the upper-layer steel bars, the lower-layer steel bars, and the bottom ground or bench. Due to the shooting angle, the plane of the top-layer steel bar mesh is the closest to the camera, the plane of the lower-layer steel bar is the second closest, and the plane of the ground or bench is the farthest. Among them, Figure 4 The left-middle figure is the point cloud generated from the depth map and the RGB map. Only for display, it can be clearly observed that the point cloud of the steel bar skeleton is divided into three layers.

[0078] Please refer to Figure 5 , when shooting the steel bar skeleton at a fixed position, the distances between the top-layer steel bars, the lower-layer steel bars, and the ground or bench from the camera can be measured, and the depth range [dmin, dmax] of the lower-layer steel bars can be obtained. When shooting at a non-fixed position, the depth range [dmin, dmax] of the lower-layer steel bars can be obtained by generating the depth value distribution range.

[0079] Then, using the depth map information, by extracting the depth range of the lower-layer steel bars and removing the pixel points of the lower-layer steel bars, the interference between the upper and lower layers of steel bars during the subsequent calculation of the steel bar spacing can be eliminated. The specific operations are as follows:

[0080] It is known that the depth map is Depth, its pixel coordinates are (x, y), the depth value at this coordinate is depth(x, y), the mask with the same size as the RGB map is Mask, and the pixel value of the RGB map at this coordinate is rgb(x, y). The pixel points of the lower-layer steel bars are removed through the following formula:

[0081]

[0082] Furthermore, the specific process is expressed by the formula as follows:

[0083] S = {(x,y)|d min ≤ depth(x,y) ≤ d max ,(x,y) ∈ Depth}

[0084]

[0085] In the formula, (x,y) are the pixel coordinates of the depth map Depth, depth(x,y) is the depth value at this coordinate, Mask is the mask with the same size as the RGB map, rgb(x,y) is the pixel value at this coordinate, d min and d max respectively represent the upper and lower limits of the depth range where the lower-layer steel bars are located.

[0086] (2) Data processing

[0087] Please refer to Figure 6, collect the image data of the steel bar skeleton, and perform preprocessing operations such as cropping and denoising on the original data. Use the labelme annotation tool to perform semantic segmentation annotation on the original data. After completing the annotation work, check the annotation accuracy. Divide the dataset into a training set, a validation set, and a test set.

[0088] For the collected original images, cropping is an important preprocessing step. Since when shooting, the images may contain a large amount of background information, such as the support structure around the steel bar skeleton, ground debris, workshop environment, etc., this information is not necessary for the quality inspection of the steel bar skeleton and may interfere with subsequent analysis. Therefore, these irrelevant areas are removed through the cropping operation, and only the core area where the steel bar skeleton is located is retained.

[0089] Specifically, an image cropping algorithm is adopted to crop according to the preset boundary conditions or the automatically recognized boundary range. For example, through the analysis of the grayscale histogram of the image and edge detection algorithms (such as Canny edge detection), the approximate range of the steel bar skeleton is identified, and then the image is precisely cropped according to this range to ensure that the cropped image only contains the main steel bar skeleton information. During the cropping process, it is necessary to consider the changes in the sizes of different precast beams and the positions of the steel bar skeletons, so it may be necessary to dynamically adjust the cropping parameters to ensure that the irrelevant areas can be effectively removed in different situations.

[0090] In some other preferred embodiments, due to the complex environment of the actual production workshop, the collected images are often interfered by noise, and the noise may come from factors such as uneven illumination, thermal noise of the camera sensor, electromagnetic interference, etc., which will reduce the image quality and affect subsequent feature extraction and model processing. In other embodiments, the Gaussian filtering algorithm is used for denoising.

[0091] Gaussian filtering is a linear smoothing filter. It smooths the noise points by performing a weighted average operation on each pixel point in the image and its neighboring pixels. For each pixel point in the image, its pixel value is updated according to the weighted average value of its neighboring pixels, and the weights are determined by the Gaussian function. The pixel points closer to the central pixel have larger weights, and the pixel points farther away have smaller weights. In this way, while removing the noise, the edge and detail information of the image can be retained as much as possible.

[0092] In some preferred embodiments, after completing the cropping and denoising processes, the annotated dataset is divided into a training set, a validation set, and a test set. The division ratio is usually determined according to the actual situation and experience. In some embodiments, the ratio of 7:2:1 is adopted.

[0093] The training set is used to train the improved Mask R-CNN model. Through a large amount of data input, the model learns the feature representations and patterns of the steel bar skeletons. During the training process, the model continuously adjusts its own parameters to minimize the loss function and improve its ability to recognize and segment steel bar skeletons.

[0094] The validation set is used to evaluate the performance of the model during training and monitor whether the model is overfitting or underfitting. In each epoch or stage of training, the validation set is input into the model. According to the performance metrics (such as accuracy, recall, F1 value, etc.) of the model on the validation set, the hyperparameters of the model, such as the learning rate, the number of network layers, the regularization parameter, etc., are adjusted to ensure that the model has good generalization ability.

[0095] The test set is used to finally evaluate the performance of the model. After the model training is completed, the test set is used to test the model and evaluate its performance on unseen data to ensure that the model can have reliable performance in the actual steel bar skeleton quality inspection task.

[0096] (3) Model construction

[0097] Please refer to Figure 7 , in this method, Mask R-CNN is used as the basic model to achieve the instance segmentation of steel bars. To detect the spacing of steel bar skeletons and the diameter of steel bars, higher precision requirements are imposed on the results of steel bar instance segmentation, and clear edge details need to be ensured. In image segmentation, edge features, as shallow features, are gradually lost during multiple convolution and pooling operations, and the edge details of the segmentation results are relatively rough.

[0098] In the selection of the backbone network, in this Embodiment 1, ResNet-FPN is used as the backbone network of Mask R-CNN, and this backbone network outputs the feature maps of P2 - P5 layers. Among them, the feature map of P2 layer has the highest resolution and can retain rich detail information of the image; the feature map of P5 layer has the lowest resolution but contains higher-level semantic information, which helps to understand the overall structure of the steel bars.

[0099] Furthermore, in this Embodiment 1, the edge detection model is further constructed. The corresponding deconvolution operations are performed on the feature maps of P2 - P5 layers respectively to restore the size of the feature maps to the original input size. Then, these 4 feature maps are merged in the second dimension, and then through a 1×1 convolution operation, the number of channels is converted to 1, and finally input into the sigmoid function, thereby outputting the edge detection result and completing the construction of the edge detection model based on ResNet-FPN.

[0100] Please refer to Figure 8 , it can be clearly seen the actual effect of the edge detection model in this Embodiment 1.

[0101] Please refer toFigure 9 , further, the remaining part of the above - constructed edge detection model except the backbone network is used as the edge detection branch and directly integrated into the original Mask R - CNN model, which can effectively avoid repeated feature extraction calculations and improve the computational efficiency of the model.

[0102] Since there is a close relationship between edges and masks, a feature fusion module is introduced in the mask branch. The shape and position features in the edge detection branch contribute to the accurate prediction of masks.

[0103] F = f(F b ) + F m

[0104] where F is the output fused feature, F b is the input mask feature, and F m is the mask branch feature. f represents a 1×1 convolution and ReLU activation operation. Through this feature fusion module, combined with the shape and position features in the edge detection branch, accurate mask prediction is achieved.

[0105] Please refer to Figure 10 , due to the influence of the depth camera accuracy and the on - site environment, the lower - layer steel bar pixels cannot be completely removed. The mask detection branch feature of the improved Mask R - CNN model combines the bottom - layer edge features and high - layer semantic features. The steel bar instance segmentation result depends more on the edge features of the complete steel bar instance. For the lower - layer steel bar instances where pixels are not completely removed, the edge features of the whole steel bar cannot be obtained, resulting in the loss of some edge features during mask branch prediction, so that the lower - layer steel bar instances are not classified as the steel bar category, that is, the introduced edge detection branch has an inhibitory effect on the segmentation results of incomplete lower - layer steel bar instances.

[0106] (4) Steel bar skeleton quality inspection

[0107] 4.1 Steel bar instance segmentation

[0108] By improving the mask branch of the Mask R - CNN model, combining the shape and position features provided by the edge detection branch, and using the feature fusion module for mask prediction, a mask for each steel bar instance is generated to separate the steel bars from the background and other interfering elements, realizing the instance segmentation of steel bars;

[0109] 4.2 Steel bar spacing detection

[0110] In this embodiment 1, after all the top - layer steel bar instances are extracted, the Zhang - Suen image thinning algorithm is used to extract the centerlines of the steel bar instances, and one of the steel bar instances is selected to calculate the intersection points of the centerline of this instance with other centerlines, obtaining the intersection pixel coordinates (x, y);

[0111] The depth value Z of this point is obtained from the depth map acquired by the binocular structured light camera. Substituting it into the following formula, the three-dimensional physical coordinates (X, Y, Z) of this point can be obtained;

[0112] Calculate the distance between two points in three-dimensional space, and the steel bar spacing information can be obtained:

[0113]

[0114] where f x and f y are the focal lengths of the camera in the x and y directions respectively, c x and c y are the principal point coordinates of the camera respectively. The focal length and principal point can be obtained from the internal parameter matrix of the camera, and u and v are the pixel coordinates of the intersection point.

[0115] In other preferred embodiments, deep learning techniques are utilized. For example, a dedicated convolutional neural network (CNN) can be trained to automatically detect the feature points on the steel bars, rather than relying solely on the Zhang-Suen image thinning algorithm. The labeled steel bar images can be used as the training set, and the feature points can be manually labeled at the key positions of the steel bars (such as the endpoints, bending points, or equally spaced marking points) of the steel bars.

[0116] The trained CNN network can be a network similar to an object detection network (such as YOLO, SSD, etc.), which is modified to be specifically used for detecting the feature points on the steel bars. The input of the network is the preprocessed steel bar image, and the output is the position and confidence of the feature points. By learning from a large amount of training data, the network can automatically learn the feature representation of the steel bars and accurately locate the feature points.

[0117] After detecting the feature points on different steel bar instances, it is necessary to match and group these feature points to determine which feature points belong to the same steel bar or the same row of steel bars. A deep learning-based feature descriptor (such as a Siamese network) can be used to generate a feature vector for each feature point.

[0118] Furthermore, for the feature vectors of different feature points, similarity metrics (such as cosine similarity, Euclidean distance, etc.) are used for matching. The feature points with high similarity are grouped to ensure that the feature points belonging to the same steel bar or the same row of steel bars are correctly associated.

[0119] For steel bars with partial occlusion or deformation, the deep learning method has better robustness because the network can learn the features of the steel bars in different states, rather than relying on a simple image thinning algorithm, thereby improving the accuracy of feature point detection and matching.

[0120] 4.3 Detection of Steel Bar Diameter

[0121] In Embodiment 1, in the calculation of the steel bar diameter, the eight-direction search method is adopted; several points are selected on the center line of each steel bar instance, and the number of steel bar mask pixel points belonging to each point is searched in the directions of 0°, 45°, 90°, 135°, 180°, 235°, 270°, and 315° until the edge of the steel bar. The steel bar edge point corresponding to the minimum value is calculated according to the following formula:

[0122] Convert the edge to three-dimensional physical coordinates to obtain the steel bar diameter at this point. Take the average value among non-zero steel bar diameters to complete the diameter detection:

[0123] d1 = n1 + n5 + 1

[0124] d2 = 2 1 / 2 (n2 + n6 + 1)

[0125] d3 = n3 + n7 + 1

[0126] d4 = 2 1 / 2 (n4 + n8 + 1)

[0127] d = min(d1, d2, d3, d4)

[0128] Among them, n1 - n8 are the number of pixels in 8 directions, d1 represents the pixel width in the directions of 0° and 180°, d2 represents the pixel width in the directions of 45° and 235°, d3 represents the pixel width in the directions of 90° and 270°, d4 represents the pixel width in the directions of 135° and 315°, and d represents the final pixel width.

[0129] Embodiment 2

[0130] Please refer to Figure 12 , Embodiment 2 provides a precast beam steel bar skeleton quality detection system based on depth vision, including:

[0131] An image acquisition unit, which can simultaneously obtain the depth map and RGB map of the steel bar skeleton based on a binocular structured light camera;

[0132] A data processing unit, which is used to collect the steel bar skeleton picture data, perform preprocessing operations such as cropping and denoising on the original data, and then complete data annotation; and divide the annotated data set into a training set, a validation set, and a test set;

[0133] A model construction unit, which is used to complete the following steps:

[0134] Adopt ResNet-FPN as the backbone network of Mask R-CNN to extract multi-scale features of the image;

[0135] Build an edge detection model based on ResNet-FPN;

[0136] Use the remaining part of the edge detection model except the backbone network as the edge detection branch and directly integrate it into the original Mask R-CNN model to avoid repeated feature extraction calculations;

[0137] Introduce a feature fusion module in the mask branch, which combines the shape and position features in the edge detection branch to assist in the accurate prediction of the mask;

[0138] The steel bar skeleton quality detection unit is used to complete the following steps:

[0139] By improving the mask branch of the Mask R-CNN model, combining the shape and position features provided by the edge detection branch, using the feature fusion module for mask prediction, generating the mask of each steel bar instance, separating the steel bars from the background and other interfering elements, and realizing the instance segmentation of the steel bars;

[0140] According to the steel bar instance segmentation results, calculate the actual diameter of the steel bars and the distance between adjacent steel bars, and evaluate according to the design standards.

[0141] Embodiment 3

[0142] Embodiment 3 of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it can implement the method for detecting the quality of the steel bar skeleton of the precast beam based on depth vision.

[0143] The computer-readable storage medium may include: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.

[0144] For the introduction of the computer-readable storage medium provided in this application, please refer to the above method embodiments, and this application will not be elaborated here.

[0145] Those skilled in the art can easily understand that the above description is only a preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for detecting the quality of prefabricated beam reinforcement skeleton based on depth vision, characterized in that: include: S1. Obtain the depth map and RGB map of the steel skeleton based on the binocular structured light camera; S2. Collect steel skeleton image data, perform pre-processing operations such as cropping and denoising on the original data, and complete data annotation; and divide the annotated data set into a training set, a validation set, and a test set; S3. Model construction, the specific steps are as follows: ResNet-FPN is used as the backbone network of Mask R-CNN to extract multi-scale features of images; Build an edge detection model based on ResNet-FPN; The edge detection model except the backbone network is used as an edge detection branch and directly integrated into the original Mask R-CNN model to avoid repeated feature extraction calculations; A feature fusion module is introduced in the mask branch, which combines the shape and position features in the edge detection branch; S4. Steel bar skeleton quality inspection, the specific steps are as follows: By improving the mask branch of the Mask R-CNN model, combining the shape and position features provided by the edge detection branch, and using the feature fusion module for mask prediction, a mask for each steel bar instance is generated to separate the steel bars from the background and other interference elements. Based on the steel bar instance segmentation results, the actual diameter of the steel bar and the distance between adjacent steel bars are calculated and evaluated according to the design standards.

2. The method for detecting the quality of prefabricated beam reinforcement skeleton based on depth vision according to claim 1 is characterized in that: The depth map described in S1 is divided into three layers, the upper steel bar, the lower steel bar, and the bottom ground or frame; when the steel frame is photographed from top to bottom, the top steel mesh plane is closest to the camera, followed by the lower steel bar plane, and the ground or frame plane is farthest.

3. The method for detecting the quality of prefabricated beam reinforcement skeleton based on depth vision according to claim 2 is characterized in that: The depth map is used to eliminate the interference between the upper and lower layers of steel bars when calculating the steel bar spacing by extracting the depth interval of the lower layer of steel bars and removing the lower layer of steel bar pixels, as shown in the following formula: S={(x,y)|d min ≤depth(x,y)≤d max ,(x,y)∈Depth} Where (x, y) is the pixel coordinate of the depth map Depth, depth(x, y) is the depth value at that coordinate, Mask(x, y) is a mask of the same size as the RGB map, rgb(x, y) is the pixel value at that coordinate, and d min With d max They respectively represent the upper and lower limits of the depth range where the lower layer of steel bars are located.

4. The method for detecting the quality of prefabricated beam reinforcement skeleton based on depth vision according to claim 1, characterized in that: The output of the ResNet-FPN backbone network in S3 is a P2-P5 layer feature map, where P2 has the highest resolution and P5 has the lowest resolution but contains higher-level semantic information.

5. The method for detecting the quality of prefabricated beam reinforcement skeleton based on depth vision according to claim 4 is characterized in that: After the corresponding deconvolution operation is performed on the feature maps of the P2-P5 layers, the feature map size is restored to the original input size, the four feature maps are merged in the second dimension, and after a 1×1 convolution operation, the number of channels is converted to 1 and input into the sigmoid function to output the edge detection result.

6. The method for detecting the quality of prefabricated beam reinforcement skeleton based on depth vision according to claim 1, characterized in that: As described in S3, a feature fusion module is introduced into the mask branch, and the specific formula is expressed as follows: F=f(F b )+F m Among them, F is the output fusion feature, F b is the input mask feature, F m is the mask branch feature, and f represents 1×1 convolution and ReLU activation operations.

7. The method for detecting the quality of prefabricated beam reinforcement skeleton based on depth vision according to claim 1, characterized in that: The distance detection method between adjacent steel bars in S4 is as follows: After all top-level steel bar instances are extracted, the Zhang-Suen image thinning algorithm is used to extract the center line of the steel bar instance, and one of the steel bar instances is selected to calculate the intersection of the center line of the instance with other center lines to obtain the pixel coordinates (x, y) of the intersection; The depth value Z of the point is obtained on the depth map obtained by the binocular structured light camera. Substituting it into the following formula, the three-dimensional physical coordinates (X, Y, Z) of the point can be obtained; By calculating the distance between two points in three-dimensional space, you can get the steel bar spacing information: Among them, f x 、f y are the focal lengths of the camera in the x and y directions, c x and c y are the principal point coordinates of the camera respectively. The focal length and principal point can be obtained from the intrinsic parameter matrix of the camera. u and v are the pixel coordinates of the intersection point.

8. The method for detecting the quality of prefabricated beam reinforcement skeleton based on depth vision according to claim 1, characterized in that: The method for calculating the actual diameter of the steel bar in S4 is as follows: In the calculation of the steel bar diameter, the eight-direction search method is used; select several points on the center line of each steel bar instance, and search for the number of steel bar mask pixels in the directions of 0 degrees, 45 degrees, 90 degrees, 135 degrees, 180 degrees, 235 degrees, 270 degrees, and 315 degrees at each point until the edge of the steel bar, and calculate the steel bar edge point corresponding to the minimum value according to the following formula; Convert the edge into three-dimensional physical coordinates, obtain the steel bar diameter at that point, take the average of the non-zero steel bar diameters, and complete the diameter detection: d1=n1+n5+1 <h2 style=";text-align:left;direction:ltr">d2=2<h2 style=";text-align:left;direction:ltr"> 1 / 2 <h2 style=";text-align:left;direction:ltr"> (n2+n6+1) d3=n3+n7+1 d4=2 1 / 2 (n4+n8+1) d=min(d1,d2,d3,d4) Among them, n1-n8 are the number of pixels in 8 directions, d1 represents the pixel width in the directions of 0 degrees and 180 degrees, d2 represents the pixel width in the directions of 45 degrees and 235 degrees, d3 represents the pixel width in the directions of 90 degrees and 270 degrees, d4 represents the pixel width in the directions of 135 degrees and 315 degrees, and d represents the final pixel width.

9. A prefabricated beam reinforcement skeleton quality inspection system based on depth vision, characterized in that: include: An image acquisition unit, used to simultaneously acquire a depth map and an RGB map of a steel bar skeleton based on a binocular structured light camera; The data processing unit is used to collect steel skeleton image data, perform pre-processing operations such as cropping and denoising on the original data, and complete data annotation; and divide the annotated data set into a training set, a validation set, and a test set; Model building unit, which is used to complete the following steps: ResNet-FPN is used as the backbone network of Mask R-CNN to extract multi-scale features of images; Build an edge detection model based on ResNet-FPN; The edge detection model except the backbone network is used as an edge detection branch and directly integrated into the original Mask R-CNN model to avoid repeated feature extraction calculations; A feature fusion module is introduced in the mask branch, which combines the shape and position features in the edge detection branch to help the accurate prediction of the mask; The steel bar skeleton quality inspection unit is used to complete the following steps: By improving the mask branch of the Mask R-CNN model, combining the shape and position features provided by the edge detection branch, and using the feature fusion module to perform mask prediction, the mask of each steel bar instance is generated, the steel bar is separated from the background and other interference elements, and the instance segmentation of the steel bar is achieved; Based on the steel bar instance segmentation results, the actual diameter of the steel bar and the distance between adjacent steel bars are calculated and evaluated according to the design standards.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: The computer program is executed by a processor to implement the method for detecting the quality of a prefabricated beam reinforcement skeleton based on depth vision as described in claims 1-8.

Citation Information

Patent Citations

  • In-service concrete pavement slab load transfer member three-dimensional space position detection system

    CN114777642A

  • Mask R-CNN mineral particle identification and particle size detection method based on improved mask

    CN114897816A

  • Intelligent reinforcement detection method and system based on convolutional neural network and binocular vision

    CN116703835A

  • Steel reinforcement framework binding quality automatic inspection device

    CN117470122A

  • Image instance detection segmentation model construction method based on edge information enhancement

    CN118762042A

Cited By

  • Multi-layer reinforcing steel bar intelligent counting method and system based on deep learning

    CN122391273A

  • A multi-layer steel bar intelligent counting method and system based on deep learning

    CN122391273B