A Reinforcing Steel Structure Data Identification Method Based on Image Recognition

By using multimodal data acquisition and a dynamic dual-domain attention optimization model, combined with BIM integration, the problems of metal reflection interference, complex structure analysis, and delay in detection results in traditional rebar inspection have been solved, achieving efficient and accurate rebar structure inspection and real-time quality control.

CN122493225APending Publication Date: 2026-07-31中电建路桥集团有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
中电建路桥集团有限公司
Filing Date
2026-04-03
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Traditional steel reinforcement structure inspection technology suffers from problems such as metal reflection interference causing texture feature distortion, single depth information being insufficient to analyze complex spatial structures, construction site pollution obscuring features, and inspection results being disconnected from BIM, resulting in low inspection efficiency, insufficient accuracy, and delayed discovery of quality problems.

Method used

By employing multimodal data acquisition technology, combined with a polarized RGB camera, a ToF depth camera, and a near-infrared spectral sensor, a dynamic dual-domain attention optimization model is trained to achieve accurate prediction of the position, diameter, and arrangement parameters of reinforcing bars. A closed-loop detection system is then constructed through BIM integration and real-time feedback.

Benefits of technology

It enables complete capture of the surface texture and internal defects of steel bars under complex working conditions, improving detection accuracy and efficiency, realizing the transformation from post-inspection to real-time process control, and reducing the cost of manual inspection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure REF-OBJ-1775202686239-000002
    Figure REF-OBJ-1775202686239-000002
  • Figure REF-OBJ-1775202686239-000003
    Figure REF-OBJ-1775202686239-000003
  • Figure REF-OBJ-1775202686239-000004
    Figure REF-OBJ-1775202686239-000004
Patent Text Reader

Abstract

This invention discloses a data recognition method for rebar structures based on image recognition, comprising the following steps: eliminating metal reflection interference using a polarized RGB camera, simultaneously acquiring millimeter-level precision rebar depth information using a ToF depth camera, and combining this with a near-infrared spectral sensor to penetrate surface dust and collect rebar corrosion features; training a dynamic dual-domain attention optimization model, which optimizes the prediction accuracy of rebar position, diameter, and arrangement parameters through a joint loss function; inputting real-time collected construction data into the trained dynamic dual-domain attention optimization model, outputting rebar position coordinates, diameter measurements, and arrangement spacing; triggering a visual alarm and generating a correction report when errors occur. This invention utilizes a polarized RGB camera to eliminate metal reflection, a ToF depth camera to acquire three-dimensional coordinates, and a near-infrared spectral sensor to penetrate dust and collect corrosion features, constructing a cross-modal dataset to achieve complete capture of rebar surface texture and internal defects under complex working conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of building engineering quality technology, and in particular to a method for identifying steel reinforcement structure data based on image recognition. Background Technology

[0002] In the field of construction engineering, the automation and accuracy improvement of steel reinforcement structure quality inspection are key development directions for the industry. Traditional inspection relies on manual measurement and sampling flaw detection, which has the following technical bottlenecks: Traditional optical images are significantly affected by the reflection of metal on the surface of steel bars, resulting in distortion of texture features and a high risk of missed detection; single depth information is insufficient to analyze complex spatial structures, and the positioning error of dense steel bar intersections exceeds 5mm.

[0003] Dust and oil stains at construction sites can easily obscure the rust characteristics on the surface of steel bars. Traditional visual inspection requires manual pretreatment, which is inefficient and has a missed detection rate of over 15%.

[0004] Traditional deep learning models lack a joint optimization mechanism for spatial location and material characteristics, making it difficult to simultaneously meet the dual accuracy requirements of rebar positioning and corrosion identification.

[0005] The test results are disconnected from the Building Information Model (BIM), making it impossible to achieve a real-time closed loop of "testing-verification-correction", and the discovery of quality problems is often delayed by more than 24 hours.

[0006] The present invention aims to solve the above problems, and to this end, proposes a method for identifying steel reinforcement structure data based on image recognition. Summary of the Invention

[0007] The technical problem to be solved by the present invention is to overcome the defects of the prior art and provide a method for identifying steel structure data based on image recognition.

[0008] To solve the above-mentioned technical problems, the present invention provides the following technical solution: This invention provides a method for identifying rebar structure data based on image recognition, comprising the following steps: Step 1. Multimodal data acquisition: The interference of metal reflection is eliminated by using a polarized RGB camera, and a ToF depth camera is used simultaneously to obtain the depth information of the steel bars with millimeter-level precision. In addition, a near-infrared spectral sensor is used to penetrate the surface dust and collect the corrosion characteristics of the steel bars. Step 2. Training of Dynamic Dual-Domain Attention Optimization Model: Based on the data collected in Step 1, a dynamic dual-domain attention optimization model is trained. The model optimizes the prediction accuracy of the rebar position, diameter, and arrangement parameters through a joint loss function. Step 3. Reinforcing steel structure reasoning: Input the real-time collected construction data into the trained dynamic dual-domain attention optimization model, and output the rebar position coordinates, diameter measurement values, and arrangement spacing; Step 4. Standard Verification and Feedback: Compare the model output results with the building information model design parameters. When the diameter deviation is greater than ±2% or the spacing error is greater than ±5 mm, trigger a visual alarm and generate a correction report.

[0009] As a preferred embodiment of the present invention, the implementation of multimodal data fusion in step 1 includes the following steps: Using ToF depth map features as a style reference, the mean and variance of RGB features are adjusted through adaptive instance normalization to align the RGB features. The formula is as follows: ; in, , These are the mean and standard deviation of the depth map features, respectively; Short-time Fourier transform was performed on the 850 nm near-infrared spectral signal to enhance the near-infrared frequency domain. The window length was 64 milliseconds and the overlap rate was 50%. Corrosion-sensitive frequency bands were screened. The processed RGB features and the short-time Fourier transform frequency domain features are concatenated along the channel dimension to generate a fused feature map.

[0010] As a preferred embodiment of the present invention, the construction of the dynamic dual-domain attention optimization model includes the following steps: Based on the fused feature map, spatial weights are calculated by improving the coordinate attention module to generate spatial domain attention, as shown in the formula: ; The fused features are decomposed using Haar wavelet decomposition, and frequency band selection gating is generated through a multilayer perceptron to extract frequency domain attention. Spatial weights are superimposed with frequency domain gating to enhance dual-domain features, and the enhanced feature map is output.

[0011] As a preferred embodiment of the present invention, the joint loss function in step 2 includes: target detection loss, diameter regression loss, and arrangement similarity loss, wherein: Target loss detection: The rebar boundary box positioning is optimized using the intersection-union ratio (IUU) loss optimization method. Diameter regression loss: The diameter prediction error is constrained by a smoothed L1 loss, and the formula is as follows: ; Arrangement similarity loss: Quantifying the consistency between the rebar layout and the design drawing based on Hausdorf distance; Wherein: the weight combination of the loss function is detection loss: diameter loss: layout loss = 6:3:1.

[0012] As a preferred embodiment of the present invention, the model training method in step 2 includes the following steps: Based on a model-independent meta-learning framework, a task distribution is constructed on 200 labeled images, and the meta-model parameters are optimized. A single-step gradient update is performed when a new sample arrives.

[0013] As a preferred embodiment of the present invention, step 3 further includes: The 3D voxel field is updated using a depth map sequence, and the symbolic distance field is reconstructed in real time using the following formula: ; in, Weights for the confidence level of depth measurements; Calculate diameter deviation and spacing error, and compare with building information model according to standards.

[0014] Compared with the prior art, the beneficial effects of the present invention are as follows: 1: This invention utilizes a multimodal data collaborative acquisition mechanism, employing a polarized RGB camera to eliminate metal reflections, a ToF depth camera to acquire three-dimensional coordinates, and near-infrared spectroscopy to penetrate dust and collect corrosion features, to construct a cross-modal dataset. This solves the technical bottleneck of single sensors being affected by environmental interference and enables the complete capture of the surface texture and internal defects of steel bars under complex working conditions.

[0015] 2: This invention uses a dynamic dual-domain attention optimization model to introduce spatial domain attention to enhance the positioning weight of dense steel bar intersections and combine it with frequency domain attention to extract rust feature frequency bands, thereby achieving joint modeling of the spatial layout and material state of steel bars. This breaks through the generalization limitation of traditional models in a single scenario and improves the detection accuracy of complex structures.

[0016] 3: This invention, through BIM integration and real-time feedback links, dynamically compares the detection results with BIM design parameters, triggers AR visualization alarms and automatically generates correction reports, constructs a complete closed loop for the entire process, and promotes the transformation of quality control from post-event sampling inspection to real-time process control. Attached Figure Description

[0017] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of the present invention; Figure 2 This is a flowchart of the construction scenario deployment of the present invention; Figure 3 This is a flowchart of the diameter regression loss of the present invention; Figure 4 This is a schematic diagram of the deployment of the on-site sensor array and edge computing terminal in the embodiment; Detailed Implementation

[0018] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention. Example 1

[0019] like Figure 1-4 As shown, this invention provides a method for identifying rebar structure data based on image recognition, comprising the following steps: Step 1. Multimodal data acquisition: The interference of metal reflection is eliminated by using a polarized RGB camera, and a ToF depth camera is used simultaneously to obtain the depth information of the steel bars with millimeter-level precision. In addition, a near-infrared spectral sensor is used to penetrate the surface dust and collect the corrosion characteristics of the steel bars. Step 2. Training of Dynamic Dual-Domain Attention Optimization Model: Based on the data collected in Step 1, a dynamic dual-domain attention optimization model is trained. The model optimizes the prediction accuracy of the rebar position, diameter, and arrangement parameters through a joint loss function. Step 3. Reinforcing steel structure reasoning: Input the real-time collected construction data into the trained dynamic dual-domain attention optimization model, and output the rebar position coordinates, diameter measurement values, and arrangement spacing; Step 4. Standardized Verification and Feedback: Compare the model output results with the Building Information Model (BIM) design parameters. When the diameter deviation is greater than ±2% or the spacing error is greater than ±5 mm, trigger a visual alarm and generate a correction report.

[0020] Furthermore, the implementation of multimodal data fusion in step 1 includes the following steps: Using ToF depth map features as a style reference, the mean and variance of RGB features are adjusted through adaptive instance normalization to align the RGB features. The formula is as follows: ; in, , These are the mean and standard deviation of the depth map features, respectively, to address the texture distortion problem caused by metallic reflections; Short-time Fourier transform was performed on the 850 nm near-infrared spectral signal to enhance the near-infrared frequency domain. The window length was 64 milliseconds and the overlap rate was 50%. Corrosion-sensitive frequency bands were screened. The processed RGB features and the short-time Fourier transform frequency domain features are concatenated along the channel dimension to generate a fused feature map.

[0021] Furthermore, the construction of the dynamic dual-domain attention optimization model includes the following steps: Based on the fused feature map, spatial weights are calculated by improving the coordinate attention module to generate spatial domain attention, as shown in the formula: ; The fused features are decomposed using Haar wavelet decomposition, and frequency band selection gating is generated through a multilayer perceptron to extract frequency domain attention. Spatial weights are superimposed with frequency domain gating to enhance dual-domain features, and the enhanced feature map is output.

[0022] Furthermore, the joint loss function in step 2 includes: target detection loss, diameter regression loss, and arrangement similarity loss, where: Target loss detection: The rebar boundary box positioning is optimized using the intersection-union ratio (IUU) loss optimization method. Diameter regression loss: The diameter prediction error is constrained by a smoothed L1 loss, and the formula is as follows: ; Arrangement similarity loss: Quantifying the consistency between the rebar layout and the design drawing based on Hausdorf distance; Wherein: the weight combination of the loss function is detection loss: diameter loss: layout loss = 6:3:1.

[0023] Furthermore, the model training method in step 2 includes the following steps: Based on a model-independent meta-learning framework, a task distribution is constructed on 200 labeled images, and the meta-model parameters are optimized. A single-step gradient update is performed when a new sample arrives.

[0024] Furthermore, step 3 also includes: The 3D voxel field is updated using a depth map sequence, and the symbolic distance field is reconstructed in real time using the following formula: in, Weights for the confidence level of depth measurements; The diameter deviation and spacing error were calculated, and the standard was compared with the building information model. The threshold setting complied with the "Code for Acceptance of Construction Quality of Concrete Structures" (GB50204-2015).

[0025] Specific implementation This embodiment is based on a multimodal sensor array and edge computing architecture, wherein the devices are as follows: Polarized RGB camera: The FLIRBFS-PGE-50S5C-C polarized camera is used, and a ring polarizing filter is configured to eliminate the interference of metal reflection on the surface of the steel bars.

[0026] ToF depth camera: Intel RealSense LiDAR L515 camera, 30fps, supports hardware synchronization of depth map and RGB map.

[0027] Near-infrared spectral sensor: Hamamatsu C12702-12 spectrometer, detection range 800-900nm, equipped with 850±10nm bandpass filter, to collect the characteristic spectrum of corrosion and solve the problem of feature loss caused by surface shading.

[0028] Edge computing terminal: NVIDIA Jetson Xavier NX, equipped with CUDA acceleration module, achieves multimodal data fusion and model inference latency ≤500ms, meeting the real-time requirements of construction site.

[0029] BIM integration platform: Autodesk Revit API interface, real-time synchronization of design parameters, allowable error of ±5mm for rebar diameter and spacing, and build a closed loop of "inspection-verification-feedback".

[0030] Implementation steps: first step: A polarized RGB camera captures the surface texture of the rebar at 30fps, simultaneously triggering a ToF camera to acquire a depth map (X, Y, Z coordinates), while a near-infrared spectrometer scans the rebar surface point by point, acquiring signals in the 850nm band. This hardware-synchronized triggering mechanism ensures spatiotemporal alignment of the three-modal data, avoiding feature misalignment caused by traditional asynchronous acquisition.

[0031] Using ToF depth map features as a reference, the mean and variance of RGB features are adjusted through adaptive instance normalization (formula below) to eliminate brightness unevenness caused by reflection: ; in, , These are the mean and standard deviation of the depth map features, respectively. , , which is an RGB feature statistic, improves the contrast of the RGB image by 35% after processing.

[0032] A short-time Fourier transform (SFT) was performed on the spectral signal with a window length of 64 ms and an overlap rate of 50%, extracting the 12-18 Hz corrosion-sensitive frequency band, which improved the signal-to-noise ratio of the frequency domain features by 20 dB. The processed 3-channel RGB features and 8-channel SFT frequency domain features were then concatenated along the channels to generate an 11-channel fused feature map, providing cross-modal enhancement information for subsequent models.

[0033] Step Two: Based on the fused feature map, spatial weights are calculated by improving the coordinate attention module, using the following formula: ; By assigning a high weight of 0.8-0.95 to the intersections of densely packed reinforcing bars, actual measurements showed that the positioning error in dense areas was reduced from 5mm to 3mm.

[0034] The fused features are decomposed using Haar wavelet decomposition, and a frequency band selection gate is generated through a multilayer perceptron to automatically filter high-frequency components related to corrosion, improving the utilization rate of frequency domain features by 50%. The spatial weights are superimposed with the frequency domain gate to output an enhanced feature map, enabling the model to capture both spatial location and material features simultaneously.

[0035] Multi-task joint loss function optimization: Detection target loss: CIoU loss, optimized rebar boundary box positioning, average intersection-to-union ratio reaches 0.92.

[0036] Diameter regression loss: smoothed L1 loss, constrained diameter prediction error ≤ 0.5 mm, actual measured root mean square error is 0.3 mm.

[0037] Arrangement similarity loss: Hausdorff distance, quantifying the consistency between steel reinforcement layout and BIM design, with 95% of samples having an error ≤2mm.

[0038] Based on a model-independent meta-learning framework, a meta-training task is constructed using 200 labeled images.

[0039] Step 3: Real-time fusion of feature maps as input to the trained model, output: Reinforcing bar center coordinates (X, Y, Z), accuracy ±2mm; Diameter measurement, accuracy ±0.5mm; The spacing between adjacent reinforcing bars is accurate to ±1mm.

[0040] The three-dimensional voxel field is updated in real time using depth map sequences and reconstructed using truncated symbolic distance field (TSDF).

[0041] Step 4: Compare the model output parameters with the BIM design values: Diameter deviation threshold: ±2% (e.g., if the design diameter is 20mm, the allowable range is 19.6-20.4mm). Spacing error threshold: ±5mm (compliant with GB50204-2015 standard).

[0042] When deviations exceed limits, the edge computing terminal triggers an AR visualization alarm via API, marking abnormal steel bars with red boxes in the construction scene. Simultaneously, a correction report containing error data and a 3D comparison image is generated and pushed to the construction management platform. This achieves a shift from "post-event sampling" to "process control," with an anomaly response time of ≤10 seconds, improving efficiency by 90% compared to traditional manual inspection.

[0043] Practical application examples: Project Background: A frame structure with two underground floors and 18 floors above ground, with a total construction area of ​​82,000 square meters and 4,200 tons of steel reinforcement. Traditional manual inspection suffers from low efficiency, insufficient accuracy, and significant interference from the construction environment.

[0044] Implementation process: Two FLIRBFS-PGE-50S5C-C polarization cameras were deployed at the construction site, covering a 200㎡ work area; one Intel RealSense L515 depth camera was installed at the end of the robotic arm; and one Hamamatsu C12702-12 spectrometer was integrated into the inspection robot. All equipment was synchronized via an NVIDIA Jetson Xavier NX edge computing terminal.

[0045] The polarization camera eliminates 92% of metal reflections and obtains clear textures on the surface of steel bars, such as the depth error of the rib pattern of rebar ≤0.1mm.

[0046] The depth camera reconstructs the 3D model of the reinforcing steel in real time. In dense areas, such as beam-column joints, the positioning error is ≤2mm.

[0047] The spectrometer can penetrate 200μm dust particles to detect the characteristic spectra of rust, with a 30% improvement in response value at the 12-18Hz frequency range.

[0048] Based on the project's rebar types (HRB400EΦ12-Φ32), a small-sample training was conducted using the MAML framework, with an initial set of 200 labeled samples. The adaptation period for new samples was ≤1 hour. The rebar diameter measurement error was ≤0.3mm, with a design value of Φ20mm and a measured mean of 20.1mm and a standard deviation of 0.2mm. The spacing error was ≤2mm, with a design spacing of 150mm and a measured mean of 151mm, conforming to GB50204-2015 standards. The corrosion level identification accuracy reached 95%. The detection data was synchronized to the Autodesk Revit platform in real time, automatically generating a "Rebar Construction Quality Comparison Report." For example, in a certain area, the measured diameter of Φ25 rebar was 24.8mm, a deviation of -0.8%, triggering a system alert and pushing it to the construction management platform.

[0049] The AR terminal marks abnormal locations with red boxes on site and simultaneously generates correction plans, such as replacing steel bars or adjusting spacing.

[0050] Implementation results: The daily testing capacity increased from 20 tons to 120 tons, improving efficiency by 6 times.

[0051] The testing cycle for a single rebar is ≤15 seconds.

[0052] The positioning error at the intersection of densely packed reinforcing bars was reduced from 5mm to 2mm, improving the positioning accuracy by 60%.

[0053] Under complex operating conditions, the feature retention rate increased from 60% to 85%, and the false detection rate decreased from 12% to 3%.

[0054] Based on reducing the average number of people from 8 to 2 per day, the cost of manual testing will be reduced by about 70%.

[0055] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for identifying steel reinforcement structure data based on image recognition, characterized in that, Includes the following steps: Step 1. Multimodal data acquisition: The interference of metal reflection is eliminated by using a polarized RGB camera, and a ToF depth camera is used simultaneously to obtain the depth information of the steel bars with millimeter-level precision. In addition, a near-infrared spectral sensor is used to penetrate the surface dust and collect the corrosion characteristics of the steel bars. Step 2. Training of Dynamic Dual-Domain Attention Optimization Model: Based on the data collected in Step 1, a dynamic dual-domain attention optimization model is trained. The model optimizes the prediction accuracy of the rebar position, diameter, and arrangement parameters through a joint loss function. Step 3. Reinforcing steel structure reasoning: Input the real-time collected construction data into the trained dynamic dual-domain attention optimization model, and output the rebar position coordinates, diameter measurement values, and arrangement spacing; Step 4. Standard Verification and Feedback: Compare the model output results with the building information model design parameters. When the diameter deviation is greater than ±2% or the spacing error is greater than ±5 mm, trigger a visual alarm and generate a correction report.

2. The method for identifying steel reinforcement structure data based on image recognition according to claim 1, characterized in that, The implementation of multimodal data fusion in step 1 includes the following steps: Using ToF depth map features as a style reference, the mean and variance of RGB features are adjusted through adaptive instance normalization to align the RGB features. The formula is as follows: ; in, , These are the mean and standard deviation of the depth map features, respectively; Short-time Fourier transform was performed on the 850 nm near-infrared spectral signal to enhance the near-infrared frequency domain. The window length was 64 milliseconds and the overlap rate was 50%. Corrosion-sensitive frequency bands were screened. The processed RGB features and the short-time Fourier transform frequency domain features are concatenated along the channel dimension to generate a fused feature map.

3. The method for identifying steel reinforcement structure data based on image recognition according to claim 2, characterized in that, The construction of the dynamic dual-domain attention optimization model includes the following steps: Based on the fused feature map, spatial weights are calculated by improving the coordinate attention module to generate spatial domain attention, as shown in the formula: ; The fused features are decomposed using Haar wavelet decomposition, and frequency band selection gating is generated through a multilayer perceptron to extract frequency domain attention. Spatial weights are superimposed with frequency domain gating to enhance dual-domain features, and the enhanced feature map is output.

4. The method for identifying steel reinforcement structure data based on image recognition according to claim 1, characterized in that, The joint loss function in step 2 includes: target detection loss, diameter regression loss, and layout similarity loss, wherein: Target loss detection: The rebar boundary box positioning is optimized using the intersection-union ratio (IUU) loss optimization method. Diameter regression loss: The diameter prediction error is constrained by a smoothed L1 loss, and the formula is as follows: ; Arrangement similarity loss: Quantifying the consistency between the rebar layout and the design drawing based on Hausdorf distance; Wherein: the weight combination of the loss function is detection loss: diameter loss: layout loss = 6:3:

1.

5. The method for identifying rebar structure data based on image recognition according to claim 1, characterized in that, The model training method in step 2 includes the following steps: Based on a model-independent meta-learning framework, a task distribution is constructed on 200 labeled images, and the meta-model parameters are optimized. A single-step gradient update is performed when a new sample arrives.

6. The method for identifying steel reinforcement structure data based on image recognition according to claim 1, characterized in that, Step 3 also includes: The 3D voxel field is updated using a depth map sequence, and the symbolic distance field is reconstructed in real time using the following formula: ; in, Weights for the confidence level of depth measurements; Calculate diameter deviation and spacing error, and compare with building information model according to standards.