Hot-dip galvanized steel pipe cage verticality measurement system based on machine vision

Through polarization coding and multi-view camera technology based on machine vision, the mirror reflection problem in the verticality detection of hot-dip galvanized steel pipe cages was solved, and efficient and accurate posture recognition and verticality assessment of multiple steel pipes were achieved.

CN120489067BActive Publication Date: 2025-09-12SHANXI INSTALLATION GRP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510969592.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-09-12
Estimated Expiration
2045-07-15

AI Technical Summary

Technical Problem

In the existing technology, the verticality detection of hot-dip galvanized steel pipe cages is subject to mirror reflection, which causes laser spot distortion and measurement errors. The traditional point measurement method is inefficient and manual measurement is highly subjective, making it difficult to achieve simultaneous evaluation of multiple steel pipes in the entire cage.

Method used

A machine vision-based method is adopted. Polarization-coded grayscale stripes are projected onto the steel pipe through a projection unit. A multi-view camera is used to capture images and perform polarization differential fusion to reconstruct a three-dimensional point cloud. The vertical deviation angle is calculated in combination with robust cylindrical axis fitting.

Benefits of technology

It achieves high-fidelity imaging and three-dimensional reconstruction of groups of hot-dip galvanized steel pipes, suppresses mirror reflection artifacts, improves detection efficiency, ensures measurement accuracy and robustness, and can synchronously identify the spatial posture of multiple steel pipes at a time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120489067B_ABST
    Figure CN120489067B_ABST
Patent Text Reader

Abstract

The present application discloses a machine vision-based verticality measurement system for hot-dip galvanized steel pipe cages, which relates to the field of physical measurement technology. The system comprises: a projection unit for projecting polarization-coded grayscale stripes onto the hot-dip galvanized steel pipe cage to be measured; an imaging unit comprising at least two cameras that form a triangulation baseline with the projection unit to synchronously capture multi-perspective images from different perspectives; and a processing unit configured to: perform polarization differential fusion on the multi-perspective images to suppress specular reflection, generate corresponding polarization differential fusion images, reconstruct a three-dimensional point cloud of the pipe body of each hot-dip galvanized steel pipe in the hot-dip galvanized steel pipe cage, generate the pipe body axial vector by robust cylindrical axis fitting, and calculate the pipe body verticality deviation angle based on the pipe body axial vector and the gravity direction. Thus, while ensuring measurement accuracy, the spatial posture of multiple steel pipes can be synchronously captured and identified at one time, eliminating the need for repeated scanning at multiple points, significantly improving detection efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of physical measurement technology, and in particular to a machine vision-based hot-dip galvanized steel pipe cage verticality measurement system. Background Art

[0002] As the core load-bearing component in foundation pit support and pile foundation reinforcement, the verticality of hot-dip galvanized steel pipe cages is directly related to the force balance and long-term stability of the entire superstructure. Therefore, rapid and accurate verticality testing during the construction or prefabrication phase has become a key step in quality control.

[0003] Currently, common methods for measuring verticality primarily include manual measurement and those assisted by electronic measuring instruments. In manual measurement, workers use hand tools such as plumb bobs, laser plumbs, or levels to estimate the vertical angle between the steel pipe and the ground. However, this method relies on the operator's experience and subjective judgment, resulting in unstable measurement results. It typically only allows for alignment of a single point, requiring repeated measurements to cover the entire component. However, simultaneous evaluation of multiple steel pipes within a cage remains difficult, and the method lacks the ability to record unified, quantitative data.

[0004] In the electronic measuring instrument-assisted method, construction companies use laser levels or laser scanners (such as laser rangefinders and structured light 3D scanners) for auxiliary measurement. These instruments can achieve high projection accuracy through a non-contact method. However, laser-assisted measurement generally uses a point measurement method, which makes it difficult to fully capture the spatial axis and posture of the entire steel pipe. In addition, the surface of hot-dip galvanized steel pipe has specular reflective properties. Strong reflections can cause distortion or severe diffraction of the laser spot, causing severe interference with the echo signal, leading to measurement errors or data loss. Summary of the Invention

[0005] The present application provides a method, system, storage medium, computer program product and electronic device for measuring the verticality of a hot-dip galvanized steel pipe cage based on machine vision, which is used to at least solve the problems in the current related technology of laser spot distortion and measurement error caused by mirror reflection on the surface of the hot-dip galvanized steel pipe cage, as well as the strong subjectivity and low efficiency of manual measurement in traditional point measurement methods.

[0006] In a first aspect, an embodiment of the present application provides a machine vision-based verticality measurement system for a hot-dip galvanized steel pipe cage, comprising: a projection unit for projecting polarization grayscale stripes with polarization coding onto the hot-dip galvanized steel pipe cage to be measured; an imaging unit comprising at least two cameras forming a triangulation baseline with the projection unit, wherein each of the cameras is arranged at a different position around the hot-dip galvanized steel pipe cage to synchronously capture multi-perspective images after projecting the polarization grayscale stripes from different perspectives; a processing unit electrically connected to the projection unit and the imaging unit, and configured to: perform polarization differential fusion on the multi-perspective images to suppress mirror reflection, and generate corresponding polarization differential fusion images; reconstruct a three-dimensional point cloud of the pipe body corresponding to each hot-dip galvanized steel pipe in the hot-dip galvanized steel pipe cage according to the polarization differential fusion image; use robust cylindrical axis fitting for each of the three-dimensional point clouds of the pipe body to generate a pipe body axial vector, and calculate the pipe body verticality deviation angle based on the pipe body axial vector and the gravity direction.

[0007] In the second aspect, an embodiment of the present application provides a method for measuring the verticality of a hot-dip galvanized steel pipe cage based on machine vision, comprising: performing polarization differential fusion on multi-perspective images to suppress mirror reflection, and generating a corresponding polarization differential fusion image; reconstructing a three-dimensional point cloud of the pipe body corresponding to each hot-dip galvanized steel pipe in the hot-dip galvanized steel pipe cage based on the polarization differential fusion image; using robust cylindrical axis fitting on each of the three-dimensional point clouds of the pipe body to generate an axial vector of the pipe body, and calculating the verticality deviation angle of the pipe body based on the axial vector of the pipe body and the direction of gravity.

[0008] The hot-dip galvanized steel pipe cage verticality measurement system provided by this application based on machine vision can produce at least the following technical effects:

[0009] (1) Through the integrated design of polarization-coded grayscale stripes and multi-view polarization differential fusion, high-fidelity imaging and three-dimensional reconstruction of a group of hot-dip galvanized steel pipes are achieved, while suppressing mirror reflection artifacts while maintaining the integrity and continuity of the point cloud data. Furthermore, with the help of robust cylindrical model fitting, noise and occlusion errors are effectively removed, and the central axial vector of each pipe body is accurately extracted. Finally, through the geometric relationship between the axial vector and the direction of gravity, a quantifiable vertical deviation angle is directly output. Thus, while ensuring measurement accuracy, the spatial posture of multiple steel pipes can be synchronously acquired and identified at one time, without the need for repeated scanning at multiple points, significantly improving detection efficiency.

[0010] (2) In the visual perception link, by introducing polarization-coded grayscale stripes as active projection structured light sources and combining them with a multi-view camera layout strategy, multi-directional imaging coverage is achieved on the surface of hot-dip galvanized steel pipes. Due to the strong specular reflection characteristics of the steel pipe surface, traditional imaging methods are difficult to obtain clear and stable structural features. However, this system effectively suppresses the overexposure areas and noise interference caused by specular reflections by performing polarization differential fusion on the collected images during the image processing stage. The fused image enhances the recognition of object boundaries while maintaining the clarity of the stripe structure, providing a better visual basis for the system to restore the actual structural morphology of the steel pipe.

[0011] (3) In terms of geometric analysis, the system performs 3D point cloud reconstruction based on the fused image. Then, to address possible occlusion, missing data, or residual noise in the point cloud, the system introduces a robust estimation algorithm to fit the cylindrical axis of each steel pipe's 3D point cloud. This algorithm can automatically identify and eliminate outliers, ensuring that only valid data segments are used during the fitting process, thereby accurately recovering the central axial vector of the pipe body. As a result, the system achieves stable measurement of the pipe body's posture parameters and maintains strong robustness in various field conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0013] Figure 1 A structural connection block diagram of an example of a machine vision-based hot-dip galvanized steel pipe cage verticality measurement system according to an embodiment of the present application is shown;

[0014] Figure 2 A schematic diagram showing a layout of an example of a structured light scanning device according to an embodiment of the present application;

[0015] Figure 3 A schematic structural diagram of an example of a projection unit according to an embodiment of the present application is shown;

[0016] Figure 4 A schematic structural diagram of an example of a stripe encoder according to an embodiment of the present application is shown;

[0017] Figure 5 A flowchart of an example of a method for measuring the verticality of a hot-dip galvanized steel pipe cage based on machine vision, performed by a processing unit according to an embodiment of the present application, is shown;

[0018] Figure 6An operation flow chart showing an example of bias self-calibration operation according to an embodiment of the present application is shown;

[0019] Figure 7 An operational flow chart of an example of performing polarization differential fusion on multi-view images according to an embodiment of the present application is shown;

[0020] Figure 8 An operational flowchart of an example of reconstructing a three-dimensional point cloud of a pipe body according to an embodiment of the present application is shown;

[0021] Figure 9 A structural connection diagram of an example of a spatiotemporal attention network according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0022] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0023] It should be noted that hot-dip galvanized steel pipe cages are core load-bearing components in foundation pit support and pile foundation reinforcement. Their verticality is directly related to the uniformity of load distribution and the overall stability of the structure. Traditional one-time manual calibration is time-consuming and affects construction progress. On-site quality inspections often require hundreds or even thousands of cage steel pipes to be quickly carried out.

[0024] At present, the common methods for measuring the verticality of ground cages mainly include:

[0025] 1) Measurements using manual levels / theodolites are widely available and easy to use. However, they suffer from single-point alignment and time-consuming repeated measurements, leading to line-of-sight deviations and reading errors that directly affect accuracy. Furthermore, simultaneous measurement of multiple tubes in an entire cage is difficult.

[0026] 2) Measurements are performed using contact-type inclination sensors, which can output inclination data in real time and do not require advanced operator experience. However, they require close contact with the pipe surface and are significantly affected by mounting position, magnetic field interference, and zero-point drift. They are also sensitive to even the slightest unevenness of the galvanized surface.

[0027] 3) Non-contact laser scanners can capture multi-point data, but the strong reflection from the hot-dip galvanized mirror surface can cause the laser spot to distort and diffract, leading to distorted echo signals. Furthermore, high-end 3D laser scanning systems are bulky and expensive, making them unsuitable for large-scale deployment.

[0028] It should be understood that the purpose of the above description of the current related art is only to facilitate the public to better understand the inventive spirit and motivation of this application, and is not to be construed as limiting this application. In addition, the technical solutions described in the above-mentioned current related art are not prior art and may also be undisclosed technical solutions, such as solutions under research or in the laboratory stage.

[0029] In the technical solutions of this application, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved shall comply with the provisions of relevant laws and regulations and shall not violate public order and good morals.

[0030] Figure 1 A structural connection block diagram of an example of a hot-dip galvanized steel pipe cage verticality measurement system based on machine vision according to an embodiment of the present application is shown.

[0031] like Figure 1 As shown, the hot-dip galvanized steel pipe cage verticality measurement system 100 based on machine vision includes a projection unit 10, an imaging unit 20 and a processing unit 30, and the processing unit 30 is electrically connected to the projection unit 10 and the imaging unit 20 respectively.

[0032] More specifically, the projection unit 10 is used to project polarization-coded grayscale stripes onto the hot-dip galvanized steel pipe cage to be measured. The imaging unit 20 includes at least two cameras that form a triangulation baseline with the projection unit. Each camera is positioned at a different position around the hot-dip galvanized steel pipe cage to simultaneously capture multi-view images of the projected grayscale stripes from different perspectives.

[0033] Figure 2 A schematic diagram showing the layout of an example of a structured light scanning device according to an embodiment of the present application is shown.

[0034] like Figure 2 As shown, the projection unit 10 includes a projector, and the imaging unit 20 includes two industrial cameras: a first camera and a second camera. The first camera, the second camera, and the projector form a triangulation baseline. The two cameras are positioned at different positions on the cage (e.g., left and right) to capture images from different perspectives. A switchable polarization filter is installed in front of each camera. When the filter is rotated to 0° or 90°, it captures images in the corresponding polarization state.

[0035] exist Figure 2 In the example shown, each camera establishes its own camera coordinate system, and a point on the surface of the ground cage to be measured directly in front of the projector is p (In the world coordinate system) is imaged by the first camera and the second camera as the image plane p 1( u 1, v 1) and p 2(u 2, v 2). The image plane coordinate system takes o as the origin, and the u and v axes point to the horizontal and vertical directions respectively. p → p 1. p → p The dotted line in 2 indicates the projection ray of the camera optical center through the pixel point.

[0036] More specifically, the projector is located between the two cameras and can rotate to project polarized grayscale stripes to provide coded light patterns for the ground cage steel pipe. The first camera and the second camera synchronously capture the coded multi-frame images to obtain two sets of corresponding pixel coordinates. and , these two pixel pairs together determine the point p Positional relationship in three-dimensional space.

[0037] It should be understood that Figure 2 The dual-camera measurement method described in the figure is only used as an example. A larger number of cameras can also be used to achieve more accurate stripe structured light analysis and angle measurement.

[0038] Figure 3 A schematic structural diagram of an example of a projection unit according to an embodiment of the present application is shown. Figure 4 A structural schematic diagram of an example of a stripe encoder according to an embodiment of the present application is shown.

[0039] like Figure 3 and 4 As shown, the projection unit 10 includes a projector, a stripe encoder, and a stepper motor. The stripe encoder includes an inner polarization encoding wheel and an outer rotating polarization wheel. More specifically, the inner polarization encoding wheel is configured with multiple grayscale stripe segments at equal intervals. The grayscale level of each grayscale stripe segment varies by a preset increment between adjacent segments and corresponds to a fixed polarization angle to achieve primary polarization encoding of the stripes. Specifically, the inner polarization encoding wheel is used to sequentially project multiple grayscale stripes, each stripe having a corresponding polarization angle. The outer rotating polarization wheel is configured to continuously rotate around the optical axis, dynamically changing the global polarization direction of the entire projected beam by the rotation angle.

[0040] In some embodiments, the visible light structured light emitted by the projector first passes through the outer rotating polarization wheel, and then enters the grayscale stripe section of the inner polarization encoding wheel. The encoded polarization grayscale stripes are finally projected onto the surface of the hot-dip galvanized steel pipe cage to be measured; the stepper motor is fixedly connected to the outer rotating polarization wheel, and drives the outer polarization wheel to rotate continuously around the optical axis through precise stepping instructions issued by the processing unit to achieve dynamic adaptive control of the global polarization direction.

[0041] For example, the inner polarization encoding wheel is evenly spaced with 10 (or more) grayscale stripe segments, and the grayscale level of each grayscale stripe segment is Changes between adjacent segments by preset increments (e.g. 16, 40, 64...240) and corresponds to a fixed polarization angle (alternating 0° / 90°) to achieve primary polarization encoding of the stripes; the outer rotating polarization wheel is used to continuously rotate around the optical axis, dynamically changing the global polarization direction of the entire projected light beam through the rotation angle, thereby suppressing specular highlights and enhancing the contrast of diffuse reflection details.

[0042] Specifically, the processing unit 30 is configured to determine the outer polarization angle based on the current ambient light reflectance estimate. Furthermore, the projection unit 10 is configured to drive the outer rotating polarization wheel to rotate to the outer polarization angle and control the inner polarization encoding wheel to switch grayscale stripe segments to sequentially project multiple grayscale stripes, thereby projecting multiple frames of polarized grayscale stripes.

[0043] For example, the processing unit 30 collects the reflected brightness value of the target surface through a built-in ambient light sensor or a pair of auxiliary cameras. , and calculate the optimal global polarization angle for the current scene based on the pre-calibrated reflectivity-optimal polarization angle mapping relationship table More specifically, the processing unit 30 collects N frames of ambient light images, calculates the grayscale mean, and determines the required angle of the outer polarization wheel by querying the reflectivity-optimal polarization angle mapping relationship table. Then, it sends a command to the stepper motor to rotate the outer polarization wheel to the angle with a minimum step accuracy of 0.1°. At the same time, the processing unit 30 sends a switching signal to drive the inner polarization coding wheel to step, switching to the first Segment, projected grayscale stripes and preserve the polarization state of the stripes Repeat the above steps and cyclically project 10 or more frames of polarization grayscale stripes to complete the projection of multi-frame polarization coded structured light.

[0044] In this embodiment, a built-in optical sensor measures the cage's surface reflectivity in real time and maps the optimal global polarization angle. This controls the outer rotating polarization wheel to precisely adjust the polarization direction of the projected light beam, effectively removing specular highlights. Simultaneously, the inner encoding wheel projects multiple frames of stripes at preset grayscale increments and alternating polarization angles, ensuring high contrast and consistency for each frame of structured light stripes on the zinc metal surface. This achieves adaptive dual-layer polarization encoding, ensuring stable output of clear, uniform grayscale stripes even in scenarios with fluctuating light intensity or significant differences in tube reflectivity.

[0045] Figure 5A flowchart of an example of a method for measuring the verticality of a hot-dip galvanized steel pipe cage based on machine vision, performed by a processing unit according to an embodiment of the present application is shown.

[0046] like Figure 5 As shown, in step S510, polarization differential fusion is performed on the multi-view images to suppress specular reflection, and a corresponding polarization differential fusion image is generated.

[0047] In some embodiments, cameras positioned at multiple spatial locations synchronously capture the cage structure, capturing images of the fringe structure formed under specific projection conditions. The projection unit emits grayscale fringe light with polarization-encoded properties. The projected fringe forms a reflection pattern on the steel pipe surface, and different cameras capture the changes in this pattern from different angles.

[0048] It should be noted that due to the high specular reflectivity of hot-dip galvanized steel pipes, strong reflective interference often occurs during ordinary imaging, resulting in overexposed areas, halo effects, or blurred structural edges in the image, which seriously affects the usability of subsequent images. Here, the system adjusts the polarization direction of the camera and projection light source, and uses images obtained at multiple polarization angles to perform polarization differential operations. That is, by performing differential calculations and weighted fusion on images under different polarization states, the system can effectively separate the specular reflection component from the diffuse reflection component. As a result, the problem of excessive brightness concentration in the specular reflection area is suppressed, allowing the structural details of the surface of the hot-dip galvanized steel pipe to be visually restored, especially at the edges of the stripe pattern, which shows higher image contrast and texture clarity.

[0049] In step S520, the three-dimensional point cloud of each hot-dip galvanized steel pipe in the hot-dip galvanized steel pipe cage is reconstructed based on the polarization difference fusion image.

[0050] In some embodiments, the system relies on the spatial geometric relationship between the projection system and the camera, that is, the triangulation principle. The same structure point in the multi-view image has spatial projection differences in different cameras. The system calculates the structure point by performing pixel-level stereo matching on these differences. p The real coordinate position in three-dimensional space. In addition, since specular reflection has been suppressed through polarization differential fusion, the mismatch rate in the point cloud reconstruction process is significantly reduced, which improves the point cloud density and accuracy. The three-dimensional point cloud not only covers the continuous surface area of ​​the steel pipe surface, but also has a high spatial resolution, which can fully describe the spatial position and overall geometric shape of each steel pipe.

[0051] In step S530, robust cylindrical axis fitting is performed on each tube 3D point cloud to generate a tube axial vector, and the tube verticality deviation angle is calculated based on the tube axial vector and the gravity direction.

[0052] In some implementations, axis extraction is performed separately for each point cloud dataset. This can be achieved based on the mathematical assumption of a cylindrical geometry model, for example, by approximating a steel pipe as a cylinder with a uniform cross-section. During the fitting process, the system employs a robust fitting algorithm to avoid the influence of local anomalies in the point cloud, such as missing points due to surface scratches, weld residue, or acquisition occlusion.

[0053] Here, through optimal fitting of spatial points, the spatial principal axis direction of each steel pipe—its axial vector—is estimated. This vector is used to represent the pipe's actual installation posture in a three-dimensional coordinate system. The angle between the two and a pre-calibrated gravity direction vector (e.g., obtained through an inertial measurement unit) is calculated to reflect the deviation between the actual installation orientation of each steel pipe and the ideal gravity direction. This system then outputs a numerical verticality deviation indicator for each steel pipe, effectively identifying pipes with abnormal verticality in the cage and enabling a refined assessment of the verticality of each pipe in the entire cage.

[0054] It should be noted that the surface of hot-dip galvanized steel pipes is prone to local high-intensity specular reflections. If the ambient light intensity and reflection characteristics are not considered during projection and imaging, problems such as discontinuous projection stripes, camera acquisition saturation, or low contrast may occur.

[0055] In view of this, in some examples of the embodiments of the present application, the processing unit 30 is further configured to perform a bias self-calibration operation before measurement, by monitoring the ambient light and providing an initial polarization angle for the polarization encoding strategy, effectively suppressing the mirror reflection noise through polarization correction, and ensuring that the polarization difference only reflects the scene diffuse reflection, thereby improving the measurement accuracy.

[0056] Figure 6 An operational flow chart illustrating an example of a bias self-calibration operation according to an embodiment of the present application is shown.

[0057] like Figure 6 As shown, in step S610, the fringe projection of the projection unit is turned off, each camera is controlled to capture multiple frames of ambient light images at 0° and 90° polarization directions, the average of the absolute values ​​of the polarization differences is calculated to estimate the scene reflectivity, and the initial angle of the polarization wheel of the projection unit is adaptively adjusted based on the reflectivity range.

[0058] For example, when the projection unit is turned off for fringe projection, different cameras each capture the fringe with the polarizer at 0° / 90°. M Frame ambient light image.

[0059] The system turns off grayscale structured light projection and only lets the two cameras collect light at the 0° / 90° position of the polarization filter. Frame environment images, thereby obtaining two sets of "pure environment" data.

[0060] Calculate the polarization difference:

[0061] , formula (1)

[0062] Where, Indicates the In the frame ambient light image, the polarizing filter is placed at an angle (value 0° or 90°), located at the horizontal pixel coordinate of the image and vertical pixel coordinates The grayscale intensity value at . Indicates the Frames in the same pixel The grayscale difference of the polarization image at 0°–90° is used to estimate the intensity of the specular reflection component of the pixel.

[0063] Take the absolute value and mean to estimate the average reflectivity of the area:

[0064] , Formula (2)

[0065] Where, It represents the average value of the absolute polarization difference of all pixels in all frames, reflecting the specular reflection intensity level of the overall environment; and Represent the number of rows and columns of the image, Indicates the total number of ambient light images.

[0066] Map the initial value of the outer polarization wheel according to the calculation results :

[0067] , Formula (3)

[0068] Where, Represents the initial rotation angle of the outer rotating polarization wheel, used to generate adaptive polarization for the specular reflection intensity. and Represent the preset low reflectivity threshold and high reflectivity threshold respectively.

[0069] Here, dual polarization difference (single frame ), remove most of the consistent components of diffuse reflection and highlight the residual of specular reflection. After taking the absolute value of the difference of each frame, average it pixel by pixel across frames to get the overall average reflectivity This process takes into account both spatial (full resolution pixels) and temporal (multi-frame time scale) characteristics to avoid misjudgment of accidental anomalies at a single point or a single frame. In addition, by setting the low and high reflection thresholds and , and linear interpolation is used to Mapping to the initial value of polarization angle The initial polarization angle is set so that it neither oversaturates nor overly suppresses details.

[0070] In step S620, the camera's intrinsic parameter matrix, distortion parameters, and extrinsic parameters are calibrated using the multi-pose checkerboard calibration plate image and a reprojection error minimization method, and the triangular baseline length is calculated using the camera's translation vector difference.

[0071] Here, in order to achieve high-precision 3D reconstruction, it is necessary to accurately obtain the intrinsic parameters (focal length, principal point, distortion parameters) and extrinsic parameters (rotation, translation) of each camera, as well as the baseline length between cameras. With minimizing the reprojection error as the core, the integrated calibration of internal and external parameters is completed through multi-view checkerboard shooting.

[0072] For example, a standard checkerboard calibration board is arranged on the site, and Z groups of images are taken at multiple orientations and postures, and the world coordinates of the checkerboard corresponding to the camera pixel points are recorded.

[0073] By minimizing the reprojection error:

[0074] , Formula (4)

[0075] Where, Represent the effective focal length in the horizontal and vertical directions respectively, Indicates the horizontal and vertical coordinates of the principal point in the image coordinate system, and Respectively represent The rotation matrix and translation vector (external parameter) corresponding to the group calibration image, represents the three-dimensional point projection function, Indicates the On the group image, The pixel coordinates corresponding to the corner points of the chessboard, Indicates the total number of corner points detected on each set of calibration plates, Indicates the number of calibration image groups;

[0076] Solve the camera intrinsic parameter matrix And distortion parameters, external parameters of each group .

[0077] After all cameras are calibrated simultaneously, the baseline is obtained by measuring the difference in their translation vectors in the same world coordinate system. , and record the baseline length .

[0078] Where, and denote the translation vectors of the first camera and the second camera respectively, represents the 3D baseline vector between the two camera centers, Indicates the Euclidean length of the baseline.

[0079] In some embodiments, the calibration plate is photographed from multiple different angles, such as looking down, looking up, and looking from the side, to improve the calibration accuracy under extreme viewing angles. Construct the reprojection error sum of squares objective function and iteratively solve the camera intrinsic parameter matrix And the external parameters of each frame , which integrates radial / tangential distortion models and can maintain distortion correction accuracy even under small-scale non-ideal chessboard and micro-displacement conditions. In addition, the extrinsic translation vectors of the two cameras are recorded in the same world coordinate system. and The difference, that is, the baseline vector . Use high-precision ruler or laser distance measurement to measure its length Perform secondary confirmation to ensure scale accuracy during stereo matching.

[0080] In step S630, a calibration target containing multiple quadrants of fixed polarization states is projected, the camera's response to different polarization states is measured, a polarization mapping matrix is ​​established, and its inverse matrix is ​​used to perform polarization mismatch compensation on subsequent captured images to eliminate the influence of polarization device errors on differential fusion.

[0081] In some embodiments, a projection unit projects a known polarization-encoded calibration target (e.g., a 4×4 quadrant with fixed polarization states of 0°, 45°, 90°, and 135° in each quadrant). For example, a four- or eight-quadrant target plate is designed, with different quadrants having known polarization states (0°, 45°, 90°, and 135°), and the target surface's ideal intensity is recorded. Each camera acquires the calibration target image by continuously switching between polarizers at 0° and 90°, sampling multiple pixels within each quadrant to measure the image intensity.

[0082] , Formula (5)

[0083] Where, is the camera filter angle, To calibrate the target polarization state, is the set of pixels in this quadrant. Indicates the camera's polarizing filter angle (0° / 90°) and the polarization state of the calibration target (such as 0°, 45°, 90°, 135°), the corresponding quadrant pixel set The average gray intensity of Indicates that the camera filter is placed at an angle Time in pixels The gray value of Indicates the calibration target A set of pixels in a quadrant, .

[0084] Assuming that the camera's response to different polarizations satisfies a linear combination, the polarization mapping matrix is ​​established

[0085] , Formula (6)

[0086] Where, The target is in polarization state The ideal projection intensity under Solve linear equations, estimate , which is the camera polarization mismatch matrix. and For ideal 0° and 90° polarization components mapped to the camera 0° response, and is the crosstalk of the ideal 90° polarization component in the camera's 0° and 90° responses.

[0087] Here, a linear mapping model is established by collecting multiple sets of polarization states and camera channel responses, and the polarization mismatch matrix is ​​solved using the least squares method.

[0088] Capture any image When , first correct the pixel intensity according to the inverse mapping:

[0089] , Formula (7)

[0090] Where, Uncorrected dual polarization state original image grayscale value, Represents the grayscale value of the corrected dual-polarization state original image.

[0091] After correction Perform subsequent differential fusion to ensure that the polarization difference reflects only the diffuse reflection of the scene. In subsequent runs, the dual polarization images collected for each frame are application Decoupling is performed to restore the ideal polarization components and then perform differentiation. This allows for rapid correction using the same mismatch matrix, even when replacing filters or cameras, and completes online self-calibration.

[0092] Figure 7 An operational flowchart of an example of performing polarization differential fusion on multi-view images according to an embodiment of the present application is shown.

[0093] like Figure 7 As shown, in step S710, for each perspective image in the multi-perspective image, the first polarization image and the second polarization image when the corresponding polarization filter is set to 0° and 90° are respectively obtained, and a differential perspective image is generated according to the normalized value of each pixel of the first polarization image and the second polarization image.

[0094] It should be noted that, unlike diffuse reflection, specular highlights exhibit significant intensity differences under specific polarization directions (primarily related to the direction of the outer polarization code projection). By simultaneously utilizing images through 0° and 90° polarization filters, a "difference / sum" analysis is performed on the two polarization state responses of the same pixel. The difference between the two polarization images is then compared to their sum. This not only eliminates absolute variations in ambient light intensity but also amplifies the polarization difference signal.

[0095] Specifically, for each camera With each frame , calculate the normalized polarization difference at the pixel level:

[0096] , formula (8)

[0097] Where, Indicates the After the frame structured light stripes are projected, they are located at the camera When the polarization filter in front is set to 0°, the image collected is at the pixel Gray intensity value at ; Indicates the After the frame structured light stripes are projected, they are located at the camera When the polarization filter in front is placed at 90°, the image collected is at the pixel Gray intensity value at ; Indicates the Frame, Camera In pixels The normalized polarization difference value at is used to measure the relative polarization difference of the pixel; A preset positive constant.

[0098] And take the absolute value to get the differential perspective image:

[0099] , formula (9)

[0100] Where, Indicates the Frame, Camera The differential perspective image is in pixels The diffuse intensity value at .

[0101] It should be understood that both positive and negative values ​​represent polarization differences, and the polarization signal strength is uniformly measured by the absolute value.

[0102] In some embodiments, after each structured light grayscale stripe projection is completed, the camera collects two images after waiting for a short stabilization delay (usually 5ms): a "first polarization image" with the filter at 0° and a "second polarization image" at 90°. This delay ensures that the stripe projection, encoder wheel switching, and stripe optical stability are completed, reducing image blur caused by mechanical vibration.

[0103] Furthermore, after completing the multi-frame projection of grayscale structured light, the system immediately acquires the final polarization difference image from the two cameras. This difference image corresponds to the grayscale contrast when the camera filters are positioned at 0° and 90°, respectively. By calculating the normalized residual ratio, the system avoids the influence of absolute light intensity variations. This ensures that the acquisition covers not only the illuminated surface of the tube, but also the shadows within the cage and the overlapping areas of adjacent tubes, fully reflecting the mirror residual under the system's current projection and imaging conditions.

[0104] In step S720 , the differential perspective images are fused to generate a polarization differential fused image.

[0105] To offset local data loss caused by occlusion or shadows in a single view, pixel-level weighting is introduced. It should be noted that simple averaging of multi-view fusion will treat poor-quality pixels, such as shadows, occlusions, or overexposed areas, as equal, affecting overall image quality. Therefore, brightness weighting is introduced to measure the signal-to-noise ratio at different camera pixels.

[0106] , formula (10)

[0107] Where, Indicates the Frame, Camera In pixels The brightness weight at is obtained by averaging the grayscale values ​​of the two polarization states and is used to measure the signal-to-noise ratio of the pixel.

[0108] More specifically, higher brightness generally results in a better signal-to-noise ratio, and for overexposed (grayscale close to maximum) or underexposed (close to zero) areas, the weights will approach the maximum or minimum values, respectively, thereby reducing the noise contribution.

[0109] The differential view images of different cameras are fused by pixel weighting to obtain Polarization difference fusion image corresponding to the frame structured light stripes:

[0110] , formula (11)

[0111] Where, Indicates the total number of cameras, Indicates the The polarization difference fusion image corresponding to the frame structured light stripes is in the pixel The final diffuse reflection intensity at is obtained by multi-view weighted averaging, taking into account the signal-to-noise ratio and visible area of ​​each view.

[0112] Here, the differential view images from multiple perspectives are weighted and averaged on a pixel-by-pixel basis, ensuring that the visible areas of each perspective complement each other. The resulting diffuse reflectance image is free of blind spots. Furthermore, by weighting the data from each perspective to determine which is more reliable, and dynamically switching the dominant contributor pixel by pixel, random noise can be effectively reduced.

[0113] Figure 8 An operational flowchart of an example of reconstructing a three-dimensional point cloud of a tube according to an embodiment of the present application is shown.

[0114] like Figure 8 As shown, in step S810, the polarization difference fusion image is subjected to continuous frame analysis through the spatiotemporal attention network, and a binary segmentation mask corresponding to each hot-dip galvanized steel pipe body is output to determine the corresponding pipe body area.

[0115] To accurately locate hot-dip galvanized steel pipes against complex backgrounds, the system employs a deep neural network based on a spatiotemporal attention mechanism to analyze polarization-differential fusion images of consecutive frames. Polarization images are resistant to strong reflections, but traditional edge- or grayscale-gradient-based segmentation methods struggle to reliably capture clear, coherent pipe outlines in situations where multiple pipes overlap, background interference is strong, and fringe projections vary significantly. Therefore, by introducing information fusion in the temporal dimension, the system focuses on significant textures and boundaries in the current image in the spatial domain, while capturing the dynamic patterns of fringe evolution and deformation between consecutive frames in the temporal domain.

[0116] For example, the spatial attention module enhances the high-responsiveness of the steel pipe area, while the temporal attention module enhances motion consistency judgment from inter-frame information, effectively reducing segmentation misjudgments caused by illumination changes, streak reflection noise, or local occlusion. Ultimately, the system generates a separate binary segmentation mask for each hot-dip galvanized steel pipe, marking the set of pixels corresponding to the pipe area and accurately locating the pixel regions within the pipe area.

[0117] Figure 9 A structural connection diagram of an example of a spatiotemporal attention network according to an embodiment of the present application is shown.

[0118] like Figure 9 As shown, the spatiotemporal attention network 900 includes a temporal feature extraction branch 910, a spatial polarization fusion branch 920 and a cross-channel self-attention fusion module 930.

[0119] The temporal feature extraction branch 910 is used to perform temporal convolution on the input polarization differential fusion image using continuous 3D convolution to capture the changing pattern of the fringe code in the time domain.

[0120] It should be noted that the input polarization differential fusion image can also be normalized beforehand. For example, mean-variance normalization is performed on the polarization differential fusion image, mapping the grayscale values ​​originally concentrated in [0, 1] or [0, 255] to an approximate standard normal distribution, so that the input of each pixel has zero mean and unit variance in each channel. This offsets the brightness distribution drift that may occur between different frames and different camera perspectives, and also keeps the activation distribution concentrated in a reasonable range, preventing gradient vanishing or explosion.

[0121] Here, structured light coding presents the dynamic switching of grayscale stripes in the temporal dimension. Utilizing 3D convolution, it not only captures the stripe texture and polarization characteristics spatially, but also detects the temporal patterns of the stripes from one frame to the next. Firstly, stripes of varying grayscale levels and polarization states are projected sequentially, and 3D convolution can identify this temporal pattern, providing additional dynamic and static contrast for tube segmentation and capturing the encoding sequence. Secondly, cross-frame filtering mitigates occasional artifacts within a single frame, enhancing the temporal coherence of the stripes under curved projection distortion and, ultimately, the coherence of temporal features.

[0122] The spatial polarization fusion branch 920 uses depthwise separable convolution to extract spatial features from each frame of the polarization difference fusion image, preserving geometric edges and polarization characteristics.

[0123] Here, for each image frame, geometric edges and polarization features are extracted in the spatial domain. Depthwise separable convolution (channel-by-channel convolution + 1x1 convolution) efficiently extracts features from single-channel polarization images. This channel-by-channel convolution learns local spatial neighborhood features pixel by pixel, preserving the texture contrast differences introduced by polarization. 1x1 convolution maps each channel to a high-dimensional feature space, combining the outputs of different spatial filters to aggregate polarization and texture information. This effectively captures polarization texture and improves segmentation capabilities at tube edges and joints.

[0124] After mapping the temporal and spatial features to a unified dimension, the cross-channel self-attention fusion module 930 calculates the self-attention weight matrix and fuses them to output a binary segmentation mask for each tube.

[0125] Here, the features output by the temporal and spatial branches each focus on different dimensions of information, but for the tube segmentation task, the two need to be deeply integrated. In cross-channel self-attention, a query / key / value fusion mechanism is used to achieve interactive perception of temporal and spatial features. The network spontaneously learns which time steps or spatial filters are more conducive to distinguishing tube edges. In addition, the attention weight matrix covers the entire spatial position, allowing each pixel to refer to the full image context, and filling local convolutional gaps by capturing global dependencies. This makes it sensitive to subtle differences between the entire surface of the tube and the surrounding background, improving the segmentation accuracy of small tubes or dense tube groups.

[0126] In step S820, for each tube body region, multiple frames of grayscale profiles are extracted from the tube body region along the stripe direction, and sub-pixel centerline decoding is performed in combination with the relative displacement law of stripes in adjacent frames. The decoded stripe center pixel pairs are input into the polarization-geometry constrained semi-global disparity matching algorithm to calculate the corresponding optimal disparity. The depth value is calculated in combination with the pre-calibrated camera focal length and baseline length, and a three-dimensional point cloud of the corresponding tube body region is generated through back projection.

[0127] Here, with the help of the grayscale stripes generated by structured light projection, the stripes on the surface of the tube are analyzed at the sub-pixel level in the camera image, and the center position of the stripes in the grayscale profile is extracted to achieve high-resolution parallax measurement.

[0128] Next, the system extracts grayscale profile information along the stripe direction in sequential frames and tracks it by combining subtle shifts in the stripes between frames. Because stripes may exhibit slight distortion on highly reflective surfaces, the system analyzes the local movement of the stripe center point across consecutive frames, performing sub-pixel decoding and extracting the stripe's central trajectory. This approach overcomes the limitations of pixel-level precision, improving image measurement accuracy to sub-pixel levels and significantly enhancing the accuracy of the final depth data.

[0129] The system then feeds the decoded stripe center pixel pairs into a semi-global disparity matching algorithm (SGM). This algorithm combines left and right image matching while incorporating polarization-geometry constraints. This strategy uses polarized light reflection characteristics to adjust the confidence level of disparity in reflective areas during disparity optimization. By adding reflection sensitivity screening to the multipath cost aggregation, the system further improves the stability and effectiveness of matching.

[0130] After matching is complete, the system converts the calculated parallax information into depth values ​​using the calibrated camera intrinsic parameters (including focal length) and the precise baseline length between the left and right cameras. Finally, the depth map is combined with the camera projection model to perform back-projection of the spatial points, generating high-precision 3D point cloud data for each steel pipe. This point cloud exhibits spatial continuity and detailed restoration capabilities, faithfully reflecting the actual 3D topography of the steel pipe surface.

[0131] This enables high-resolution, high-fidelity spatial reconstruction, significantly improving the system's ability to restore and control errors for curved structures (such as cylindrical steel pipes) while maintaining measurement accuracy. This is particularly true for metal materials with high reflectivity and surface interference, effectively avoiding the noise accumulation and blurred outlines common in traditional matching algorithms, providing a strong data foundation for subsequent spatial posture analysis and verticality calculations.

[0132] Regarding the specific implementation details of sub-pixel centerline decoding, it should be noted that relying on a single set of stripe profiles within a single frame is susceptible to noise and occlusion. However, the tube stripes only undergo small pixel-level movement in consecutive frames. Leveraging this temporal coherence significantly improves profile quality. Profiles are taken only within the segmentation mask to avoid background clutter. Furthermore, profiles taken perpendicular to the stripe direction yield clearer grayscale peaks. Therefore, the stripe tilt angle in pixel space is estimated, and then uniform sampling is performed along the perpendicular direction.

[0133] In some examples of the embodiments of the present application, this can be achieved by:

[0134] Define two mutually orthogonal unit vectors:

[0135] , formula (12)

[0136] Where, Indicates the The slope of the frame structured light stripe relative to the horizontal axis of the pixel, For the The stripes of the frame move towards the unit vector, For the The unit vector of the frame's profile sampling direction.

[0137] Here, an orthogonal vector pair is constructed by using the known stripe tilt angle to ensure that the profile is strictly perpendicular to the stripe direction, thereby obtaining a profile section with the minimum width and the maximum grayscale gradient.

[0138] More specifically, the initial center pixel coordinates of each stripe are roughly estimated based on the projected stripe index map inside the tube body area. The cross-section is extracted inside the segmentation mask of each tube to avoid interference from background or adjacent tubes and ensure that the cross-section samples only come from the same cylindrical tube surface.

[0139] The initial center pixel coordinates of each stripe As the center, along the cross-section direction Discretely sample multiple profile frames:

[0140] , Formula (13)

[0141] Where, is the half width of the section; is the profile sampling point index, and its value range is .

[0142] It should be noted that within sub-pixel accuracy, millimeter-level 3D accuracy corresponds to a pixel-level decoding accuracy of less than 0.1px, and the corresponding angular deviation can be controlled to ±0.05°. The grayscale profile approximates a Gaussian peak, and quadratic curve fitting is sufficient to approximate the peak. Incorporating gradient weights further suppresses noise.

[0143] Specifically, in the continuous In each frame, the grayscale is sampled at the same section center and direction.

[0144] It should be noted that under ideal conditions, the grayscale profile has a nearly bell-shaped or parabolic distribution, but noise and speckle can distort the peak shape. Using gradient-weighted quadratic least squares, we can automatically focus on the peak region with the largest grayscale variation within a weighted framework, suppressing the influence of flat or noisy areas on the fit and achieving sub-pixel localization of the fringe center.

[0145] , formula (14)

[0146] Where, Indicates the Frame No. The grayscale value of each profile sampling point, Indicates the Grayscale intensity value of the pixel in the frame polarization difference fusion image.

[0147] right The corresponding gray profile sequence can be obtained from the frame .

[0148] In the highly reflective and densely packed hot-dip galvanized steel pipe cage scene, it is difficult to obtain a continuous and high signal-to-noise ratio stripe grayscale curve by relying on a single frame grayscale profile, which is often affected by local noise, occlusion, and unevenness of the pipe body. The profile is extracted at the same stripe position in the frame, and the stability of the profile center is ensured by time alignment or cross-validation.

[0149] The weight of each profile sampling point is defined as the absolute value of the grayscale value difference between two adjacent points:

[0150] , formula (15)

[0151] Where, Indicates the The weighting coefficient of each profile sampling point is used to emphasize the fitting weight at the grayscale peak.

[0152] By taking advantage of the fact that stripes only shift a few pixels in adjacent frames and extracting and aligning the cross-sections of multiple frames, the overall curve smoothness can be improved and occasional noise can be suppressed.

[0153] Constructing a weighted quadratic model

[0154] , formula (16)

[0155] Where, They represent the quadratic term coefficient, linear term coefficient and constant term of the quadratic fitting model respectively;

[0156] Solving for parameters by weighted least squares :

[0157] , formula (17)

[0158] Solve for the vertex position of the quadratic curve:

[0159] , formula (18)

[0160] Where, Represents the continuous offset of the quadratic curve vertex in the section coordinate system.

[0161] Here, the absolute value of the gradient of each sample point is introduced as a weight in the quadratic fitting target, which significantly enhances the fitting weight at the fringe turning point (i.e., the central peak), so that the fitting fidelity and positioning accuracy are both taken into account. In addition, the coefficients are obtained by solving the normal equations Finally, the vertices are calculated directly, avoiding searching or iteration, improving efficiency and ensuring the stability of the solution.

[0162] Calculate the pixel coordinates of the center of each stripe in pixel space:

[0163] , formula (19)

[0164] Where, Indicates the sub-pixel coordinates of the stripe center.

[0165] Here, the sub-pixel centerline coordinates It is naturally aligned with the original segmentation mask and polarization fusion map, and can be directly used as the internal parameter for subsequent semi-global matching and depth back projection without additional conversion, thereby improving decoding efficiency.

[0166] Regarding the details of calculating the optimal disparity, it should be noted that in traditional stereo matching, grayscale consistency alone is often difficult to deal with highly reflective metal surfaces because specular highlights destroy the photometric assumption. In the embodiments of the present application, a comprehensive cost is adopted that simultaneously incorporates photometric consistency, polarization consistency, and disparity priors. Photometric consistency is used to measure the difference in the grayscale images of the same physical point on the left and right cameras. Based on polarization consistency, the diffuse reflection intensity obtained by suppressing the specular component using the polarization difference fusion image is used to compare the polarization response difference between the two perspectives. Based on the disparity prior, the geometric prior of the tube body (tube diameter, installation height) is combined to provide initial guidance within the search space. Therefore, by adopting the three-item hybrid cost of "photometric + polarization + prior", a robust metric with the ability to distinguish between highly reflective and low-texture areas is constructed.

[0167] In some examples of the embodiments of the present application, for each stripe center pixel pair obtained by sub-pixel decoding, The comprehensive matching cost is defined on the frame:

[0168] , formula (20)

[0169] , formula (21)

[0170] , formula (22)

[0171] , Formula (23)

[0172] Where, and is the fringe center pixel pair, represents the coordinates of the center point of the stripe obtained after sub-pixel decoding of the first camera; is a candidate disparity value, which represents the difference in pixel positions between the center points of the stripes of the first camera and the second camera on the same scan line. Indicates in On the frame, for the first camera point and The comprehensive matching cost defined by the corresponding second camera point; 、 and Represent the photometric consistency cost term, polarization consistency cost term and parallax prior cost term respectively, Respectively represent the corresponding weight coefficients; Indicates the The original grayscale image of the first camera of the frame is at point The gray value of Indicates the The original grayscale image of the second camera of the frame is at point Gray value of Indicates the The first camera polarization difference fusion map of the frame is at point The intensity value of Indicates the The second camera polarization difference fusion map of the frame is at point The intensity value of Indicates the The frame is based on the initial disparity value obtained by prior estimation of the tube geometry.

[0173] In the photometric consistency cost, the grayscale images of different cameras are usually integer coordinates, and bilinear interpolation is required at sub-pixel locations to ensure the continuity of the photometric difference, which significantly improves the continuity and matching stability of stripes. The polarization consistency cost reflects the diffuse reflection intensity map of each pixel. The prior parallax cost is used to infer the depth range based on the known tube radius and tube height range. In areas where tubes intersect or are partially occluded, the temporal continuity and geometric window constraints brought by the parallax prior effectively avoid the generation of invalid or jumpy parallax points.

[0174] For each aggregation direction, calculate the corresponding cumulative path cost:

[0175] , Formula (24)

[0176] , formula (25)

[0177] , formula (26)

[0178] , formula (27)

[0179] Where, , Indicates the Pixel offset vector of the strip aggregation direction, 8 directions are enumerated in semi-global matching; Indicates along direction In pixels Candidate disparity is accumulated at The minimum total cost; 、 and They represent the same parallax penalty term, small parallax jump penalty term and large parallax jump penalty term respectively. and They respectively represent the preset small-scale parallax jump penalty coefficient and the preset large-scale parallax jump penalty coefficient.

[0180] In some embodiments, eight propagation directions, namely horizontal, vertical, and two diagonal directions, are selected, and the cumulative cost of each direction can be executed in parallel on the GPU, requiring only lightweight minimum and addition operations on adjacent pixels and disparity planes, resulting in high processing efficiency.

[0181] Sum the cumulative costs of all aggregation directions and take the candidate disparity with the minimum total cost as the optimal disparity:

[0182] , formula (28)

[0183] Where, represents the geometric prior range of disparity search, represents the optimal disparity that minimizes the total cost.

[0184] In some embodiments, the geometric prior range can be pre-set, such as a window calculated based on camera intrinsic parameters, baseline length, and tube depth prior, to limit the parallax search range, significantly reduce the amount of calculation, and avoid matching results that are far away from the true depth of the object.

[0185] It should be noted that the classic SGM balances global consistency and local flexibility by performing dynamic programming in multiple directions. In this embodiment, three types of path smoothing are considered in each direction: Isoparallax continuation allows pixels to be consistent with their neighbors at the same disparity; small-scale jumps allow ±1px disparity variations with a small penalty to accommodate object edges and parallax fine-tuning; and large-scale jumps impose heavier penalties for larger parallax jumps to suppress mismatches and jump distortion. Thus, after multi-directional aggregation, the combined total cost guides the system to find a disparity distribution across the entire image that balances edge preservation and global smoothness.

[0186] Regarding the process of calculating depth and reconstructing a three-dimensional point cloud based on parallax, camera focal length, and baseline length, reference may be made to the implementation methods in current related technologies and no limitation is imposed here.

[0187] In some embodiments, before depth calculation, the center pixel of the stripe is corrected from the distorted image coordinates to the ideal pinhole model coordinates to eliminate the impact of radial and tangential lens distortion on back-projection accuracy. Then, utilizing the inverse relationship between parallax and depth in stereo geometry, the sub-pixel parallax is directly converted to physical depth.

[0188] For example, let the corrected pixel coordinates be , and the final parallax is , the camera's equivalent focal length is , the principal point coordinates are , the baseline length is . Then the 3D point corresponding to the pixel is Can be written uniformly:

[0189] , formula (29)

[0190] Since the point cloud only comes from reliable tube segmentation areas and high-confidence parallax, no interference points are introduced through the above point cloud mapping, and the generated 3D point cloud is dense and continuous, directly supporting subsequent tube cylindrical fitting and tilt angle calculation without additional filling or filtering.

[0191] Regarding the fitting details of the tube body axial vector, in some examples of the embodiments of the present application, it can be achieved by using a two-stage RANSAC (Random Sample Consensus) algorithm to fit the tube body axis of each tube body three-dimensional point cloud.

[0192] It should be noted that when measuring the verticality of hot-dip galvanized steel pipe cages, the spatial axis of each pipe reflects its installation tilt. Point clouds reconstructed using multi-view structured light often contain numerous outliers, such as measurement noise, residual occlusions, pipe end openings, and pipe connections. Accurately determining the pipe axis requires reliably distinguishing between inliers that conform to the cylindrical surface model and various anomalous outliers, and then accurately fitting a straight line based on these inliers.

[0193] More specifically, in the first stage of processing, a minimum sample set is randomly extracted from the 3D point cloud of a single pipe body, a least squares straight line fitting is performed, and the outliers with the highest proportion are eliminated to obtain a preliminary axial model.

[0194] It should be noted that traditional RANSAC classifies inliers and outliers based on a residual threshold, requiring a prior estimate of the outlier rate. In ground cage point clouds, the proportion of outliers often exceeds 30% due to interference from occlusions, fasteners, and other factors, and their spatial distribution is uneven. Here, a fixed rejection ratio strategy is employed: after fitting a line between two random points, the residuals of the perpendicular distances from all points to the fitted line are calculated. A fixed q% of the residuals with the largest residuals are considered outliers. This residual sorting and rejection mechanism eliminates the need for a pre-set threshold and can adapt to varying outlier densities.

[0195] In the second stage of processing, weights are assigned to the remaining internal points based on normal consistency, and weighted least squares optimization is performed to obtain the corresponding tube axial vector.

[0196] Here, the normal direction of each point on the cylindrical surface should be orthogonal to the axial direction. The axial direction given by the preliminary model can be used to evaluate the consistency of the normal direction of each point with the model and use this weighting to highlight points that meet cylindrical characteristics and suppress non-cylindrical areas such as pipe ends or connectors.

[0197] It should be pointed out that traditional symmetric least squares straight-line fitting treats all interior points equally. In the embodiment of the present application, weighted least squares is used to give higher consistency points a greater influence, that is, the influence of interior points that do not conform to the cylindrical characteristics is automatically weakened through weights, resulting in a more robust and accurate axial estimate.

[0198] It should be noted that, for the aforementioned method or system embodiments, for the sake of simplicity of description, they are all expressed as a series of combined actions, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application. In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0199] The system embodiment described above is merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units. That is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.

[0200] Through the description of the above embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the relevant technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for enabling a computer device (such as a personal computer, server, or network device) to execute the method or system described in each embodiment or certain portions of the embodiment.

[0201] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A hot-dip galvanized steel pipe cage verticality measurement system based on machine vision, characterized in that: The system comprises: A projection unit, used for projecting polarization grayscale stripes with polarization coding onto the hot-dip galvanized steel pipe cage to be measured; An imaging unit, comprising at least two cameras forming a triangulation baseline with the projection unit, wherein the cameras are respectively arranged at different positions around the hot-dip galvanized steel pipe cage to synchronously capture multi-view images after projecting the polarized grayscale stripes from different viewing angles; a processing unit electrically connected to the projection unit and the imaging unit, and configured to: Performing polarization differential fusion on the multi-view images to suppress specular reflection, and generating corresponding polarization differential fusion images; Reconstructing a three-dimensional point cloud of each hot-dip galvanized steel pipe in the hot-dip galvanized steel pipe cage according to the polarization difference fusion image; Robust cylindrical axis fitting is adopted for each of the three-dimensional point clouds of the tube body to generate an axial vector of the tube body, and the verticality deviation angle of the tube body is calculated according to the axial vector of the tube body and the direction of gravity.

2. The system according to claim 1, wherein: The processing unit is further configured to perform a bias self-calibration operation before measurement, including: Turning off the fringe projection of the projection unit, controlling each of the cameras to capture multiple frames of ambient light images at polarization directions of 0° and 90°, calculating the mean of the absolute values ​​of polarization differences to estimate the scene reflectivity, and adaptively adjusting the initial angle of the polarization wheel of the projection unit based on the reflectivity range; The camera's intrinsic matrix, distortion parameters, and extrinsic parameters are calibrated using a multi-pose checkerboard calibration plate image and a reprojection error minimization method. The triangular baseline length is calculated using the camera's translation vector difference. A calibration target containing multiple quadrants of fixed polarization states is projected, the camera's response to different polarization states is measured, and a polarization mapping matrix is ​​established. Its inverse matrix is ​​used to compensate for polarization mismatch in subsequent acquired images to eliminate the impact of polarization device errors on differential fusion.

3. The system according to claim 1 or 2, characterized in that The projection unit includes: An inner polarization encoding wheel for sequentially projecting a plurality of grayscale stripes, each stripe having a corresponding polarization angle; wherein the inner polarization encoding wheel is configured with a plurality of grayscale stripe segments arranged at equal intervals, the grayscale level of each grayscale stripe segment varying by a predetermined increment between adjacent segments and corresponding to a fixed polarization angle to achieve primary polarization encoding of the stripes; The outer rotating polarization wheel is used to continuously rotate around the optical axis and dynamically change the global polarization direction of all projected light beams by the rotation angle; Wherein, the processing unit is used to determine the outer polarization angle according to the current ambient light reflectivity estimation value; The projection unit is used to drive the outer rotating polarization wheel to rotate to the outer polarization angle, and control the inner polarization encoding wheel to switch the grayscale stripe segments to sequentially project multiple grayscale stripes, thereby realizing the projection of multi-frame polarization grayscale stripes.

4. The system according to claim 1, wherein: The performing polarization differential fusion on the multi-view images to suppress specular reflection and generating corresponding polarization differential fusion images includes: For each perspective image in the multi-perspective images, respectively obtaining a first polarization image and a second polarization image when the corresponding polarization filter is set at 0° and 90°, and generating a differential perspective image according to the normalized value of each pixel of the first polarization image and the second polarization image, including: For each camera With each frame , calculate the normalized polarization difference at the pixel level: , Where, Indicates the After the frame structured light stripes are projected, the camera When the polarization filter in front is set to 0°, the image collected is at the pixel Gray intensity value at ; Indicates the After the frame structured light stripes are projected, the camera When the polarization filter in front is placed at 90°, the image collected is at the pixel Gray intensity value at ; Indicates the Frame, Camera In pixels The normalized polarization difference value at is used to measure the relative polarization difference of the pixel; is a preset positive constant; And take the absolute value to get the differential perspective image: , Where, Indicates the Frame, Camera The differential perspective image is in pixels The diffuse reflection intensity value at ; The differential perspective images are fused to generate a polarization differential fused image, comprising: In order to offset the local data loss caused by occlusion or shadow in a single view, pixel-level weights are introduced: , Where, Indicates the Frame, Camera In pixels The brightness weight at is obtained by averaging the grayscale values ​​of the two polarization states and is used to measure the signal-to-noise ratio of the pixel. The differential view images of different cameras are fused by pixel weighting to obtain Polarization difference fusion image corresponding to the frame structured light stripes: , Where, Indicates the total number of cameras, Indicates the The polarization difference fusion image corresponding to the frame structured light stripes is in the pixel The final diffuse reflection intensity at is obtained by multi-view weighted averaging, taking into account the signal-to-noise ratio and visible area of ​​each view.

5. The system according to claim 1, wherein The step of reconstructing the three-dimensional point cloud of each hot-dip galvanized steel pipe in the hot-dip galvanized steel pipe cage according to the polarization difference fusion image includes: Performing continuous frame analysis on the polarization difference fusion image through a spatiotemporal attention network, and outputting a binary segmentation mask corresponding to each hot-dip galvanized steel pipe body to determine the corresponding pipe body area; For each of the tube body regions, multiple frames of grayscale profiles are extracted from the tube body region along the stripe direction, and sub-pixel centerline decoding is performed in combination with the relative displacement law of stripes in adjacent frames. The decoded stripe center pixel pairs are input into a polarization-geometry constrained semi-global disparity matching algorithm to calculate the corresponding optimal disparity. The depth value is calculated in combination with the pre-calibrated camera focal length and baseline length, and a three-dimensional point cloud of the corresponding tube body region is generated through back projection.

6. The system according to claim 5, wherein: The method of extracting multiple frames of grayscale profiles from the tube body region along the stripe direction and performing sub-pixel centerline decoding based on the relative displacement pattern of stripes in adjacent frames includes: Define two mutually orthogonal unit vectors: , Where, Indicates the The slope of the frame structured light stripe relative to the horizontal axis of the pixel, For the The stripes of the frame move towards the unit vector, For the The unit vector of the frame's profile sampling direction; Inside the tube area, the initial center pixel coordinates of each stripe are roughly estimated based on the projected stripe index map ; The initial center pixel coordinates of each stripe As the center, along the cross-section direction Discretely sample multiple profile frames: , Where, is the half width of the section; is the profile sampling point index, and its value range is ; Continuous extraction in the cross section In each frame, the grayscale is sampled at the same section center and direction: , Where, Indicates the Frame No. The grayscale value of each profile sampling point, Indicates the Gray intensity value of the pixel of the frame polarization difference fusion image; right The corresponding gray profile sequence can be obtained from the frame ; The weight of each profile sampling point is defined as the absolute value of the grayscale value difference between two adjacent points: , Where, Indicates the The weighting coefficient of each profile sampling point is used to emphasize the fitting weight at the grayscale peak; Constructing a weighted quadratic model , Where, They represent the quadratic term coefficient, linear term coefficient and constant term of the quadratic fitting model respectively; Solving for parameters by weighted least squares : , Solve for the position of the quadratic curve vertex: , Where, Represents the continuous offset of the quadratic curve vertex in the section coordinate system; Calculate the pixel coordinates of the center of each stripe in pixel space: , Where, Indicates the sub-pixel coordinates of the stripe center.

7. The system according to claim 5, wherein: The decoded stripe center pixel pairs are input into the polarization-geometry constrained semi-global disparity matching algorithm to calculate the corresponding optimal disparity: For each stripe center pixel pair obtained by sub-pixel decoding, The comprehensive matching cost is defined on the frame: , , , , Where, and is the fringe center pixel pair, represents the coordinates of the center point of the stripe obtained after sub-pixel decoding of the first camera; is the candidate disparity value, which represents the difference in pixel position between the center points of the stripes of the first camera and the second camera on the same scan line; Indicates in On the frame, for the first camera point and The comprehensive matching cost defined by the corresponding second camera point; 、 and Represent the photometric consistency cost term, polarization consistency cost term and parallax prior cost term respectively, Respectively represent the corresponding weight coefficients; Indicates the The original grayscale image of the first camera of the frame is at point The gray value of Indicates the The original grayscale image of the second camera of the frame is at point Gray value of Indicates the The first camera polarization difference fusion map of the frame is at point The intensity value of Indicates the The second camera polarization difference fusion map of the frame is at point The intensity value of Indicates the The frame is based on the initial disparity value obtained by prior estimation of the tube geometry; For each aggregation direction, calculate the corresponding cumulative path cost: , , , , Where, , Indicates the Pixel offset vector of the strip aggregation direction, 8 directions are enumerated in semi-global matching; Indicates along direction In pixels Candidate disparity is accumulated at The minimum total cost; 、 and They represent the same parallax penalty term, small parallax jump penalty term and large parallax jump penalty term respectively. and represent the preset small-scale parallax jump penalty coefficient and the preset large-scale parallax jump penalty coefficient respectively; Sum the cumulative costs of all aggregation directions and take the candidate disparity with the minimum total cost as the optimal disparity: , Where, represents the geometric prior range of disparity search, represents the optimal disparity that minimizes the total cost.

8. The system according to claim 5, wherein: The spatiotemporal attention network includes: The temporal feature extraction branch is used to perform temporal convolution on the input polarization difference fusion image using continuous 3D convolution to capture the changing pattern of fringe coding in the time domain; The spatial polarization fusion branch uses depthwise separable convolution to extract spatial features from each frame of polarization difference fusion image, preserving geometric edges and polarization characteristics; The cross-channel self-attention fusion module maps the temporal and spatial features to a unified dimension, calculates the self-attention weight matrix and fuses them to output the binary segmentation mask of each tube.

9. The system according to claim 1, wherein: The step of using robust cylindrical axis fitting on each of the three-dimensional point clouds of the tube body to generate the tube body axial vector comprises: The two-stage RANSAC algorithm is used to fit the tube axis of each tube three-dimensional point cloud: In the first stage of processing, a minimum sample set is randomly extracted from the 3D point cloud of a single tube, a least squares straight line fit is performed, and the outliers with the highest proportion are eliminated to obtain a preliminary axial model; In the second stage of processing, weights are assigned to the remaining internal points based on normal consistency, and weighted least squares optimization is performed to obtain the corresponding tube axial vector.

Citation Information

Patent Citations

  • Binocular stereo infrared salient target detection method and system

    CN108460794A

  • High reflective surface three-dimensional reconstruction method and device based on polarization structured light camera

    CN115876124A