Workpiece three-dimensional reconstruction and size measurement method based on multi-view polarization binocular imaging
By combining multi-view polarization binocular imaging and neural implicit representation with polarization information and normal Gaussian modeling, the problems of local geometric deviation and reconstruction distortion in non-contact measurement of workpieces with high reflectivity and weak texture are solved, and high-precision 3D reconstruction and dimensional measurement are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HOHAI UNIV
- Filing Date
- 2026-01-30
- Publication Date
- 2026-05-12
AI Technical Summary
Existing technologies are prone to local geometric deviations and reconstruction distortions in non-contact measurements of workpieces with high reflectivity and weak texture, making it difficult to achieve accurate 3D reconstruction and dimensional measurement.
A multi-view polarization binocular imaging method is adopted, which combines polarization information and neural implicit representation. An imaging device is constructed by using a polarization binocular camera, a ring light source and a rotating platform to acquire multi-view polarization images. Stereo matching is performed by utilizing polarization azimuth consistency and depth prior. Combined with normal Gaussian modeling and adaptive ray sampling, the three-dimensional surface of the workpiece is restored and geometric primitives are automatically identified to achieve dimensional measurement.
It effectively suppresses the problems of inconsistent light intensity and missing texture caused by specular reflection, improves the reconstruction stability and measurement accuracy of workpieces with high reflectivity and weak texture, and can automatically calculate the key dimensions of workpieces such as thickness, diameter, and hole center distance.
Smart Images

Figure CN122015644A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer vision and 3D reconstruction technology, specifically relating to a method for 3D reconstruction and dimensional measurement of workpieces based on multi-view polarization binocular imaging. Background Technology
[0002] The dimensional and shape inspection of industrial parts typically employs contact or non-contact measurement methods. While contact measurement offers high accuracy, it is prone to causing surface wear on the workpiece and is limited by geometric shape and probe contact, making it unsuitable for complex or fragile workpieces. Non-contact measurement avoids the risks of physical contact and is more widely used in industry. Key methods include machine vision measurement, laser triangulation, structured light measurement, and laser scanning measurement. Among these, machine vision measurement has become a widely adopted non-contact measurement technology in industry due to its low cost, high flexibility, and ease of integration into automated production lines. These methods typically rely on image intensity or color information for two-dimensional edge detection or obtain three-dimensional information through binocular and multi-view stereo matching. However, two-dimensional vision measurement can only obtain the projected features of the workpiece, making it difficult to fully describe complex surface morphology or step heights and other three-dimensional geometric structures. To achieve accurate inspection of the complete three-dimensional structure of the workpiece, it is necessary to introduce three-dimensional reconstruction technology based on multi-view images to recover a high-fidelity geometric model from multiple perspectives.
[0003] Neural implicit representation, as an emerging 3D reconstruction method, can model spatial geometry through continuous functions and has the potential to recover complex geometric structures under multi-view conditions. However, existing neural implicit reconstruction methods mostly rely on the assumption of image photometric consistency and lack explicit modeling of the reflective properties of the workpiece surface. On highly reflective and weakly textured workpieces, such as metals, they are easily affected by specular interference and the loss of effective features, leading to local geometric errors and reconstruction distortion, which in turn affects the accuracy of workpiece size measurement.
[0004] To overcome the challenges of measuring workpieces with high reflectivity and weak texture, polarization information can be introduced to assist in 3D reconstruction. This information reflects the physical characteristics of the interaction between light and the workpiece surface, providing additional surface geometric constraints for 3D reconstruction in areas with high reflectivity and weak texture. This effectively suppresses photometric inconsistencies caused by specular reflection and unreliable matching problems caused by texture loss, thereby improving the stability and measurement accuracy of the reconstruction results. Summary of the Invention
[0005] The purpose of this invention is to provide a method for three-dimensional reconstruction and dimensional measurement of workpieces based on multi-view polarization binocular imaging, which solves the problem that existing technologies are prone to local geometric deviations and reconstruction distortions in non-contact measurement of high-reflectivity, weak-texture workpieces.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions.
[0007] A method for 3D reconstruction and dimensional measurement of a workpiece based on multi-view polarization binocular imaging includes:
[0008] Step 1: Construct a multi-view polarization binocular imaging device, which includes a polarization binocular camera 2, a ring light source 1, and a rotating platform. The binocular camera is calibrated using a calibration plate 6 to obtain intrinsic parameters and the fixed relative pose relationship between the binocular cameras.
[0009] Step 2: Fix the workpiece 5 to be tested on the rotating platform, control the imaging device to take pictures around the workpiece 5 at a preset angle, and continuously acquire multi-view polarized binocular image pairs to obtain an image sequence with high overlap around the workpiece 5.
[0010] Step 3: Input the left camera image sequence into the structure-of-motion algorithm to recover the spatial pose of the left camera in the world coordinate system from each viewpoint, and calculate the pose of the right camera by combining the binocular relative pose, thereby obtaining the binocular camera pose information from all viewpoints.
[0011] Step 4: Under camera pose constraints, extract polarization information and introduce polarization constraints in stereo matching to obtain multi-view depth information with confidence; perform depth back projection based on camera intrinsic parameters and camera pose to form multi-view depth prior.
[0012] Step 5: Based on multi-view depth prior, polarization information and camera pose, construct a polarization neural implicit representation model. Guide adaptive sampling through depth prior, supervise geometric learning by using polarization volume rendering, and combine normal Gaussian modeling constraints to suppress local normal fluctuations. Use the signed distance field to characterize the three-dimensional surface of the workpiece under test.
[0013] Step 6: Extract the zero isosurface mesh from the symbolic distance field to restore the true scale; identify planar, cylindrical, and freeform surface structures under the constraints of normal consistency and neighborhood relationship; realize size measurement based on the geometric primitive fitting results; and automatically calculate key dimensions such as thickness, diameter, hole center distance, and coaxiality.
[0014] In the aforementioned method, the polarization binocular camera 2 is fixed to the rotating platform by a rigid structure, and the relative pose of the polarization binocular camera 2 remains unchanged during the acquisition process; the polarization binocular camera 2 and the ring light source 1 move synchronously around the workpiece 5 to be measured fixed on the rotating platform, and acquire a multi-view polarization binocular image sequence of about 60 viewpoints in a preset angle step of 6° to ensure sufficient overlap between adjacent viewpoints.
[0015] In the aforementioned method, the total matching cost of stereo matching in step four is composed of the photometric matching cost. Cost of consistency with polarization azimuth angle According to the preset weighting coefficient Weighted summation representation:
[0016] ;
[0017] in, and These are the left pixel and the right image candidate matching point, respectively.
[0018] Cost of polarization azimuth consistency The definition is as follows:
[0019] ;
[0020] in, and These are the surface normal azimuth angles corresponding to the pixels in the left and right images, respectively.
[0021] In the aforementioned method, in step four, the confidence level... By matching cost confidence Left and right parallax confidence and local smoothness confidence The weighted composition is used to dynamically adjust the expansion of the sampling interval in subsequent adaptive ray sampling:
[0022] ;
[0023] in, Fixed experience weights;
[0024] The confidence level of the matching cost Described by the prominence of the optimal match in the candidate set. For pixels At optimal parallax The minimum total cost of matching at point , For all candidate parallax within the search space The sum of the corresponding costs is given by the following formula:
[0025] ;
[0026] The confidence level of left and right disparity Determined by the independently calculated difference in left and right disparity. The optimal disparity was found using the left image as a reference. The formula for the optimal disparity obtained by reverse matching based on the right image is as follows:
[0027] ;
[0028] in This is the scaling adjustment parameter.
[0029] In the aforementioned method, step five, the neural implicit representation model incorporates a volume rendering calculation process based on a polarization physics model, wherein the outgoing Stokes vector... Based on the polarization bidirectional reflection distribution function model, the diffuse reflection component... With specular reflection component Linear superposition yields:
[0030] ;
[0031] Diffuse reflection polarization characteristic vector Constructed based on Fresnel's transmission law, its polarization azimuth angle From surface normal Determine the transmission coefficient. and Based on the material's refractive index and the direction of observation With surface normal The included angle was calculated as follows:
[0032] ;
[0033] Specular reflection polarization characteristic vector Constructed based on the microsurface reflection theory, its polarization azimuth angle From half-length vector Determined, where the half-range vector By aligning the incident light direction with the observation direction The sum vector is normalized to obtain the reflection coefficient. and Based on the material's refractive index and half-range vector With the direction of observation The included angle was calculated as follows:
[0034] ;
[0035] The polarization distribution characteristics determined by the physical laws described above need to be combined with the corresponding energy intensity term to characterize the final emission state. Diffuse reflection intensity By neural network According to spatial location Predicted specular reflection intensity By neural network Combined with spatial location Observation direction and surface normals The predicted outgoing Stokes vector is ultimately formed through linear superposition. Based on predicted values Compared with observed values The difference between them can be defined as the polarization uniformity loss. .
[0036] The aforementioned method, in step five, further includes a normal Gaussian modeling method, comprising:
[0037] At a point in space Selecting within the neighborhood sampling points The unit normal vector is calculated based on the gradient of the signed distance field at each sampling point. Construct a three-dimensional Gaussian distribution:
[0038] ;
[0039] in Represents the spatial point The mean vector of the normal line, The corresponding normal covariance matrix is calculated as follows:
[0040] ;
[0041] .
[0042] The aforementioned method, considering that polarization information only constrains the direction of the surface normal in the imaging plane, projects the three-dimensional Gaussian distribution of the normal into a two-dimensional Gaussian distribution, and constructs a normal Gaussian consistency loss based on the two-dimensional Gaussian distribution and polarization information. :
[0043] ;
[0044] in and The two-dimensional Gaussian mean and covariance obtained by projection. The normal projection direction is a priori, constructed from the polarization azimuth angle; This is used to constrain the dispersion of the local normal distribution.
[0045] In the aforementioned method, step five, the depth-guided adaptive sampling refers to the sampling interval of the ray. It is based on depth confidence. Make dynamic adjustments:
[0046] ;
[0047] in, For the sake of deep priors The parameter positions of the 3D points obtained by back projection on the corresponding camera ray; For The half-width of the ray parameter sampling interval centered on the ray; This is the preset maximum value of the interval half-width under scene normalization conditions.
[0048] In the aforementioned method, step six, the dimensional measurement includes: on a three-dimensional model that restores the true scale, based on the normal consistency and spatial neighborhood relationship of surface sampling points, automatically identifying regular structural regions, and fitting geometric primitives of planes, cylinders and freeform surfaces, so as to achieve dimensional measurement of the workpiece 5 to be measured without manual point selection;
[0049] Plane fitting is performed on the planar region to obtain the unit normal vector and displacement parameters characterizing the spatial position and orientation of the plane, and the plane equation is determined accordingly.
[0050] By minimizing the distance error from the sampling point to the cylinder axis, the axial direction, points on the axis, and radius parameters of the cylinder are obtained by fitting the data for the cylindrical region.
[0051] For the freeform surface region, a parametric surface modeling method is used to represent the freeform surface as a continuous mapping function of parametric coordinates, and the surface parameters of the freeform surface are automatically obtained by minimizing the distance error from the sampling point to the parametric surface.
[0052] The aforementioned method, wherein the calculation of the critical dimension includes:
[0053] Thickness calculation: based on displacement parameters of a pair of parallel planes , ,according to calculate;
[0054] Diameter calculation: based on cylinder radius ,according to calculate;
[0055] Hole center distance calculation: based on points on different cylinder axes and ,according to calculate;
[0056] Coaxiality calculation: The direction of the principal axis is ,according to calculate;
[0057] Calculation of surface curve length: Select a parametric curve within the surface parameter domain. The corresponding space curve length is ;
[0058] Surface curvature radius calculation: Based on the principal curvature of the surface, at the parameter points The principal curvature at that point is The corresponding radius of curvature is .
[0059] The advantages of this invention compared to the prior art are as follows:
[0060] This invention provides a method for 3D reconstruction and dimensional measurement of workpieces based on multi-view polarization binocular imaging, solving the problems of local geometric deviations and reconstruction distortions that easily occur in non-contact measurement of highly reflective and weakly textured workpieces in existing technologies. This invention utilizes binocular stereo vision to obtain a depth prior of the true scale and performs adaptive ray sampling based on this prior, concentrating sampling points closer to the surface of the real object, thereby significantly improving the stability and convergence efficiency of implicit geometry in weakly textured and highly reflective regions. This invention combines Stokes vectors calculated from polarization images to establish polarization photometric consistency loss, which can recover normals and local geometric details in specular and specular regions, improving the reconstruction quality of highly reflective and weakly textured workpiece surfaces. This invention introduces Gaussian modeling of normals, imposing constraints on the normals obtained from the gradient of the signed distance function, effectively suppressing normal direction fluctuations caused by reflection noise, ensuring the implicit surface remains continuous and reliable. This invention reconstructs the 3D model and recovers the true scale through a signed distance field, and combines fitting with geometric primitives such as planes, cylinders, and freeform surfaces to directly obtain key dimensions such as aperture, thickness, and distance between opposite edges. Attached Figure Description
[0061] Figure 1 This is a schematic diagram of the workpiece three-dimensional reconstruction and dimensional measurement method of the present invention;
[0062] Figure 2 This is a schematic diagram of the multi-view polarization binocular imaging device constructed according to the present invention;
[0063] Figure 3 A schematic diagram illustrating the construction of neural implicit geometric constraints based on polarization information;
[0064] Figure 4 A schematic diagram of the ray sampling interval based on depth prior;
[0065] Figure 5 A schematic diagram of the surface normal distribution obtained by volume rendering;
[0066] Figure 6 This is an example of the surface effect of a 3D reconstructed workpiece.
[0067] Explanation of reference numerals in the attached diagram: 1-ring light source; 2-polarizing binocular camera; 3-dual-camera gimbal; 4-extension plate bracket; 5-workpiece to be tested; 6-calibration plate. Detailed Implementation
[0068] The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0069] Example 1:
[0070] like Figure 1 As shown, this embodiment provides a method for workpiece 3D reconstruction and dimensional measurement based on multi-view polarization binocular imaging, including:
[0071] Step 1: Construct a multi-view polarization binocular imaging device, which includes a polarization binocular camera 2, a ring light source 1, and a rotating platform. The binocular camera is calibrated using a calibration plate 6 to obtain intrinsic parameters and the fixed relative pose relationship between the binocular cameras.
[0072] Step 2: Fix the workpiece 5 to be tested on the rotating platform, control the imaging device to take pictures around the workpiece 5 at a preset angle, and continuously acquire multi-view polarized binocular image pairs to obtain an image sequence with high overlap around the workpiece 5.
[0073] Step 3: Input the left camera image sequence into the structure-of-motion algorithm to recover the spatial pose of the left camera in the world coordinate system from each viewpoint, and calculate the pose of the right camera by combining the binocular relative pose, thereby obtaining the binocular camera pose information from all viewpoints.
[0074] Step 4: Under camera pose constraints, extract polarization information and introduce polarization constraints in stereo matching to obtain multi-view depth information with confidence; perform depth back projection based on camera intrinsic parameters and camera pose to form multi-view depth prior.
[0075] Step 5: Based on multi-view depth prior, polarization information and camera pose, construct a polarization neural implicit representation model. Guide adaptive sampling through depth prior, supervise geometric learning by using polarization volume rendering, and combine normal Gaussian modeling constraints to suppress local normal fluctuations. Use the signed distance field to characterize the three-dimensional surface of the workpiece under test.
[0076] Step 6: Extract the zero isosurface mesh from the symbolic distance field to restore the true scale; identify planar, cylindrical, and freeform surface structures under the constraints of normal consistency and neighborhood relationship; realize size measurement based on the geometric primitive fitting results; and automatically calculate key dimensions such as thickness, diameter, hole center distance, and coaxiality.
[0077] Example 2:
[0078] This embodiment describes the method of the present invention based on Embodiment 1, with specific steps as follows:
[0079] This invention constructs a multi-view polarization binocular imaging device for acquiring binocular depth prior and multi-view polarization images, the structure of which is as follows: Figure 2As shown, the system consists of polarization binocular cameras 2, a ring light source 1, a dual-camera gimbal 3, an extension plate bracket 4, and a 360° rotating imaging platform. The two polarization binocular cameras 2 are rigidly fixed to the dual-camera gimbal 3, maintaining a constant baseline length and relative orientation throughout the acquisition process. The dual-camera gimbal 3 and the extension plate bracket 4 are integrally mounted on the rotating platform column, allowing the polarization binocular cameras 2 to form multi-view imaging paths around the workpiece 5 under test as the platform rotates. During the imaging process, the workpiece 5 remains stationary, while the rotating platform rotates gradually at a preset angle of 6°, causing the binocular camera assembly to sequentially acquire polarization images of the workpiece 5 from various angles along a circular trajectory, acquiring approximately 60 pairs of polarization binocular images to ensure consistency and sufficient overlap of the polarization image results from different perspectives.
[0080] To improve the illumination stability during polarization image imaging, a ring light source 1 is installed in front of the binocular polarization camera 2. The ring light source 1 is mounted on the same mounting structure as the binocular polarization camera 2 and moves synchronously around the workpiece 5 under test during multi-view acquisition. The emitting surface of the ring light source 1 faces the workpiece area to provide stable and uniform illumination conditions for the workpiece surface, thereby reducing the influence of ambient light and viewing angle changes on polarization image acquisition and ensuring the consistency of polarization image results under different viewing angles.
[0081] Before acquiring multi-view polarized images, the binocular system is calibrated to obtain the imaging parameters required for subsequent depth estimation and camera pose determination. Calibration is performed by acquiring multi-pose calibration board images within the shared field of view of both cameras, yielding the internal parameters of the left and right cameras (including focal length, principal point coordinates, distortion coefficients, etc.) and the fixed relative pose between the two cameras. Since the binocular components have a rigid mounting structure, their relative pose remains unchanged throughout the multi-view acquisition process; therefore, calibration only needs to be performed once and does not need to be recalculated with each viewing angle.
[0082] After calibrating the binocular system, the workpiece is placed in the center of the rotating imaging platform's stage, maintaining its position and orientation throughout the acquisition process. The polarization binocular camera, supported by a gimbal, rotates around the workpiece at preset angles, acquiring intensity images in four polarization directions at each angle to obtain a multi-view polarization binocular image sequence covering the entire orientation. To ensure the stability of subsequent pose acquisition and spatial reconstruction, adjacent viewpoints should maintain visual overlap to enhance the robustness of feature matching.
[0083] After multi-view image acquisition, the left camera image sequence is input into a 3D reconstruction algorithm based on motion recovery structure. Through feature extraction, feature matching, and incremental solving, the extrinsic parameter matrix of the left camera in a unified world coordinate system is obtained. Since the relative pose between the left and right cameras is determined during the calibration phase and remains unchanged throughout the acquisition process, this invention can obtain the camera parameters of the right camera by transforming the extrinsic parameters obtained from calibration after obtaining the left camera's extrinsic parameters, without needing to solve them repeatedly. This provides the camera pose of the binocular system under all shooting views, offering a unified and robust geometric reference for subsequent depth prior calculations and neural implicit reconstruction.
[0084] Since the relative poses of the left and right cameras are determined during the calibration phase, to ensure that the disparity search range falls on the same scan line, epipolar correction is first performed on the left and right images to improve the stability of disparity estimation. For the pixels of the left image after epipolar correction... Select several candidate parallaxes within its horizontal parallax search range. Each disparity candidate corresponds to a pixel position in the right image. .
[0085] To ensure stable matching between high-reflectivity and low-texture regions, this invention introduces polarization azimuth consistency as a geometric constraint. This constraint assumes that for pixels in the left and right views to match, their corresponding surface normals should have a consistent directionality in physical space, which can be characterized by the azimuth angle extracted from the polarization image. Specifically, based on the polarization images of the left and right views, pixel... and The polarization angle at that point. The polarization angle is calculated based on the Stokes vector. Intensity images from four polarization directions constitute:
[0086]
[0087] From this, the corresponding polarization angle can be further obtained. :
[0088]
[0089] Since the principal polarization direction of the specular reflection is perpendicular to the incident plane, the azimuth angle With polarization angle There are Phase difference; using this relationship, the left pixel can be estimated. Candidate matching points in the right image The corresponding surface normal azimuth angle and The deviation of the azimuth angle between the two points is calculated based on this, and the polarization azimuth angle consistency cost function is defined. for:
[0090]
[0091] Simultaneously, the photometric matching cost is obtained by calculating the sum of the absolute differences in gray levels between the corresponding pixel windows of the two points. The final total cost of matching. Defined as the weighted sum of the two:
[0092]
[0093] in These are the preset weighting coefficients.
[0094] In practical depth estimation, the algorithm will evaluate each candidate disparity. Calculate the total cost And select the disparity value that minimizes its cost. This optimal matching yields a disparity estimate that approximates the true depth of the workpiece surface. The optimal disparity value is then obtained. Then, based on the triangulation principle of the binocular system, combined with the baseline length obtained from calibration... With camera focal length The corresponding depth value is calculated. :
[0095]
[0096] depth This describes the distance from the pixel's corresponding spatial point to the camera center. To make depth usable for implicit geometric constraints in 3D space, it needs to be back-projected into a 3D point in the left camera coordinate system. Let the homogeneous coordinates of the pixel be... Based on the pinhole imaging model, its three-dimensional coordinates in the left camera coordinate system can be obtained:
[0097]
[0098] in The inverse matrix of the camera intrinsic parameter matrix obtained from the calibration of the left camera. Used to convert pixel coordinates into the line-of-sight direction in the camera coordinate system.
[0099] To ensure that multi-view depth data falls within a unified coordinate framework, it is necessary to use the extrinsic parameter matrix obtained from calibration. This maps 3D points from the camera coordinate system to the world coordinate system. The extrinsic parameter matrix gives the rotation and translation relationship between the camera coordinate system and the world coordinate system; therefore, the mapping relationship is:
[0100]
[0101] Through the aforementioned back projection and coordinate transformation, multi-view depth data can be unified into the same world coordinate system, providing a consistent spatial reference for subsequent ray sampling and implicit field training.
[0102] During depth estimation, due to the potential for strong reflections and weak textures on the surface of metal workpieces, some parallax values may become unstable or unreliable. To distinguish between reliable and unreliable depths, this invention constructs a depth confidence score. Its value is determined by the confidence level of the matching cost. Left and right parallax confidence and local smoothness confidence Weighted composition. Confidence level. The calculation form is:
[0103]
[0104] in To establish fixed experience weights, their proportional relationships can be set according to the application scenario.
[0105] confidence in matching cost , For pixels At optimal parallax The minimum total cost of matching at point , For all candidate parallax The sum of the corresponding costs. This confidence level is described by the prominence of the optimal match in the candidate set, and its expression is:
[0106]
[0107] Confidence level for left and right disparity , The optimal disparity was found using the left image as a reference. This represents the optimal disparity obtained through reverse matching using the right image as a baseline. The confidence level is determined by the independently calculated difference between the left and right disparities, expressed as:
[0108]
[0109] in This is the scaling adjustment parameter.
[0110] For local smoothness confidence ,in To match the total cost function At optimal parallax The discrete second-order difference at the point. This confidence level is described by the continuity of the matching cost curve near the optimal match, and is expressed as:
[0111]
[0112] in For smooth control parameters.
[0113] depth confidence It only reflects the reliability of pixel-level depth values and has no spatial geometric meaning, therefore it does not require 3D backprojection or coordinate transformation like depth. This confidence level is only used as a weight to adjust subsequent sampling intervals. After the above processing, the depth information has achieved scale uniformity and multi-view alignment in the world coordinate system, and comes with a confidence level, which can be used to guide ray sampling of the neural implicit field.
[0114] After obtaining multi-view images and their corresponding camera poses, this invention employs a neural implicit representation model to uniformly model the three-dimensional geometric structure and polarization reflection characteristics of the workpiece. The model construction block diagram is shown below. Figure 3 As shown, the neural implicit representation model predicts the geometric and appearance information of any location in space by treating the neural network as a continuous three-dimensional spatial function. This model establishes the geometric correspondence between image pixels and three-dimensional spatial rays using multi-view images and corresponding camera poses. It generates predicted images through volume rendering and iteratively optimizes network parameters by comparing them with real-world images, gradually approximating the actual workpiece shape.
[0115] However, existing neural implicit representation models mainly rely on consistency constraints of multi-view image information during geometric learning. When the workpiece surface exhibits high reflectivity or weak texture, the same surface area in the image often shows significantly different brightness and appearance features under different viewpoints, making it difficult for the model to determine whether these changes are caused by the true geometric shape or by the surface reflectivity. This can easily lead to local geometric deviations or instability during reconstruction.
[0116] To address this problem, this invention introduces polarization information as an additional physical constraint. Compared to intensity information such as brightness, polarization features are determined by the surface normal and exhibit stable variation patterns. This provides consistent physical constraints for the same surface region under different viewing angles, effectively distinguishing between appearance changes caused by surface reflection and actual geometric shape changes, thus enhancing the geometric modeling stability of the neural implicit representation model under complex reflection conditions. This invention refers to this method as "polarization neural implicit representation," which compensates for the shortcomings of traditional visual constraints by introducing polarization physical properties, thereby achieving stable acquisition of the three-dimensional structure of the workpiece.
[0117] (1) Construction of implicit geometric representation and acquisition of normals
[0118] Among various implicit neural models, this invention employs a method based on the symbolic distance field, utilizing its zero isosurface to directly define the boundary of an object's surface, thereby extracting information about the workpiece's surface and normals. Based on this, the polarization implicit neural representation model uses spatial points in the world coordinate system... Construct a continuous implicit signed distance function as input. The sign distance representing the point from the workpiece surface is defined in this invention. For spatial points The shortest geometric distance to the workpiece surface is determined, and a sign function is introduced. Used to reflect spatial points Spatial relationship with the workpiece surface. Continuous implicit signed distance function. The expression is as follows:
[0119]
[0120] Among them, the sign function The geometric meaning of is as follows:
[0121]
[0122] Applying the signed distance function to the three-dimensional space where the workpiece is located creates a continuous global signed distance field, whose zero isosurface (i.e., satisfying...) The point set constitutes the implicit representation of the workpiece surface. Because... For a continuously differentiable function, the unit normal vector at a point in space can be obtained through the gradient:
[0123]
[0124] (2) Joint constraint of polarization volume rendering and Stokes vector consistency
[0125] To supervise the geometric learning of neural implicit fields using polarization information, this invention introduces polarization information as a supervisory constraint within the volume rendering framework of the neural implicit representation model; this process is called polarization volume rendering. Polarization volume rendering uses a polarization bidirectional reflectance distribution function model to convert the surface normals and polarization reflection components predicted by the implicit field into corresponding outgoing Stokes vectors, and compares them with the observed Stokes vectors calculated from the input polarization image. This utilizes polarization consistency error to constrain the learning process of 3D geometry.
[0126] The polarization bidirectional reflection distribution function model will output the Stokes vector. Modeled as diffuse reflection component With specular reflection component A linear superposition of . Where, Represents a point in space in the world coordinate system. It is the unit surface normal obtained from the gradient of the sign distance field. It is the observation direction vector, determined by the camera pose, pointing from the camera's optical center to a point in space.
[0127] In diffuse reflection modeling, diffuse reflection feature vector Constructed based on Fresnel's law of transmission. Its polarization azimuth angle. From surface normal Determine the polarization transmission coefficient. and Then, based on the material's refractive index and the direction of observation... With surface normal The included angle is calculated using the Fresnel transmission equation, and this constitutes the diffuse reflection polarization characteristic vector. :
[0128]
[0129] In specular reflection modeling, the specular feature vector Constructed based on microsurface reflection theory and Fresnel's law of reflection. Its polarization azimuth angle... From half-length vector Determined, where the half-range vector By comparing the direction of incident light with the direction of observation The sum vector is normalized to obtain the polarization reflection coefficient. and Based on the material's refractive index and the half-range vector... With the direction of observation The determined microscopic incident angle is obtained by solving the Fresnel reflection equation, thus forming the specular reflection polarization characteristic vector. :
[0130]
[0131] The polarization distribution characteristics determined by the physical laws described above need to be combined with the corresponding energy intensity term to characterize the final emission state. Diffuse reflection component intensity By neural network According to spatial location Prediction, i.e. The intensity of the specular reflection component; Then by neural network Combined with spatial location Observation direction and surface normals Joint forecasting, i.e. The final predicted outgoing Stokes vector is obtained by multiplying the analytically determined polarization feature vectors by the intensity terms predicted by the neural network and then linearly superimposing them. :
[0132]
[0133] To obtain pixel-level predicted Stokes vectors For subsequent loss monitoring, it is necessary to follow each line from the camera's optical center. Departure, direction The ray is weighted and integrally applied over the outgoing Stokes vectors at multiple sampling points, where the position of the ray in 3D space is represented as... To rationally allocate the contribution of each sampling point on the ray to the final pixel, a volume density function is introduced to describe the influence of each sampling point on the observation value. Based on the distribution characteristics of the symbolic distance field, a monotonically differentiable mapping function is used. Construct density function ,in Represents ray parameters The corresponding three-dimensional spatial point; the mapping function uses a bell-shaped probability density function. Its output decreases as the absolute value of the input increases, resulting in a larger volume density near the zero isosurface, which gradually decreases further away from the implicit surface. Volume density function It can be represented as:
[0134]
[0135] in and It is a learnable or fixed hyperparameter used to adjust the density mapping of the symbolic distance function so that the zero-level set corresponds to a higher density value in volume rendering.
[0136] Based on the aforementioned volume density distribution, the ray is defined in the parameters Transmittance function at This is used to represent the cumulative attenuation of light before it reaches the current sampling position from the camera; wherein, the integration starting position corresponding to the transmittance is determined by the depth-guided adaptive sampling interval described later. The transmittance... The expression is:
[0137]
[0138] Combined with the local outgoing Stokes vector determined by the polarization bidirectional reflection distribution function model The observed Stokes vector at the pixel can be obtained. The integral expression for continuum rendering is:
[0139]
[0140] The volume rendering integral shown indicates that the final observation at a pixel is jointly determined by the local polarization reflection response of each sampling point along the ray direction, where the volume density function... Determines transmittance Distribution, transmittance It describes the visibility attenuation process, thereby enabling the joint reconstruction of the three-dimensional geometry of the workpiece surface and its polarization reflection characteristics without the need for an explicit mesh structure.
[0141] By minimizing pixels Predicted value of Stokes vector at location Extracting observations from input images The residuals between them are used to construct the polarization uniformity loss. :
[0142]
[0143] This loss function is used to constrain the three-dimensional geometry and surface properties of the neural implicit field, so that the predicted Stokes vector is consistent with the input polarization image observation, thereby improving the reconstruction accuracy and stability of complex reflective surfaces.
[0144] (3) Gaussian modeling based on local normals and normal consistency constraints
[0145] To further reduce fluctuations in normal direction caused by noise or changes in viewpoint, and to ensure that implicit surfaces maintain detail while possessing local consistency, this invention constructs a Gaussian distribution model based on local normals in three-dimensional space. Using spatial points... Centered on a given point, the variation of the local surface normal direction within its neighborhood is examined, and its direction distribution is characterized by a three-dimensional Gaussian distribution. This Gaussian distribution is defined as:
[0146]
[0147] in, This represents the mean value of the local normal direction. To describe the dispersion and anisotropy of the normal direction The covariance matrix. To obtain the above parameters, at point... Selecting within the neighborhood sampling points And using the signed distance function The gradient calculation corresponds to the unit normal vector. Then the mean With covariance It can be estimated using the following formula:
[0148]
[0149]
[0150] After obtaining the aforementioned three-dimensional Gaussian normal distribution, since polarization images can only observe the projection direction of the normals onto the imaging plane, this invention projects the three-dimensional Gaussian normal distribution into a two-dimensional Gaussian distribution to establish consistency constraints with polarization priors. The mean of the projected two-dimensional Gaussian normal distribution is... Covariance corresponding to the normal projection direction This reflects the anisotropy of the local normal distribution. Since the polarization azimuth represents the direction of the normal projection, and the mean of a two-dimensional Gaussian distribution is geometrically the vector of the normal projection, a Gaussian normal consistency loss function can be established by comparing the differences in mean and covariance between the predicted two-dimensional Gaussian distribution and the direction prior constructed from the polarization azimuth. The definition is as follows:
[0151]
[0152] in This represents the set of pixels or ray sampling points that participate in the Gaussian consistency constraint of the normal. Used to constrain the discreteness of the local normal distribution. These are the weighting coefficients.
[0153] In the volume rendering process of neural implicit representation models, if no external geometric prior is introduced, rays typically lie within a uniformly defined global range. Uniform or layered sampling is performed within the area. This area needs to cover the entire spatial range of the workpiece, but in areas with high reflection or weak texture, the sampling points are prone to deviating from the actual surface position, which is not conducive to the stable convergence of implicit surface geometry.
[0154] To improve the ability to focus on real surface positions during volume rendering and enhance the training stability of the symbolic distance field in areas with weak textures and strong reflections, this invention uses depth information obtained from binocular stereo matching during the ray sampling stage. and its confidence level The sampling interval is adaptively adjusted to concentrate the sampling points in the neighborhood of the real surface. The sampling interval is constructed as follows: Figure 4 As shown.
[0155] Specifically, for any pixel Its depth The three-dimensional spatial position of the pixel in the world coordinate system can be determined. Based on this location and the camera's optical center Based on this relationship, its ray parameters in the current line-of-sight direction can be calculated. :
[0156]
[0157] Based on this reference ray parameter Constructing an adaptive sampling interval ,in Indicates The half-width of the interval centered on the depth confidence level, the size of which is determined by the depth confidence level. Linear control:
[0158]
[0159] In the formula Indicates the half-width of the sampling interval The maximum allowed value under scene normalization conditions. When the confidence level... At higher levels, The corresponding reduction concentrates the sampling points at the depth estimation location; when the confidence level is low, The sampling interval is increased to cover potential depth deviations and avoid geometric omissions due to an excessively narrow sampling interval. This adaptive strategy effectively reduces invalid sampling and improves the efficiency and stability of reconstructing the real surface.
[0160] Within a volume rendering framework, the implicit geometric distribution along the ray direction can be determined using volume density. With transmittance By modeling, the predicted depth corresponding to the ray can be defined as:
[0161]
[0162] This invention compares With binocular depth Construct depth-consistency constraints to make the signed distance function Maintaining consistency with the true depth on a spatial scale avoids local depressions, bulges, or scale shifts in areas with weak textures and high reflectivity. At locations with high depth confidence, constraints with near-zero signed distance can be added near the corresponding 3D points to ensure the reconstructed surface remains continuous and stable near the actual measurement location.
[0163] After the aforementioned deep-guided sampling and optimization training, the surface normal distribution of the workpiece output by the neural implicit field is optimized, as shown in the schematic diagram below. Figure 5 As shown.
[0164] In this invention, workpiece size measurement is based on three-dimensional reconstruction completed by neural implicit representation. By automatically identifying and parametrically describing the regular structures in the three-dimensional model, key size indicators are further automatically calculated, thereby realizing non-contact size detection and evaluation of various types of industrial workpieces.
[0165] (1) 3D model extraction and measurement data preparation
[0166] After the neural implicit representation model is trained, a 3D mesh model of the workpiece is obtained by extracting the zero isosurface of the symbolic distance field. The reconstructed 3D surface structure of the workpiece is as follows: Figure 6As shown. Since neural implicit representations are usually trained in a normalized coordinate system, in order to make the subsequent measurement results meaningful for engineering purposes, based on the scale factor determined by the camera calibration results and depth priors before model training, the 3D mesh model reconstructed by the neural implicit representation is mapped back from the normalized coordinate system to the real physical coordinate system, thereby obtaining a 3D model of the workpiece with actual size units.
[0167] On the restored 3D mesh model, its surface is uniformly sampled, and the corresponding unit normal vector is calculated for each sampling point. The discrete mesh representation is transformed into a point set form containing spatial location and normal information, providing a unified data foundation for subsequent geometric structure analysis and parametric modeling.
[0168] (2) Automatic identification of workpiece surface regions based on normal consistency and spatial neighborhood
[0169] Considering that the key dimensions of typical industrial workpieces (such as fasteners, shafts, flanges and plate workpieces) are usually composed of regular geometry such as end faces, hole walls and shaft structures, and the above structures can be stably abstracted into two basic geometric models, plane or cylinder, this invention identifies the structural regions of the workpiece surface based on the surface sampling points and their corresponding normal information obtained in step (1).
[0170] Specifically, within the surface sampling point set, local point sets are selected based on spatial neighborhood relationships, and the geometric characteristics of local regions are analyzed by considering the consistency of the normal directions of the sampling points. Based on this, a model estimation method based on random sampling consistency can be used to detect point sets that satisfy planar or cylindrical model constraints within the sampling point set, and sampling points that satisfy the model constraints and are spatially interconnected are grouped into the same candidate region. Point sets that do not satisfy planar or cylindrical model constraints but are spatially continuous and have smooth normal changes are retained as candidate regions for freeform surfaces for subsequent surface parametric modeling and dimensional measurement.
[0171] (3) Fit geometric primitives such as planes and cylinders to the candidate region and obtain the parameters.
[0172] For each candidate region obtained in step (2), regular geometric model verification and parametric fitting modeling are performed respectively. Considering that the key measurement dimensions in industrial workpieces can usually be described by plane, cylinder and freeform surface geometry, the present invention performs fitting of three types of geometric primitives of plane, cylinder and freeform surface on the candidate regions.
[0173] For a planar region, the unit plane normal vector is obtained by fitting. With displacement parameters The plane equation is expressed as:
[0174]
[0175] For regions that satisfy the characteristics of a cylindrical structure, the geometric model parameters of the cylinder are obtained by fitting the sampled points and minimizing the consistency error of the distance from the cylinder axis. Let the direction of the cylinder axis be a unit vector. A point on the axis is The radius of the cylinder is The cylinder fitting process can then be expressed as:
[0176]
[0177] in These are the sampling points within the candidate cylindrical region.
[0178] For candidate regions of free-form surfaces that cannot be accurately described by planes or cylinders, this invention employs a parametric surface modeling method, representing the surface as a continuous mapping function of parametric coordinates:
[0179]
[0180] in The coordinates are in the parametric domain of the surface. By minimizing the distance error from the sampling point to the parametric surface, this invention automatically identifies candidate regions for free-form surfaces that cannot be accurately described by a plane or cylinder. It employs a parametric surface modeling method, representing the surface as a continuous mapping function of parametric coordinates. :
[0181]
[0182] in These are the coordinates in the surface parameter domain. The surface parameters are automatically calculated by minimizing the distance error from the sampling points to the parametric surface.
[0183]
[0184] in, For sampling points within the candidate region of the freeform surface, These are the coordinates of the surface parameters corresponding to the sampling points.
[0185] (4) Establish the measurement reference direction and unified coordinate definition based on the fitting primitives.
[0186] After obtaining regular geometric primitives such as planes and cylinders, this invention automatically establishes a measurement benchmark based on the workpiece structure. For workpieces with obvious shaft or hole structures, the axial direction of the main cylindrical geometric primitive is selected. The main measurement direction for the workpiece is selected as the main plane normal; for plate-type workpieces with end faces as the main measurement direction, the main plane normal is selected as the main plane normal. As a reference direction.
[0187] (5) Automatically calculate key dimension indicators based on primitive parameters and spatial relationships
[0188] After establishing a unified measurement benchmark, this invention transforms the measurement of workpiece dimensions into the calculation of geometric element parameters and their spatial relationships, thereby achieving automatic measurement without the need for manual point selection.
[0189] When the workpiece contains a pair of relatively parallel end faces, the displacement parameters of the corresponding planar model are used. , The distance between end faces is calculated to characterize the thickness, height, or axial length of a workpiece. The calculation method is as follows:
[0190]
[0191] When a workpiece contains cylindrical structures such as holes, shafts, or sleeves, the radius parameter of the cylindrical geometric model is used. Calculate the corresponding diameter directly:
[0192]
[0193] When a workpiece contains multiple holes or multiple cylindrical sections, it can be constructed from points on different cylindrical axes. and The spatial distance between holes is calculated using dimensional parameters such as hole spacing or center distance, for example, the center distance between two holes:
[0194]
[0195] When a workpiece contains both inner and outer cylinders or multiple coaxial cylindrical structures, coaxiality or eccentricity can be calculated based on the radial offset between the cylinder axes. Let the spindle direction be... The coaxiality index can then be expressed as:
[0196]
[0197] For dimensions along a certain direction of the surface (such as chamfer width or fillet arc length), select a parametric curve within the surface parameter domain. The corresponding space curve length is:
[0198]
[0199] For surface regions with rounded corners or transitional features, the local radius of curvature is calculated based on the principal curvatures of the surface. Let the surface at the parameter point... The principal curvature at that point is The corresponding radius of curvature is:
[0200]
[0201] The above-mentioned dimensional indicators are all directly derived from the fitted geometric primitive parameters, with clear geometric origins and engineering significance. They can cover the measurement of inner and outer diameters and thicknesses of fasteners and bushings, the measurement of diameters and axial lengths of shafts, and the measurement requirements of thickness, hole diameters, and hole spacing of flanges and plate parts. Furthermore, they can be used to measure the curve lengths, radii of curvatures, and other surface geometric features of free-form surface workpieces.
Claims
1. A method for three-dimensional reconstruction and dimensional measurement of a workpiece based on multi-view polarization binocular imaging, characterized in that, include: Step 1: Construct a multi-view polarization binocular imaging device, which includes a polarization binocular camera (2), a ring light source (1) and a rotating platform. The binocular camera is calibrated by a calibration plate (6) to obtain the intrinsic parameters and the fixed relative pose relationship between the binocular cameras. Step 2: Fix the workpiece (5) to be tested on the rotating platform, control the imaging device to take pictures around the workpiece (5) at a preset angle, and continuously collect multi-view polarized binocular image pairs to obtain an image sequence with high overlap around the workpiece (5) once. Step 3: Input the left camera image sequence into the structure-of-motion algorithm to recover the spatial pose of the left camera in the world coordinate system from each viewpoint, and calculate the pose of the right camera by combining the binocular relative pose, thereby obtaining the binocular camera pose information from all viewpoints. Step 4: Under camera pose constraints, extract polarization information and introduce polarization constraints in stereo matching to obtain multi-view depth information with confidence; perform depth back projection based on camera intrinsic parameters and camera pose to form multi-view depth prior. Step 5: Based on multi-view depth prior, polarization information and camera pose, construct a polarization neural implicit representation model. Guide adaptive sampling through depth prior, supervise geometric learning by polarization volume rendering, and combine normal Gaussian modeling constraints to suppress local normal fluctuations. Use the sign distance field to characterize the workpiece under test (5) three-dimensional surface. Step 6: Extract the zero isosurface mesh from the symbolic distance field to restore the true scale; Under the constraints of normal consistency and neighborhood relationship, it identifies planar, cylindrical and free-form surface structures, realizes dimension measurement based on the geometric primitive fitting results, and automatically calculates key dimensions such as thickness, diameter, hole center distance, and coaxiality.
2. The method according to claim 1, characterized in that, The polarization binocular camera (2) is fixed to the rotating platform by a rigid structure, and the relative pose of the polarization binocular camera (2) remains unchanged during the acquisition process. The polarization binocular camera (2) and the ring light source (1) move synchronously around the workpiece (5) to be measured fixed on the rotating platform, and acquire a multi-view polarization binocular image sequence of about 60 viewpoints in a preset angle step of 6° to ensure sufficient overlap between adjacent viewpoints.
3. The method according to claim 1, characterized in that, The total matching cost of stereo matching in step four is composed of the photometric matching cost. Cost of consistency with polarization azimuth angle According to the preset weighting coefficient Weighted summation representation: ; in, and These are the left pixel and the right image candidate matching point, respectively. Cost of polarization azimuth consistency The definition is as follows: ; in, and These are the surface normal azimuth angles corresponding to the pixels in the left and right images, respectively.
4. The method according to claim 1, characterized in that, In step four, confidence level By matching cost confidence Left and right parallax confidence and local smoothness confidence The weighted composition is used to dynamically adjust the expansion of the sampling interval in subsequent adaptive ray sampling: ; in, Fixed experience weights; The confidence level of the matching cost Described by the prominence of the optimal match in the candidate set. For pixels At optimal parallax The minimum total cost of matching at point , For all candidate parallax within the search space The sum of the corresponding costs is given by the following formula: ; The confidence level of left and right disparity Determined by the independently calculated difference in left and right disparity. The optimal disparity was found using the left image as a reference. The formula for the optimal disparity obtained by reverse matching based on the right image is as follows: ; in This is the scaling adjustment parameter.
5. The method according to claim 1, characterized in that, In step five, the neural implicit representation model introduces a volume rendering calculation process based on a polarization physics model, wherein the outgoing Stokes vector... Based on the polarization bidirectional reflection distribution function model, the diffuse reflection component... With specular reflection component Linear superposition yields: ; Diffuse reflection polarization characteristic vector Constructed based on Fresnel's transmission law, its polarization azimuth angle From surface normal Determine the transmission coefficient. and Based on the material's refractive index and the direction of observation With surface normal The included angle was calculated as follows: ; Specular reflection polarization characteristic vector Constructed based on the microsurface reflection theory, its polarization azimuth angle From half-length vector Determined, where the half-range vector By aligning the incident light direction with the observation direction The sum vector is normalized to obtain the reflection coefficient. and Based on the material's refractive index and half-range vector With the direction of observation The included angle was calculated as follows: ; The polarization distribution characteristics determined by the physical laws described above need to be combined with the corresponding energy intensity term to characterize the final emission state. Diffuse reflection intensity By neural network According to spatial location Predicted specular reflection intensity By neural network Combined with spatial location Observation direction and surface normals The predicted outgoing Stokes vector is ultimately formed through linear superposition. Based on predicted values Compared with observed values The difference between them can be defined as the polarization uniformity loss. .
6. The method according to claim 1, characterized in that, Step five also includes a Gaussian normal modeling method, including: At a point in space Selecting within the neighborhood sampling points The unit normal vector is calculated based on the gradient of the signed distance field at each sampling point. Construct a three-dimensional Gaussian distribution: ; in Represents the spatial point The mean vector of the normal line, The corresponding normal covariance matrix is calculated as follows: ; 。 7. The method according to claim 6, characterized in that, Considering that polarization information only constrains the direction of the surface normal in the imaging plane, the three-dimensional Gaussian distribution of the normal is projected into a two-dimensional Gaussian distribution, and a normal Gaussian consistency loss is constructed based on the two-dimensional Gaussian distribution and the polarization information. : ; in and The two-dimensional Gaussian mean and covariance obtained by projection. The normal projection direction is a priori, constructed from the polarization azimuth angle; This is used to constrain the dispersion of the local normal distribution.
8. The method according to claim 1, characterized in that, In step five, the depth-guided adaptive sampling refers to the sampling interval of the ray. It is based on depth confidence. Make dynamic adjustments: ; in, For the sake of deep priors The parameter positions of the 3D points obtained by back projection on the corresponding camera ray; For The half-width of the ray parameter sampling interval centered on the ray; This is the preset maximum value of the interval half-width under scene normalization conditions.
9. The method according to claim 1, characterized in that, In step six, the size measurement includes: on the three-dimensional model that restores the true scale, based on the normal consistency and spatial neighborhood relationship of the surface sampling points, automatically identifying regular structural regions and fitting geometric primitives of planes, cylinders and freeform surfaces, so as to realize the size measurement of the workpiece to be measured without manual selection of points (5); Plane fitting is performed on the planar region to obtain the unit normal vector and displacement parameters characterizing the spatial position and orientation of the plane, and the plane equation is determined accordingly. By minimizing the distance error from the sampling point to the cylinder axis, the axial direction, points on the axis, and radius parameters of the cylinder are obtained by fitting the data for the cylindrical region. For the freeform surface region, a parametric surface modeling method is used to represent the freeform surface as a continuous mapping function of parametric coordinates, and the surface parameters of the freeform surface are automatically obtained by minimizing the distance error from the sampling point to the parametric surface.
10. The method according to claim 9, characterized in that, The calculation of the critical dimensions includes: Thickness calculation: based on displacement parameters of a pair of parallel planes , ,according to calculate; Diameter calculation: based on cylinder radius ,according to calculate; Hole center distance calculation: based on points on different cylinder axes and ,according to calculate; Coaxiality calculation: The direction of the principal axis is ,according to calculate; Calculation of surface curve length: Select a parametric curve within the surface parameter domain. The corresponding space curve length is ; Surface curvature radius calculation: Based on the principal curvature of the surface, at the parameter points The principal curvature at that point is The corresponding radius of curvature is .