A multi-view small target fast three-dimensional reconstruction method based on stereo calibration
By combining a stereo calibration object with a multi-view camera, and using a "3 side view + 1 top view" topological layout and epipolar geometric constraints, the problem of 3D reconstruction of clustered small targets at low resolution was solved, achieving high-precision and low-ambiguity 3D coordinate reconstruction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHANGCHUN UNIV OF SCI & TECH
- Filing Date
- 2026-07-03
- Publication Date
- 2026-07-31
AI Technical Summary
Existing technologies struggle to achieve high-precision, low-ambiguity 3D coordinate reconstruction of clustered small targets under low-resolution conditions, especially in long-distance, large-field-of-view measurement scenarios where issues such as coplanar degeneracy, tangential distortion, occlusion, and matching redundancy exist.
By employing a stereo calibration object combined with a "3 side view + 1 top view" multi-camera topology layout, and using a multi-view joint reconstruction method that combines epipolar constraints and uncertainty weighted fusion, stable matching and high-precision positioning are achieved by utilizing sub-pixel centroid extraction and epipolar geometric constraints.
It improves the accuracy of solving extrinsic parameters in the depth direction, reduces occlusion interference and matching ambiguity, realizes high-precision 3D reconstruction of small targets under low-resolution conditions, and enhances the stability and accuracy of reconstruction results.
Smart Images

Figure CN122492953A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of optical measurement and 3D reconstruction technology, specifically to a fast 3D reconstruction method for small targets from multiple perspectives based on stereo calibration. Background Technology
[0002] Spatial three-dimensional coordinate reconstruction of clustered small targets plays a crucial metrological support role in fields such as industrial precision manufacturing, experimental mechanical measurement, and condition monitoring of micro-devices. Research on multi-view visual reconstruction of clustered small targets helps reveal the correlation between the geometry, spatial pose, and physical properties of micro-components, and has significant engineering application value for precision machining quality control, automated optical inspection, and material mechanical property analysis.
[0003] Existing methods for measuring clustered small targets mostly employ single-view observation or contact-based point-by-point measurement, which makes it difficult to synchronously and efficiently acquire the precise coordinates of multiple targets in three-dimensional space. Furthermore, for low-texture, densely arranged small targets, feature matching ambiguities are easily generated, making it difficult to meet the requirements of high-precision and high-efficiency three-dimensional reconstruction.
[0004] Current research has proposed 3D reconstruction methods based on multi-view cameras. By integrating multi-view geometric constraints and triangulation principles, these methods transform spatial targets from 2D image coordinates to 3D spatial coordinates. Furthermore, by combining multi-view joint calibration and bundle adjustment, they improve the accuracy of camera pose estimation and the completeness of 3D reconstruction to some extent. However, for special scenarios involving "small, low-texture, and multiple targets," existing technologies still have the following problems: 1. Planar target calibration leads to coplanar degeneracy and amplification of tangential errors. Existing multi-camera systems commonly use planar Zhang Zhengyou calibration plates for camera calibration. These planar targets only provide two-dimensional geometric constraints within the calibration plane, lacking effective constraints in the depth direction perpendicular to the calibration plane. Especially in long-distance, large-field-of-view measurement scenarios, this easily leads to "coplanar degeneracy," resulting in significantly lower accuracy in solving the translation parameters in the depth direction (Z-axis) of the camera's extrinsic parameters compared to the X / Y directions.
[0005] Meanwhile, the traditional Zhang Zhengyou two-step method has limited ability to compensate for tangential distortion in the lens edge region during nonlinear optimization. When imaging rays pass through the lens edge region, the tangential distortion coefficient is prone to coupling with parameters such as principal point offset and image plane tilt, leading to unstable estimation of tangential distortion parameters, which in turn affects the consistency and accuracy of the 3D reconstruction results across the entire field of view.
[0006] 2. Inappropriate multi-view topology layout leads to occlusion and matching redundancy. For 3D reconstruction tasks involving clustered small targets, existing multi-view camera systems typically lack topological layout designs that consider spatial occlusion relationships.
[0007] For example, while the traditional single-layer, uniformly arranged ring can provide some horizontal field of view coverage, it cannot effectively solve the problems of target occlusion and projection overlap in the depth direction. When multiple targets are distributed along the camera's optical axis, foreground targets can easily occlude background targets, creating blind spots.
[0008] 3. Feature extraction and matching are difficult under low resolution conditions. Under low-resolution and low-texture conditions, traditional feature point algorithms (such as SIFT and SURF) struggle to extract stable keypoints and effective local descriptors, leading to a significant decrease in the accuracy and recall of feature point matching. Especially in dense, small-object scenes, objects are prone to mutual occlusion and projection overlap, further increasing feature matching ambiguity and causing 3D reconstruction failure or a significant decrease in reconstruction accuracy.
[0009] Therefore, how to achieve stable recognition, high-precision matching, and robust 3D coordinate reconstruction of clustered, low-texture, small targets under low-resolution conditions remains a pressing technical problem in the field of multi-view visual measurement. Summary of the Invention
[0010] To address the aforementioned problems, this invention proposes a rapid 3D reconstruction method for small targets from multiple perspectives based on stereo calibration. By constructing a stereo-coded calibration object, combining a "3 side views + 1 top view" multi-camera topology layout, and a multi-perspective joint reconstruction method based on epipolar constraints and uncertainty weighted fusion, high-precision, low-ambiguity 3D coordinate reconstruction of clustered small targets under low-resolution conditions is achieved.
[0011] The method includes the following steps: S1. Set up 4 cameras to simultaneously photograph the stereo calibration object and obtain the sub-pixel coordinates of the feature corner points of each calibration surface of the stereo calibration object; S2. Based on the sub-pixel coordinates obtained in step S1, solve for the intrinsic and extrinsic parameters of each camera respectively, and construct the relative rotation matrix and relative translation vector between any two cameras based on the extrinsic parameter matrix of each camera. S3. Simultaneously photograph the target object with 4 cameras to obtain the sub-pixel coordinates of the center of the target object photographed by each camera. Then, based on the relative rotation matrix and relative translation vector between any two cameras, perform epipolar matching on the sub-pixel coordinates of the center of the target object photographed by each camera to obtain several pairs of two-dimensional points that are successfully epipolar matched. S4. Based on several pairs of two-dimensional points that have been successfully matched with epipolar lines, solve for the initial three-dimensional coordinates of several target objects; S5. Based on several successfully matched epipolar point pairs, obtain the angle between the corresponding camera's line-of-sight vector and the target object's line-of-sight vector. ; S6, based on Given the initial 3D coordinates of several target objects, introduce fusion weights to reconstruct the 3D coordinates of the target objects.
[0012] Furthermore, in step S1, the first camera, the second camera, and the third camera are positioned on the side of the stereo calibration object, and the fourth camera is positioned directly above the stereo calibration object. Each calibration surface of the three-dimensional calibration object is provided with a checkerboard pattern and a corresponding ArUco code mark; The four cameras simultaneously acquire several images of the stereo calibration object. Each image is traversed, and the corresponding calibration surface is identified through ArUco encoding. Then, the sub-pixel coordinates of the feature corner points of the corresponding calibration surface are extracted using a sub-pixel corner detection algorithm. Furthermore, a world coordinate system is constructed to obtain the three-dimensional coordinates of the characteristic corner points of each calibration surface of the three-dimensional calibration object; For any feature corner point on any calibration surface of a stereo calibration object, establish perspective projection constraint equations for sub-pixel coordinates and three-dimensional coordinates, thereby obtaining several 3D-2D matching point pairs; Based on several 3D-2D matching point pairs, solve for the intrinsic parameter matrix of each camera. and extrinsic parameter matrix , Indicates the first The rotation matrix of the camera relative to the world coordinate system. Indicates the first The translation vector of the camera relative to the world coordinate system; The relative rotation matrix between any two cameras is: ,in, Indicates the first The index value of the camera, Indicates the first The index value of the camera, , , , Indicates the first Camera coordinate system to the first The rotation matrix of the camera coordinate system. Indicates the first The rotation matrix of the camera relative to the world coordinate system. No. The transpose of the rotation matrix of the camera relative to the world coordinate system; The relative translation vector between any two cameras is ,in, Indicates the first Camera coordinate system to the first The translation vector of the camera coordinate system. Indicates the first The translation vector of the camera relative to the world coordinate system. Indicates the first The translation vector of the camera relative to the world coordinate system.
[0013] Furthermore, the polar matching includes the following steps: S31, Calculate the first Subpixel coordinates of the center of the target object captured by the camera In the Polar relation on a camera ,in, Indicates camera and camera The fundamental matrix between them; S32, Calculate the first Subpixel coordinates of the center of the target object captured by the camera arrive set distance ; S33, if Then judge and A pair of two-dimensional points that are successfully matched by epipolar lines is considered a successful match; otherwise, the match is unsuccessful. This represents the polar distance residual; S34. Repeat steps S31 to S33 to traverse all camera combinations and obtain several two-dimensional point pairs with successful epipolar matching.
[0014] Furthermore, based on several successfully matched two-dimensional point pairs, the initial three-dimensional coordinates of several target objects are solved by the following steps; S41. Based on each pair of two-dimensional points that are successfully matched with epipolar lines, establish an overdetermined system of equations using the direct linear transformation method. S42. Using the singular value decomposition method, solve the overdetermined system of equations to obtain the initial three-dimensional coordinates of several target objects. , This represents the index value of a two-dimensional point pair that has successfully matched the epipolar line. For the first The initial three-dimensional coordinates of the two-dimensional point pairs that have been successfully matched with the epipolar lines.
[0015] Furthermore, the fusion weights are specifically as follows: .
[0016] Furthermore, the formula for reconstructing the three-dimensional coordinates of the target object is: ,in, These are the final three-dimensional coordinates of the target object.
[0017] The beneficial effects of the method described in this invention are as follows: (1) This invention overcomes the coplanar degeneracy problem of traditional planar calibration by combining three-dimensional calibration objects with multi-view spatial constraints, and improves the accuracy of solving extrinsic parameters in the depth direction.
[0018] (2) This invention reduces occlusion interference and matching ambiguity in clustered small target scenarios by combining side-view and top-view topology layout and spatial topology sorting constraints, thereby improving the reliability of multi-target matching.
[0019] (3) The present invention uses sub-pixel centroid extraction combined with epipolar geometric constraints to achieve stable matching and high-precision positioning of small targets under low resolution and low texture conditions.
[0020] (4) The present invention introduces an uncertainty weighted fusion mechanism based on the line of sight angle and the epipolar residual, which effectively suppresses the problem of depth error amplification caused by small intersection angle and improves the stability and accuracy of the three-dimensional reconstruction results. Attached Figure Description
[0021] Figure 1 This is a flowchart of the method described in this invention; Figure 2 This is a schematic diagram of the asymmetric spatial topology layout and included angle theory of the four cameras described in this invention; Figure 3 This is a schematic diagram of a single side of the cubic checkerboard marker described in this invention; Figure 4 This is a schematic diagram of the weighted fusion of uncertainties in the multi-view reconstructed three-dimensional coordinates and the convergence of the error ellipse described in this invention. Detailed Implementation
[0022] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0023] Example 1 This embodiment provides a fast 3D reconstruction method for small targets from multiple perspectives based on stereo calibration. The flowchart of the method is as follows: Figure 1 As shown, the method includes the following steps: S1. Set up 4 cameras to simultaneously photograph the stereo calibration object and obtain the sub-pixel coordinates of the feature corner points of each calibration surface of the stereo calibration object; The relevant operations in step S1 will be introduced with specific examples: In this embodiment, four cameras with a resolution of 640×480 are selected, and their physical field of view is [missing information]. , ,in, The current horizontal field of view of the camera. The average working distance of the four cameras to the center of the cube; like Figure 2 As shown, based on the principle of spatial complementarity, the four cameras adopt an asymmetric topological layout of "3 side-view + 1 top-view": Top-down camera positioning: The fourth camera (camera 4) looks vertically downwards at the top surface (upper surface) of the stereo calibration object, which is specifically used to capture the global two-dimensional topological distribution of clustered small targets in the XY plane, as a way to prevent target confusion.
[0024] Derivation of the side-view camera angle: The remaining three cameras (cameras 1, 2, and 3) are arranged in a side-view ring. To achieve the best trade-off between "reconstructed depth accuracy" and "image feature deformation," the space between the two side-view cameras will form an angle. Strictly limited to Within this range, the included angle ensures that the error ellipse in the reconstructed space is as close to a perfect circle as possible, thereby minimizing the uncertainty in the depth direction.
[0025] In view of the physical field of view of the above four cameras, the design of the stereo calibration object in this embodiment is as follows: Based on the principle of field of view coverage, the side length of the stereo calibration object... The cube should occupy 1 / 3 to 1 / 2 of the image. Since side-view cameras (cameras 1, 2, and 3) will see two or three faces, the cube's projection will appear wider. Therefore, the side length of a single face of the cube is set to 1 / 4 to 1 / 3 of the physical field of view width. In this embodiment, The formula for calculation is: .
[0026] In multi-view computer vision systems, obtaining the precise rotation and translation matrices between each camera is a core prerequisite. Traditional two-dimensional planar checkerboard patterns have limitations in multi-camera joint calibration: when cameras are distributed at different angles, such as in a ring arrangement, a single plane cannot be completely observed by all cameras simultaneously, leading to a cumbersome calibration process and the introduction of accumulated errors. The six-sided 3D coded checkerboard pattern, with its cubic structure and six faces covered by a checkerboard pattern with a globally encoded ID, enables simultaneous observation from multiple perspectives by identifying the three-dimensional spatial coordinates and unique identifier of each feature point. This simplifies the calibration process and improves global calibration accuracy.
[0027] In this embodiment, the camera resolution is 640×480. To accurately extract ArUco corner points on the calibration object, the visual algorithm needs to achieve sub-pixel accuracy. However, under multi-view shooting, due to perspective distortion, the calibration board will tilt and deform, causing the original feature points to appear as blurred or degraded small areas in the image, potentially leading to decoding failure. To ensure reliable feature point recognition even under tilt distortion, this invention designs each square to retain at least 20 pixels of space in the image. This ensures that corner point information is fully preserved even with perspective distortion, meeting the sub-pixel accuracy requirement for extraction.
[0028] The total width of the images captured by the four cameras is 640 pixels, and the cube's face width occupies approximately 210 pixels (1 / 3). To meet the requirement that each square has at least 20 pixels, the maximum number of squares that can be arranged on a single face is 210 / 20 ≈ 10. Therefore, the physical dimensions of a single square should be: The “white space” refers to the blank area reserved at the edge of the cube. Its function is to ensure that the outer squares and the inner squares form sufficient gray-scale gradient contrast, so that OpenCV can accurately extract the corner points of the internal checkerboard and ArUco codes even when there is reflection or overexposure at the edge when looking for high-frequency gradients in local images.
[0029] In this embodiment, the side length of the 3D target is 400 mm. To ensure stable extraction of corner points and to account for perspective distortion allowance, the size of each square in the checkerboard is set to 40 mm. The side length of the effective pattern area is calculated as follows: Effective pattern area side length = total side length 2 × blank space = 400mm 2 × 40mm = 320mm Effective pattern area side length = Total side length - 2 White space = 400 mm - 2 40 mm = 320 mm; Therefore, the blank area around the cube is 40 mm, which is used to ensure the gradient difference between the edge of the checkerboard and the inner squares, and to ensure the robustness and sub-pixel accuracy of the visual calibration.
[0030] Based on the above design, such as Figure 3 The result is a high-precision alumina cube calibrator with sides of 300mm×300mm×300mm. Each face is printed with a 30mm×30mm square and a unique ArUco code. The ArUco code is a QR code-like visual reference mark based on binary square markings.
[0031] In this embodiment, four 640×480 cameras (cameras 1-3 surround the object from the side, and camera 4 faces the top surface of the stereo calibration object from above) are simultaneously triggered to acquire 20 sets of synchronous images of the stereo calibration object under different translation and tilt states. The line-of-sight angle between two adjacent cameras in cameras 1, 2, and 3 is adjusted to 60°. Each set of images consists of 4 images taken by 4 cameras at the same time for the stereo calibration object in the same posture.
[0032] Each set of synchronized images corresponds to a spatial pose of the stereo calibration object. The images captured by each camera in each set of images are processed as follows: First, the ArUco marker detection algorithm (cv2.aruco.detectMarkers) is used to identify the ArUco code (marker) ID in the image, and the cube calibration face corresponding to the current image is determined according to the ArUco code ID. Then, the checkerboard feature corner points on the corresponding calibration face are detected, and the sub-pixel corner optimization algorithm (cv2.cornerSubPix) is used to obtain the sub-pixel coordinates of each feature corner point, thereby obtaining the two-dimensional pixel observation data of the feature corner points of each calibration face under different poses.
[0033] By using the above method, the feature corner points of each calibration surface of the stereo calibration object can be repeatedly observed from multiple perspectives and in various spatial postures, increasing the number of correspondences between two-dimensional pixel coordinates and three-dimensional world coordinates, thereby improving the stability and calibration accuracy of solving the camera's intrinsic and extrinsic parameters.
[0034] S2. Based on the sub-pixel coordinates obtained in step S1, solve for the intrinsic and extrinsic parameters of each camera respectively, and construct the relative rotation matrix and relative translation vector between any two cameras based on the extrinsic parameter matrix of each camera. The relevant operations in step S2 will be introduced with specific examples: This embodiment establishes a unified global world coordinate system. The origin of the coordinate system is the geometric center of the top surface (upper surface) of the hexagonal cube calibration object. The plane containing the top surface is The plane is used as the X-axis and Y-axis, respectively. A right-handed Cartesian coordinate system is established.
[0035] In this embodiment, all cameras uniformly adopt... It means that, among them: Each of the four cameras corresponds to one of them.
[0036] definition Indicates the first A camera, Indicates the first A camera, , , ; In this embodiment, a unique ArUco-coded marker (ID: 0~5) is added to the white margin area (white space) around each face of the cube. By recognizing the ArUcoID, the local target face of the cube corresponding to the current image can be automatically determined, thus completing the automatic identification and matching of different target faces without manual intervention.
[0037] The outer white space is used to ensure a stable grayscale gradient between the edge area and the inner checkerboard. Even if there is reflection or local overexposure at the edge, the stable parsing of the corner points of the inner checkerboard and ArUco encoding can still be guaranteed.
[0038] This embodiment uses a cube calibration object to establish a 3D-2D correspondence between three-dimensional points in space and two-dimensional points in an image.
[0039] The “3D-2D correspondence” refers to the one-to-one mapping relationship between the three-dimensional coordinates of the same physical feature point in the world coordinate system and the two-dimensional pixel coordinates of the feature point in the image pixel coordinate system after being imaged by the camera.
[0040] 3D World Coordinates: Because the six-sided cube calibration object is manufactured with high precision, the three-dimensional coordinates of each corner point of its checkerboard pattern in the world coordinate system are as follows: ; , , These represent the spatial coordinates of the feature corner point along the X-axis, Y-axis, and Z-axis in the world coordinate system, respectively.
[0041] 2D pixel coordinates: Images of the cube calibration object are simultaneously acquired by multiple cameras, and the coordinates of the checkerboard corner points are extracted using a sub-pixel corner detection algorithm. Simultaneously, the ArUco encoded ID is used to determine the target surface to which the corresponding corner point belongs, thus obtaining the two-dimensional pixel coordinates of the corresponding corner point in the image plane. ,in, and These represent the horizontal and vertical pixel coordinates of the feature corner point in the image pixel coordinate system, respectively.
[0042] For any feature corner point on a 3D calibration object, a perspective projection constraint equation is established for its sub-pixel coordinates and 3D coordinates, thereby obtaining several 3D-2D matching point pairs.
[0043] Establish perspective projection constraint equations: Based on the pinhole camera model, the projection relationship from any 3D feature corner point to the 2D pixel coordinates is expressed as: ,in, The scale factor is the depth value that varies for different points in 3D space due to their different distances from the camera. The values of also differ; for the same spatial point, during a single projection process, its corresponding scale factor varies. It is a fixed value, and this value is equal to the depth coordinate of the spatial point in the camera coordinate system. For the first A camera Rotation matrix, For the first A camera Translation vector.
[0044] Solving for camera intrinsic and extrinsic parameters: Intrinsic parameter initialization: Treat each face of the 3D calibration object as an independent planar calibration face; For each plane calibration surface, perspective projection constraint equations are established using two-dimensional pixels collected under various poses and corresponding three-dimensional corner points; For the above perspective projection constraint equations, the orthogonal constraint relationships between their column vectors are extracted. These constraints are directly related to the camera intrinsic parameters.
[0045] By combining the constraint information of multiple planar mapping matrices, a system of linear constraint equations is constructed. The above linear constraint equations are solved using singular value decomposition to obtain the intrinsic parameter matrix for each camera. .
[0046] Solving for extrinsic parameters: in the intrinsic parameter matrix Given the conditions, for each camera Using it to capture N sets of 3D-2D corresponding points of feature corner points in the image The PnP algorithm or the direct linear transformation algorithm is used to solve the problem of the camera relative to the world coordinate system. extrinsic parameter matrix: . Due to optical distortion inherent in camera lenses, this embodiment further introduces radial and tangential distortion models. In the normalized image plane: , This represents the radial distance from the current image point to the center of the normalized image plane. and These represent the horizontal and vertical coordinates in the normalized image plane, respectively. The distortion model is represented as: ; in, This represents the horizontal pixel coordinates after lens distortion. This represents the vertical pixel coordinates after lens distortion. and The radial distortion coefficient is used to describe the barrel or pincushion distortion of a lens. and The tangential distortion coefficient is used to describe the asymmetric distortion caused by lens mounting misalignment or sensor non-parallelism.
[0047] No. Distortion parameter vector corresponding to each camera Recorded as:
[0048] The distortion parameter vector Used to describe nonlinear geometric distortions caused by the optical system during lens imaging.
[0049] After obtaining the initial intrinsic and extrinsic parameters of the camera, normalize the image plane coordinates. Substituting the distortion model, the ideal projected coordinates are nonlinearly corrected to obtain the actual distorted image coordinates: .
[0050] In this embodiment, , , and Instead of being preset fixed values, the parameters to be estimated are obtained through joint optimization based on the 3D-2D correspondence during the calibration process.
[0051] Relative pose transformation between any two cameras Seamless matrix transfer simultaneous equations are achieved through their absolute extrinsic parameters:
[0052] Its corresponding relative rotation matrix Relative translation matrix . After obtaining the initial intrinsic and extrinsic parameters, this embodiment also constructs a global nonlinear optimization function to jointly optimize all visible corner points in all cameras, with the optimization objective being to minimize the total reprojection error:
[0053] in, This represents the sum of squared total weight projection errors. Total number of cameras This represents the total number of three-dimensional points in space, specifically the total number of all extracted corner points on the cube. Represents the visibility indicator function, if the first... The spatial point can be... If the camera sees and extracts the data successfully, then... If the view is obstructed or the back is not visible, then . Indicates the observer, that is, the one in the... The actual extracted from the camera image. Two-dimensional pixel coordinates of each corner point . Indicates the first The three-dimensional homogeneous coordinates of the corner points in the cubic coordinate system , This represents a nonlinear projection function. It describes the theoretical calculation process of mapping 3D points to 2D pixels. This represents the square of the L2 norm.
[0054] For any camera The world coordinates of a corner point in space Coordinates transformed to this camera coordinate system Its matrix transfer formula is:
[0055] in, It is a 3×3 rotation matrix. It is a 3×1 translation vector. and This constitutes a camera The extrinsic parameter matrix relative to the world origin.
[0056] Retrieve the reference origin image and calculate the relative extrinsic parameters between any two cameras using the matrix transfer formula: for example, the relative rotation matrix from camera 1 (side view) to camera 4 (top view). and translation vector Establish pairwise joint relationships.
[0057] The reference origin image refers to: after the system is rigidly fixed, placing the stereo calibration object at the center of the common field of view of multiple cameras, and selecting a certain geometric feature on the calibration object as the globally unique origin of the world coordinate system. At that time, the first set of four images were captured simultaneously by the four cameras.
[0058] Based on the transformation relationship of the single camera above, we can deduce the world coordinates:
[0059] Because the rotation matrix is an orthogonal matrix ,so:
[0060] Substituting this formula into the transformation formula of camera 4 In the middle, we get:
[0061] Therefore, we derived the coordinate transformation relationship from camera 1 to camera 4: Relative rotation matrix:
[0062] Relative translation vector:
[0063] By using the relative pose relationship, any two cameras (whether they are two side-view cameras or side-view and top-view cameras) can solve for the relative rotation matrix and relative translation vector by determining their respective extrinsic parameters relative to the cube.
[0064] S3. Simultaneously photograph the target object with 4 cameras to obtain the sub-pixel coordinates of the center of the target object photographed by each camera. Then, based on the relative rotation matrix and relative translation vector between any two cameras, perform epipolar matching on the sub-pixel coordinates of the center of the target object photographed by each camera to obtain several pairs of two-dimensional points that are successfully epipolar matched. The relevant operations in step S3 will be introduced with specific examples: In this embodiment, the target object consists of 16 gray steel columns, each 16mm × 16mm in diameter, which are simultaneously captured by four cameras. Due to the low resolution, the gray steel columns appear as gray circular spots with a diameter of approximately 12 to 15 pixels in the image.
[0065] First, the background is removed by binarization and connected component analysis. Then, the pixel coordinates of the target center are accurately extracted by the sub-pixel gray-scale centroid method, with an extraction accuracy of 0.05 pixels.
[0066] Using the pairwise relative external parameters (relative rotation) of the cameras established in step S2 and relative translation ), calculate the fundamental matrix F between each pair of cameras, specifically: 1. Construct the inverse symmetric matrix of the translation vector. relative translation vector Convert to its corresponding 3×3 inverse symmetric matrix :
[0067] 2. Calculate the essential matrix Q The essential matrix Q represents the epipolar geometric relationship between two cameras on the normalized image plane, and depends only on the relative extrinsic parameters of the cameras. According to the epipolar constraint equation, the essential matrix... The formula for calculation is:
[0068] 3. Derive the fundamental matrix F The fundamental matrix F is based on the essential matrix Q and further introduces the camera's intrinsic parameters to describe the epipolar geometric relationship between the two actual pixel coordinate systems.
[0069] The formula for calculating the fundamental matrix F is: ,in, Indicates camera and camera The fundamental matrix between them and Indicates camera and The intrinsic parameter matrix, The expansion is as follows: .
[0070] For cameras The target point on the camera is calculated. The upper polar line .
[0071] Target object in camera The matching point in the equation must fall on the polar line. Nearby. By combining the XY-axis spatial topology sorting provided by the overhead camera, false correspondences caused by the dense and overlapping of small targets are eliminated, achieving high-speed and error-free matching of 16 to 20 small targets. The process of obtaining the spatial topological sort is as follows: The optical axis of the top-view camera is approximately perpendicular to the XY plane of the world coordinate system. After synchronous trigger acquisition, the image from the top-view camera (camera 4) is first processed. Using the sub-pixel centroid method, the two-dimensional pixel coordinate set of all M small targets in the field of view is extracted. The coordinates of the midpoint are
[0072] By performing a topological sort of the point coordinates, the coordinates of the top-view image directly map to the actual spatial arrangement of the target in the XY plane in the physical world, allowing for the sorting of the point set. Set the sorting rules: sort in ascending order based on the pixel horizontal coordinate u; if the u values are similar, sort by the vertical coordinate v.
[0073] After sorting, the serialized points are numbered (ID: 1-N).
[0074] The process of using spatial topological sorting is as follows: After obtaining the IDs (ID: 1-N), the fundamental matrix between the top-view camera 4 and the side-view camera m, which was obtained through pre-calibration, is used. ,Will A directed epipolar line is generated by projecting the image onto the side-view camera m. The target in the side-view camera m will be on this polar line.
[0075] If we look at the polar line of camera m from the side... If only one target is extracted, the matching is completed directly. If the epipolar line... If multiple overlapping targets are crossed, epipolar ambiguity occurs, and the system will adopt the "sequence consistency constraint" rule in topological sorting coding.
[0076] In this embodiment, the fundamental matrix of camera 1 and camera 4 is used. The first one extracted from camera 1 The center point of each steel column is projected onto the image of camera 4, generating an epipolar line. The corresponding real steel column in the image of camera 4 must fall within a strip-shaped area of ±1 pixel on both sides of this epipolar line. If there are multiple overlapping steel columns on this epipolar line, the global XY axis topology recorded by camera 4 (view from above) is retrieved and sorted for filtering, successfully locking in the unique and correct matching point pair.
[0077] In the stereo matching process, epipolar geometric constraints are used to determine the matching of candidate points. The specific steps are as follows: (1) Generation of polar lines: Set up a camera Points in In the camera The polar lines generated in the middle are ,set up The parameters of the linear equation are ,in , and The equation of the straight line representing the polar line is automatically obtained through matrix multiplication. .
[0078] (2) Calculate the polar distance residual: Candidate Points to the poles The orthogonal geometric distance d must satisfy: , (3) Match success judgment: If Then it is believed and Match successful; The polar distance residual. In an ideal mathematical model, the corresponding point should fall precisely on the epipolar line. However, in the physical world, the nonlinear distortion residuals of the camera lens and the centroid shift caused by reflections from the target will cause the real point to deviate from the epipolar line. This is equivalent to widening the epipolar lines to ensure successful matching. However, it cannot be too wide, as this will lead to extreme ambiguity. In this embodiment, a resolution of 640×480 is set... Pixel.
[0079] S4. Based on several pairs of two-dimensional points that have been successfully matched with epipolar lines, solve for the initial three-dimensional coordinates of several target objects; The relevant operations in step S4 will be introduced with specific examples: For each successfully matched pair of two-dimensional points (such as cameras) of and camera of The overdetermined equation system is established using the Direct Linear Transform (DLT) method: for each pair of two-dimensional points with successfully matched epipolar lines, the position of the target object on the camera is recorded. and camera The subpixel coordinates of the center of the photographed target object. Using the intrinsic and extrinsic parameters of each camera, the target object is positioned within the camera's field of view. and camera A linear constraint relationship is established between the sub-pixel coordinates of the center of the captured target object and the three-dimensional position of the target object in the world coordinate system. Each pair of matching points provides multiple constraints, which are combined to form an overdetermined system of linear equations.
[0080] The above overdetermined equations were solved using the Singular Value Decomposition (SVD) method to obtain the initial three-dimensional coordinates of several target objects in the world coordinate system. , This represents the index value of a two-dimensional point pair that successfully matches the epipolar line. For the first The initial three-dimensional coordinates of the two-dimensional point pairs that are successfully matched with epipolar lines are calculated; for each set of matched point pairs, a set of initial three-dimensional coordinates is obtained by solving.
[0081] S5. Based on several successfully matched epipolar point pairs, obtain the angle between the corresponding camera's line-of-sight vector and the target object's line-of-sight vector. ; The relevant operations in step S5 will be introduced with specific examples: The method described in this invention uses four cameras to generate multiple sets of simultaneous three-dimensional coordinates for the same target object. .
[0082] Because the intersection angles of different camera combinations are different, the error uncertainties also differ. Based on several successfully matched epipolar 2D point pairs, the angle between the line-of-sight vectors from the corresponding cameras to the target object is obtained. .
[0083] In this embodiment, for the target object (steel column), three different three-dimensional spatial coordinate values were calculated by combining three sets of cameras: camera 1-camera 4 (angle approximately 45°), camera 2-camera 4 (angle approximately 45°), and camera 1-camera 2 (angle between side views 60°). .
[0084] S6, based on Given the initial 3D coordinates of several target objects, introduce fusion weights to reconstruct the 3D coordinates of the target objects.
[0085] The relevant operations in step S6 will be described using specific examples: The method described in this invention introduces a weighted fusion mechanism based on uncertainty to calculate the final three-dimensional coordinates of the target object. : Among them, the fusion weight This strategy dynamically suppresses the high-noise coordinates caused by ill-conditioned intersection perspectives, enabling convergence of reconstruction accuracy.
[0086] In this embodiment, since the angle between camera 1 and camera 2 is 60°, which is closest to the optimal intersection angle (sin60°≈0.866), it is given the highest weight; while the intersection angle of the side view and top view combination is smaller (sin45°≈0.707), and is given a lower weight. The weights are determined as follows: 1. Spatial geometric field of view weights:
[0087] Representing the The spatial intersection angle of the principal optical axes of the two cameras at the target point determines the "structural uncertainty" during triangulation reconstruction.
[0088] Based on the properties of the sine function, when the intersection angle of the two cameras... As it gets closer to 90°, The closer the value is to 1, the greater the weight, and at this point the error ellipse converges to a circle with a minimum value; when the two cameras are almost parallel... →0° or 180°, →0, at which point a very large deep blind zone will be generated, and the system will automatically reduce its weight to the minimum.
[0089] 2. Image bottom-level matching confidence weights:
[0090] Let the absolute error of the polar distance be... The maximum allowable polar tolerance threshold of the system is Then the normalization error When the extraction is very precise ,but Preserve geometric weights; if the error is extremely large ,but Remove this set of coordinates to prevent the introduction of noise.
[0091] Specific values in this embodiment In this embodiment, the four cameras are paired up to produce a maximum of [number missing] images. Solution .
[0092] 1. Geometric intersection angle Optimal value Adjacent side-view camera combinations (e.g., side-view 1 and side-view 2): The physical angle constraint is 60°. At this time... Its geometric weight It provides stable base weights.
[0093] For a pair of side-view cameras (such as Side View 1 and Side View 3): the physical angle spans both sides, which is 120°. At this time... Its geometric weight .
[0094] Combination of top-down and side-view views (e.g., top-down view 4 and any side-view): The camera is set up vertically downwards for the top-down view and tilted for the side-view view; their intersection angle is typically between 60° and 90°. At this time... It remained stable in the range of 0.866 to 1.0.
[0095] 2. Normalized residuals Value constraints of the implementation examples In this embodiment, the selected absolute threshold for epipolar distance tolerance is... Pixel.
[0096] For the The orthogonal distance of the polar lines calculated by the group : Its normalization formula is: (and when) (At that time, it will not participate in the integration).
[0097] In this embodiment, the epipolar distance of high-quality matching pairs is used in the actual sub-pixel extraction of a gray steel column (6mm long, 6mm in diameter). It's usually around 0.2 to 0.3 pixels. Substituting into the formula gives... Then the confidence weight .
[0098] Will Multidimensional fusion was performed using a weighted fusion formula. The absolute physical reconstruction errors of the final output steel column's 3D center coordinates in the X, Y, and Z directions were measured using standard gauge blocks. This achieved efficient, rapid, and high-precision 3D reconstruction closed-loop verification of small clustered targets at low resolution.
[0099] The error convergence effect after the above uncertainty weighted fusion is as follows: Figure 4 As shown, by Figure 4It is known that relying solely on binocular combinations will result in inaccurate depth geometric information during triangulation. Due to the small angle between the lines of sight, the minute noise extracted from pixels is amplified, causing the probability distribution of the reconstructed coordinates to exhibit an elongated "pre-fusion spatial uncertainty ellipse," with the major axis of the ellipse pointing in the depth direction along the Z-axis.
[0100] Convergence effect of the error ellipse: Multi-view dynamic weighting ensures optimal constraints on uncertainty in the three orthogonal dimensions of X, Y, and Z space. The elongated error ellipse eventually converges to an "optimal weighted fusion center point" at the real target location. The system avoids the depth blind spot caused by a single edge viewpoint and outputs optimal three-dimensional coordinates with isotropy and minimal variance. .
Claims
1. A fast 3D reconstruction method for small targets from multiple perspectives based on stereo calibration, characterized in that, The method includes the following steps: S1. Set up 4 cameras to simultaneously photograph the stereo calibration object and obtain the sub-pixel coordinates of the feature corner points of each calibration surface of the stereo calibration object; S2. Based on the sub-pixel coordinates obtained in step S1, solve for the intrinsic and extrinsic parameters of each camera respectively, and construct the relative rotation matrix and relative translation vector between any two cameras based on the extrinsic parameter matrix of each camera. S3. Simultaneously photograph the target object with 4 cameras to obtain the sub-pixel coordinates of the center of the target object photographed by each camera. Then, based on the relative rotation matrix and relative translation vector between any two cameras, perform epipolar matching on the sub-pixel coordinates of the center of the target object photographed by each camera to obtain several pairs of two-dimensional points that are successfully epipolar matched. S4. Based on several pairs of two-dimensional points that have been successfully matched with epipolar lines, solve for the initial three-dimensional coordinates of several target objects; S5. Based on several successfully matched epipolar 2D point pairs, obtain the angle between the corresponding camera's line-of-sight vector and the target object's line-of-sight vector. ; S6, based on Given the initial 3D coordinates of several target objects, introduce fusion weights to reconstruct the 3D coordinates of the target objects.
2. The method for rapid 3D reconstruction of small targets from multiple perspectives based on stereo calibration according to claim 1, characterized in that, In step S1, the first camera, the second camera, and the third camera are positioned on the side of the stereo calibration object, and the fourth camera is positioned directly above the stereo calibration object. Each calibration surface of the three-dimensional calibration object is provided with a checkerboard pattern and a corresponding ArUco code mark; The four cameras simultaneously acquire several images of the stereo calibration object. Each image is traversed, and the corresponding calibration surface is identified through ArUco encoding. Then, the sub-pixel coordinates of the feature corner points of the corresponding calibration surface are extracted using a sub-pixel corner detection algorithm.
3. The method for rapid 3D reconstruction of small targets from multiple perspectives based on stereo calibration according to claim 2, characterized in that, Construct a world coordinate system and obtain the three-dimensional coordinates of the characteristic corner points of each calibration surface of the three-dimensional calibration object; For any feature corner point on any calibration surface of a stereo calibration object, establish perspective projection constraint equations for sub-pixel coordinates and three-dimensional coordinates, thereby obtaining several 3D-2D matching point pairs; Based on several 3D-2D matching point pairs, solve for the intrinsic parameter matrix of each camera. and extrinsic parameter matrix , Indicates the first The rotation matrix of the camera relative to the world coordinate system. Indicates the first The translation vector of the camera relative to the world coordinate system; The relative rotation matrix between any two cameras is: ,in, Indicates the first The index value of the camera, Indicates the first The index value of the camera, , , , Indicates the first Camera coordinate system to the first The rotation matrix of the camera coordinate system. Indicates the first The rotation matrix of the camera relative to the world coordinate system. No. The transpose of the rotation matrix of the camera relative to the world coordinate system; The relative translation vector between any two cameras is ,in, Indicates the first Camera coordinate system to the first The translation vector of the camera coordinate system. Indicates the first The translation vector of the camera relative to the world coordinate system. Indicates the first The translation vector of the camera relative to the world coordinate system.
4. The method for rapid 3D reconstruction of small targets from multiple perspectives based on stereo calibration according to claim 3, characterized in that, The polar matching includes the following steps: S31, Calculate the first Subpixel coordinates of the center of the target object captured by the camera In the Polar relation on a camera ,in, Indicates camera and camera The fundamental matrix between them; S32, Calculate the first Subpixel coordinates of the center of the target object captured by the camera arrive set distance ; S33, if Then judge and A pair of two-dimensional points that are successfully matched by epipolar lines is considered a successful match; otherwise, the match is unsuccessful. This represents the polar distance residual; S34. Repeat steps S31 to S33 to traverse all camera combinations and obtain several two-dimensional point pairs with successful epipolar matching.
5. The method for rapid 3D reconstruction of small targets from multiple perspectives based on stereo calibration according to claim 4, characterized in that, The initial three-dimensional coordinates of several target objects are obtained by solving for several pairs of two-dimensional points that have been successfully matched with epipolar lines, including the following steps; S41. Based on each pair of two-dimensional points that are successfully matched with epipolar lines, establish an overdetermined system of equations using the direct linear transformation method. S42. Using the singular value decomposition method, solve the overdetermined system of equations to obtain the initial three-dimensional coordinates of several target objects. , This represents the index value of a two-dimensional point pair that has successfully matched the epipolar line. For the first The initial three-dimensional coordinates of the two-dimensional point pairs that have been successfully matched with the epipolar lines.
6. The method for rapid 3D reconstruction of small targets from multiple perspectives based on stereo calibration according to claim 5, characterized in that, The fusion weights are specifically as follows: .
7. The method for rapid 3D reconstruction of small targets from multiple perspectives based on stereo calibration according to claim 6, characterized in that, The formula for calculating the 3D coordinates of the reconstructed target object is: ,in, These are the final three-dimensional coordinates of the target object.