Single-frame high-precision multi-view reconstruction system and method based on curvature perception adaptive window
The single-frame high-precision multi-view reconstruction system based on curvature-aware adaptive windows solves the shortcomings of existing technologies in high-precision and high-completeness reconstruction, and achieves high-precision measurement in high and low curvature regions, especially in high-precision measurement in industrial fields, achieving an accuracy of about 0.06mm.
Patent Information
- Application Number
- CN202511446500.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-11
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2045-10-11
AI Technical Summary
Existing multi-view matching methods have shortcomings in high-precision and high-completeness reconstruction, especially in high-precision measurements in industrial fields. Existing methods are difficult to achieve high-precision measurements, and existing speckle projection profilometry and digital image correlation algorithms have significant errors in areas with large curvature.
A high-precision multi-view reconstruction system based on curvature-aware adaptive windows is adopted, which includes modules for multi-camera acquisition, camera calibration, multi-view matching, and curvature-aware window size selection. The system calculates a comprehensive score by considering local plane fitting error, the consistency of assumed oblique plane normal vectors, and the cost of multi-view aggregation matching, and then adaptively selects the window size for multi-view reconstruction.
It achieves non-contact measurement with high accuracy and speed, and can achieve high measurement accuracy in both high and low curvature areas. It is also robust to occlusion and shadows, and can achieve high integrity measurement of about 0.06 mm at a physical resolution of 4 pixels/mm.
Smart Images

Figure CN120912686B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of optical reconstruction technology, specifically relating to a single-frame high-precision multi-view reconstruction system and method based on curvature-aware adaptive windows. Background Technology
[0002] Optical measurement is a key technology in geometric measurement research and industrial applications, including biomedical devices, shape measurement, and industrial manufacturing. In optical reconstruction tasks, the goal is to achieve high-precision and high-completeness point clouds while ensuring the adaptability of the reconstruction algorithm. Existing multi-view matching methods mostly focus on high-precision and high-completeness reconstruction of large-scale scenes. Direct matching based on object surface texture significantly reduces reconstruction accuracy, making it difficult to use for high-precision measurements in industrial settings. Existing speckle projection profilometry for shape reconstruction is inefficient, and digital image correlation algorithms with small matching windows lead to significant errors. In regions with high curvature, digital image correlation algorithms with large matching windows result in larger reconstruction errors. Therefore, a new matching method and a curvature-aware adaptive window size selection algorithm are needed to achieve high-precision measurement. Many researchers have studied cost-summarization methods and window size selection techniques. Most methods rely on empirical selection of the matching window size using auxiliary information in the image, and few are developed based on the surface morphological characteristics of the object, which limits the improvement of accuracy.
[0003] In applications where high measurement accuracy is required, such as industrial workpiece measurement and reverse engineering, existing methods are lacking in terms of accuracy and completeness in surface morphology measurement. Summary of the Invention
[0004] To address the problems existing in the prior art, this invention proposes a single-frame high-precision multi-view reconstruction system and method based on curvature-aware adaptive windows, which achieves high-precision and high-completeness measurement of a single frame.
[0005] To achieve the above objectives, the present invention provides the following solution:
[0006] A single-frame high-precision multi-view reconstruction system based on curvature-aware adaptive windows includes:
[0007] The multi-camera acquisition module is used to acquire single-frame images of the workpiece under test using multiple cameras;
[0008] The camera calibration module is used to acquire the intrinsic and extrinsic parameters of the multi-view camera and to perform distortion correction on the single-frame image.
[0009] The multi-view matching module is used to calculate the multi-view aggregation matching cost of a single frame image after distortion correction using the multi-view stereo pipeline algorithm, and obtain the pixel-by-pixel hypothetical oblique plane of each single frame image;
[0010] The curvature-aware window size selection module is used to obtain a comprehensive score for curvature-aware adaptive window size selection based on local plane fitting error, consistency of local hypothetical oblique plane normal vector, and the multi-view aggregation matching cost.
[0011] The multi-view reconstruction module is used to allocate matching window sizes based on the comprehensive score, and to iterate the multi-view stereo pipeline algorithm using the allocation results of the matching window sizes to obtain a depth map and complete the multi-view reconstruction.
[0012] Preferably, the camera calibration module includes:
[0013] The monocular camera calibration unit is used to capture the calibration plate image of the monocular camera and obtain the intrinsic parameters of the monocular camera and the initial radial and tangential distortion coefficients using the monocular camera calibration algorithm.
[0014] The multi-camera calibration unit is used to construct a multi-camera group using a monocular camera and simultaneously capture multiple frames of calibration board images of the multi-camera group. The extrinsic parameters of the multi-camera group are obtained through calibration. Then, the intrinsic parameters, extrinsic parameters, and initial radial and tangential distortion coefficients of the monocular camera in the multi-camera group are jointly optimized using the beam adjustment algorithm to obtain the optimal solution.
[0015] The distortion correction unit is used to perform distortion correction on the single frame image based on the optimal solution.
[0016] Preferably, the multi-view matching module includes:
[0017] An initialization unit is used to randomly initialize an initial hypothetical inclined plane with normal vector and depth values for each pixel of a single frame image after distortion correction, within a preset depth range and normal vector direction range.
[0018] The matching cost calculation unit is used to take a single frame image as a reference image and the remaining single frame images as source images, and use the parameters of the initial assumed oblique plane to calculate the homography matrix between the current pixel of the reference image and the pixel of the source image to obtain the multi-view aggregation matching cost.
[0019] An optimization unit is used to optimize the multi-view aggregation matching cost by utilizing the geometric consistency between the reference image and the source image, and to take the hypothetical oblique plane with the minimum multi-view aggregation matching cost as the final hypothetical oblique plane.
[0020] Preferably, in the matching cost calculation unit, the process of calculating the multi-view aggregation matching cost includes:
[0021] The calculated weights of each source image are obtained based on the weights obtained from the binocular matching cost between the pixels with the optimal matching cost in the multiple neighborhoods around the pixels of the reference image and the corresponding pixels in each source image, as well as the preset good view selection set and the set of pixel indices in the source image that meet the preset matching cost.
[0022] The multi-view aggregation matching cost is obtained based on the calculated weights of each source image and the binocular matching cost between each source image and the reference image at the current reference pixel.
[0023] Preferably, the curvature perception window size selection module includes:
[0024] The sampling unit is used to perform local neighborhood adaptive sampling on the black pixels of the reference image using an eight-directional ring sequential sampling strategy to obtain a set of local neighborhood spatial points.
[0025] The scoring calculation unit is used to calculate a comprehensive score for curvature-aware adaptive window size selection by utilizing the fitting plane error of the local neighborhood spatial point set, the angle between the assumed oblique plane normal vector corresponding to each pixel in the neighborhood and the assumed oblique plane normal vector corresponding to the center pixel, and the difference between the minimum and maximum matching window multi-view aggregation matching costs.
[0026] Preferably, the scoring calculation unit includes:
[0027] The first scoring subunit is used to perform spatial plane fitting using the local neighborhood spatial point set, calculate the fitting standard deviation of the spatial plane, and obtain a window size score based on the standard deviation using the fitting standard deviation.
[0028] The second scoring subunit is used to smooth the normal vector of the sampling point by using the assumed slope plane normal vector around the sampling point, and to calculate the average normal vector angle between the smoothed sampling point normal vector and the assumed slope plane normal vector corresponding to the center pixel; and to calculate the window size score based on the average normal vector angle using the average normal vector angle.
[0029] The third scoring subunit is used to aggregate the matching cost using a multi-view method with a predefined minimum and maximum matching window to obtain a window size score based on the matching cost.
[0030] The comprehensive score calculation subunit is used to obtain a comprehensive score for curvature-aware adaptive window size selection by utilizing window size scores based on standard deviation, window size scores based on the average normal vector angle, and window size scores based on matching cost.
[0031] This invention also provides a single-frame high-precision multi-view reconstruction method based on curvature-aware adaptive windows for implementing the system, comprising:
[0032] A multi-view camera is used to acquire a single-frame image of the workpiece under test;
[0033] Obtain the intrinsic and extrinsic parameters of the multi-view camera and perform distortion correction on the single-frame image;
[0034] The multi-view aggregation matching cost of a single frame image after distortion correction is calculated using the multi-view stereo pipeline algorithm, and the pixel-wise hypothetical oblique plane of each single frame image is obtained.
[0035] Based on the local plane fitting error, the consistency of the local hypothetical oblique plane normal vector, and the multi-view aggregation matching cost, a comprehensive score for curvature-aware adaptive window size selection is obtained.
[0036] Based on the comprehensive score, the matching window size is allocated, and the allocation result of the matching window size is used to iterate the multi-view stereo pipeline algorithm to obtain the depth map and complete the multi-view reconstruction.
[0037] Preferably, the method for distortion correction of the single-frame image includes:
[0038] Capture the calibration plate image of the monocular camera, and use the monocular camera calibration algorithm to obtain the intrinsic parameters of the monocular camera and the initial radial and tangential distortion coefficients;
[0039] A multi-camera group is constructed using a monocular camera, and multiple frames of calibration plate images of the multi-camera group are captured synchronously. The extrinsic parameters of the multi-camera group are calibrated and obtained. Then, the intrinsic parameters, extrinsic parameters, and initial radial and tangential distortion coefficients of the monocular camera in the multi-camera group are jointly optimized using the beam adjustment algorithm to obtain the optimal solution.
[0040] Based on the optimal solution, distortion correction is performed on the single-frame image.
[0041] Compared with existing technologies, the beneficial effects of this invention are as follows: The technical solution of this invention achieves non-contact measurement with high accuracy, high speed, good versatility, and the ability to perceive the curvature of the object under test and perform single-frame measurement. It achieves a high measurement accuracy that is not found in existing methods, adapting to both high and low curvature regions, while also being robust to occlusion and shadows (multi-view vision). It achieves high-completeness measurement with an accuracy of approximately 0.06 mm within a field of view with a physical resolution of 4 pixels / mm. Attached Figure Description
[0042] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0043] Figure 1This is a schematic diagram of an assumed inclined plane in an embodiment of the present invention;
[0044] Figure 2 This is a schematic diagram of the two-color checkerboard grid division and iteration mode according to an embodiment of the present invention;
[0045] Figure 3 This is a schematic diagram of local neighborhood adaptive sampling in an embodiment of the present invention;
[0046] Figure 4 This is a schematic diagram of the single-frame high-precision multi-view reconstruction system based on curvature-aware adaptive windows according to an embodiment of the present invention. Detailed Implementation
[0047] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0048] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0049] Example 1:
[0050] like Figure 4 As shown, the single-frame high-precision multi-view reconstruction system based on curvature-aware adaptive window includes: a multi-camera acquisition module, a camera calibration module, a multi-view matching module, a curvature-aware window size selection module, and a multi-view reconstruction module.
[0051] The multi-camera acquisition module is used to acquire single-frame images of the workpiece under test using multiple cameras. Specifically, the camera's exposure time is adjusted to a reasonable value, and a single-frame image of the workpiece under test is acquired using the multi-camera acquisition module with photolithographic digital speckle projection (the number of cameras is variable, this is a general method).
[0052] The camera calibration module is used to acquire the intrinsic and extrinsic parameters of the multi-view camera and to perform distortion correction on the single-frame image.
[0053] A further embodiment includes a camera calibration module comprising:
[0054] The monocular camera calibration unit is used to capture calibration plate images of the monocular camera and obtain the intrinsic parameters and initial radial and tangential distortion coefficients of each monocular camera using a monocular camera calibration algorithm. In this embodiment, 30 calibration plate images are captured for each camera, covering various imaging positions and working distances.
[0055] The specific calibration algorithm for a monocular camera is as follows: The calibration board is either a checkerboard or a circular spot calibration board. In the acquired images, the black and white corner points of the checkerboard or the center points of the circular spots are obtained using a corner detection algorithm (checkerboard) and an ellipse edge extraction algorithm (circular spots), respectively. This point set is of sub-pixel accuracy. A world coordinate system (xyz) is constructed using the predefined corner size information on the calibration board. Calibration is performed using the same world coordinate system for all poses. The world coordinates of the corner points and the detected sub-pixel coordinates of the image are jointly input into the `calibrateCamera` function in OpenCV for monocular calibration to obtain the camera's intrinsic parameters. With distortion coefficient .
[0056] The multi-camera calibration unit is used to construct a multi-camera group using a monocular camera and capture multiple calibration board images for each multi-camera group. Similar to monocular calibration, a world coordinate system for the calibration board is established. Simultaneously, sub-pixel sets of markers on the calibration board are extracted from the multi-camera images, and a one-to-one correspondence between sub-pixel points is established between cameras. One camera is considered as the reference camera, and the stereoCalibration algorithm in OpenCV is used to calculate the binocular extrinsic parameters of the other cameras and the reference camera (the intrinsic parameters of the monocular camera and the initial radial and tangential distortion coefficients are used as initial inputs). The beam adjustment algorithm is then used to jointly optimize the intrinsic and extrinsic parameters, as well as the initial radial and tangential distortion coefficients of the multi-camera group, to obtain the optimal solution. The specific cost function formula for the beam adjustment algorithm is as follows:
[0057] ,
[0058] By iteratively optimizing the camera's intrinsic and extrinsic parameters and distortion parameters, the above cost function is solved to minimize the overall reprojection error, thereby achieving beam adjustment. The spatial coordinates are considered constant (unchanging) because the calibration plate is manufactured with sufficiently high precision. For reprojected pixels, For the detected pixels, The number of calibration board poses, where n is the number of cameras. This represents the set of rotation matrices for each camera in the world coordinate system. This represents the set of translation vectors for each camera in the world coordinate system. The spatial coordinates of the marker point reconstructed at the i-th position of the calibration plate are represented. represent The components in represent The components in.
[0059] Specifically, multiple cameras are treated as a group, and 30 sets of images of the calibration board pose (determined by the camera arrangement) are captured using the method in the monocular camera calibration unit (the calibration board is in the common field of view of at least two cameras).
[0060] The distortion correction unit is used to correct distortion in a single frame of an image based on the optimal solution.
[0061] The multiview matching module is used to calculate the multiview aggregation matching cost of a single frame image after distortion correction using the multiview stereo pipeline algorithm, and obtain the pixel-by-pixel hypothetical oblique plane of each single frame image.
[0062] A further implementation method includes a multi-view matching module comprising:
[0063] The initialization unit is used to initialize an initial hypothetical oblique plane with normal vector and depth values for each pixel of a single frame image after distortion correction, within a preset depth range and normal vector direction range; specifically, pixel-by-pixel random spatial plane initialization, such as... Figure 1 As shown.
[0064] The matching cost calculation unit uses a single-frame image as a reference image and the remaining single-frame images as source images. It calculates the homography matrix between the current pixel of the reference image and the pixels of the source images using parameters of the initially assumed oblique plane to achieve sub-pixel bounding box correspondence, thus obtaining the multi-view aggregated matching cost. A further implementation involves the following process in the matching cost calculation unit:
[0065] The calculated weights of each source image are obtained based on the weights obtained from the binocular matching cost between the pixels with the optimal matching cost in the multiple neighborhoods around the reference image pixels and the corresponding pixels in each source image, as well as the preset good view selection set and the pixel index set in the source image that meets the preset matching cost.
[0066] Based on the calculated weights of each source image and the stereo matching cost between each source image and the reference image at the current reference pixel, the multi-view aggregation matching cost is obtained. :
[0067] ,
[0068] Where n-1 is the number of source images, Represents the source image. The weights for each source image can be expressed as: ,in This represents the pixel with the optimal matching cost within its multiple neighborhoods surrounding the reference pixel. The weights obtained from the stereo matching cost between each source image, This represents the set of good view selections. A good view is a set of pixels in each neighborhood that has the best matching cost and satisfies more than [a certain condition]. The cost of matching is less than ,less than The cost of matching is greater than (The lower the matching cost, the more accurate the match). and The cost threshold constant is set in advance. Indicates less than The total number of matching costs, Let represent the set of pixel indices with lower matching cost in the j-th source image. The stereo matching cost at the current reference pixel is calculated using a bilinear normalized cross-correlation algorithm, which is applied between each reference image and the source image. Finally, the hypothetical oblique plane parameters of the center pixel are updated to the plane parameters with the lowest multi-view aggregation matching cost. For example... Figure 2 As shown.
[0069] The optimization unit is used to optimize the multi-view aggregation matching cost by utilizing the geometric consistency between the reference image and the source image, and takes the hypothetical slope plane with the minimum multi-view aggregation matching cost as the final hypothetical slope plane.
[0070] Specifically, after multiple iterations, a geometric consistency check between the reference image and the source image is used to improve the consistency among multiple views, as shown in the following formula:
[0071] ,
[0072] in This represents the reprojection error between the reference image and the source image at the reference pixel. This represents the balance constant between geometric consistency and photometric consistency in space. Finally, a planar parameter dithering method is used to calculate more multi-view aggregation matching costs, obtaining more accurate oblique plane parameters in a larger solution space. Ultimately, the oblique plane parameter with the minimum multi-view aggregation matching cost is taken as the hypothetical oblique plane corresponding to that pixel.
[0073] The curvature-aware window size selection module is used to obtain a comprehensive score for curvature-aware adaptive window size selection based on the assumed oblique plane and the cost of multi-view aggregation matching.
[0074] A further implementation method includes a curvature-sensing window size selection module comprising:
[0075] The sampling unit is used to perform local neighborhood adaptive sampling of black pixels in the reference image using an eight-directional circular sequential sampling strategy to obtain a set of local neighborhood spatial points; specifically, the eight-directional circular sequential sampling strategy is used, such as... Figure 3As shown. Specifically, starting from the first black pixel at the top, black pixels are sampled sequentially in eight directions clockwise. High-confidence matching is achieved through the following condition, thus realizing high-confidence spatial points. reconstruction:
[0076] ,
[0077] in This represents the total number of reference images plus source images that satisfy the reprojection error threshold at the sampling points. This represents a 3×3 matrix formed by the first three columns of the projection matrix of the k-th camera. This represents a 3×1 matrix formed by the last column of the projection matrix of the k-th camera. Represents the k-th image in pixels The depth value at that location. Satisfies... Pixels meeting the condition greater than 1 will be recorded, while pixels not meeting the condition will be discarded. The number of validly recorded pixels will reach a threshold. The search stops when the center point is reached. Using the above formula, the calculation is performed including the center point. The spatial coordinates of each pixel, using the reference camera coordinate system. Each spatial coordinate represents a structured region of the neighborhood surrounding the center point, and is fitted into a spatial tilted plane using the least squares method.
[0078] The scoring calculation unit is used to calculate a comprehensive score for curvature-aware adaptive window size selection by utilizing the fitting plane error of the local neighborhood spatial point set, the angle between the hypothetical oblique plane normal vector corresponding to each pixel in the neighborhood and the hypothetical oblique plane normal vector corresponding to the center pixel, and the difference between the minimum and maximum matching window multi-view aggregation matching costs.
[0079] A further embodiment of the invention includes a scoring calculation unit comprising:
[0080] The first scoring subunit is used to fit a spatial plane using a set of local neighborhood points, calculate the standard deviation of the fit, and obtain a window size score based on the standard deviation. Specifically, it is assumed that the standard deviation of the plane fit is used... This represents the window size score derived from the standard deviation of the plane fit. As shown in the following formula:
[0081] ,
[0082] in Let be the physical resolution of the camera in the current field of view. To ensure the universality of the method, These are predefined constant coefficients.
[0083] The second scoring subunit, in order to reduce the impact of single-point noise on the calculation of the normal vector angle, uses the assumed inclined plane normal vector around the sampling point to smooth the normal vector of the sampling point, and calculates the average normal vector angle between the smoothed sampling point normal vector and the assumed inclined plane normal vector corresponding to the center pixel; The angle values are arranged in ascending order, and the window size score is calculated based on the average normal vector angle.
[0084] The average angle between the normal vectors is given by the following formula:
[0085] ,
[0086] in Let be the angle between the smoothed normal vector of the o-th sampling point and the assumed oblique plane normal vector corresponding to the center pixel, expressed in radians. Based on this, the window size score derived from the average normal vector angle can be expressed by the following formula:
[0087] ,
[0088] in It is a constant coefficient.
[0089] The third scoring subunit is used to aggregate the matching costs across multiple views using predefined minimum and maximum matching windows to obtain a window size score based on the matching cost:
[0090] ,
[0091] in and These are the multi-view aggregation matching costs calculated using predefined minimum and maximum matching windows, respectively. and These are predefined constant coefficients.
[0092] The comprehensive score calculation subunit is used to obtain a comprehensive score for curvature-aware adaptive window size selection by utilizing window size scores based on standard deviation, average normal vector angle, and matching cost. After calculating the above three scores, the comprehensive score formula can be expressed as follows:
[0093] , The preset constant weight representing the first scoring sub-unit. The preset constant weight representing the second scoring sub-unit. The preset constant weight represents the third scoring sub-unit.
[0094] when and One of them is higher than the preset matching cost threshold constant. The comprehensive scoring formula is rewritten as follows:
[0095] , This represents the preset constant weight of the first scoring sub-unit in this case. This represents the preset constant weight of the second scoring sub-unit under this condition.
[0096] The multi-view reconstruction module is used to allocate matching window sizes based on comprehensive scores, and to iterate the multi-view stereo pipeline algorithm using the allocation results of matching window sizes to obtain depth maps and complete multi-view reconstruction.
[0097] Specifically, set the minimum matching window size. With maximum matching window size and predefined Window size levels, setting a lower score limit With the upper limit of scores When the score is below the lower limit, the minimum matching window size is assigned to that pixel; when the score is above the upper limit, the maximum matching window size is assigned to that pixel. When the score is between the two, the window size level is proportional to the curvature perception comprehensive score.
[0098] After one or two iterations of the multi-view stereo pipeline algorithm using an adaptive window size distribution obtained through curvature-aware scoring, multiple depth maps corresponding to multiple images are obtained. Depth map fusion is achieved by using a consistency check method between the reference view and each source image at pixel-by-pixel, including reprojection error, relative depth value difference, and normal vector consistency, resulting in a high-precision and highly complete point cloud.
[0099] This invention achieves non-contact measurement with high accuracy, speed, and versatility. It can perceive the curvature of the object under test and perform single-frame measurement, achieving a high measurement accuracy that is not found in existing methods, adapting to both high and low curvature regions. It is also robust to occlusion and shadows (multi-view vision). With a field of view of physical resolution of 4 pixels / mm, it achieves high-completeness measurement with an accuracy of approximately 0.06mm.
[0100] Example 2:
[0101] A single-frame high-precision multi-view reconstruction method based on curvature-aware adaptive windows, used in the system applying Embodiment 1, includes:
[0102] A multi-view camera is used to acquire a single-frame image of the workpiece under test.
[0103] Obtain the intrinsic and extrinsic parameters of the multi-view camera and perform distortion correction on the single-frame image.
[0104] The multi-view aggregation matching cost of a single frame image after distortion correction is calculated using the multi-view stereo pipeline algorithm, and the pixel-by-pixel hypothetical oblique plane of each single frame image is obtained.
[0105] Based on the local plane fitting error, the consistency of the local hypothetical oblique plane normal vector, and the multi-view aggregation matching cost, a comprehensive score for curvature-aware adaptive window size selection is obtained.
[0106] Based on the comprehensive score, the matching window size is allocated, and the allocation result of the matching window size is used to iterate the multi-view stereo pipeline algorithm to obtain the depth map and complete the multi-view reconstruction.
[0107] A further embodiment of the method for distortion correction of the single-frame image includes:
[0108] Capture the calibration plate image of the monocular camera, and use the monocular camera calibration algorithm to obtain the intrinsic parameters of the monocular camera and the initial radial and tangential distortion coefficients.
[0109] A multi-camera group is constructed using a monocular camera, and multiple frames of calibration plate images of the multi-camera group are captured synchronously. The extrinsic parameters of the multi-camera group are calibrated and obtained. Then, the intrinsic parameters, extrinsic parameters, and initial radial and tangential distortion coefficients of the monocular camera in the multi-camera group are jointly optimized using a beam adjustment algorithm to obtain the optimal solution.
[0110] Based on the optimal solution, distortion correction is performed on the single-frame image.
[0111] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.
Claims
1. A single-frame high-precision multi-view reconstruction system based on curvature-aware adaptive windows, characterized in that, include: The multi-camera acquisition module is used to acquire single-frame images of the workpiece under test using multiple cameras; The camera calibration module is used to acquire the intrinsic and extrinsic parameters of the multi-view camera and to perform distortion correction on the single-frame image. The multi-view matching module is used to calculate the multi-view aggregation matching cost of a single frame image after distortion correction using the multi-view stereo pipeline algorithm, and obtain the pixel-by-pixel hypothetical oblique plane of each single frame image; The curvature-aware window size selection module is used to obtain a comprehensive score for curvature-aware adaptive window size selection based on local plane fitting error, consistency of local hypothetical oblique plane normal vector, and the multi-view aggregation matching cost. The multi-view reconstruction module is used to allocate matching window sizes based on the comprehensive score, and to iterate the multi-view stereo pipeline algorithm using the allocation results of the matching window sizes to obtain a depth map and complete the multi-view reconstruction. The multi-view matching module includes: An initialization unit is used to randomly initialize an initial hypothetical inclined plane with normal vector and depth values for each pixel of a single frame image after distortion correction, within a preset depth range and normal vector direction range. The matching cost calculation unit is used to take a single frame image as a reference image and the remaining single frame images as source images, and use the parameters of the initial assumed oblique plane to calculate the homography matrix between the current pixel of the reference image and the pixel of the source image to obtain the multi-view aggregation matching cost. An optimization unit is used to optimize the multi-view aggregation matching cost by utilizing the geometric consistency between the reference image and the source image, and to take the hypothetical oblique plane with the minimum multi-view aggregation matching cost as the final hypothetical oblique plane.
2. The system according to claim 1, characterized in that, The camera calibration module includes: The monocular camera calibration unit is used to capture the calibration plate image of the monocular camera and obtain the intrinsic parameters of the monocular camera and the initial radial and tangential distortion coefficients using the monocular camera calibration algorithm. The multi-camera calibration unit is used to construct a multi-camera group using a monocular camera and simultaneously capture multiple frames of calibration board images of the multi-camera group. The extrinsic parameters of the multi-camera group are obtained through calibration. Then, the intrinsic parameters, extrinsic parameters, and initial radial and tangential distortion coefficients of the monocular camera in the multi-camera group are jointly optimized using the beam adjustment algorithm to obtain the optimal solution. The distortion correction unit is used to perform distortion correction on the single frame image based on the optimal solution.
3. The system according to claim 1, characterized in that, The matching cost calculation unit includes the following process for calculating the multi-view aggregation matching cost: The calculated weights of each source image are obtained based on the weights obtained from the binocular matching cost between the pixels with the optimal matching cost in the multiple neighborhoods around the pixels of the reference image and the corresponding pixels in each source image, as well as the preset good view selection set and the set of pixel indices in the source image that meet the preset matching cost. The multi-view aggregation matching cost is obtained based on the calculated weights of each source image and the binocular matching cost between each source image and the reference image at the current reference pixel.
4. The system according to claim 1, characterized in that, The curvature-sensing window size selection module includes: The sampling unit is used to perform local neighborhood adaptive sampling on the black pixels of the reference image using an eight-directional ring sequential sampling strategy to obtain a set of local neighborhood spatial points. The scoring calculation unit is used to calculate a comprehensive score for curvature-aware adaptive window size selection by utilizing the fitting plane error of the local neighborhood spatial point set, the angle between the assumed oblique plane normal vector corresponding to each pixel in the neighborhood and the assumed oblique plane normal vector corresponding to the center pixel, and the difference between the minimum and maximum matching window multi-view aggregation matching costs.
5. The system according to claim 4, characterized in that, The scoring calculation unit includes: The first scoring subunit is used to perform spatial plane fitting using the local neighborhood spatial point set, calculate the fitting standard deviation of the spatial plane, and obtain a window size score based on the standard deviation using the fitting standard deviation. The second scoring subunit is used to smooth the normal vector of the sampling point by using the assumed slope plane normal vector around the sampling point, and to calculate the average normal vector angle between the smoothed sampling point normal vector and the assumed slope plane normal vector corresponding to the center pixel; and to calculate the window size score based on the average normal vector angle using the average normal vector angle. The third scoring subunit is used to aggregate the matching cost using a multi-view method with a predefined minimum and maximum matching window to obtain a window size score based on the matching cost. The comprehensive score calculation subunit is used to obtain a comprehensive score for curvature-aware adaptive window size selection by utilizing window size scores based on standard deviation, window size scores based on the average normal vector angle, and window size scores based on matching cost.
6. A single-frame high-precision multi-view reconstruction method based on curvature-aware adaptive windows, used to implement the system described in any one of claims 1-5, characterized in that, include: A multi-view camera is used to acquire a single-frame image of the workpiece under test; Obtain the intrinsic and extrinsic parameters of the multi-view camera and perform distortion correction on the single-frame image; The multi-view aggregation matching cost of a single frame image after distortion correction is calculated using the multi-view stereo pipeline algorithm, and the pixel-wise hypothetical oblique plane of each single frame image is obtained. Based on the local plane fitting error, the consistency of the local hypothetical oblique plane normal vector, and the multi-view aggregation matching cost, a comprehensive score for curvature-aware adaptive window size selection is obtained. Based on the comprehensive score, the matching window size is allocated, and the allocation result of the matching window size is used to iterate the multi-view stereo pipeline algorithm to obtain the depth map and complete the multi-view reconstruction.
7. The method according to claim 6, characterized in that, The method for distortion correction of the single-frame image includes: Capture the calibration plate image of the monocular camera, and use the monocular camera calibration algorithm to obtain the intrinsic parameters of the monocular camera and the initial radial and tangential distortion coefficients; A multi-camera group is constructed using a monocular camera, and multiple frames of calibration plate images of the multi-camera group are captured synchronously. The extrinsic parameters of the multi-camera group are calibrated and obtained. Then, the intrinsic parameters, extrinsic parameters, and initial radial and tangential distortion coefficients of the monocular camera in the multi-camera group are jointly optimized using the beam adjustment algorithm to obtain the optimal solution. Based on the optimal solution, distortion correction is performed on the single-frame image.
Citation Information
Patent Citations
High-dynamic surface single-frame measurement method based on multi-view digital speckle correlation method
CN119268592A
Systems and Methods for Targetless Auto-calibration and Depth Estimation
US20250173886A1