Multi-frame superposition point cloud super-resolution system for three-dimensional reconstruction
By generating multi-frame point cloud sets and point-level confidence sets, pose registration and complementary sampling label output are performed, solving the problem of retaining complementary sampling points in multi-frame super-resolution point cloud, and achieving a more continuous point cloud distribution and more stable reconstruction results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-03
- Publication Date
- 2026-03-31
AI Technical Summary
In 3D reconstruction, existing technologies struggle to effectively preserve complementary sampling points during multi-frame super-resolution of point clouds, resulting in insufficient resolution improvement and even the smoothing out of edges and microstructures.
By generating a multi-frame point cloud set and a point-level confidence set, pose registration and complementary sampling labels are performed, and complementary point sets, outlier point sets, and point sets to be corrected are output. Based on the labels, super-resolution point cloud results are generated through overlay fusion and consistency verification.
It achieves cleaner, denser point clouds with more continuous point distribution in detailed areas such as edges and grooves, reducing the number of rework scans. The output super-resolution point cloud results are suitable for direct entry into the size inspection and surface reconstruction process.
Smart Images

Figure CN121767571A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of structured light 3D scanning, and more specifically, to a multi-frame overlay point cloud super-resolution system for 3D reconstruction. Background Technology
[0002] In industrial measurement and 3D reconstruction using structured light 3D scanning, a common practice to make point clouds finer and closer to the real surface is to repeatedly scan the same target, align the point clouds obtained from multiple scans in the same coordinate system, and then superimpose and fuse them. This utilizes the minute sampling differences between different scans to supplement details. Existing technologies, such as the "A 3D Point Cloud Acquisition Method and Structured Light Scanning Device" (application number CN202411803899.0), propose acquiring images at multiple time points and performing pose estimation and matching. The results from adjacent time points are then used for connected component partitioning and collision detection to filter point clouds, eliminating false point clouds caused by mismatches and thus improving point cloud reliability.
[0003] However, when the goal shifts from obtaining reliable point clouds to achieving point cloud super-resolution through multi-frame overlay, the above screening approach raises a more subtle but crucial issue: super-resolution overlay truly requires complementary sampling points that appear in a particular frame, do not completely overlap with adjacent frames, but genuinely originate from real surfaces. However, adjacent-time screening often treats inconsistencies with neighboring frames directly as errors, eliminating them through connected component sorting, collision detection, deletion, or merging. As a result, while the remaining point clouds are more consistent and smoother, they primarily contain repetitive information, lacking new samples that can improve resolution. Ultimately, this manifests as an increase in the number of overlays without a clearer understanding of details, and even the gradual smoothing of edges and microstructures. This problem is not easily detected initially because, intuitively, deleting inconsistencies is usually correct in traditional point cloud processing. However, in multi-frame overlay point cloud super-resolution systems, it is precisely necessary to retain truly valuable complementary sampling points while eliminating false point clouds; otherwise, the overlay technique will struggle to achieve the desired resolution improvement.
[0004] To address the aforementioned problems, a technical solution is provided. Summary of the Invention
[0005] To overcome the aforementioned deficiencies of the prior art, embodiments of the present invention provide a multi-frame overlay point cloud super-resolution system for 3D reconstruction. This system generates a multi-frame point cloud set by acquiring multiple frames of structured light point clouds of the same target and simultaneously generates a point-level confidence set. The multi-frame point cloud set is then registered to obtain pose transformation results, generating a unified coordinate point cloud set and a point-level registration residual set. Complementary sampling labels are generated on the unified coordinate point cloud set based on fringe phase reproducibility parameters and line-of-sight layer separation parameters, and complementary point sets, outlier set, and point set to be corrected are output. Adhesion and constrained fusion are performed based on the complementary sampling labels to obtain the super-resolution point cloud result. Finally, the point-level confidence set and complementary sampling labels are updated through back-projection consistency verification, and the verified and constrained super-resolution point cloud result is output, thus solving the problems mentioned in the background art.
[0006] To achieve the above objectives, the present invention provides the following technical solution: Acquisition and construction unit: Acquire multiple frames of structured light point cloud of the same target, generate a set of multiple frame point cloud, and simultaneously generate a set of point-level confidence scores corresponding to each point in the set of multiple frame point cloud; Pose registration unit: Registers a multi-frame point cloud set to obtain the pose transformation results from each frame to the reference frame, generates a unified coordinate point cloud set, and simultaneously generates a point-level registration residual set corresponding to each point in the unified coordinate point cloud set. Complementary discrimination unit: Performs complementary sampling discrimination on a unified coordinate point cloud set, generates complementary sampling labels based on imaging stability and geometric consistency, and outputs complementary point set, outlier point set, and point set to be corrected; Layered fusion unit: Fusion is performed based on complementary sampling labels; outlier sets are stopped from entering the fusion input; complementary point sets participate in densification fusion; and point sets to be corrected participate in restricted fusion to obtain super-resolution point cloud results. The feedback verification unit performs consistency verification on the super-resolution point cloud results, updates the point-level confidence set and complementary sampling labels based on the verification results, and outputs the verified and constrained super-resolution point cloud results.
[0007] Furthermore, each frame of structured light image sequence for the same target is decoded to generate depth information, phase information, validity information, and calculate fringe modulation. Based on the validity information, 3D points are generated by combining depth information with camera intrinsic parameters and calibration extrinsic parameters, written into a multi-frame point cloud set, and bound to frame identifiers and pixel coordinates. Based on the fringe modulation and phase information, decoded stable components are generated and geometrically continuous components are generated by combining spatial neighborhood. These components are fused to obtain a point-level confidence set and correspond point-by-point with the multi-frame point cloud set.
[0008] Furthermore, the reliability score of each frame is calculated based on the point-level confidence set, and the reference frame point cloud is determined; the registration point set and the reference point set are selected from the point cloud of the frame to be registered and the reference frame point cloud based on the spatial partition and the point-level confidence set; the initial coarse alignment value is generated based on the centroid alignment and principal direction alignment of the registration point set, and the initial pose relationship between the point cloud of the frame to be registered and the reference frame point cloud is formed.
[0009] Furthermore, based on the initial coarse alignment value, a set of corresponding point pairs that are closest to each other is constructed; a local plane is fitted in the corresponding neighborhood of the reference frame point cloud and the point-to-plane residual is calculated; the set of interior points is determined based on the abrupt change position of the residual sorting and the pose transformation result is iteratively solved; a unified coordinate point cloud set is generated based on the pose transformation result and the normal projection distance is calculated and written into the point-level registration residual set.
[0010] Furthermore, in the unified coordinate point cloud set, spatial nearest neighbor points are constructed for each point to be judged, and a set of support points is selected based on the point-level confidence set. A local reference surface is fitted based on the set of support points. Based on the pose transformation results, the point to be judged is projected back onto the multi-frame imaging plane, and phase information, fringe modulation, and validity information are extracted along the fringe direction. The phase coherence is calculated, and fringe phase reproducibility parameters are generated.
[0011] Furthermore, in the unified coordinate point cloud set, a line-of-sight bucket is constructed based on the line-of-sight direction of the reference frame and a depth sequence is formed. Based on the layer segmentation position of the depth sequence, the thickness length of the main layer and the length of the interlayer gap are calculated and the line-of-sight interlayer separation parameters are generated. Based on the maximum difference boundary of the neighborhood between the fringe phase reproducibility parameter and the line-of-sight interlayer separation parameter, complementary sampling labels are generated and complementary point sets, outlier point sets, and sets of points to be corrected are output.
[0012] Furthermore, a fused surface set is established based on the grid cell division of the unified coordinate point cloud set. The fused surface set includes the center coordinates and normals of the fused surface cells. Each grid cell is divided by maximum difference according to the point-level confidence set and a stable point set is selected. The stable point set is used to calculate the center coordinates and normals of the fused surface cells and write them into the fused surface set.
[0013] Furthermore, after the complementary point set is mapped to the fused surface set, the normal signed distance is calculated and the attached coordinates are generated. The attached coordinates are segmented according to the maximum gap of the normal signed distance and the fused surface set is updated with zero-crossing principal clusters. The point set to be corrected is segmented into close clusters and conflict clusters and conflict buffer records are generated. The abnormal point set stops entering the fusion input, and the fused surface set is used to export the super-resolution point cloud results.
[0014] Furthermore, the super-resolution point cloud results are projected back onto the multi-frame imaging plane based on the pose transformation results to generate a set of projected pixel landing points. The set of projected pixel landing points is combined with imaging range determination, validity information determination, and occlusion determination by comparing predicted depth and observed depth to generate observable markers. Depth information is read at the corresponding positions of the observable markers to form the line-of-sight difference and difference rate. The fringe modulation and phase information are read to form phase coherence and combined with validity information to generate decoding consistency evidence.
[0015] Furthermore, the single-frame support score is synthesized from the decoded consistent evidence and the geometric suppression term corresponding to the difference rate. After sorting the single-frame support scores of multiple frames, a stable support set is screened out by the position of the largest adjacent difference and a consistent support is generated. At the same time, the geometric conflict frequency is generated from the difference rate sequence and a verification result is formed. The verification result is used to update the point-level confidence set and complementary sampling labels and output the super-resolution point cloud result with verification constraints.
[0016] The technical effects and advantages of this invention for a multi-frame overlay point cloud super-resolution system for 3D reconstruction: 1. During the acquisition phase, a point-level confidence set is generated synchronously from the multi-frame point cloud set. During the registration phase, a point-level registration residual set is generated synchronously. During the complementary sampling discrimination phase, complementary sampling labels are generated using the fringe phase reproducibility parameter and the line-of-sight layer separation parameter. In the field operation, the decoding jump points of reflective boundaries and the mismatched ghost points of occluded edges can be pre-separated into abnormal point sets and point sets to be corrected. The complementary point sets are concentrated for supplementary sampling and densification. Scanners do not need to repeatedly delete points based on experience to obtain a cleaner densified point cloud. The point distribution in detailed areas such as edges and grooves is more continuous, and the number of rework scans is reduced with the marking of conflict areas.
[0017] 2. In the overlay and fusion stage, the fused surface set carries the surface skeleton. The complementary point set is attached and merged, and then the fused surface set is updated. The point set to be corrected is written into the conflict buffer record in a restricted fusion manner without pulling the surface. In the consistency verification stage, the super-resolution point cloud results are projected back onto the multi-frame imaging plane to form verification results and update the point-level confidence set and complementary sampling labels. In field use, the densified points are spread out on the surface without bulging a thick shell. The conflict area retains clear markings of the point set to be corrected for easy verification. The output super-resolution point cloud results with empirical constraints are more suitable for direct entry into the size inspection, defect observation and surface reconstruction process. Operators can still maintain stable reconstruction results when facing complex materials. Attached Figure Description
[0018] Figure 1 This is a schematic diagram of the multi-frame overlay point cloud super-resolution system for 3D reconstruction according to the present invention. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] Example 1: Figure 1 The present invention provides a multi-frame overlay point cloud super-resolution system for 3D reconstruction, comprising: Acquisition and construction unit: Acquire multiple frames of structured light point cloud of the same target, generate a set of multiple frame point cloud, and simultaneously generate a set of point-level confidence scores corresponding to each point in the set of multiple frame point cloud; Pose registration unit: Registers a multi-frame point cloud set to obtain the pose transformation results from each frame to the reference frame, generates a unified coordinate point cloud set, and simultaneously generates a point-level registration residual set corresponding to each point in the unified coordinate point cloud set. Complementary discrimination unit: Performs complementary sampling discrimination on a unified coordinate point cloud set, generates complementary sampling labels based on imaging stability and geometric consistency, and outputs complementary point set, outlier point set, and point set to be corrected; Layered fusion unit: Fusion is performed based on complementary sampling labels; outlier sets are stopped from entering the fusion input; complementary point sets participate in densification fusion; and point sets to be corrected participate in restricted fusion to obtain super-resolution point cloud results. The feedback verification unit performs consistency verification on the super-resolution point cloud results, updates the point-level confidence set and complementary sampling labels based on the verification results, and outputs the verified and constrained super-resolution point cloud results.
[0021] When structured light 3D reconstruction enters the multi-frame overlay point cloud super-resolution process, the multi-frame point cloud set needs to simultaneously carry geometric coordinates and acquisition semantics, while the point-level confidence set needs to simultaneously carry decoding confidence and geometric confidence. The structured light acquisition chain naturally includes grayscale changes in the projection sequence, the coherence of the phase decoding process, and depth validity information. If this information is directly solidified into point-level markers during the acquisition stage, subsequent steps can achieve stable registration and filtering without adding extra sensors or complex parameter adjustments. However, structured light point clouds are prone to local jumps and flying points at reflections, occlusions, and edges. Without a point-level confidence set, multiple frames of point clouds will experience an equal mixing effect when overlaying, affecting the convergence of subsequent pose transformation results and the usability of the unified coordinate point cloud set.
[0022] S101 structured light decoding and validity information generation.
[0023] The structured light image sequence consists of multiple projection numbers of the same frame. The projection numbers correspond to different projection patterns or phase steps, and the pixel coordinates maintain a one-to-one correspondence in the projection sequence.
[0024] Structured light decoding first extracts the grayscale intensity sequence along the projection sequence direction for each pixel location. This grayscale intensity sequence is used to generate phase information, which is generated using the decoding rules of the projection sequence to obtain the phase value, and phase quality cues are output simultaneously. Depth information is obtained by combining the phase information with the calibration relationship, and the output is the depth value corresponding to each pixel location. Validity information is generated simultaneously with depth information. Validity information indicates whether the pixel location meets the decoding conditions, which include the grayscale intensity sequence exhibiting significant fringe changes and not being in a saturated or near-saturated state, and the phase information showing no obvious abrupt changes. Fringe modulation is calculated at the same pixel location. The calculation method is to take the maximum grayscale intensity minus the minimum grayscale intensity in the projection sequence as the brightness difference, then take the sum of the maximum and minimum grayscale intensities as the reference brightness, and add a positive number close to zero to the reference brightness to avoid division by zero. Finally, the brightness difference is divided by the processed reference brightness to obtain the fringe modulation. The fringe modulation is written as a dimensionless quantity into a record table consistent with the pixel coordinates, and can be referenced by the point-level confidence set to maintain a consistent physical meaning across frames. Stripe modulation can reflect stripe imaging quality without relying on geometric neighborhood. Pixels that fall on low texture or edges can still obtain effective decoding cues, thus reducing the dependence of subsequent steps on geometric statistics.
[0025] S102 generates 3D points and writes them into a multi-frame point cloud set.
[0026] Depth information, validity information, and camera intrinsic parameters are used together to determine the 3D coordinate reconstruction relationship. 3D point generation is only performed on pixel positions marked as valid by the validity information to avoid invalid depth information from entering the multi-frame point cloud set.
[0027] The pixel x-coordinate is first subtracted from the principal point x-coordinate, then divided by the lateral focal length to obtain the normalized lateral imaging scale. This normalized lateral imaging scale is then multiplied by the depth value to obtain the lateral coordinate in the camera coordinate system. Similarly, the pixel y-coordinate is first subtracted from the principal point y-coordinate, then divided by the lateral focal length to obtain the normalized lateral imaging scale. This normalized lateral imaging scale is then multiplied by the depth value to obtain the lateral coordinate in the camera coordinate system. The depth value is directly used as the forward coordinate in the camera coordinate system. The calibration extrinsic parameters are used to transform the 3D coordinates of the camera coordinate system to the point cloud output coordinate system. The transformation result, along with the frame identifier and pixel coordinates, is written into a multi-frame point cloud set. Each point in the multi-frame point cloud set retains its frame identifier and pixel coordinates, enabling the pose registration unit to perform pose solving for different frames, and the complementary discrimination unit to re-project the image frame by frame onto the decoding plane to read phase information and fringe modulation.
[0028] The 3D point generation stage is limited to pixel locations with valid information, which can reduce the dense mixing of invalid points caused by saturation, occlusion, and phase jump from the source, and improve the usable density of multi-frame point cloud sets.
[0029] S103 decodes stable clues and writes them into a point-level confidence set.
[0030] The point-level confidence set needs to provide confidence clues on the decoding side before the geometric neighborhood is stably established, and the fringe modulation and phase information coherence of structured light just meet this requirement.
[0031] Decoding stable components uses validity information as a gate to avoid invalid pixels generating artificially high confidence levels. Phase coherence is calculated within the fringe direction neighborhood, determined by the projection pattern direction. Neighboring pixels are selected along the fringe direction, maintaining consistent lengths. The phase coherence calculation first calculates the phase difference by subtracting the phase value of the center pixel from the phase value of the neighboring pixels. The phase difference is then cosine-shifted, resulting in a near-one response when the phase difference approaches zero and a near-negative-one response when the phase difference approaches a jump. The cosine response is then shifted and scaled to a range of zero to one, resulting in high values in continuous regions and low values in fragmented regions. Finally, the zero-to-one responses within the fringe direction neighborhood are averaged to obtain the phase coherence of the center pixel.
[0032] The decoding stable component is obtained by multiplying validity information, fringe modulation, and phase coherence. When the validity information is invalid, the decoding stable component is zero; when the validity information is valid, the decoding stable component changes along with both fringe modulation and phase coherence. The decoding stable component is written into the point-level confidence set, corresponding one-to-one with the points in the multi-frame point cloud set. Each entry retains both the frame identifier and pixel coordinates, ensuring that the pose registration unit and the complementary discrimination unit can reuse the same index relationship. The decoding stable component merges fringe imaging quality and phase coherence at the point level, enabling the priority identification of unstable points on the decoding side when processing edges, reflective areas, and occluded regions, thereby reducing false deletions and false inclusions caused by relying solely on geometric neighborhood judgments.
[0033] S104 geometrically reasonable clues are fused to generate point-level confidence set entries.
[0034] In addition to stable cues on the decoding side, the point-level confidence set also requires geometric continuity cues to identify isolated flying points and local surface discontinuities. The calculation of geometrically continuous components is completed within the current frame's point cloud. The neighborhood point set is obtained through spatial nearest neighbor search, which is performed within the same frame's point cloud and excludes depth-invalid points. Local normals are generated using a multi-candidate consistency screening method to avoid normal jitter caused by a single three-point selection. The multi-candidate normals are generated by repeatedly selecting two non-collinear neighborhood points from the neighborhood point set, forming two vectors with the center point respectively, performing a cross product on the two vectors to obtain candidate normals, and then normalizing the candidate normals to unit length.
[0035] The consistency screening method for multiple candidate normals is to project the difference vector between each neighboring point in the neighborhood point set and the center point onto the candidate normal direction, and the absolute value of the projection is used as the normal deviation of the point from the local surface; the neighboring points with smaller normal deviations are regarded as support points; the candidate normal with the most support points is selected as the local normal.
[0036] The neighborhood scale length is calculated by calculating the Euclidean distance from the center point to each neighboring point and taking the median of all distances.
[0037] The normal deviation scale length is calculated by taking the median of the normal deviations of all neighboring points.
[0038] The geometric continuous component is calculated by taking the neighborhood scale length as the numerator and the neighborhood scale length plus the normal deviation scale length plus a positive number close to zero as the denominator. The result is a geometric continuous component in the range of zero to one. The smaller the normal deviation scale length, the closer the geometric continuous component is to one. The larger the normal deviation scale length, the closer the geometric continuous component is to zero.
[0039] Point-level confidence is obtained by multiplying the decoded stable component and the geometrically continuous component. The point-level confidence is written into the point-level confidence set and maintains a one-to-one correspondence with the points in the multi-frame point cloud set. The geometrically continuous component achieves cross-scene adaptation through length dimension scaling. When the neighborhood density changes, the neighborhood scale length automatically changes. The normal deviation scale length of isolated flying points becomes significantly larger, thus naturally being suppressed in the point-level confidence. This facilitates the pose registration unit to prioritize geometrically continuous and decoded stable points for pose solving during registration.
[0040] In one embodiment, after a multi-frame structured light scanner is aligned with the same target, the operator keeps the relative position of the scanner and the target unchanged, allowing the structured light pattern to be continuously projected onto the target surface. The camera simultaneously acquires each frame of structured light image sequence and completes decoding. The decoding process directly outputs depth information and validity information. The depth information, combined with camera intrinsic parameters and calibration extrinsic parameters, is converted into three-dimensional coordinates to form a frame of point cloud. Multiple consecutive frame point clouds are sequentially written into a multi-frame point cloud set. Simultaneously, each point in each frame of point cloud records a point-level confidence set. The point-level confidence set writes validity information, fringe modulation, and phase coherence as decoding stability cues, and local neighborhood continuity and isolation as geometric rationality cues. The operator sees a denser point cloud in the target's edges and groove areas, and a decrease in valid points at the boundaries of reflective areas. These phenomena correspond to the recording differences of different points in the point-level confidence set.
[0041] After the multi-frame point cloud set is completed, each point in the set carries both a frame identifier and pixel coordinates. Each entry in the point-level confidence set corresponds point-to-point with the multi-frame point cloud set and carries both a decoded stable component and a point-level confidence score. The pose registration unit reads the multi-frame point cloud set to generate pose transformation results, and reads the point-level confidence set for reference frame selection and point set filtering for registration constraints. The index relationship between these two types of objects is fixed in the acceptance construction unit and will not undergo meaning drift in subsequent steps. The structured light multi-frame overlay point cloud super-resolution process thus obtains a set of directly callable point-level reliability markers, providing stable input semantics for complementary sampling discrimination and hierarchical fusion.
[0042] After the data acquisition and construction unit completes the multi-frame point cloud set and the point-level confidence set, the multi-frame information from repeated structured light scanning already has point-level reliability markers. However, the risk of superposition misalignment due to inconsistencies in inter-frame coordinates still exists. Multi-frame super-resolution of the point cloud relies on minute sampling differences to supplement details. Pose errors can amplify these minute differences into ghosting and thick shells, and complementary sampling labels can also be misled, leading to misclassification. However, the point-level confidence set already provides the decoding stability and geometric continuity of each point. Based on this, the pose registration unit selects reference frames and constructs stable correspondences, outputting pose transformation results, a unified coordinate point cloud set, and a point-level registration residual set, providing consistent coordinates and residual evidence for the complementary discrimination unit.
[0043] S201 Reference Frame Determination and Reference Frame Point Cloud Establishment.
[0044] The stripe imaging quality and decoding stability of different frames in a multi-frame point cloud set vary, and the reference frame needs to represent a relatively stable observation state.
[0045] The frame reliability score is calculated once for each frame. The calculation process first takes all point-level confidence scores belonging to that frame from the point-level confidence set and sums them. Then, the number of points in the frame's point cloud is used as the divisor, and the sum is divided by the number of points to obtain the frame reliability score. The reference frame is determined by the frame with the highest frame reliability score. The reference frame point cloud is obtained by indexing multiple frame point clouds according to the reference frame identifier. The reference frame point cloud has a higher proportion of stable points, more continuous local plane fitting, more stable subsequent point-to-surface residual calculations, and makes it easier to maintain detail consistency in pose transformation results.
[0046] S202 Confidence-guided point set selection and coarse alignment initial value generation.
[0047] The initial coarse alignment value needs to avoid unstable and isolated points in the decoding process, and it also needs to cover the target surface to avoid establishing alignment relationships only in local areas.
[0048] The registration point set selection process first divides the point cloud of the frame to be registered into multiple spatial partitions according to spatial location. The spatial partitions are obtained by uniformly dividing the bounding range of the target point cloud. Within each spatial partition, the point with the highest point confidence is selected from the point-level confidence set as the representative point. The reference point set selection process adopts the same spatial partitioning rules and selects representative points within the reference frame point cloud.
[0049] The initial value generation process for coarse alignment first calculates the centroids of the registration point set and the reference point set. The centroid is obtained by summing the point coordinates item by item and dividing by the number of points. The initial value for coarse translation is obtained by subtracting the centroid of the registration point set from the centroid of the reference point set. The initial value for coarse rotation is obtained by aligning the principal directions. The process of obtaining the principal directions involves first subtracting the centroid of the point set from the coordinates of all points in the point set to form centroid-free coordinates, then constructing a scatter matrix. Each element of the scatter matrix is obtained by averaging the product of the corresponding components of the centroid-free coordinates over the entire point set. The scatter matrix is then decomposed to obtain three mutually orthogonal principal directions. The initial value for coarse rotation aligns the principal directions of the registration point set with the principal directions of the reference point set one by one. Spatial partitioning ensures that the registration point set covers the target surface, making the initial value for coarse alignment less susceptible to local anomalies, and allowing the fine registration iteration to enter the stable convergence region more quickly.
[0050] S203 uses the closest correspondences to construct and finely register the pose transformation results iteratively.
[0051] Super-resolution of multi-frame overlaid point clouds is sensitive to minor mismatches. Fine registration requires suppressing spurious correspondences caused by occlusion boundary mismatches and thick shell ghosting. The process of constructing the corresponding point pair set first transforms the point cloud of the frame to be registered to the reference coordinate system according to the initial coarse alignment value. Then, for each transformed point, the nearest point is searched in the reference frame point cloud to form a one-way correspondence. The one-way correspondence reverse verification process searches for the nearest point in the transformed point cloud for each reference frame point and checks whether the nearest point returns to the original one-way corresponding point. Point pairs that satisfy the mutual nearest relationship are retained as the corresponding point pair set.
[0052] The point-to-surface residual calculation process involves constructing a neighborhood point set for each corresponding point's reference endpoint in the reference frame point cloud. The neighborhood point set is obtained by spatial nearest neighbor search and excludes obviously isolated points. The local plane is obtained by least squares fitting of the neighborhood point set, and the local normal is calculated and normalized from the local plane. The point-to-surface residual is obtained by subtracting the reference endpoint from the transformed registration endpoint to obtain the difference vector, and then the projection of the difference vector on the local normal direction is used as the residual value.
[0053] The process of determining the interior point set involves sorting the absolute values of the residuals of all corresponding point pairs in ascending order, and then calculating the ratio sequence of the absolute values of adjacent residuals. To avoid the ratio becoming unstable due to the denominator being close to zero, a positive number close to zero is added to the denominator before calculating the ratio. The adjacent position with the largest ratio is regarded as the abrupt change position, and the corresponding point pairs before the abrupt change position are taken as the interior point set.
[0054] The pose update process minimizes the sum of squares of the point-to-surface residuals on the inlier set to obtain the pose increment. This increment is obtained through linearization. The pose update adds the pose increment to the current pose. Iteration terminates when the translation and rotation components of the pose increment stabilize during continuous iterations. The pose transformation result consists of the final rotation and the final translation, and is written into the pose transformation result. Nearest-to-nearest correspondences suppress one-to-many mismatches, and abrupt position segmentation excludes thick-shell ghosting correspondences from the inlier set, resulting in more consistent pose transformation results in the detail region.
[0055] S204 generates a unified coordinate point cloud set and a point-level registration residual set.
[0056] A unified coordinate point cloud set needs to unify all frame point clouds of a multi-frame point cloud set to a reference coordinate system. A point-level registration residual set needs to fix the fitting error of each point as a point-by-point record to facilitate subsequent complementary sampling and discrimination.
[0057] The unified coordinate point cloud set generation process performs a rigid body transformation on each frame of the point cloud in the multi-frame point cloud set according to the pose transformation result. The rigid body transformation first multiplies the point coordinates by rotation and then adds translation. The transformed points are written into the unified coordinate point cloud set and the original frame identifier and pixel coordinates are retained.
[0058] The point-level registration residual set generation process searches for a neighborhood point set in the reference frame point cloud for each point in the unified coordinate point cloud set. If the neighborhood point set satisfies the conditions of sufficient number of points and non-degenerate point distribution, local plane fitting is performed to obtain the unit normal. The normal projection distance is obtained by subtracting the neighborhood geometric center from the point coordinates to obtain the difference vector. The absolute value of the projection of the difference vector onto the unit normal direction is then written as the residual value into the point-level registration residual set. If the neighborhood point set is insufficient or degenerate, the support marker is recorded as unsupported and the residual value is recorded as zero in the point-level registration residual set. The point-level registration residual set expresses the local fitting error using a length quantity. Complementary sampling discrimination can directly utilize the residuals to locate mismatched regions and distinguish between complementary differences and registration deviations.
[0059] In one embodiment, after the multi-frame point cloud set enters the registration stage, the point-level confidence set is used to select the reference frame. The reference frame is the frame with a more complete distribution of effective points and a more coherent point-level confidence set. Each frame to be registered is first downsampled based on the point-level confidence set, retaining points with uniform coverage and higher point-level confidence, and removing obviously isolated points and obviously invalid points. Then, initial alignment is performed, which comes from scanner calibration extrinsic parameters or geometric feature matching from stable regions. After the initial alignment is completed, fine registration is performed. Fine registration forms local surface constraints on the reference frame point cloud, and the point cloud of the frame to be registered searches for matches in the corresponding regions and completes the pose transformation result solution. The pose transformation result transforms the point cloud of the frame to be registered to the reference coordinate system and writes it into a unified coordinate point cloud set. The operator observes on the screen that the overall contour of the target is aligned frame by frame, and the originally scattered contours gradually overlap; at the same time, the point-level registration residual set is recorded. The registration residual is more obvious at occluded edges and in areas with weak texture, and the registration residual is more stable in flat areas.
[0060] After the pose transformation is completed, the unified coordinate point cloud set converts the sampling differences of multiple frames into superimposed information within the same coordinate system, and the point-level registration residual set fixes the local fitting error as point-by-point residual evidence. The complementary discrimination unit reads the unified coordinate point cloud set to complete the local support construction, and reads the point-level registration residual set to identify thick shells and mismatched regions. The complementary discrimination unit then combines the point-level confidence set to calculate the fringe phase reproducibility parameter and the line-of-sight layer separation parameter and generate complementary sampling labels. The registration process of the pose registration unit avoids dependence on manual thresholds, and the nearest correspondence and abrupt exponential segmentation form a stable set of interior points, making the pose transformation results more consistent in the accumulation of details for multi-frame super-resolution point cloud super-resolution.
[0061] After the pose transformation results are unified, the inconsistencies across frames in the unified coordinate point cloud set may originate from true complementary sampling or from ghosting caused by local mismatches. The point-level confidence set has incorporated the stability and geometric continuity of structured light decoding into the point-level tags, and the point-level registration residual set has incorporated the point-by-point fitting error into the residual signal. The complementary discrimination unit introduces two directional pieces of evidence into the unified coordinate point cloud set: fringe direction evidence to characterize imaging reproduction stability, and line-of-sight evidence to characterize layer structure conflicts. These two pieces of evidence are then converted into complementary sampling labels, outputting a complementary point set, an outlier point set, and a point set to be corrected.
[0062] S301 Local Support Construction and Surface Reference Determination.
[0063] The unified coordinate point cloud set has point density variations and a small number of flying points in local areas. If the local reference surface is fitted by random points, the local normal is prone to shift, and the fringe phase reproducibility parameter and the line-of-sight layer separation parameter will drift accordingly.
[0064] In the unified coordinate point cloud set, each point to be judged first retrieves a set of spatial nearest neighbors. The number of spatial nearest neighbors is automatically determined by the point density. The point-level confidence set of each spatial nearest neighbor is read to form a confidence sequence. The confidence sequence is then sorted in ascending order of value. Subsequently, the difference sequence between adjacent confidences is calculated, and the position with the largest difference sequence is used as the split position.
[0065] The nearest neighbor points above the segmentation position form a set of support points, which is used to fit a local reference surface. The fitting process first calculates the geometric center of the support point set, and then calculates the scattering intensity of the support points relative to the geometric center. The weakest scattering intensity in the three directions corresponds to the local normal. After the local normal is normalized, it and the geometric center jointly define the local reference surface.
[0066] The local reference surface provides unified surface orientation constraints in subsequent calculations, and the local support is formed by high confidence points. The local reference surface is more stable at occlusion edges and reflective areas.
[0067] Calculation of phase reproducibility parameters for S302 fringes.
[0068] In structured light, complementary sampling points often exhibit slight positional differences rather than quality jumps across multiple frames. The fringe phase reproducibility parameter needs to characterize the reproduction stability of the same point in the multi-frame decoding link to avoid misjudging real complementary sampling points as outliers.
[0069] For each point to be judged, the pixel landing point is first obtained by using the mapping relationship between the frame identifier and the pixel coordinates to return to the imaging plane of the reference frame. Then, the pixel landing point is projected back to the imaging plane of each frame using the pose transformation result to obtain the sequence of projected pixel landing points. For each frame of projected pixel landing points, a continuous pixel band is taken along the stripe direction. The phase information and stripe modulation degree are read from the continuous pixel band.
[0070] Phase coherence is calculated for each frame by comparing the phase of each pixel in a continuous pixel band with the center phase. The closer the difference is to zero, the higher the coherence response; when the difference jumps, the lower the coherence response. The phase coherence is obtained by averaging the coherence response over the continuous pixel band.
[0071] For each frame, the intra-frame decoding stability score is calculated by using validity information as a gate. When the gate is invalid, the intra-frame decoding stability score is recorded as the lowest. When the gate is valid, the intra-frame decoding stability score is obtained by multiplying the fringe modulation degree and the phase coherence.
[0072] The stable scores of multi-frame intra-frame decoding form a stable score sequence. The stable score sequence is sorted in ascending order of value. Then, the difference sequence between adjacent stable scores is calculated, and the position with the largest difference sequence is taken as the stable frame segmentation position. Frames above the stable frame segmentation position form a stable frame set.
[0073] The fringe phase reproducibility parameter is synthesized from two parts: one part is the proportion of the stable frame set to all frames, and the other part is the average value of the stable scores within the stable frame set. The two parts are multiplied together to obtain the fringe phase reproducibility parameter.
[0074] The stripe phase reproducibility parameter compresses the ability to consistently reproduce multiple frames into a point-level index. The set of stable frames is adaptively given by the maximum difference, and it can still maintain consistent discrimination when the material and exposure change.
[0075] S303 line-of-sight separation parameter calculation.
[0076] Unified coordinate point cloud sets tend to form thick shells or double layers in registration conflict areas. The line of sight is most likely to expose this layer structure separation. The line of sight layer separation parameter needs to be described by length to describe the degree of separation between the main layer and the sub-layer, so as to facilitate its use in conjunction with local surface reference.
[0077] For each point to be judged, the line of sight is first determined based on the imaging model of the reference frame. The line of sight is obtained by normalizing the direction vector from the optical center of the reference frame camera to the point to be judged. The optical center of the reference frame camera is determined by the calibration extrinsic parameters.
[0078] A line-of-sight bucket is constructed around the point to be judged. The line-of-sight bucket is obtained by filtering the spatial nearest neighbor points of the point to be judged. The filtering criterion uses the quantile of the nearest neighbor distance sequence to determine the spatial range. The spatial range is automatically contracted or expanded by the local point density.
[0079] The depth projection distance of each point within the line of sight is calculated along the line of sight direction. The depth projection distances are sorted from near to far to form a depth sequence. The gap sequence between adjacent depth projection distances is calculated, and the position with the largest gap sequence is used as the layer segmentation position.
[0080] The layer segmentation position divides the depth sequence into two layers. The layer with more points is defined as the main layer, and the other layer is defined as the sublayer. The thickness of the main layer is obtained by subtracting the projection distance of the nearest depth from the projection distance of the farthest depth of the main layer. The length of the gap between layers is obtained by taking the positive value of the difference in the projection distance of the depths of the adjacent boundaries of the two layers.
[0081] The line-of-sight separation parameter is formed using a proportional method. The calculation method is to divide the interlayer gap length by the main layer thickness length and add a positive number close to zero to obtain the dimensionless separation degree. Then, the separation degree is corrected for consistency by the ratio of the number of main layer points to the number of line-of-sight bucket points. The corrected result is multiplied back by the main layer thickness length to obtain the line-of-sight separation parameter while maintaining the length dimension.
[0082] The line-of-sight separation parameter expresses the degree of separation between the thick shell and the bilayer as an indicator with clear length significance, and still maintains interpretability when there are local point density changes.
[0083] S304 generates complementary sampling labels and outputs point sets by partitioning two-dimensional evidence.
[0084] The stripe phase reproducibility parameter and the line-of-sight separation parameter need to be jointly decided under the local adaptive boundary. If the boundary is given by a fixed threshold, label drift is likely to occur when the material reflection and occlusion change.
[0085] For each point to be judged, a neighborhood sequence of fringe phase reproducibility parameters and a neighborhood sequence of line-of-sight layer separation parameters are formed within the range of spatial nearest points. The two neighborhood sequences are sorted in ascending order of value, and then adjacent difference sequences are calculated. The positions with the largest difference sequences are used as the imaging stability boundary and the geometric conflict boundary, respectively.
[0086] Complementary sampling label discrimination adopts two-dimensional partitioning. When the fringe phase reproducibility parameter is higher than the imaging stability boundary and the line-of-sight layer separation parameter is lower than the geometric conflict boundary, the complementary sampling label is assigned as a complementary point; when the fringe phase reproducibility parameter is lower than the imaging stability boundary, the complementary sampling label is assigned as an anomaly point; when the fringe phase reproducibility parameter is higher than the imaging stability boundary and the line-of-sight layer separation parameter is higher than the geometric conflict boundary, the complementary sampling label is assigned as a point to be corrected.
[0087] The point-level registration residual set is used to refine the partition of the points to be corrected. The refinement method is to sort the point-level registration residual set within the spatial nearest neighbor range of the points to be corrected and take the position of the maximum difference as the residual boundary. Points with residuals higher than the residual boundary remain as points to be corrected, while points with residuals lower than the residual boundary and whose fringe phase reproducibility parameters meet the imaging stability boundary are transferred to complementary points.
[0088] The complementary point set, the outlier point set, and the point set to be corrected are generated by merging the complementary sampling labels point by point.
[0089] The boundaries of the two-dimensional partitions are all adaptively obtained by the maximum difference of the neighborhood. The complementary point set retains the real complementary sampling information, the abnormal point set contains unstable pseudo-points in the decoding, and the point set to be corrected contains registration conflict area points. The point set division is more in line with the needs of point cloud super-resolution overlay.
[0090] In one embodiment, after the unified coordinate point cloud set enters the complementary sampling discrimination stage, the complementary sampling label no longer relies on a single distance for judgment, but instead processes imaging stability and geometric consistency separately. Each point to be discriminated first selects its spatial nearest neighbors in the unified coordinate point cloud set and filters out a set of support points based on the point-level confidence set. The support point set is then fitted to form a local reference surface. Subsequently, the point to be discriminated is projected back onto the imaging plane of each frame according to the pose transformation result. Phase information and fringe modulation are read along the fringe direction near the projected pixel landing point and combined with validity information to form a fringe phase reproducibility parameter. Simultaneously, a viewing bucket is constructed along the viewing direction of the reference frame for the point to be discriminated. Points within the viewing bucket are sorted by depth and layer segmentation is completed. The thickness length of the main layer and the interlayer gap length are combined to form the viewing layer separation parameter. The fringe phase reproducibility parameter and the viewing layer separation parameter generate complementary sampling labels under the neighborhood segmentation boundary. The complementary sampling labels divide the points into a complementary point set, an anomaly point set, and a point set to be corrected set. Operators observed complementary point clusters near the target edge, with the point cloud becoming finer at the edge details; anomaly point clusters increased at the boundary of the reflective area, and these anomaly point clusters no longer participated in subsequent overlays; and a set of points to be corrected appeared near the occlusion edge, and these points were marked as locations requiring verification.
[0091] The complementary discrimination unit outputs three types of point sets: complementary point set, outlier point set, and point set to be corrected set. These three point sets are directly derived from the point-by-point merging results of the complementary sampling labels. The fringe phase reproducibility parameter and the line-of-sight layer separation parameter maintain consistent spatial semantics under the constraint of the local reference plane. The point-level registration residual set is only used for fine-grained partitioning of the point to be corrected areas, avoiding the direct incorporation of ghosting and mismatched points into the complementary point set. The complementary sampling discrimination results maintain local adaptive boundaries and are supported by both structured light and geometric evidence. When reading the three types of point sets during the layered fusion stage, it is not necessary to re-estimate the reliable attributes of the points.
[0092] After the complementary point set, outlier point set, and point set to be corrected are output by the complementary point discrimination unit, the densification information and conflict information have been separated. However, the complementary point set still needs to be incorporated into the stable surface in a controlled manner. If a surface-supporting structure is lacking during the point cloud overlay stage, the point position update will be driven by local outliers; if a conflict isolation structure is lacking, the point set to be corrected will write the two-layer structure into the final point cloud. The hierarchical fusion unit uses a fusion element set to support the surface skeleton. The complementary point set is incorporated into and updates the fusion element set through attachment. The point set to be corrected enters restricted fusion and forms a conflict buffer record. The outlier point set stops entering the fusion input.
[0093] S401 fusion element set establishment and stable point selection.
[0094] The unified coordinate point cloud set exhibits spatial density differences. The fusion support structure needs to fix the surface orientation within a local area and suppress outlier dragging. The unified coordinate point cloud set is first divided into grid cells in three-dimensional space, and the grid cells are formed by covering cubes of uniform scale. Each grid cell collects the points falling into the grid cell and retains the frame identifier and pixel coordinates.
[0095] Each grid cell reads a set of point-level confidence scores to form a confidence score sequence, which is then sorted from low to high. Adjacent confidence scores are subtracted sequentially to obtain a difference sequence, with the position of the largest difference sequence serving as the split position. Points above the split position form a set of stable points.
[0096] The set of stable points is used to calculate the center coordinates of the merged surface element. The calculation method involves adding the coordinates of all points in the stable point set point by point and then dividing by the number of stable points. The set of stable points is also used to calculate the normal vector of the merged surface element. This is done by repeatedly selecting three points from the stable point set to form two difference vectors; the cross product of these two difference vectors yields candidate normal vectors; the magnitude of these candidate normal vectors is first calculated and normalized; candidate normal vectors with a magnitude lower than the area threshold are discarded. The area threshold is determined by the nearest neighbor distance distribution of the stable point set and given by the quantile points; the remaining candidate normal vectors are then filtered based on their angle consistency with the first candidate normal vector. If the angle exceeds the consistency threshold, the direction is reversed before accumulation. The accumulated normal vectors are then normalized again to obtain the normal vector of the merged surface element.
[0097] The center coordinates of the merged surface element, the normal of the merged surface element, and the grid cell index are all written into the merged surface element set. The merged surface element set provides a stable reference direction in a local range, and subsequent attachment and incorporation are not easily distorted by a small number of outliers.
[0098] S402 Complementary Point Set Attachment and Merging Surface Element Update.
[0099] Directly superimposing complementary point sets introduces decoding noise into the surface along the normal direction. Attaching and incorporating the complementary point sets limits their densification contribution to the local tangential direction, resulting in a smoother surface skeleton update. The complementary point sets are first mapped to the fusion element set according to the mesh cell index, with each complementary point entering the fusion element candidate point set of its corresponding mesh cell; the outlier point set terminates during the mapping stage and enters the fusion input.
[0100] For each candidate point in the merged surface element set, the signed normal distance is calculated point by point. The signed normal distance is obtained by subtracting the projection of the center coordinates of the merged surface element onto the normal direction of the merged surface element from the projection of the complementary point coordinates onto the normal direction of the merged surface element. A positive signed normal distance indicates that the complementary point is located on the side pointed to by the normal direction of the merged surface element, and a negative signed normal distance indicates that the complementary point is located on the opposite side. The attachment coordinates of the complementary points are obtained by subtracting the normal direction of the merged surface element from the coordinates of the complementary points and multiplying by the signed normal distance. The attachment coordinates fall on the plane defined by the center coordinates and the normal direction of the merged surface element.
[0101] The attachment coordinates are grouped into thick-shell suppression groups, which are based on the normal signed distance sequence. The normal signed distances are sorted in ascending order to obtain a sorting sequence; adjacent elements in the sorting sequence are subtracted sequentially to obtain a gap sequence; the position with the maximum gap in the gap sequence is used as the layer split position. The layer split position divides the attachment points into two groups. The main cluster is defined as the group containing normal signed distances that cross zero. Crossing zero means that the sorting sequence contains both positive and negative values within the group. The main cluster is more common on single-layer surfaces.
[0102] The primary cluster attachment coordinates are used to update the center coordinates of the fused surface elements. The update method is to add the primary cluster attachment coordinates point by point and then divide by the number of primary clusters. The primary cluster attachment coordinates are also used to update the normals of the fused surface elements. The update method follows the candidate selection and consistency accumulation process of the three-point cross product of the stable point set. The fused surface element set is also updated with source frame identifier statistics. The source frame identifier statistics record the number of times the primary cluster attachment coordinates appear on different frame identifiers. Subsequent consistency verification can be directly grouped and re-projected according to frame identifiers.
[0103] S403 Restricted fusion and conflict buffer record generation of the set of points to be corrected.
[0104] The set of points to be calibrated is close to the set of complementary points in terms of imaging stability, but exhibits layer structure conflicts in terms of geometric consistency. When the set of points to be calibrated enters the position update, the conflicts are written into the surface. Constrained fusion leaves the conflicts as a buffer rather than forcibly incorporating them. The set of points to be calibrated is mapped to the fusion surface set according to the grid cell index to form the fusion surface buffer point set. The normal signed distance of the fusion surface buffer point set is calculated point by point and sorted by value. Then, the adjacent gaps are calculated and the position of the largest gap is taken as the layer split position.
[0105] The layer segmentation location divides the buffer points into proximity clusters and conflict clusters. Proximity clusters are defined as those containing points whose signed normal distance crosses zero, while conflict clusters are defined as the other group. Proximity clusters do not participate in the update of the fused surface cell center coordinates or the update of the fused surface cell normals; they are only used for sample density recording. Sample density recordings are grouped by source frame identifiers and extracted alternately. The alternating extraction order is determined by the chronological order of the frame identifiers. Each frame identifier retains only a fixed number of representative points in the current fused surface cell. These representative points are determined by sorting the point-level confidence set and selecting points above the segmentation location.
[0106] Conflict clusters are written to conflict buffer records. Conflict buffer records are bound to fusion surface element indexes and retain the residual signals corresponding to the point-level registration residual sets. Conflict buffer records can be directly accessed when locating conflict regions during the consistency verification phase. Conflicts will not be solidified into thick-shell surfaces during the fusion phase.
[0107] S404 super-resolution point cloud result organization and output.
[0108] After the fused element set is updated, it has formed a surface skeleton. It is still necessary to organize the skeleton points and detail enhancement points into a super-resolution point cloud result and maintain traceability. The fused element set outputs the center coordinates of the fused element as skeleton points for each element, the normal of the fused element is written into the skeleton point attributes, and the skeleton point set is written into the super-resolution point cloud result.
[0109] The coordinates of the main cluster of complementary points are used as detail augmentation points and entered into the super-resolution point cloud result. Before the detail augmentation points enter, a consistency screening is performed. The consistency screening reads the point-level confidence set in each fused surface element to form a confidence sequence. After sorting the confidence sequence, the adjacent differences are calculated and the position with the largest difference is taken as the segmentation position. Detail augmentation points above the segmentation position are retained and entered into the super-resolution point cloud result.
[0110] The super-resolution point cloud results simultaneously output source frame identifier statistics and fusion surface element index mapping table. The source frame identifier statistics are used for re-projection groups for consistency verification, and the fusion surface element index mapping table is used to locate the re-projection inconsistencies to specific fusion surface elements. The super-resolution point cloud results maintain a balance between shape stability and detail enhancement while retaining the verification entry point.
[0111] In one embodiment, the overlay and fusion stage transforms the unified coordinate point cloud set into a fused surface set. The fused surface set is built by grid cell, and stable point sets are selected in each grid cell based on point-level confidence sets. The stable point sets calculate the center coordinates and normals of the fused surface cells and write them into the fused surface set. After the complementary point set is mapped to the fused surface set, the signed normal distance is calculated. The complementary point coordinates are attached to the local surface along the normal of the fused surface cell to form attached coordinates. The attached coordinates are segmented according to the signed normal distance to obtain zero-crossing principal clusters. The principal cluster attached coordinates update the fused surface set and are written into the super-resolution point cloud result as detail augmentation points. Abnormal point sets are no longer included in the fusion input. After mapping, the point set to be corrected is divided into proximity clusters and conflict clusters. Proximity clusters are entered into the sample density record, while conflict clusters are written into the conflict buffer record and bound to the point-level registration residual set. Operators observe that the augmented points on the target surface mainly spread along the surface, and no obvious thick shell appears; conflict markers are still retained near the occlusion edges, but the surface position is not pulled out into a double layer.
[0112] After the fusion phase, the set of outliers has stopped entering the fusion input, the set of complementary points is incorporated through attachment to form detail-enhancing points, and the set of points to be corrected is formed into conflict buffer records through constrained fusion. Under the constraints of stable point selection and zero-major-cluster updates, the fused surface set maintains the single-layer surface morphology. The super-resolution point cloud result retains both skeleton points and detail-enhancing points and has traceable index information, resulting in a more stable point cloud super-resolution stacking effect and a clear verification path.
[0113] The super-resolution point cloud results already include the fused surface center sample points and detail-enhancing points, and the splitting results of the complementary point set and the point set to be corrected are also reflected in the output organization. However, structured light is prone to phase jumps at reflection and occlusion edges, and registration is prone to local mismatches in areas with weak texture. Even after multiple frames are superimposed, false details may still be introduced into the super-resolution point cloud results. The re-projection verification unit puts the super-resolution point cloud results back into the multi-frame observation link, generates verification results based on the pose transformation results and the original decoding information, and then writes the verification results back to the point-level confidence set and complementary sampling labels, outputting the verified and constrained super-resolution point cloud results.
[0114] S501 return positioning and observable marker generation.
[0115] Multi-frame consistency verification requires first confirming whether each point has observational conditions in each frame; otherwise, invisible points will be mistakenly identified as inconsistencies. The 3D coordinates of the super-resolution point cloud results are read point by point, and the 3D coordinates are transformed to the camera coordinates of each frame according to the pose transformation results. Then, the camera coordinates are projected onto the imaging plane using the camera intrinsic parameters to obtain the landing point of the returned pixel. When the landing point of the returned pixel falls outside the imaging range, an unobservable mark is written.
[0116] When the projected pixel falls within the imaging range, validity information is read. If the reading result is invalid, an unobservable marker is written; if the reading result is valid, a candidate observable marker is written. Candidate observable markers are then used for occlusion determination, which is performed by comparing the predicted depth with the observed depth. The predicted depth comes from the transformed camera coordinate depth component, while the observed depth comes from the depth information value at the projected pixel's location. If the observed depth is significantly less than the predicted depth, an occlusion-unobservable marker is written; if the observed depth is close to the predicted depth, an observable marker is written.
[0117] Observable markers limit the verification range to the real visible area, making the verification results less susceptible to interference from points outside the field of view and occlusion.
[0118] S502 Single-Frame Consistency Evidence Calculation.
[0119] Single-frame consistency requires simultaneous coverage of geometric consistency and decoding consistency. Geometric consistency reflects whether the point position matches the frame observation, while decoding consistency reflects whether the fringe decoding is stably reproduced. For each observable marker, the depth information of the returned pixel is read to form the observation depth, and the fringe modulation and phase information are read to form phase coherence. Phase coherence is calculated by taking a continuous pixel band along the fringe direction and comparing the difference between the phase of each pixel in the continuous pixel band and the center phase. The closer the difference is to zero, the higher the coherence response. The phase coherence is obtained by averaging the coherence response over the continuous pixel band.
[0120] Evidence of geometric consistency consists of line-of-sight difference and difference rate. Line-of-sight difference is calculated by subtracting the predicted depth from the observed depth and taking the absolute value. Difference rate is calculated by dividing the line-of-sight difference by the absolute value of the predicted depth and adding a small positive number to avoid the denominator being zero. Difference rate is a dimensionless quantity, and the smaller the difference rate, the stronger the geometric consistency.
[0121] The decoding consistency evidence adopts a gated product method. The gate comes from the validity information. When the gate is invalid, the decoding consistency evidence is recorded as zero. When the gate is valid, the fringe modulation and phase coherence are multiplied to obtain the decoding consistency evidence. Both the fringe modulation and phase coherence are dimensionless quantities, and the decoding consistency evidence is also a dimensionless quantity.
[0122] The single-frame support score fuses the difference rate with the decoded consistent evidence. The fusion method is to take the inverse of the difference rate and then perform a natural exponential transformation to obtain a geometric suppression term. The decoded consistent evidence is then multiplied by the geometric suppression term to obtain the single-frame support score. When the single-frame support score is close to zero, it means that the single frame does not support the data. When the single-frame support score is close to the gating level, it means that the single frame supports the data.
[0123] Single-frame support scores unify and compress geometric conflicts and decoding jumps to the same scale, and cross-frame aggregation makes it easier to maintain consistent discrimination criteria.
[0124] S503 cross-frame aggregation generates verification results.
[0125] In multi-frame observations, there are occasional occlusions and local reflections. Directly averaging the results can be biased by a few abnormal frames. The aggregation process needs to suppress extreme values and retain the trend of most frames. For each point, the single-frame support scores corresponding to all observable markers are collected and sorted in ascending order of value. The difference between adjacent values in the sorted sequence is calculated and the position of the maximum difference is found. The position of the maximum difference divides the sorted sequence into a low support segment and a high support segment. The high support segment is defined as the stable support set.
[0126] Consistent support is obtained by averaging the stable support set. Consistent support remains dimensionless and within the effective range; if the stable support set is empty, the consistent support is recorded as the lowest. Geometric conflict frequency is formed by the difference rate sequence. After sorting the difference rate sequence by value, the conflict boundary is formed by the position of the largest adjacent difference. Frames above the conflict boundary are included in the conflict frame set. The proportion of the conflict frame set to the observable frame set forms the geometric conflict frequency.
[0127] The validation results are jointly provided by consistent support and geometric conflict frequency, and the decision boundary is derived from local neighborhood adaptive segmentation. Local neighborhoods are determined under the constraints of the fusion surface index mapping table, and points with the same fusion surface are grouped into the same neighborhood; the consistent support neighborhood sequence forms the support boundary through the position of the largest adjacent difference, and the geometric conflict frequency neighborhood sequence forms the conflict boundary through the position of the largest adjacent difference.
[0128] When the consistent support is higher than the support boundary and the geometric conflict frequency is lower than the conflict boundary, the verification result is recorded as a consistent support state; when the consistent support is lower than the support boundary, the verification result is recorded as a consistent missing state; when the consistent support is higher than the support boundary and the geometric conflict frequency is higher than the conflict boundary, the verification result is recorded as a conflict-preserved state.
[0129] The verification results divide the points into three states and maintain local adaptive boundaries, and can still maintain stable judgment under changes in material and exposure.
[0130] S504 confidence and label updates, and outputs the super-resolution point cloud results with validated constraints.
[0131] The validation results need to be stored as point-level labels to avoid encountering the same type of pseudo-details repeatedly in the next round of processing. The point-level confidence set is updated using log-odds summation. The update process first converts the point-level confidence before the update into log-odds, which is calculated by dividing the confidence by one and then taking the natural logarithm. Then, the consistent support is converted into log-odds, which is also calculated by dividing the consistent support by one and then taking the natural logarithm. The two log-odds are added together to obtain the updated log-odds. The updated log-odds are then converted back into confidence by taking the natural exponent of the updated log-odds, dividing the natural exponent by one, and then adding the natural exponent.
[0132] When the uniform support is close to the boundary value, the conversion will result in numerical instability. Before the conversion, the uniform support is clamped to extreme values. The clamping method is to restrict the uniform support to an open interval far from zero and far from one. The boundary of the open interval is determined by the minimum and maximum values of the stable support set to avoid introducing fixed constants.
[0133] The complementary sampling label update executes migration rules based on the verification result status. Points in the complementary point set that are in a consistent missing state are migrated to the abnormal point set, and points in the point set to be corrected that are in a consistent support state are migrated to the complementary point set. Points in the complementary point set and the point set to be corrected that are in a conflict retention state maintain the point set to be corrected and are written into the conflict buffer record. The conflict buffer record also retains the residual signal of the point-level registration residual set, which is used to locate the conflict concentration area.
[0134] The super-resolution point cloud results output after verification and constraint are labeled and conflict marked. Points corresponding to the abnormal point set are deleted from the super-resolution point cloud results, while points corresponding to the point set to be corrected are retained in the super-resolution point cloud results, and the point set to be corrected and the conflict buffer record index are retained. The fused surface element index mapping table and the source frame identifier statistics are synchronously written to the output to ensure the continuity of the re-projection verification path.
[0135] The output point cloud forms a constraint relationship between detail enhancement and consistency support, making it more difficult for pseudo-details to remain in the output for a long time, and retaining clear verification entry points in conflict areas.
[0136] In one embodiment, during the consistency verification phase, the super-resolution point cloud results are projected back onto the imaging plane of each frame according to the pose transformation results, and observable markers are generated. Points projected outside the imaging range, with invalid validity information, or failing the occlusion determination are written as unobservable markers. The depth information of the projected pixel landing point corresponding to the observable marker is read to form the observation depth. The predicted depth comes from the camera coordinate depth component before projection, and the difference between the two forms the line-of-sight difference and is converted into a difference rate. At the same time, phase information is read along the fringe direction to calculate the phase coherence. The fringe modulation and phase coherence form decoded consistency evidence under the validity information gating. The single-frame support score is obtained by multiplying the decoded consistency evidence by the geometric suppression term. After sorting the single-frame support scores of multiple frames, the stable support set is screened out by the position of the largest adjacent difference and a consistent support is formed. At the same time, the geometric conflict frequency is obtained and a verification result is formed. The verification result is written back to the point-level confidence set and the complementary sampling label is updated. Points corresponding to the outlier set are deleted from the super-resolution point cloud results, and points corresponding to the point set to be corrected retain the conflict marker and maintain the conflict buffer record index. The operator sees the scattered spikes of the reflective boundary gradually disappear on the screen, while the details of the edges and grooves are still preserved. Marked areas are retained near the occluded edges for easy rescanning or verification, and the final output is a super-resolution point cloud result with verified constraints.
[0137] The super-resolution point cloud results are verified by re-projection comparison and cross-frame aggregation. The verification results separate consistent support, decoding jumps and geometric conflicts into stable states and write back the point-level confidence set and complementary sampling labels. The super-resolution point cloud results with verification constraints delete the corresponding points of the abnormal point set and retain the conflict markers of the point set to be corrected. The output point cloud has both increased detail and multi-frame consistent support constraints.
[0138] Specifically, the above are merely preferred embodiments of this application and are not intended to limit this application.
[0139] In the description of this specification, references to terms such as "an embodiment," "an example," and "a specific example" indicate that a particular feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0140] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to any specific implementation. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.
Claims
1. A multi-frame overlay point cloud super-resolution system for 3D reconstruction, characterized in that, include: Acquisition and construction unit: Acquire multiple frames of structured light point cloud of the same target, generate a set of multiple frame point cloud, and simultaneously generate a set of point-level confidence scores corresponding to each point in the set of multiple frame point cloud; Pose registration unit: Registers a multi-frame point cloud set to obtain the pose transformation results from each frame to the reference frame, generates a unified coordinate point cloud set, and simultaneously generates a point-level registration residual set corresponding to each point in the unified coordinate point cloud set. Complementary discrimination unit: Performs complementary sampling discrimination on a unified coordinate point cloud set, generates complementary sampling labels based on imaging stability and geometric consistency, and outputs complementary point set, outlier point set, and point set to be corrected; Layered fusion unit: Fusion is performed based on complementary sampling labels; outlier sets are stopped from entering the fusion input; complementary point sets participate in densification fusion; and point sets to be corrected participate in restricted fusion to obtain super-resolution point cloud results. The feedback verification unit performs consistency verification on the super-resolution point cloud results, updates the point-level confidence set and complementary sampling labels based on the verification results, and outputs the verified and constrained super-resolution point cloud results.
2. The multi-frame overlay point cloud super-resolution system for 3D reconstruction according to claim 1, characterized in that, The acceptance building unit is used for: Decoding is performed on each frame of structured light image sequence for the same target to generate depth information, phase information, validity information, and calculate fringe modulation. Based on the validity information, 3D points are generated by combining depth information with camera intrinsic parameters and calibration extrinsic parameters, written into a multi-frame point cloud set, and frame identifiers and pixel coordinates are bound. Based on the fringe modulation and phase information, decoded stable components are generated and geometrically continuous components are generated by combining spatial neighborhood. These components are fused to obtain a point-level confidence set and correspond point-by-point with the multi-frame point cloud set.
3. The multi-frame overlay point cloud super-resolution system for 3D reconstruction according to claim 2, characterized in that, The pose registration unit is used for: Calculate the frame reliability score for each frame based on the point-level confidence set and determine the reference frame point cloud; select the registration point set and reference point set from the point cloud of the frame to be registered and the reference frame point cloud based on the spatial partition and the point-level confidence set; generate coarse alignment initial values based on the centroid alignment and principal direction alignment of the registration point set, and form the initial pose relationship between the point cloud of the frame to be registered and the point cloud of the reference frame.
4. The multi-frame overlay point cloud super-resolution system for 3D reconstruction according to claim 3, characterized in that, The pose registration unit is also used for: Based on the initial coarse alignment value, construct a set of corresponding point pairs that are closest to each other; fit a local plane in the corresponding neighborhood of the reference frame point cloud and calculate the point-to-plane residual; determine the set of interior points based on the abrupt change position of the residual sorting and iteratively solve the pose transformation result; generate a unified coordinate point cloud set based on the pose transformation result and calculate the normal projection distance and write it into the point-level registration residual set.
5. The multi-frame overlay point cloud super-resolution system for 3D reconstruction according to claim 4, characterized in that, Complementary discriminant units are used for: In a unified coordinate point cloud set, spatial nearest neighbor points are constructed for each point to be judged, and a set of support points is selected based on the point-level confidence set. A local reference surface is fitted based on the set of support points. Based on the pose transformation results, the point to be judged is projected back onto the multi-frame imaging plane, and phase information, fringe modulation, and validity information are extracted along the fringe direction. The phase coherence is calculated, and fringe phase reproducibility parameters are generated.
6. The multi-frame overlay point cloud super-resolution system for 3D reconstruction according to claim 5, characterized in that, The complementary discrimination unit is also used for: In a unified coordinate point cloud set, a line-of-sight bucket is constructed based on the line-of-sight direction of the reference frame and a depth sequence is formed. Based on the layer segmentation position of the depth sequence, the thickness length of the main layer and the length of the interlayer gap are calculated and the line-of-sight layer separation parameters are generated. Complementary sampling labels are generated based on the maximum difference boundary between the fringe phase reproducibility parameter and the line-of-sight layer separation parameter, and complementary point sets, outlier set sets, and sets of points to be corrected are output.
7. The multi-frame overlay point cloud super-resolution system for 3D reconstruction according to claim 6, characterized in that, The hierarchical fusion unit is used for: A fused surface set is established based on the grid cells of the unified coordinate point cloud set. The fused surface set contains the center coordinates and normals of the fused surface cells. Each grid cell is divided into maximum difference segments based on the point-level confidence set and a stable point set is selected. The stable point set is used to calculate the center coordinates and normals of the fused surface cells and write them into the fused surface set.
8. The multi-frame overlay point cloud super-resolution system for 3D reconstruction according to claim 7, characterized in that, The hierarchical fusion unit is also used for: After the complementary point set is mapped to the fused surface set, the normal signed distance is calculated and the attached coordinates are generated. The attached coordinates are divided according to the maximum gap of the normal signed distance and the fused surface set is updated by crossing the zero principal cluster. The set of points to be corrected is divided into proximity clusters and conflict clusters, and a conflict buffer record is generated. The set of outliers stops entering the fusion input, and the set of fused surface elements is used to export the super-resolution point cloud result.
9. The multi-frame overlay point cloud super-resolution system for 3D reconstruction according to claim 8, characterized in that, The feedback verification unit is used for: The super-resolution point cloud results are back-projected onto the multi-frame imaging plane based on the pose transformation results to generate a set of back-projected pixel landing points. The set of back-projected pixel landing points is combined with imaging range determination, validity information determination, and occlusion determination by comparison of predicted depth and observed depth to generate observable markers. Depth information is read from the location corresponding to the observable marker to form the line-of-sight difference quantity and difference rate. The fringe modulation and phase information are read to form the phase coherence. Combined with the validity information, decoding consistency evidence is generated.
10. The multi-frame overlay point cloud super-resolution system for 3D reconstruction according to claim 9, characterized in that, The feedback verification unit is also used for: The single-frame support score is synthesized from the decoded consistent evidence and the geometric suppression term corresponding to the difference rate. After sorting the single-frame support scores of multiple frames, a stable support set is screened out by the position of the largest adjacent difference and a consistent support is generated. At the same time, the geometric conflict frequency is generated from the difference rate sequence and a verification result is formed. The verification result is used to update the point-level confidence set and complementary sampling labels and output the super-resolution point cloud result with verification constraints.
Citation Information
Patent Citations
Three-dimensional point cloud acquisition method and structured light scanning equipment
CN119289864A
Point cloud fusion three-dimensional reconstruction method and system based on multivariate confidence coefficient filtering
CN115375836A
Pavement crack identification method based on point cloud-RGB (Red, Green, Blue) heterogenous image multistage registration mapping
CN117036300A
Indoor real scene three-dimensional reconstruction method based on improved 3D Gaussian sputtering
CN120279159A
Three-dimensional environment reconstruction optimization method based on multi-sensor fusion data
CN120931829A