Point cloud segmentation method and device

By extracting the two-dimensional mask of image frames frame by frame and the polar line correction technology of multi-view phase encoding diagram, the problem of point cloud segmentation in complex and dynamic scenarios is solved, and efficient and accurate three-dimensional reconstruction is achieved, suitable for complex and dynamic environments.

CN120388040APending Publication Date: 2025-07-29SHENZHEN UNIV +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510235687.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The prior art is difficult to effectively separate the target point cloud from the background point cloud in complex or dynamic scenarios, resulting in reduced three-dimensional reconstruction accuracy and data redundancy, and depends on specific environment settings or requires a large amount of labeled data, which lacks flexibility and real-timeness.

Method used

By extracting the two-dimensional mask of image frames frame by frame, combining multi-view phase coding diagrams and polar line correction technology, we automatically segment the objects and backgrounds to be reconstructed to generate high-precision three-dimensional point clouds, suitable for complex and dynamic scenarios.

Benefits of technology

It realizes efficient and accurate point cloud segmentation in complex and dynamic scenarios, reduces the amount of computing, improves the quality and efficiency of three-dimensional reconstruction, reduces dependence on specific environments, and does not require a large amount of labeled data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388040A_ABST
    Figure CN120388040A_ABST
Patent Text Reader

Abstract

The invention discloses a point cloud segmentation method, which comprises the following steps that: two-dimensional masks of image frames are extracted frame by frame, the image frames are pictures of an object to be reconstructed shot from different visual angles, and the two-dimensional mask extracted from the previous image frame can be transmitted to the subsequent image frame; the method comprises the following steps: acquiring fringe image sequences of a to-be-reconstructed object at different visual angles, extracting phase information, generating a multi-visual-angle phase coding diagram of the to-be-reconstructed object, performing epipolar correction on the phase coding diagrams at different visual angles to eliminate geometric distortion at different visual angles, and calculating parallax of the same name point at different visual angles to obtain the multi-visual-angle phase coding diagram of the to-be-reconstructed object. Calculating a depth value of an object point corresponding to the homonymy point according to the parallax, and obtaining depth value data of the object to be reconstructed; and obtaining a two-dimensional contour according to the two-dimensional mask of the to-be-reconstructed object, calculating the depth value of the two-dimensional contour of the to-be-reconstructed object pixel by pixel according to the depth value data of the to-be-reconstructed object, mapping the two-dimensional contour information of the to-be-reconstructed object to a three-dimensional space, generating the three-dimensional contour of the to-be-reconstructed object of the current frame, and realizing point cloud segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to three-dimensional scanning technology, and particularly to a point cloud segmentation method and device. Background Art

[0002] In a hand-held three-dimensional scanning system, the point cloud data collected by a sensor usually contains the object to be reconstructed and the background environment or other irrelevant objects. These unnecessary point cloud data will interfere with subsequent processing and modeling, resulting in data redundancy, reduced reconstruction accuracy, loss of model details, and even inability to obtain accurate three-dimensional reconstruction results. Especially in complex or dynamic scenarios, such as outdoor environments with cluttered backgrounds, mixing of multiple targets in industrial inspections, or real-time background changes in dynamic inspections, traditional methods are difficult to effectively separate the target point cloud from the background point cloud, thus severely limiting the applicable scope and reliability of three-dimensional reconstruction technology.

[0003] With the rapid development of three-dimensional reconstruction technology, the structured light reconstruction system, as a key three-dimensional imaging means, has been widely used in fields such as object reconstruction, environmental modeling, and robot navigation. During the reconstruction process, the point cloud data obtained by the system usually contains the object to be reconstructed and the background environment or other irrelevant objects. These background point clouds not only bring data redundancy but may also interfere with subsequent reconstruction processing, thereby affecting the accuracy of object modeling. For example, in a complex scene, such as an environment with a cluttered background or many objects, the point clouds of the background and the target object may overlap, making it difficult to separate the object from the background, thus affecting the final three-dimensional reconstruction effect.

[0004] Currently, the following techniques are commonly used for point cloud segmentation:

[0005] Point cloud segmentation method based on environmental constraints: By setting constraint conditions (such as black cloth, calibration board, mask, or auxiliary frame) in a specific physical environment to reduce the interference of background point clouds, thereby achieving the separation of the object from the background. Although these methods are effective in scenarios with a relatively simple background and controllable environment, their limitation lies in being highly dependent on specific environmental settings, lacking flexibility, and having poor effects in the face of complex and dynamic backgrounds.

[0006] Point cloud segmentation method based on geometric features: By extracting geometric features (such as planes, curved surfaces, edges, etc.) in the point cloud for segmentation. Common algorithms include region-growing-based segmentation methods, random sample consensus algorithm (RANSAC), etc. They identify the boundary between the object and the background by analyzing the local structure of the point cloud data and are applicable to simple scenarios. However, in complex or dynamic scenarios, these methods may be affected by noise and occlusion, resulting in inaccurate segmentation results.

[0007] Point cloud segmentation method based on deep learning: By training models such as convolutional neural networks (CNNs) and graph convolutional networks (GCNs), it is possible to automatically classify and segment point cloud data. Common deep learning algorithms include PointNet, PointNet++, etc. These methods can handle relatively complex scenes and backgrounds and are suitable for large-scale and highly complex point cloud data. However, these methods still require a large amount of labeled data for training and pose challenges in terms of real-time performance and computational overhead.

[0008] It can be seen that the existing technologies mainly have the following disadvantages:

[0009] Strong environmental constraint dependence: Point cloud segmentation methods based on environmental constraints rely on specific environmental settings, such as black cloth, calibration board, mask, or auxiliary framework. These methods are only effective in simple and controllable backgrounds and lack flexibility. When the scene is complex or the background is dynamic, the effectiveness of these methods drops significantly. For example, in outdoor environments with cluttered backgrounds or industrial inspection scenes where multiple target objects coexist, these constraints cannot work effectively, resulting in inaccurate separation of objects from the background.

[0010] Limitations of geometric feature methods: Point cloud segmentation methods based on geometric features (such as RANSAC and region growing) are susceptible to noise, occlusion, or object overlap in complex or dynamic backgrounds, leading to inaccurate segmentation results. In highly complex scenes, it is difficult for these methods to achieve precise separation of the target from the background, especially when the object boundaries are blurred.

[0011] Point cloud segmentation methods based on deep learning (such as PointNet, PointNet++) have the following challenges: high requirements for data quality, as noise or missing data can affect the model's performance; poor generalization ability, with unstable performance in different application scenarios; low segmentation accuracy in complex scenes such as object overlap or occlusion; and the need for a large amount of labeled data, resulting in high costs for obtaining high-quality data. Summary of the Invention

[0012] Based on the above situation, the main objective of the present invention is to provide an efficient and accurate point cloud segmentation method that can automatically distinguish the object to be reconstructed from the background area and extract pure target point cloud data. It can improve the reconstruction quality and efficiency of the 3D scanning system, reduce the limitations of environmental conditions on the application of the technology, enhance the applicability in complex and dynamic scenes, and further broaden the application fields of 3D reconstruction technology.

[0013] To achieve the above objective, the technical solution adopted by the present invention is as follows:

[0014] A point cloud segmentation method, comprising the steps:

[0015] S100, extract the two-dimensional mask of the image frame frame by frame, where the image frame is a picture of the object to be reconstructed taken from different perspectives. Among them, the two-dimensional mask extracted from the previous image frame can be passed to the subsequent image frames;

[0016] S200, collect the fringe image sequence of the object to be reconstructed from different perspectives and extract the phase information to generate the multi-view phase encoding map of the object to be reconstructed. Perform epipolar correction on the phase encoding maps of different perspectives to eliminate geometric distortion under different perspectives, so that corresponding points in the phase encoding maps of different perspectives are located on the same scan line. The corresponding points are the image points projected by the same object point in the object to be reconstructed on the imaging planes of different perspectives. Then calculate the disparity of the corresponding points under different perspectives, and calculate the depth value of the object point corresponding to the corresponding point according to the disparity to obtain the depth value data of the object to be reconstructed;

[0017] S300, obtain the two-dimensional contour of the object to be reconstructed according to the two-dimensional mask of the object to be reconstructed, and calculate the depth value of the two-dimensional contour of the object to be reconstructed pixel by pixel according to the depth value data of the object to be reconstructed, so as to map the two-dimensional contour information of the object to be reconstructed into three-dimensional space and generate the three-dimensional contour of the object to be reconstructed in the current frame to achieve point cloud segmentation.

[0018] Preferably, extract the two-dimensional mask of the image frame frame by frame based on the segmentation arbitrary model.

[0019] Preferably, the two-dimensional mask of the object to be reconstructed is input in the form of prompt point coordinates or a rectangular frame for the first frame in the image frame.

[0020] Preferably, the S100 includes the steps of:

[0021] S101, perform normalization processing on the image frame so that the resolutions of all image frames are the same to obtain the normalized image frame;

[0022] S102, perform multi-scale feature extraction on the normalized image frame. At the same time, perform position encoding on the contour of the object to be reconstructed identified in the image frame to retain the spatial position information. Decode the current image frame according to the multi-scale features, and cache the object pointers of the first frame image, the current image frame, and the previous M frames of images;

[0023] S103, calculate and cache the mask memory features and mask memory position encodings of the previous N frames of the current frame for the normalized image frames other than the first frame;

[0024] S104, calculate the low-resolution image features of the current image frame according to the mask memory features and mask memory position encodings, and the object pointers of the first frame image, the current image frame, and the previous M frames cached in step S102 and the position encoding of the current image frame;

[0025] In S105, decode the low-resolution image to obtain the low-resolution mask of the current image frame, upsample the low-resolution mask to obtain the two-dimensional mask of the current image frame, and continue the calculation until all image frames are completed to obtain the two-dimensional mask of the object to be reconstructed.

[0026] Preferably, in step S102, for the first frame image, directly decode the multi-scale features of the first frame image, and determine the two-dimensional contour of the object to be reconstructed by identifying the mask of the object to be reconstructed in the first frame image.

[0027] Preferably, in step 103, extract pixel-by-pixel features from the normalized image frames other than the first frame, and encode the mask features and positions of the object to be reconstructed according to the pixel-by-pixel features to obtain mask memory features and mask memory position encodings.

[0028] Preferably, M is 14 and N is 7.

[0029] Preferably, in step S200, extract phase information from the fringe image sequence through a phase unwrapping algorithm to generate a multi-view phase encoding map of the object to be reconstructed.

[0030] Preferably, the epipolar rectification of the phase encoding maps from different views in step S200 includes: adjusting the projection matrices of multiple cameras through geometric transformation to align their optical axes and make the pixel coordinates satisfy the epipolar constraint condition, so that corresponding points in the phase encoding maps from different views are located on the same scan line.

[0031] Preferably, calculate the disparity by calculating the horizontal distance difference between the corresponding points.

[0032] Preferably, calculate the depth value of the object point corresponding to the corresponding point according to the disparity, the calibration parameters of the camera, and the focal length.

[0033] Preferably, the calibration parameter of the camera is the baseline distance between the optical centers of multiple cameras.

[0034] The present invention also discloses a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed, it can implement the point cloud segmentation method described in any one of the present invention.

[0035] The present invention also discloses a three-dimensional scanning system, which uses the point cloud segmentation method described in any one of the present invention to perform point cloud segmentation on the object to be reconstructed to achieve three-dimensional reconstruction of the object.

[0036] The present invention also discloses a point cloud segmentation device, including a mask extraction module, a depth value calculation module, and a point cloud calculation module.

[0037] The mask extraction module is used to extract the two-dimensional mask of each image frame, where the image frames are pictures of the object to be reconstructed taken from different perspectives. Among them, the two-dimensional mask extracted from the previous image frame can be passed to subsequent image frames;

[0038] The depth value calculation module is used to collect the fringe image sequences of the object to be reconstructed from different perspectives and extract phase information, generate a multi-perspective phase-encoded map of the object to be reconstructed, perform epipolar correction on the phase-encoded maps from different perspectives to eliminate geometric distortions under different perspectives, so that corresponding points in the phase-encoded maps from different perspectives are located on the same scan line. The corresponding points are the image points projected by the same object point in the object to be reconstructed on the imaging planes of different perspectives. Then, calculate the disparity of the corresponding points under different perspectives, and calculate the depth value of the object point corresponding to the corresponding point according to the disparity to obtain the depth value data of the object to be reconstructed;

[0039] The point cloud calculation module is used to obtain the two-dimensional contour of the object to be reconstructed according to the two-dimensional mask of the object to be reconstructed, and calculate the depth value of each pixel of the two-dimensional contour of the object to be reconstructed according to the depth value data of the object to be reconstructed, so as to map the two-dimensional contour information of the object to be reconstructed into three-dimensional space, generate the three-dimensional contour of the object to be reconstructed in the current frame, and achieve point cloud segmentation.

[0040] Preferably, the mask extraction module includes a segmentation arbitrary model unit, and the segmentation arbitrary model unit is used to extract the two-dimensional mask of each image frame.

[0041] Preferably, the mask extraction module includes a normalization processing unit, an image frame feature extraction unit, an image decoding unit, a mask feature calculation unit, a low-resolution image feature calculation unit, and a two-dimensional mask calculation unit.

[0042] The normalization processing unit is used to perform normalization processing on the image frame so that the resolutions of all image frames are the same, and obtain a normalized image frame;

[0043] The image frame feature extraction unit is used to perform multi-scale feature extraction on the normalized image frame. At the same time, perform position encoding on the contour of the object to be reconstructed identified in the image frame to retain spatial position information;

[0044] The image decoding unit is used to decode the current image frame according to the multi-scale features, and cache the object pointers of the first frame image, the current image frame, and its first M previous image frames;

[0045] The mask feature calculation unit is used to calculate and cache the mask memory features and mask memory position encodings of the first N previous frames of the current frame for the normalized image frames other than the first frame;

[0046] The low-resolution image feature calculation unit is used to calculate the low-resolution image features of the current image frame according to the mask memory features, mask memory position encodings, the first-frame image cached in step S102, the current image frame, the object pointers of the previous M frames, and the position encodings of the current image frame;

[0047] The two-dimensional mask calculation unit is used to decode the low-resolution image to obtain the low-resolution mask of the current image frame, upsample the low-resolution mask to obtain the two-dimensional mask of the current image frame, and continue the calculation until all image frames are completed, so as to obtain the two-dimensional mask of the object to be reconstructed.

[0048] Preferably, for the first-frame image, the image decoding unit is further used to directly decode the multi-scale features of the first-frame image, and determine the two-dimensional contour of the object to be reconstructed by identifying the mask of the object to be reconstructed in the first-frame image.

[0049] Preferably, the mask feature calculation unit is used to extract per-pixel features from the normalized image frames except the first frame, and encode the mask features and positions of the object to be reconstructed according to the per-pixel features to obtain mask memory features and mask memory position encodings.

[0050] Preferably, the depth value calculation module further includes a phase encoding map calculation unit, which is used to extract phase information from the fringe image sequence through a phase unwrapping algorithm to generate a multi-view phase encoding map of the object to be reconstructed.

[0051] Preferably, the depth value calculation module further includes an epipolar rectification calculation unit, which is used to adjust the projection matrices of multiple cameras through geometric transformation to align their optical axes and make the pixel coordinates satisfy the epipolar constraint condition, so that corresponding points in the phase encoding maps of different views are located on the same scan line.

[0052] Preferably, the depth value calculation module further includes a depth value calculation unit, which is used to calculate the depth value of the object point corresponding to the corresponding point according to the parallax, the calibration parameters of the camera, and the focal length.

[0053] Existing point cloud segmentation methods either highly rely on specific environmental settings, lack flexibility, and have poor performance in the face of complex and dynamic backgrounds, or are affected by noise and occlusion in complex or dynamic scenarios, resulting in inaccurate segmentation results, or require a large amount of labeled data for training, presenting challenges in terms of real-time performance and computational overhead. The point cloud segmentation method provided by the present invention extracts the two-dimensional mask (image segmentation technology) of the image frame frame by frame, enables the two-dimensional mask to be transmitted between frames, and combines the fringe image sequences of the object to be reconstructed from different perspectives and extracts phase information (structured light reconstruction), stereo matching (determining corresponding points), and epipolar rectification technology to achieve automatic target segmentation and background removal, thereby improving the efficiency and accuracy of point cloud data processing and avoiding errors and redundant data in traditional methods. Extracting the two-dimensional mask of the image frame frame by frame is used to extract the two-dimensional contour of the object, and the consistency processing of the object to be reconstructed is maintained through the mask transmission between multiple frames to ensure accurate object segmentation in complex backgrounds and dynamic scenarios. At the same time, combining stereo matching of fringe projection (determining the positions of corresponding points in multi-view images) and epipolar rectification technology can accurately obtain the depth information of the object and generate high-precision three-dimensional point clouds. This method breaks through the limitations of traditional point cloud segmentation technology, does not rely on specific environmental conditions, does not require model training, and at the same time ensures real-time performance, reduces the amount of calculation, and is applicable to more complex and dynamic environments.

[0054] Other beneficial effects of the present invention will be elaborated in the specific implementation manners through the introduction of specific technical features and technical solutions. Those skilled in the art should be able to understand the beneficial technical effects brought by the said technical features and technical solutions through these introductions. Brief Description of the Drawings

[0055] The following will describe the preferred embodiments of the point cloud segmentation method and device according to the present invention with reference to the drawings. In the figures:

[0056] Figure 1 is a flowchart of a point cloud segmentation method according to a preferred embodiment of the present invention;

[0057] Figure 2 is a flowchart of a method for extracting the two-dimensional mask of the image frame frame by frame according to a preferred embodiment of the present invention;

[0058] Figure 3 is a schematic diagram of the effect of extracting the two-dimensional mask of the image frame frame by frame in a single-object scene on a plane according to a preferred embodiment of the present invention;

[0059] Figure 4 is a schematic diagram of the effect of extracting the two-dimensional mask of the image frame frame by frame in a multi-object scene on a plane according to a preferred embodiment of the present invention;

[0060] Figure 5Schematic diagram of two - dimensional mask effect for frame - by - frame extraction of image frames in a multi - object overlapping scene on a plane according to a preferred embodiment of the present invention;

[0061] Figure 6 Schematic diagram of two - dimensional mask effect for frame - by - frame extraction of image frames in a dynamic occlusion scene according to a preferred embodiment of the present invention;

[0062] Figure 7 Schematic diagram of binocular vision principle according to a preferred embodiment of the present invention;

[0063] Figure 8 Schematic diagram of corresponding points according to a preferred embodiment of the present invention;

[0064] Figure 9 Schematic diagram of epipolar rectification principle according to a preferred embodiment of the present invention;

[0065] Figure 10 Schematic diagram of corresponding - point search principle based on epipolar constraint and disparity constraint according to a preferred embodiment of the present invention;

[0066] Figure 11 Schematic diagram of triangulation principle according to a preferred embodiment of the present invention;

[0067] Figure 12 Effect diagram of single - frame point - cloud segmentation according to a preferred embodiment of the present invention;

[0068] Figure 13 Effect diagram of three - dimensional point - cloud model of an object reconstructed from point clouds after multi - frame segmentation according to a preferred embodiment of the present invention;

[0069] Figure 14 Block diagram of the device for the point - cloud segmentation method according to a preferred embodiment of the present invention. Detailed implementation manner

[0070] Figure 1 Flowchart of the point - cloud segmentation method according to a preferred embodiment of the present invention, including steps:

[0071] S100, extract the two - dimensional mask of the image frame frame - by - frame. The image frame is a picture of the object to be reconstructed taken from different perspectives. Among them, the two - dimensional mask extracted from the previous image frame can be passed to the subsequent image frame. Usually, for different objects to be reconstructed, the number of pictures of the object to be reconstructed at different angles used is different, and pictures of the object to be reconstructed at different angles can be taken specifically according to requirements.

[0072] S200, collect the fringe image sequences of the object to be reconstructed from different perspectives and extract the phase information to generate multi-perspective phase encoding maps of the object to be reconstructed. Perform epipolar correction on the phase encoding maps from different perspectives to eliminate geometric distortions under different perspectives, so that corresponding points in the phase encoding maps from different perspectives are located on the same scan line. The corresponding points are the image points projected by the same object point in the object to be reconstructed on the imaging planes of different perspectives. Then, calculate the disparity of the corresponding points under different perspectives, and calculate the depth value of the object point corresponding to the corresponding point according to the disparity to obtain the depth value data of the object to be reconstructed;

[0073] S300, obtain the two-dimensional contour of the object to be reconstructed according to the mask of the object to be reconstructed, and calculate the depth value of the two-dimensional contour of the object to be reconstructed pixel by pixel according to the depth value data of the object to be reconstructed, so as to map the two-dimensional contour information of the object to be reconstructed into the three-dimensional space and generate the three-dimensional contour of the object to be reconstructed in the current frame, realizing point cloud segmentation.

[0074] Existing point cloud segmentation methods either highly rely on specific environmental settings, lack flexibility, and have poor effects when facing complex and dynamic backgrounds, or are affected by noise and occlusion in complex or dynamic scenes, resulting in inaccurate segmentation results, or require a large amount of labeled data for training, presenting challenges in terms of real-time performance and computational overhead. The point cloud segmentation method provided by the present invention realizes automatic target segmentation and background elimination by extracting the two-dimensional mask (image segmentation technology) of the image frame frame by frame, enabling the two-dimensional mask to be transferred between frames, and combining the fringe image sequences of the object to be reconstructed from different perspectives and extracting the phase information (structured light reconstruction), stereo matching (determining corresponding points), and epipolar correction techniques, thereby improving the efficiency and accuracy of point cloud data processing and avoiding errors and redundant data in traditional methods. The two-dimensional mask of the image frame is extracted frame by frame to extract the two-dimensional contour of the object, and the target consistency processing of the object to be reconstructed is maintained through mask transfer between multiple frames to ensure accurate object segmentation in complex backgrounds and dynamic scenes. At the same time, by combining stereo matching of fringe projection (determining the positions of corresponding points in multi-perspective pictures) and epipolar correction techniques, the depth information of the object can be accurately obtained to generate high-precision three-dimensional point clouds. This method breaks through the limitations of traditional point cloud segmentation techniques, does not rely on specific environmental conditions, does not require model training, and also ensures real-time performance, reduces the computational amount, and is applicable to more complex and dynamic environments.

[0075] In a preferred embodiment, a two-dimensional mask of an image frame can be extracted frame by frame based on the Segment Anything Model 2 (SAM2). The SAM2 network model was proposed by the research team of Meta AI (formerly Facebook). SAM stands for Segment Anything Model, aiming to solve the promptable visual segmentation tasks in images and videos. By constructing a unified image and video segmentation model, SAM2 introduces a streaming memory mechanism based on the Transformer architecture, enabling the effective transmission of two-dimensional masks of cross-frame information, thus significantly improving the robustness and accuracy of object segmentation in dynamic scenarios.

[0076] In a preferred embodiment, the first frame in the image frame inputs the mask of the object to be reconstructed in the form of prompt point coordinates or a rectangular box. Thus, taking this as the initialization reference, the transmission and automatic generation of the mask among multiple image frames are realized.

[0077] In a preferred embodiment, as Figure 2 shown, the extraction of the two-dimensional mask of the image frame frame by frame includes:

[0078] S101, Normalize the image frame so that all image frames have the same resolution, obtaining a normalized image frame. In a specific embodiment, after the image frame is normalized, its resolution can be uniformly adjusted to 1024×1024.

[0079] S102, Extract multi-scale features from the normalized image frame. At the same time, perform position encoding on the contour of the object to be reconstructed identified in the image frame to retain spatial position information. Decode the current image frame according to the multi-scale features, and cache the object pointers of the first frame image, the current image frame, and its first M previous frame images. Extract multi-scale feature maps from the normalized image frame to capture the features of the object to be reconstructed at different scales, thereby enhancing the adaptability to multi-scale targets and improving the robustness of the model in complex scenarios. In a specific embodiment, the multi-scale features can include high-resolution image features and medium-resolution image features.

[0080] S103, Calculate and cache the mask memory features and mask memory position encodings of the first N previous frames of the current frame for the normalized image frames other than the first frame;

[0081] S104, Calculate the low-resolution image features of the current image frame according to the mask memory features, mask memory position encodings, the object pointers of the first frame image, the current image frame, and the first M previous frames cached in step S102, and the position encoding of the current image frame. In this way, the two-dimensional mask of the object to be reconstructed can be automatically generated and transmitted between subsequent frames, ensuring the continuity and consistency of the contour of the object to be reconstructed among multiple frames.

[0082] S105, decode the low-resolution image to obtain the low-resolution mask of the current image frame, upsample the low-resolution mask to obtain the two-dimensional mask of the current image frame, and continue the calculation until all image frames are completed to obtain the two-dimensional mask of the object to be reconstructed.

[0083] By effectively combining the mask memory feature and the mask memory position encoding with the first-frame image, the current image frame, the object pointers of the first M frames, and the position encoding of the current image frame, the accurate transmission of mask information in the time series is achieved, thereby improving the automation and accuracy of the mask generation of the object to be reconstructed and providing technical guarantee for the target consistency among multiple frames.

[0084] To verify the robustness and accuracy of the two-dimensional mask generation of the object to be reconstructed in the present invention, the following typical experimental scenarios were tested. The experiments covered complex conditions such as single planar target, overlapping interference of multiple similar planar targets, and dynamic occlusion. The experimental results are as follows:

[0085] (1) Single-target scenario on a plane

[0086] As Figure 3 shown, the target is a white plaster statue on a plane. Under the condition of no obvious background interference, the two-dimensional mask of the plaster statue can be accurately generated, with clear mask edges and high coincidence with the actual contour of the target, verifying the high-precision segmentation ability of this method in a single-target simple scenario.

[0087] (2) Multi-target scenario on a plane

[0088] The target is multiple objects placed on a plane. In the case of multi-target interference, the target to be segmented can be selected through hint points or rectangular frames, and the corresponding mask of the target can be correctly generated, excluding the influence of other objects, showing strong interference robustness and target selectivity.

[0089] (3) Multi-object overlapping scenario on a plane

[0090] The target is multiple similar Crayon Shin-chan ornaments placed on a plane, and the selected mask is a Crayon Shin-chan at the bottom that is blocked. The experimental results show that the mask of the target object (such as Crayon Shin-chan) can be accurately generated in a complex scenario, and the mask contour is clear and complete. This indicates strong robustness and superiority in dealing with the segmentation task of complex backgrounds and overlapping objects.

[0091] (4) Dynamic occlusion scenario (finger occluding the white plaster statue)

[0092] The target is two-frame dynamic scenes of fingers on a white plaster statue. Experiments have found that it can accurately generate the two-dimensional mask of the fingers and clearly separate the occlusion relationship between the fingers and the background object (the plaster statue). Even when the occlusion degree is relatively high, it can still retain the complete mask boundary of the object. This indicates that it has significant target segmentation advantages in dynamic occlusion scenes.

[0093] The experimental results show that the method for extracting the two-dimensional mask of the object to be reconstructed in the present invention exhibits excellent target segmentation capabilities in various complex scenes: in a single-target simple scene, it can generate the target mask with high precision; in a multi-target interference scene, it shows good target selectivity; in a scene with multiple object overlaps, it can accurately generate the mask contour of the target; in dynamic occlusion conditions, it can clearly separate the occluder and the background object. The above results verify its applicability and robustness under multi-scene, multi-target, and multi-occlusion conditions, providing reliable support for target point cloud extraction in point cloud segmentation tasks.

[0094] In a preferred embodiment, for the first-frame image in step 102, directly decode the multi-scale features of the first-frame image, and determine the two-dimensional contour of the object to be reconstructed by identifying the mask of the object to be reconstructed in the first-frame image.

[0095] In a preferred embodiment, in step 103, extract per-pixel features from the normalized image frames other than the first frame, and encode the mask features and positions of the object to be reconstructed according to the per-pixel features to obtain the mask memory feature and the mask memory position encoding.

[0096] In a preferred embodiment, after testing during the design process and considering both accuracy and time efficiency, M in steps S102 and S104 can be selected as 14, and N in step S103 can be selected as 7. That is to say, cache the object pointers of the current image frame and its previous 14 image frames, and cache the mask memory features and mask memory position encodings of the previous 7 frames of the current frame.

[0097] In a preferred embodiment, in step S200, phase information can be extracted from the fringe image sequence through a phase unwrapping algorithm to generate multi-view phase-encoded maps of the object to be reconstructed. In a specific embodiment, the fringe projection phase stereo matching method can be used. A digital projection device projects a series of sine or cosine fringe patterns onto the object to be measured, and at the same time, a camera is used to collect the fringe image sequence modulated by the object from different perspectives. Subsequently, phase information is extracted through a phase unwrapping algorithm to generate multi-view phase-encoded maps of the target object. Since the phase information of the fringe pattern is less sensitive to ambient light interference, and phase encoding performs high-density and continuous encoding on the object surface, each pixel point can independently obtain a specific phase value, thereby achieving pixel-level stereo matching (determining the positions of corresponding points in multi-view images). Compared with traditional stereo matching techniques that rely on object surface texture, feature information, or gray values, phase stereo matching performs excellently in terms of anti-light interference ability, matching point density, accuracy, and applicability, providing a reliable guarantee for 3D reconstruction in complex environments.

[0098] In a preferred embodiment, the epipolar rectification of the phase-encoded maps from different perspectives in step S200 includes: adjusting the projection matrices of multiple cameras through geometric transformation to align their optical axes and make the pixel coordinates satisfy the epipolar constraint condition, so that the corresponding points in the phase-encoded maps from different perspectives are located on the same scan line.

[0099] The implementation of the fringe projection phase stereo matching method has diversity under different hardware configurations. According to the system requirements and application scenarios, various hardware combinations can be selected to optimize the accuracy and efficiency of 3D reconstruction. Common hardware configurations include monocular systems, binocular systems, and multi-camera systems, etc. Each configuration has different advantages and applicable ranges under different conditions. The monocular system combines a single camera with a projection device and is suitable for scanning objects at close range or with relatively simple geometric structures; the binocular system obtains the fringe images of the object from different angles through two cameras and calculates the depth information more accurately through disparity calculation, which is suitable for 3D reconstruction of complex objects or at relatively long distances; while the multi-camera system uses multiple camera perspectives to improve the accuracy and coverage of 3D reconstruction, especially suitable for large-scale scanning or situations where there are occlusions on the object surface.

[0100] The point cloud segmentation method proposed by the present invention is applicable to all the above hardware configurations. Whether it is a monocular, binocular, or multi-camera system, it can effectively utilize the fringe projection phase stereo matching technology to obtain high-quality 3D reconstruction data. Taking the binocular vision system as an example, this system simultaneously captures images of the same scene from different perspectives by two cameras, simulating the working principle of human vision. By calculating the disparity of these two images, the depth information of each point in the scene can be obtained, that is, the distance from each pixel point to the camera.

[0101] Taking the binocular vision system as an example, the principle of binocular vision is as follows Figure 7 As shown, points L, M, and N of the same measured object at different depth positions are projected onto camera 1 as L1, M1, and N1 respectively, and these points are located at the same point on the image of the left camera. This indicates that a single camera can only obtain the position of an object in a certain direction, but cannot determine the depth of the object in that direction. However, in the right camera, the projection positions of these points will change, such as L2, M2, and N2. By comparing the disparities between the left and right camera images, the system can accurately calculate the depth information of each point, thereby achieving accurate three-dimensional reconstruction of the measured object.

[0102] As Figure 8 Shown, the left and right cameras acquire two pictures A and B of the same scene from different perspectives. A point P on the surface of the measured object in the scene is projected onto the imaging planes of the two imaging devices as an image point, labeled as P' and P" respectively in the figure. These two projection points are both images of the three-dimensional object point P, and are called corresponding points or homologous points. The purpose of stereo matching is to determine the positions of the corresponding points in the multi-perspective pictures A and B. In the ideal state where the left and right cameras are parallel and coplanar, the corresponding relationships of the two perspectives of the cameras can be directly found using matching algorithms. However, due to certain distortions and perspective differences in camera imaging, the matching points may not be strictly located on the same scan line. Therefore, it is necessary to perform epipolar rectification on the images to constrain the matching points of the two images on the same scan line to improve the matching accuracy.

[0103] Figure 9 Shows the basic principle of epipolar rectification. Let a point P in the line of sight of the left camera be projected onto the point P1 on the image plane A of the left camera, and its corresponding matching point on the image plane B of the right camera is P2. In theory, point P2 should be located on the epipolar line m2 on the right camera image plane, while point P1 is located on the epipolar line m1 on the left camera image plane. However, due to factors such as the camera's perspective and imaging distortion, the distribution of epipolar lines in the original image plane is usually irregular, resulting in the search for matching points needing to be carried out in a two-dimensional plane, increasing the matching complexity and computational cost.

[0104] Through epipolar rectification, the image planes of the two cameras can be remapped to a new plane where the epipolar lines are parallel and horizontally arranged, making the epipolar lines parallel to the baseline and moving the epipoles to infinity. After rectification, the search area for matching points is reduced from two dimensions to one dimension, and both P1' and P2' are restricted to the same horizontal line, greatly reducing the search complexity and improving the matching efficiency and accuracy. The essence of epipolar rectification is to adjust the projection matrices of the two cameras through geometric transformation to align their optical axes and make the pixel coordinates satisfy the epipolar constraint conditions, thereby achieving efficient stereo matching.

[0105] After epipolar rectification, as Figure 10As shown, P1 and P2 are the image points obtained by projecting the spatial point P onto the left and right camera imaging planes A and B respectively. P1 and P2 satisfy the following relationship:

[0106]

[0107] where F is the fundamental matrix, K1 and K2 are the internal parameter matrices of the left and right cameras respectively, s x is the skew-symmetric matrix of the translation vector between the left and right cameras, and R is the rotation matrix. The phase provides a constraint in one direction, so P1 only needs to find the point with equal phase on the epipolar line m2 to determine the corresponding point P2. In actual measurement, since the depth of field of the camera is fixed, the measurement range of the system is limited. The measurement space of the system is defined as the region bounded by the farthest measurement depth H max and the nearest measurement depth H min , which is located between the two dashed lines and represents the depth range that the system can effectively measure.

[0108] Figure 10 In A , the projection ray formed by the spatial point P and the projection center O of the left camera is intercepted by the measurement space range as the line segment D1D2. The line segment D1D2 is projected onto the right camera as the line segment d1d2. Combining the epipolar constraint, it can be known that the corresponding point P2 of the image point P1 must be on the line segment d1d2. The disparity constraint restricts the search range and only searches for the local area of the corresponding point on one epipolar line, thereby reducing the search space on the epipolar line. This constraint not only improves the matching accuracy but also significantly reduces the computational amount, thereby optimizing the efficiency of the stereo matching algorithm. By constraining the disparity, the geometric relationship between corresponding points can be determined more precisely, thereby enhancing the quality and accuracy of point cloud reconstruction.

[0109] In a preferred embodiment, the disparity is obtained by calculating the horizontal distance difference between the corresponding points in step S200.

[0110] In a preferred embodiment, the depth value of the object point corresponding to the corresponding point is calculated according to the disparity, the calibration parameters of the camera, and the focal length. In a specific embodiment, as Figure 11 shown, assume that the position of a certain point P on the object in the left camera image is P'(X R , y), and the position in the right camera image is P"(X T , y). By calculating the horizontal distance difference between the matching points in the left and right images (the disparity d is the difference in the horizontal distance of the projection points of the same point in the left and right images), combined with the calibration parameter B (the baseline distance between the optical centers of the left and right cameras, the baseline length) and f (the focal length) of the camera, the depth Z of this point can be obtained.

[0111] The calculation formula for the depth Z is as follows:​

[0112]

[0113] In a specific embodiment, the calibration parameter of the camera is the baseline distance between multiple camera optical centers.

[0114] After completing the depth calculation, the depth value of the contour of the object to be reconstructed in the segmented two-dimensional image is calculated pixel by pixel. In this way, the two-dimensional contour information on the object surface is mapped into the three-dimensional space pixel by pixel, generating a single-frame three-dimensional point cloud of the object contour. Thus, the point cloud segmentation is completed.

[0115] The present invention can efficiently and accurately separate the object to be reconstructed from the background in complex backgrounds and dynamic scenes, demonstrating significant application advantages. When dealing with point cloud segmentation of complex backgrounds and overlapping objects, it shows excellent effects; in dynamic occlusion scenes, it can accurately identify the mask of the object to be reconstructed and effectively remove the occluder. Compared with traditional methods, the present invention not only broadens the application scenarios of three-dimensional reconstruction technology, enabling it to work stably in complex backgrounds and dynamic environments, but also significantly reduces the computational cost, improves the processing efficiency, reduces the dependence on the environmental background, can more widely adapt to different scenarios, and has strong practicality.

[0116] Compared with the prior art, the present invention has the following significant advantages and technical effects:

[0117] Ability to adapt to complex backgrounds and dynamic environments: The present invention can efficiently and accurately separate the object to be reconstructed from the background in complex backgrounds and dynamic scenes. Especially when dealing with the point cloud segmentation task of overlapping objects, the present invention can accurately identify and extract the mask of the object to be reconstructed, significantly reducing the influence of the interfering background. In addition, for dynamic occlusion problems, such as finger occlusion, the present invention ensures the scanning accuracy through an efficient background removal method, thereby improving the reliability and stability of three-dimensional reconstruction.

[0118] Broadening of application scenarios: The technology of the present invention can operate stably in dynamic environments and complex backgrounds, significantly expanding the application scope of three-dimensional scanners. Different from the existing environment-constrained methods (such as relying on physical settings such as black cloth, calibration plates or masks), the present invention gets rid of the dependence on specific environments and can adapt to a wider range of application scenarios without these auxiliary devices, especially suitable for the operation mode of synchronous scanning of a handheld scanner and an object, which cannot be achieved in the prior art.

[0119] Efficient resource utilization: By combining image segmentation and using a structured light reconstruction system to obtain depth information, the present invention can accurately separate the point cloud of the object to be reconstructed from the background point cloud in complex backgrounds and dynamic scenes. At the same time, compared with methods that solely rely on deep learning, the present invention significantly reduces the need for a large amount of data annotation and the consumption of computing resources, optimizing the resource utilization efficiency. Therefore, while improving efficiency, this method also has higher flexibility and can adapt to more diverse actual application requirements.

[0120] In summary, compared with the prior art, the present invention has stronger adaptability, lower computational overhead, higher processing efficiency when dealing with complex environments and dynamic scenes, and significantly broadens the application scope and practical application value of 3D scanners.

[0121] The entire 3D point cloud segmentation process includes multiple steps from image acquisition, processing, epipolar rectification to depth calculation based on disparity constraints and triangulation method, and finally realizes point cloud segmentation by combining object contour extraction. First, in the initial stage of 3D reconstruction, the structured light reconstruction system is used to acquire and process images of the object to be reconstructed, and then epipolar rectification is performed to eliminate the geometric distortion caused by the position and attitude differences between cameras, ensuring that corresponding points (representing the same object point in 3D space) in the left and right images are located at the same position on the horizontal scan line, facilitating the subsequent solution of corresponding points. Through epipolar rectification, the uncertainty in disparity calculation can be reduced, and the accuracy of point-to-point matching can be improved.

[0122] Based on the epipolar-rectified images, the corresponding points are searched using disparity constraints. Based on the known epipolar geometry relationship, the disparity constraints limit the matching search space to a small range along the epipolar line, thus significantly reducing the computational amount and improving the matching accuracy. After obtaining the corresponding points, the depth values of each pair of matching points are calculated by the triangulation principle. Finally, the object boundaries in the image are identified, and the contours of the object to be reconstructed are extracted. Combining with the depth map data, only the point cloud corresponding to the object contour region is retained, and the point cloud of the background and other interfering objects is removed to achieve point cloud segmentation, ensuring the accuracy and purity of the point cloud data and providing high-quality input data for subsequent 3D reconstruction and registration.

[0123] Figure 12 The segmentation effect of a single-frame point cloud is shown, and the results indicate that the segmentation algorithm successfully extracts the object point cloud, and the distinction between the object and the background is significant, demonstrating good segmentation quality.

[0124] Figure 13Shows the three-dimensional point cloud model of the object reconstructed from the point cloud after multi-frame segmentation, effectively restoring the details of the object surface and demonstrating high geometric accuracy. The overall shape of the model is clear, and the surface texture and structure are accurately reconstructed, showing strong geometric consistency. Among them, the object to be reconstructed is on the left, the complete three-dimensional point cloud of the object is in the middle, and the Poisson reconstruction of the object is on the right.

[0125] In a multi-object scene, the point cloud segmentation method proposed by the present invention can accurately segment the three-dimensional point cloud of the selected object from the interference of the background and other objects. After selecting the object to be reconstructed, an object mask is automatically generated between consecutive frames.

[0126] In a dynamic occlusion scene, objects such as fingers may dynamically occlude the object to be reconstructed in each frame, causing different degrees of occlusion. To address this problem, the point cloud segmentation method proposed by the present invention can accurately segment the point cloud of the object to be reconstructed in the context of background changes in each frame. The point cloud segmentation method proposed by the present invention can effectively remove dynamic occluders and background noise and accurately separate the point cloud of the target object.

[0127] The present invention also discloses a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed, it can implement the point cloud segmentation method described in any one of the present invention.

[0128] The present invention also discloses a three-dimensional scanning system, and the system uses the point cloud segmentation method described in any one of the present invention to perform point cloud segmentation on the object to be reconstructed to achieve three-dimensional reconstruction of the object.

[0129] The present invention also discloses a point cloud segmentation device Figure 14The block diagram of a point cloud segmentation method device according to a preferred embodiment of the present invention includes a mask extraction module 10, a depth value calculation module 20, and a point cloud calculation module 30. The mask extraction module 10 is used to extract the two-dimensional mask of the image frame frame by frame. The image frame is a picture of the object to be reconstructed taken from different perspectives. Among them, the two-dimensional mask extracted from the previous image frame can be transmitted to the subsequent image frame. The depth value calculation module 20 is used to collect the fringe image sequence of the object to be reconstructed from different perspectives and extract the phase information to generate a multi-perspective phase encoding map of the object to be reconstructed. Epipolar correction is performed on the phase encoding maps from different perspectives to eliminate geometric distortions under different perspectives, so that corresponding points in the phase encoding maps from different perspectives are located on the same scan line. The corresponding points are the image points projected by the same object point in the object to be reconstructed on the imaging planes from different perspectives. Then, the disparity of the corresponding points from different perspectives is calculated, and the depth value of the object point corresponding to the corresponding point is calculated according to the disparity to obtain the depth value data of the object to be reconstructed. The point cloud calculation module 30 is used to obtain the two-dimensional contour of the object to be reconstructed according to the two-dimensional mask of the object to be reconstructed, and calculate the depth value of the two-dimensional contour of the object to be reconstructed pixel by pixel according to the depth value data of the object to be reconstructed, so as to map the two-dimensional contour information of the object to be reconstructed into three-dimensional space and generate the three-dimensional contour of the object to be reconstructed in the current frame, realizing point cloud segmentation.

[0130] In a preferred embodiment, the mask extraction module includes a segmentation arbitrary model unit, and the segmentation arbitrary model unit is used to extract the two-dimensional mask of the image frame frame by frame.

[0131] In a preferred embodiment, the mask extraction module includes a normalization processing unit, an image frame feature extraction unit, an image decoding unit, a mask feature calculation unit, a low-resolution image feature calculation unit, and a two-dimensional mask calculation unit. The normalization processing unit is configured to perform normalization processing on the image frames so that all image frames have the same resolution, obtaining normalized image frames; the image frame feature extraction unit is configured to perform multi-scale feature extraction on the normalized image frames. Meanwhile, position encoding is performed on the contour of the object to be reconstructed identified in the image frames to retain spatial position information; the image decoding unit is configured to decode the current image frame according to the multi-scale features and cache the object pointers of the first frame image, the current image frame, and its first M previous frames; the mask feature calculation unit is configured to calculate and cache the mask memory features and mask memory position encodings of the first N previous frames of the current frame for the normalized image frames other than the first frame; the low-resolution image feature calculation unit is configured to calculate the low-resolution image features of the current image frame according to the mask memory features, mask memory position encodings, the object pointers of the first frame image, the current image frame, and its first M previous frames cached in step S102, and the position encoding of the current image frame; the two-dimensional mask calculation unit is configured to decode the low-resolution image to obtain the low-resolution mask of the current image frame, and perform upsampling on the low-resolution mask to obtain the two-dimensional mask of the current image frame until all image frames are completely calculated, obtaining the two-dimensional mask of the object to be reconstructed.

[0132] In a preferred embodiment, for the first frame image, the image decoding unit is further configured to directly decode the multi-scale features of the first frame image, and determine the two-dimensional contour of the object to be reconstructed by identifying the mask of the object to be reconstructed in the first frame image.

[0133] In a preferred embodiment, the mask feature calculation unit is configured to extract pixel-by-pixel features for the normalized image frames other than the first frame, and encode the mask features and positions of the object to be reconstructed according to the pixel-by-pixel features, obtaining mask memory features and mask memory position encodings.

[0134] In a preferred embodiment, the depth value calculation module further includes a phase encoding map calculation unit, and the phase encoding map calculation unit is configured to extract phase information from the fringe image sequence through a phase unwrapping algorithm, generating a multi-view phase encoding map of the object to be reconstructed.

[0135] In a preferred embodiment, the depth value calculation module further includes an epipolar rectification calculation unit, and the epipolar rectification calculation unit is configured to adjust the projection matrices of multiple cameras through geometric transformation to align their optical axes, and the pixel coordinates satisfy the epipolar constraint condition, so that corresponding points in the phase encoding maps of different views are located on the same scan line.

[0136] In a preferred embodiment, the depth value calculation module further includes a depth value calculation unit, and the depth value calculation unit is configured to calculate the depth value of the object point corresponding to the homologous point according to the parallax, the calibration parameters of the camera, and the focal length.

[0137] It should be noted that in the present invention, step numbers (letter or number numbers) are used to refer to certain specific method steps, only for the purpose of convenience and brevity of description, and by no means to limit the order of these method steps by letters or numbers. Those skilled in the art can understand that the order of relevant method steps should be determined by the technology itself and should not be unduly restricted by the existence of step numbers.

[0138] Those skilled in the art can understand that on the premise of no conflict, the above preferred solutions can be freely combined and superimposed.

[0139] It should be understood that the above embodiments are merely exemplary and not restrictive. Without departing from the basic principles of the present invention, various obvious or equivalent modifications or substitutions made by those skilled in the art to the above details will be included within the scope of the claims of the present invention.

Claims

1. A point cloud segmentation method, characterized in that, Including the steps: S100, extracting the two-dimensional mask of each image frame frame by frame, where the image frames are pictures of the object to be reconstructed taken from different perspectives. Among them, the two-dimensional mask extracted from the previous image frame can be passed to subsequent image frames; S200, collecting the fringe image sequences of the object to be reconstructed from different perspectives and extracting phase information to generate a multi-view phase encoding map of the object to be reconstructed. Perform epipolar correction on the phase encoding maps of different perspectives to eliminate geometric distortions under different perspectives, so that corresponding points in the phase encoding maps of different perspectives are located on the same scan line. The corresponding points are the image points projected by the same object point in the object to be reconstructed on the imaging planes of different perspectives. Then calculate the disparity of the corresponding points under different perspectives, and calculate the depth value of the object point corresponding to the corresponding point according to the disparity to obtain the depth value data of the object to be reconstructed; S300, obtaining the two-dimensional contour of the object to be reconstructed according to the two-dimensional mask of the object to be reconstructed, and calculating the depth value of each pixel of the two-dimensional contour of the object to be reconstructed according to the depth value data of the object to be reconstructed, so as to map the two-dimensional contour information of the object to be reconstructed into three-dimensional space, generate the three-dimensional contour of the object to be reconstructed in the current frame, and realize point cloud segmentation.

2. The point cloud segmentation method according to claim 1, wherein Extract the two-dimensional mask of the image frame frame by frame based on an arbitrary segmentation model.

3. The point cloud segmentation method according to claim 1, wherein The two-dimensional mask of the object to be reconstructed is input in the first frame of the image frame in the form of cue point coordinates or a rectangular frame.

4. The point cloud segmentation method according to claim 1, characterized in that The S100 includes the steps: S101, normalizing the image frame so that the resolutions of all image frames are the same, obtaining a normalized image frame; S102, performing multi-scale feature extraction on the normalized image frame. At the same time, perform position encoding on the contour of the object to be reconstructed identified in the image frame to retain spatial position information. Decode the current image frame according to the multi-scale features, and cache the object pointers of the first frame image, the current image frame, and its first M frames of images; S103, calculating and caching the mask memory features and mask memory position encodings of the first N frames of the current frame for the normalized image frames other than the first frame; S104, calculating the low-resolution image features of the current image frame according to the mask memory features and mask memory position encodings, and the object pointers of the first frame image, the current image frame, and the first M frames cached in step S102 and the position encoding of the current image frame; S105, decoding the low-resolution image to obtain the low-resolution mask of the current image frame, and upsampling the low-resolution mask to obtain the two-dimensional mask of the current image frame until all image frames are calculated to obtain the two-dimensional mask of the object to be reconstructed.

5. The point cloud segmentation method according to claim 4, characterized in that, In step S102, for the first frame image, directly decode the multi-scale features of the first frame image, and determine the two-dimensional contour of the object to be reconstructed by identifying the mask of the object to be reconstructed in the first frame image.

6. The point cloud segmentation method according to claim 1, characterized in that In step 103, extract the per-pixel features for the normalized image frames other than the first frame, and encode the mask features and positions of the object to be reconstructed according to the per-pixel features to obtain the mask memory features and mask memory position encodings.

7. The point cloud segmentation method according to claim 1, wherein The M is 14 and the N is 7.

8. The point cloud segmentation method according to claim 1, wherein In step S200, phase information is extracted from the fringe image sequence through a phase unwrapping algorithm to generate multi-view phase encoded maps of the object to be reconstructed.

9. The point cloud segmentation method according to claim 1, characterized in that, Epipolar rectification of the phase encoded maps from different views in step S200 includes: adjusting the projection matrices of multiple cameras through geometric transformation to align their optical axes and make the pixel coordinates satisfy the epipolar constraint condition, so that corresponding points in the phase encoded maps from different views are located on the same scan line.

10. The point cloud segmentation method according to claim 1, wherein, In step S200, the disparity is obtained by calculating the horizontal distance difference of the corresponding points.

11. The point cloud segmentation method according to claim 10, characterized in that, The depth value of the object point corresponding to the corresponding point is calculated based on the disparity, the calibration parameters of the camera, and the focal length.

12. The point cloud segmentation method according to claim 11, characterized in that The calibration parameter of the camera is the baseline distance between the optical centers of multiple cameras.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed, it can implement the point cloud segmentation method described in any one of claims 1-8.

14. A three-dimensional scanning system, characterized in that, The system uses the point cloud segmentation method described in any one of claims 1-12 to perform point cloud segmentation on the object to be reconstructed, so as to realize three-dimensional reconstruction of the object.

15. A point cloud segmentation device, characterized in that, It includes a mask extraction module, a depth value calculation module, and a point cloud calculation module. The mask extraction module is used to extract the two-dimensional mask of the image frame frame by frame. The image frame is a picture of the object to be reconstructed taken from different views. Among them, the two-dimensional mask extracted from the previous image frame can be transmitted to the subsequent image frame. The depth value calculation module is used to collect the fringe image sequences of the object to be reconstructed from different views and extract phase information to generate multi-view phase encoded maps of the object to be reconstructed. Epipolar rectification is performed on the phase encoded maps from different views to eliminate geometric distortion under different views, so that corresponding points in the phase encoded maps from different views are located on the same scan line. The corresponding points are the image points projected by the same object point in the object to be reconstructed on the imaging planes of different views. Then, the disparity of the corresponding points under different views is calculated, and the depth value of the object point corresponding to the corresponding point is calculated based on the disparity to obtain the depth value data of the object to be reconstructed. The point cloud calculation module is used to obtain the two-dimensional contour of the object to be reconstructed according to the two-dimensional mask of the object to be reconstructed, and calculate the depth value of each pixel of the two-dimensional contour of the object to be reconstructed according to the depth value data of the object to be reconstructed, so as to map the two-dimensional contour information of the object to be reconstructed into three-dimensional space and generate the three-dimensional contour of the object to be reconstructed in the current frame, realizing point cloud segmentation.

16. The point cloud segmentation device according to claim 15, wherein The mask extraction module includes an arbitrary model segmentation unit, and the arbitrary model segmentation unit is used to extract the two-dimensional mask of the image frame frame by frame.

17. The point cloud segmentation device according to claim 15, wherein The mask extraction module includes a normalization processing unit, an image frame feature extraction unit, an image decoding unit, a mask feature calculation unit, a low-resolution image feature calculation unit, and a two-dimensional mask calculation unit. The normalization processing unit is used to perform normalization processing on the image frame so that the resolutions of all image frames are the same, obtaining a normalized image frame. The image frame feature extraction unit is used to perform multi-scale feature extraction on the normalized image frame. At the same time, position encoding is performed on the contour of the object to be reconstructed identified in the image frame to retain spatial position information. The image decoding unit is used to decode the current image frame according to the multi-scale features, and cache the object pointers of the first frame image, the current image frame, and the previous M frame images; The mask feature calculation unit is used to calculate and cache the mask memory features and mask memory position encodings of the previous N frames of the current frame for the normalized image frames except the first frame; The low-resolution image feature calculation unit is used to calculate the low-resolution image features of the current image frame according to the mask memory features, the mask memory position encodings, the object pointers of the first frame image, the current image frame, and the previous M frame images cached in step S102, and the position encoding of the current image frame; The two-dimensional mask calculation unit is used to decode the low-resolution image to obtain the low-resolution mask of the current image frame, and upsample the low-resolution mask to obtain the two-dimensional mask of the current image frame until all image frames are completely calculated, so as to obtain the two-dimensional mask of the object to be reconstructed.

18. The point cloud segmentation device according to claim 17, characterized in that, For the first frame image, the image decoding unit is further used to directly decode the multi-scale features of the first frame image, and determine the two-dimensional contour of the object to be reconstructed by identifying the mask of the object to be reconstructed in the first frame image.

19. The point cloud segmentation device according to claim 17, characterized in that, The mask feature calculation unit is used to extract per-pixel features from the normalized image frames except the first frame, and encode the mask features and positions of the object to be reconstructed according to the per-pixel features to obtain mask memory features and mask memory position encodings.

20. The point cloud segmentation device according to claim 15, wherein The depth value calculation module further includes a phase encoding map calculation unit, and the phase encoding map calculation unit is used to extract phase information from the fringe image sequence through a phase unwrapping algorithm to generate a multi-view phase encoding map of the object to be reconstructed.

21. The point cloud segmentation device according to claim 15, wherein The depth value calculation module further includes an epipolar rectification calculation unit, and the epipolar rectification calculation unit is used to adjust the projection matrices of multiple cameras through geometric transformation to align their optical axes and make the pixel coordinates satisfy the epipolar constraint condition, so that corresponding points in the phase encoding maps of different views are located on the same scan line.

22. The point cloud segmentation device according to claim 15, characterized in that, The depth value calculation module further includes a depth value calculation unit, and the depth value calculation unit is used to calculate the depth value of the object point corresponding to the corresponding point according to the parallax, the calibration parameters of the camera, and the focal length.

Citation Information

Cited By

  • Automatic structured light point cloud substrate removing method based on frame-by-frame analysis

    CN121725161A

  • Automatic Basis Removal Method for Structured Light Point Clouds Based on Frame-by-Frame Analysis

    CN121725161B