A Zero-Shot Point Cloud Instance Segmentation Method and System Based on Multi-View Consistency

The method addresses view consistency and long-distance projection errors in 3D point cloud instance segmentation by using key frames, SAM2 model segmentation, and iterative graph clustering to enhance segmentation accuracy and precision.

CN120070901BActive Publication Date: 2025-07-15UNIV OF SCI & TECH BEIJING
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510546772.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-07-15
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

The existing three-dimensional point cloud instance segmentation method relies on high-quality labeled data, which is expensive to label. The viewing angle consistency problem and long-distance projection error affect the segmentation accuracy, especially in complex scenarios and large-scale point cloud data processing.

Method used

Using the zero-sample point cloud instance segmentation method based on multi-view consistency, the keyframes are extracted through the SAM2 model to generate a 2D mask, a hyperpoint map is constructed and the mask consistency weight is assigned. Multi-level hyperpoint iterative map clustering is used for layered merge to solve the perspective consistency and long-distance projection errors.

Benefits of technology

The accuracy and robustness of three-dimensional point cloud instance segmentation are significantly improved. Through the complex memory mechanism of the SAM2 model and the mask consistency weight correction, the consistency and accuracy of the segmentation are enhanced, especially in complex scenarios and large-scale point cloud data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070901B_ABST
    Figure CN120070901B_ABST
Patent Text Reader

Abstract

The present invention provides a zero-shot point cloud instance segmentation method and system based on multi-view consistency, including: extracting multiple key frames from a two-dimensional image sequence corresponding to a three-dimensional point cloud; using the 2D masks generated by segmenting the key frames as cues to segment and track the two-dimensional image sequence to obtain 2D masks of all objects; performing super-segmentation on the three-dimensional point cloud to obtain a series of super points, performing 3D-to-2D projection to obtain the 2D masks of each super point on each image, representing the edges between super points by the affinity scores between the 2D masks, and constructing a super point graph; using mask consistency weights to correct the affinity scores, giving low-quality masks low weights to obtain a corrected super point graph; and through multi-level super point iterative graph clustering, hierarchically merging super points with different levels of view consistency to obtain a 3D instance segmentation result. The present invention can perform instance segmentation on 3D point clouds.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of zero-shot point cloud instance segmentation, and particularly to a zero-shot point cloud instance segmentation method and system based on multi-view consistency. Background Art

[0002] The 3D point cloud instance segmentation task refers to dividing different points in three-dimensional point cloud data according to certain rules or objectives, so as to realize the recognition and separation of different objects, regions or categories. Point cloud instance segmentation has important applications in fields such as autonomous driving, robot navigation, and augmented reality. Compared with two-dimensional images, three-dimensional data (such as point cloud data, lidar scan data, etc.) provides richer spatial information, which helps to improve the accuracy and reliability of segmentation. However, at the same time, three-dimensional perception faces a series of new challenges and problems.

[0003] First of all, existing three-dimensional perception methods often rely on a large amount of high-quality labeled data, especially for the labeling of three-dimensional point clouds, which usually requires a large amount of manual intervention and has a high labeling cost. In some application scenarios, it is very difficult to obtain high-quality labeled data, which makes traditional supervised learning methods inapplicable. To address this problem, researchers have begun to explore zero-shot learning methods, hoping to achieve the segmentation of new categories through existing models and data. In addition, researchers have also begun to explore transferring the powerful generalization ability of two-dimensional vision models to three-dimensional tasks, attempting to use the generalization ability of these models in different scenarios and categories to improve the effect of three-dimensional object detection. However, the application of zero-shot learning in point cloud instance segmentation still faces great challenges.

[0004] Combining two-dimensional image information with three-dimensional point cloud data has become an effective solution. Two-dimensional images have rich detail information and can provide higher resolution, which has an important auxiliary role in fine-grained object segmentation. By fusing the two-dimensional image segmentation results with the three-dimensional point cloud data, the accuracy and robustness of point cloud instance segmentation can be improved.

[0005] An urgent problem to be solved is the perspective consistency problem. In 3D point cloud instance segmentation, existing methods usually introduce 2D image sequences to assist in completing point cloud instance segmentation. Image segmentation is completed through multi-frame data, and then the segmentation results are projected onto the point cloud. However, due to the appearance changes and occlusion phenomena of the target under different perspectives, the target segmentation results among multi-frame data are inconsistent. This perspective inconsistency directly affects the final segmentation accuracy. In addition, in 3D perception tasks, especially in the process of projecting from 2D images to 3D point clouds, as the object moves away from the camera, the error introduced by long-distance projection becomes a significant challenge. Existing research has not effectively solved this problem. Especially in the process of matching 2D images to 3D point clouds, distant objects and nearby objects are given the same weight during projection. This can lead to the incorrect matching of 3D objects with the 2D masks of another object, resulting in incorrect 3D segmentation. This problem worsens as the object moves away from the camera, ultimately leading to a decrease in segmentation accuracy. This challenge, especially in the process of processing complex scenes and large-scale point cloud data, seriously affects the accuracy of the point cloud instance segmentation task. Therefore, how to effectively solve the error problem caused by long-distance projection has also become a key difficult problem in current research. Summary of the Invention

[0006] To solve the above technical problems existing in the prior art, the present invention provides a zero-shot point cloud instance segmentation method and system based on multi-perspective consistency. The technical solutions are as follows:

[0007] On the one hand, a zero-shot point cloud instance segmentation method based on multi-perspective consistency is provided. The method includes:

[0008] S1. Obtain a 3D point cloud to be subjected to instance segmentation and the corresponding 2D image sequence;

[0009] S2. Extract multiple key frames from the 2D image sequence;

[0010] S3. Use the 2D masks generated by segmenting the key frames as prompts, and use the SAM2 model to segment and track the 2D image sequence to obtain the 2D masks of all objects in the 2D image sequence;

[0011] S4. Perform super-segmentation on the 3D point cloud to obtain a series of super points, perform 3D to 2D projection, obtain the 2D masks of each super point on each image, represent the edges between super points by the affinity scores between the 2D masks corresponding to the super points, and construct a super point graph based on the super points and the edges between the super points. The super points serve as the nodes of the super point graph, and the edges between the super points serve as the edges between the nodes;

[0012] S5. Use the mask consistency weight to correct the affinity scores, assign low weights to low-quality masks, and obtain the corrected super point graph;

[0013] S6. Through multi-level superpoint iterative graph clustering, hierarchically merge superpoints with different levels of perspective consistency to obtain the final 3D instance segmentation result.

[0014] Optionally, S2 specifically includes:

[0015] S21. Treat the two-dimensional image sequence as a video, and calculate the difference between two adjacent frames , and the calculation method is as follows:

[0016]

[0017] where represents the pixel matrix of the current frame, represents the pixel matrix of the previous frame, and represent the height and width of the image respectively;

[0018] S22. Use the Hanning window to smooth the frame difference array, and the calculation formula of the window weight is as follows:

[0019]

[0020] where N is the length of the window, n is the index within the window, is the value of the window function at position n, and define the smoothed array as S. The value at the j-th position in S is expressed as:

[0021]

[0022] where represents the value at the j - n position in . The index corresponding to the local maximum value in the array S represents the key frame, which is the frame where a new object appears.

[0023] Optionally, S3 specifically includes:

[0024] S31. Use the image predictor of the SAM2 model to automatically segment each key frame and generate a 2D mask of the object, where the object includes new objects that appear in each key frame;

[0025] S32. Treat the image sequence between every two key frames as a video clip. The first frame and the last frame of the video clip are key frames respectively. Use the video predictor of the SAM2 model, and use the 2D masks generated by segmenting each key frame as hints to track the object in the video clip using a two-way verification tracking strategy;

[0026] S33. At the edges of adjacent video segments, use IoU to connect the same objects in different segments, and finally obtain the 2D masks of all objects in the entire 2D image sequence.

[0027] Optionally, the two-way verification tracking strategy includes forward tracking and backward tracking. The targets missed in the forward tracking process will be identified in the backward tracking, and the targets missed in the backward tracking will be identified in the forward tracking. The tracking results from the two directions complement each other, and the 2D masks from the two directions are merged.

[0028] Optionally, the S4 specifically includes:

[0029] S41. Use the graph cut algorithm to perform super-segmentation on the 3D point cloud and generate a series of 3D superpoints based on 3D geometric properties;

[0030] S42. Use a general pinhole camera matrix for 3D to 2D projection, and use the corresponding camera internal parameters and pose parameters to project the i-th 3D superpoint onto the k-th image and obtain its corresponding 2D mask:

[0031]

[0032] where is the projection operator;

[0033] S43. Calculate the normalized histogram based on the 2D mask and convert it into a vector, denoted as . This histogram vector represents the 2D instance features of the 3D superpoint. Obtain the affinity score by calculating the cosine similarity between the 2D instance features of different superpoints on the k-th image, as follows:

[0034]

[0035] S44. Sum the affinity scores of each image to obtain the affinity score between all 2D masks corresponding to the i-th superpoint and the j-th superpoint, which is used as the edge between the i-th superpoint and the j-th superpoint, as follows:

[0036] .

[0037] Optionally, the S5 specifically includes:

[0038] S51. Calculate and obtain the visibility edge weight:

[0039] Project the points in the superpoint onto the image, filter out the points outside the image, and filter out the occluded points according to the depth of the points in the superpoint to obtain visible points. Define as the superpoint The ratio of the number of visible points on the k-th image to the total number of points in the superpoint. Define the visibility edge weight between the i-th superpoint and the j-th superpoint as:

[0040]

[0041] S52. Calculate the distance edge weight:

[0042] Define as the distance weighted average based on the k-th image, which is used to weight average the 2D mask :

[0043]

[0044] where represents the distance from the visible point in the i-th superpoint to the k-th image

[0045] and represent the maximum and minimum distances from all visible points to the k-th image respectively. Define the distance edge weight between the i-th superpoint and the j-th superpoint as:

[0046]

[0047] S53. Calculate the purity edge weight:

[0048] The number of 2D masks corresponding to the superpoint is defined as its purity. Define as the maximum ratio of the superpoint projected onto different 2D masks in the k-th image:

[0049]

[0050] where represents the total number of visible points of the superpoint in the k-th image, represents the number of points belonging to the mask label value of 1, represents the number of points belonging to the mask label value of 2, represents the number of points belonging to the mask label value of n, and n is the total number of 2D masks. Define the purity edge weight between the i-th superpoint and the j-th superpoint as:

[0051]

[0052] S54. Combine the visibility edge weight, distance edge weight, and purity edge weight to form a mask consistency weight:

[0053]

[0054] S55. Obtain the final affinity score according to the mask consistency weight, and assign a low weight to the low-quality mask. The low-quality mask includes 2D masks with lower visibility, greater distance, and higher purity. The calculation formula is:

[0055] 。

[0056] Optionally, the multi-level superpoint iterative graph clustering specifically includes:

[0057] In each iteration of clustering, set different thresholds according to the affinity score to merge superpoints. Give priority to processing superpoints with high view consistency. Superpoints with high view consistency whose affinity score is greater than the threshold are merged first, and then gradually lower the threshold to merge superpoints with low view consistency.

[0058] On the other hand, a zero-shot point cloud instance segmentation system based on multi-view consistency is provided. The system includes:

[0059] An acquisition module for acquiring a three-dimensional point cloud to be subjected to instance segmentation and a corresponding two-dimensional image sequence;

[0060] A key frame extraction module for extracting multiple key frames from the two-dimensional image sequence;

[0061] A 2D mask acquisition module for using the 2D masks generated by segmenting the key frames as prompts, segmenting and tracking the two-dimensional image sequence using the SAM2 model, and obtaining 2D masks of all objects in the two-dimensional image sequence;

[0062] A superpoint graph construction module for performing super-segmentation on the three-dimensional point cloud to obtain a series of superpoints, performing 3D to 2D projection, obtaining the 2D masks of each superpoint on each image, representing the edges between superpoints by the affinity scores between the 2D masks corresponding to the superpoints, and constructing a superpoint graph based on the superpoints and the edges between the superpoints. The superpoints are used as the nodes of the superpoint graph, and the edges between the superpoints are used as the edges between the nodes;

[0063] A superpoint graph correction module for correcting the affinity scores using the mask consistency weight, assigning a low weight to the low-quality masks, and obtaining a corrected superpoint graph;

[0064] A 3D instance segmentation module for hierarchically merging superpoints with different view consistency levels through multi-level superpoint iterative graph clustering to obtain the final 3D instance segmentation result.

[0065] On the other hand, an electronic device is provided, which includes a processor and a memory. At least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the above zero-shot point cloud instance segmentation method based on multi-view consistency.

[0066] On the other hand, a computer-readable storage medium is provided. At least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement the above zero-shot point cloud instance segmentation method based on multi-view consistency.

[0067] The beneficial effects brought by the technical solution provided by the present invention at least include:

[0068] 1) The zero-shot point cloud instance segmentation method based on multi-view consistency of the present invention utilizes the complex memory mechanism of the SAM2 model, and proposes a segmented processing method based on key frames. The image sequence between every two key frames is regarded as an independent video segment. The key frames are automatically segmented by the image predictor of SAM2 to generate 2D masks (including 2D masks of new objects appearing in each key frame), and these are used as cues. Even when an object is occluded or re-enters the scene (i.e., a new object appears), the video predictor is used to achieve continuous segmentation and tracking of all objects between video frames (existing methods use the objects in the first frame as cues for tracking, and when new objects appear later, they cannot be tracked). And the target omission problem is solved through a two-way verification tracking strategy to maintain the consistency of segmentation and tracking. This mechanism significantly enhances the segmentation accuracy under multiple views.

[0069] 2) The mask consistency weight designed by the present invention not only effectively solves the long-distance projection error problem, but also assigns low weights to low-quality masks from multiple angles, thereby correcting the relationship (edges) between superpoints in the superpoint map, so as to obtain a more accurate segmentation result when iteratively clustering and merging superpoints at the end (the super-segmentation of the three-dimensional point cloud used in the present invention to obtain a series of superpoints (a superpoint is a set of multiple points with similar features) is an existing method, and the super-segmentation of this method is over-segmentation, that is, different parts of an instance may be segmented into different superpoints. For example, the table instance is over-segmented, and the table legs and the table top are segmented into different superpoints. Subsequently, it is necessary to perform combined segmentation on the over-segmented superpoints belonging to the same instance).

[0070] 3) The multi-level superpoint iterative graph clustering designed by the present invention dynamically adjusts the affinity score threshold according to the view consistency level during the iterative clustering process, and preferentially processes superpoints with high view consistency. Through the iterative graph clustering and hierarchical merging strategy, the point cloud instance segmentation performance is significantly enhanced. Description of the Drawings

[0071] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0072] Figure 1 is a flowchart of a zero-shot point cloud instance segmentation method based on multi-view consistency provided by an embodiment of the present invention;

[0073] Figure 2 is another flowchart of a zero-shot point cloud instance segmentation method based on multi-view consistency provided by an embodiment of the present invention;

[0074] Figure 3 is a flowchart of two-dimensional image sequence segmentation and tracking provided by an embodiment of the present invention;

[0075] Figure 4 is a block diagram of a zero-shot point cloud instance segmentation system based on multi-view consistency provided by an embodiment of the present invention;

[0076] Figure 5 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0077] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.

[0078] An embodiment of the present invention provides a zero-shot point cloud instance segmentation method based on multi-view consistency. This method can be implemented by an electronic device, which can be a terminal or a server. Figure 1 and Figure 2 As shown in the flowchart of this method, the processing flow can include the following steps:

[0079] S1. Obtain a three-dimensional point cloud to be segmented into instances and a corresponding two-dimensional image sequence;

[0080] S2. Extract multiple key frames from the two-dimensional image sequence;

[0081] Optionally, step S2 specifically includes:

[0082] S21. Regard the two-dimensional image sequence as a video, calculate the difference between two adjacent frames , and the calculation method is as follows:

[0083]

[0084] where represents the pixel matrix of the current frame, represents the pixel matrix of the previous frame, and respectively represent the height and width of the image;

[0085] When the scene changes or a new object appears, the pixel difference between the previous frame and the current frame will increase significantly. Embodiments of the present invention use the pixel differences between these frames to identify key frames. In this way, the frames where new objects are located can be identified and marked, facilitating subsequent segmentation and tracking.

[0086] S22. Smooth the frame difference array using a Hanning window. The calculation formula for the window weight is as follows:

[0087]

[0088] where N is the length of the window, n is the index within the window, is the value of the window function at position n. Define the smoothed array as S. The value at the j-th position in S is expressed as:

[0089]

[0090] where, represents the value at the j - n position in . The index corresponding to the local maximum value in the array S represents the key frame, indicating the frame where a new object appears.

[0091] S3. Use the 2D mask generated by segmenting the key frame as a prompt, and use the SAM2 model to segment and track the two-dimensional image sequence to obtain the 2D masks of all objects in the two-dimensional image sequence;

[0092] Optionally, as Figure 3 shown, S3 specifically includes:

[0093] S31. Use the image predictor of the SAM2 model to automatically segment each key frame and generate the 2D mask of the object, where the object includes the new objects that appear in each key frame;

[0094] S32. Regard the image sequence between every two key frames as a video segment. The first frame and the last frame of the video segment are key frames respectively. Use the video predictor of the SAM2 model, and use the 2D masks generated by segmenting each key frame as prompts to track the object in the video segment using a two-way verification tracking strategy;

[0095] S33. At the edges of adjacent video segments, use IoU to connect the same objects in different segments, and finally obtain the 2D masks of all objects in the entire 2D image sequence.

[0096] Existing methods use the objects in the first frame as cues for tracking. When new objects appear later, they cannot be tracked. However, in the embodiments of the present invention, the 2D masks generated by segmenting each key frame are used as cues, so that tracking of all objects (including new objects that appear in each key frame) can be achieved.

[0097] Optionally, the two-way verification tracking strategy includes forward tracking and backward tracking. The targets missed during forward tracking will be identified during backward tracking, and the targets missed during backward tracking will be identified during forward tracking. The tracking results from the two directions complement each other, and the 2D masks from the two directions are merged.

[0098] Although the embodiments of the present invention can obtain new objects by segmenting key frames, the key frames are not always the exact frames where new objects appear (there may be incorrect segmentation). If new objects appear before or after the key frames, they may be missed or omitted. Therefore, the embodiments of the present invention design a two-way verification tracking strategy to solve this problem.

[0099] S4. Perform super-segmentation on the 3D point cloud to obtain a series of super points, perform 3D-to-2D projection, obtain the 2D masks of each super point on each image, represent the edges between super points by the affinity scores between the 2D masks corresponding to the super points, and construct a super point graph based on the super points and the edges between super points. The super points are used as the nodes of the super point graph, and the edges between super points are used as the edges between nodes (the super point graph is represented as , where v represents a super point and E represents the edges between super points);

[0100] Optionally, S4 specifically includes:

[0101] S41. Use a graph cut algorithm to perform super-segmentation on the 3D point cloud and generate a series of 3D super points based on 3D geometric attributes;

[0102] The method used in the present invention to perform super-segmentation on the 3D point cloud to obtain a series of super points (a super point is a set of multiple points with similar characteristics) is an existing method. The super-segmentation of this method is over-segmentation, that is, different parts of an instance may be segmented into different super points. For example, if a table instance is over-segmented, the table legs and the table top are segmented into different super points. Subsequently, it is necessary to merge the over-segmented super points belonging to the same instance to obtain an accurate instance segmentation result and complete the point cloud instance segmentation.

[0103] S42. Use a general pinhole camera matrix for 3D-to-2D projection and utilize the corresponding camera internal parameters and pose parameters Project the i-th 3D superpoint onto the k-th image and obtain its corresponding 2D mask:

[0104]

[0105] where is the projection operator;

[0106] S43. Based on the 2D mask Calculate the normalized histogram and convert it into a vector, denoted as , and this histogram vector represents the 2D instance feature of the 3D superpoint. The affinity score is obtained by calculating the cosine similarity between the 2D instance features of different superpoints on the k-th image , as follows:

[0107]

[0108] S44. Sum the affinity scores for each image to obtain the affinity score between all 2D masks corresponding to the i-th superpoint and the j-th superpoint , as the edge between the i-th superpoint and the j-th superpoint, as follows:

[0109] .

[0110] S5. Use the mask consistency weight to correct the affinity score, assign a low weight to the low-quality mask, and obtain the corrected superpoint graph;

[0111] Optionally, S5 specifically includes:

[0112] S51. Calculate and obtain the visibility edge weight:

[0113] Project the points in the superpoint onto the image, filter out the points outside the image, and filter out the occluded points according to the depth of the points in the superpoint to obtain the visible points. Define as the ratio of the number of visible points of the superpoint on the k-th image to the total number of points in the superpoint. Define the visibility edge weight between the i-th superpoint and the j-th superpoint as:

[0114]

[0115] S52. Calculate and obtain the distance edge weight:

[0116] Define as the distance weighted average based on the k-th image, used to weight and average the 2D mask :

[0117]

[0118] Among them, represents the distance from the visible points in the i-th superpoint to the k-th image,

[0119] and respectively represent the maximum and minimum distances from all visible points to the k-th image. The distance edge weight between the i-th superpoint and the j-th superpoint is defined as:

[0120]

[0121] S53. Calculate and obtain the purity edge weight:

[0122] The number of 2D masks corresponding to the superpoint is defined as its purity. Define as the maximum ratio of the superpoint projected onto different 2D masks in the k-th image:

[0123]

[0124] Among them, represents the total number of visible points of the superpoint in the k-th image, represents the number of points belonging to the mask label value of 1, represents the number of points belonging to the mask label value of 2, represents the number of points belonging to the mask label value of n, where n is the total number of 2D masks. The purity edge weight between the i-th superpoint and the j-th superpoint is defined as:

[0125]

[0126] S54. Combine the visibility edge weight, distance edge weight, and purity edge weight to form a mask consistency weight:

[0127]

[0128] S55. According to the mask consistency weight, obtain the final affinity score, and assign a low weight to the low-quality masks. The low-quality masks include 2D masks with lower visibility, farther distance, and higher purity. The calculation formula is:

[0129] .

[0130] S6. Through multi-level superpoint iterative graph clustering, hierarchically merge the superpoints with different levels of perspective consistency to obtain the final 3D instance segmentation result.

[0131] Optionally, the multi-level superpoint iterative graph clustering specifically includes:

[0132] In each iterative clustering, different thresholds are set according to the affinity scores to merge superpoints. Superpoints with high view consistency are processed first. Superpoints with high view consistency and affinity scores greater than the threshold are merged first, and then the threshold is gradually reduced to merge superpoints with low view consistency.

[0133] The multi-level superpoint iterative graph clustering designed by the present invention dynamically adjusts the affinity score threshold according to the view consistency level during the iterative clustering process, gives priority to processing superpoints with high view consistency, and significantly enhances the point cloud instance segmentation performance through the iterative graph clustering and hierarchical merging strategy.

[0134] As Figure 4 shown, an embodiment of the present invention also provides a zero-shot point cloud instance segmentation system based on multi-view consistency. The system includes:

[0135] An acquisition module 410, configured to acquire a three-dimensional point cloud to be subjected to instance segmentation and a corresponding two-dimensional image sequence;

[0136] A key frame extraction module 420, configured to extract multiple key frames from the two-dimensional image sequence;

[0137] A 2D mask acquisition module 430, configured to use the 2D mask generated by segmenting the key frames as a prompt, and use the SAM2 model to segment and track the two-dimensional image sequence to obtain 2D masks of all objects in the two-dimensional image sequence;

[0138] A superpoint graph construction module 440, configured to perform super-segmentation on the three-dimensional point cloud to obtain a series of superpoints, perform 3D to 2D projection, obtain 2D masks of each superpoint on each image, represent the edges between superpoints by the affinity scores between the 2D masks corresponding to the superpoints, and construct a superpoint graph according to the superpoints and the edges between the superpoints. The superpoints are used as the nodes of the superpoint graph, and the edges between the superpoints are used as the edges between the nodes;

[0139] A superpoint graph correction module 450, configured to correct the affinity scores using the mask consistency weights, assign low weights to low-quality masks, and obtain a corrected superpoint graph;

[0140] A 3D instance segmentation module 460, configured to perform hierarchical merging on superpoints with different view consistency levels through multi-level superpoint iterative graph clustering to obtain a final 3D instance segmentation result.

[0141] The function structure of a zero-shot point cloud instance segmentation system based on multi-view consistency provided by an embodiment of the present invention corresponds to a zero-shot point cloud instance segmentation method provided by an embodiment of the present invention, and will not be described in detail herein.

[0142] Figure 5 It is a schematic structural diagram of an electronic device 500 provided by an embodiment of the present invention. The electronic device 500 may vary greatly due to different configurations or performances, and may include one or more processors (central processing units, CPUs) 501 and one or more memories 502. Among them, at least one instruction is stored in the memory 502, and the at least one instruction is loaded and executed by the processor 501 to implement the steps of the above-mentioned zero-shot point cloud instance segmentation method based on multi-view consistency.

[0143] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions, and the above instructions can be executed by a processor in a terminal to complete the above-mentioned zero-shot point cloud instance segmentation method based on multi-view consistency. For example, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0144] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware, or can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk, or an optical disc, etc.

[0145] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A zero-shot point cloud instance segmentation method based on multi-view consistency, characterized in that The method includes: S1. Obtain a three-dimensional point cloud to be subjected to instance segmentation and a corresponding two-dimensional image sequence; S2. Extract multiple key frames from the two-dimensional image sequence; S3. Use the 2D masks generated by segmenting the key frames as prompts, and use the SAM2 model to segment and track the two-dimensional image sequence to obtain the 2D masks of all objects in the two-dimensional image sequence; S4. Perform super-segmentation on the three-dimensional point cloud to obtain a series of super points, perform 3D-to-2D projection, obtain the 2D masks of each super point on each image, represent the edges between super points by the affinity scores between the 2D masks corresponding to the super points, and construct a super point graph based on the super points and the edges between super points, where the super points are used as the nodes of the super point graph and the edges between super points are used as the edges between nodes; S5. Use the mask consistency weight to correct the affinity scores, assign low weights to low-quality masks, and obtain a corrected super point graph; S6. Through multi-level super point iterative graph clustering, hierarchically merge the super points with different levels of perspective consistency to obtain the final 3D instance segmentation result; The S3 specifically includes: S31. Use the image predictor of the SAM2 model to automatically segment each key frame and generate the 2D masks of the objects, where the objects include new objects that appear in each key frame; S32. Regard the image sequence between every two key frames as a video segment, where the first frame and the last frame of the video segment are key frames respectively. Use the video predictor of the SAM2 model, use the 2D masks generated by segmenting each key frame as prompts, and track the objects in the video segment using a two-way verification tracking strategy; S33. At the edges of adjacent video segments, use IoU to connect the same objects in different segments, and finally obtain the 2D masks of all objects in the entire two-dimensional image sequence.

2. The method according to claim 1, wherein The S2 specifically includes: S21. Regard the two-dimensional image sequence as a video and calculate the difference between two adjacent frames , and the calculation method is as follows: ; Among them, represents the pixel matrix of the current frame, represents the pixel matrix of the previous frame, and respectively represent the height and width of the image; S22. Use a Hanning window to smooth the frame difference array, and the calculation formula of the window weight is as follows: ; where N is the length of the window, and n is the index within the window, is the value of the window function at position n. Define the smoothed array as S, and the value at the j-th position in S is expressed as: ; Among them, represents the value at the j-n position in The index corresponding to the local maximum value in the array S represents the key frame, which indicates the frame where a new object appears.

3. The method according to claim 1, wherein The two-way verification tracking strategy includes forward tracking and backward tracking. The targets missed in the forward tracking process will be identified in the backward tracking, and the targets missed in the backward tracking will be identified in the forward tracking. The tracking results from the two directions complement each other, and the 2D masks from the two directions are merged.

4. The method according to claim 1, characterized in that, The S4 specifically includes: S41. Use a graph cut algorithm to perform super-segmentation on the three-dimensional point cloud and generate a series of 3D super points based on 3D geometric attributes; S42. Perform 3D-to-2D projection using a general pinhole camera matrix, and utilize the corresponding camera intrinsics and pose parameters to project the i-th 3D superpoint onto the k-th image and obtain its corresponding 2D mask: ; Among them, is a projection operator; S43. Based on 2D mask Calculate the normalized histogram and convert it into a vector, denoted as , this histogram vector represents the 2D instance features of 3D superpoints, and the affinity score is obtained by calculating the cosine similarity between the 2D instance features of different superpoints on the k-th image , as the following formula: ; S44. Sum the affinity scores for each image to obtain the affinity score between all 2D masks corresponding to the $i$-th superpoint and the $j$-th superpoint, which is used as the edge between the $i$-th superpoint and the $j$-th superpoint, as shown in the following formula: , as the edge between the $i$-th superpoint and the $j$-th superpoint, in the following formula: 。 5. The method according to claim 4, characterized in that, The S5 specifically includes: S51. Calculate and obtain the visibility edge weight: Project the points in the superpoint onto the image, filter out the points outside the image, and filter out the occluded points according to the depth of the points in the superpoint to obtain visible points. Define as the superpoint The ratio of the number of visible points on the k-th image to the total number of points in the superpoint. Define the visibility edge weight between the i-th superpoint and the j-th superpoint as: ; S52. Calculate and obtain the distance edge weight: Definition is the distance weighted average based on the k-th image, used for weighted averaging the 2D mask : ; Among them, represents the distance from the visible points in the \(i\)-th superpoint to the \(k\)-th image, and respectively represent the maximum and minimum distances from all visible points to the \(k\)-th image. The edge weight of the distance between the \(i\)-th superpoint and the \(j\)-th superpoint is defined as: ; S53. Calculate and obtain the purity edge weight: 2D mask corresponding to the superpoint The number of which is defined as its purity, and defined as the superpoint in the k-th image as the maximum ratio projected onto different 2D masks: ; Among them, represents the total number of visible points of the superpoint in the k-th image, represents the number of points belonging to the mask label value of 1, represents the number of points belonging to the mask label value of 2, represents the number of points belonging to the mask label value of n, where n is the total number of 2D masks. Define the purity edge weight between the i-th superpoint and the j-th superpoint as: ; S54. Merge the visibility edge weight, the distance edge weight, and the purity edge weight to form a mask consistency weight: ; S55. According to the mask consistency weight, obtain the final affinity scores, and assign low weights to low-quality masks. The low-quality masks include 2D masks with lower visibility, greater distance, and higher purity. The calculation formula is: 。 6. The method according to claim 1, wherein The multi-level super point iterative graph clustering specifically includes: In each iterative clustering, different thresholds are set according to the affinity scores to merge the superpoints. The superpoints with high view consistency are processed first. The superpoints with high view consistency and affinity scores greater than the threshold are merged first, and then the threshold is gradually decreased to merge the superpoints with low view consistency.

7. A zero-shot point cloud instance segmentation system based on multi-view consistency, characterized in that, The system includes: An acquisition module, configured to acquire a three-dimensional point cloud to be subjected to instance segmentation and a corresponding two-dimensional image sequence; A key frame extraction module, configured to extract a plurality of key frames from the two-dimensional image sequence; A 2D mask obtaining module, configured to use the 2D masks generated by segmenting the key frames as prompts, and use the SAM2 model to segment and track the two-dimensional image sequence to obtain the 2D masks of all objects in the two-dimensional image sequence; A superpoint graph construction module, configured to perform super-segmentation on the three-dimensional point cloud to obtain a series of superpoints, perform 3D to 2D projection, obtain the 2D masks of each superpoint on each image, represent the edges between the superpoints by the affinity scores between the 2D masks corresponding to the superpoints, and construct a superpoint graph according to the superpoints and the edges between the superpoints, with the superpoints as the nodes of the superpoint graph and the edges between the superpoints as the edges between the nodes; A superpoint graph correction module, configured to correct the affinity scores using the mask consistency weights, assign low weights to the low-quality masks, and obtain a corrected superpoint graph; A 3D instance segmentation module, configured to perform hierarchical merging on the superpoints with different view consistency levels through multi-level superpoint iterative graph clustering to obtain a final 3D instance segmentation result; The 2D mask obtaining module is specifically configured to: Use the image predictor of the SAM2 model to automatically segment each key frame and generate the 2D masks of the objects, where the objects include the new objects appearing in each key frame; Regard the image sequence between every two key frames as a video segment, with the first frame and the last frame of the video segment being key frames respectively. Use the video predictor of the SAM2 model, use the 2D masks generated by segmenting each key frame as prompts, and adopt a two-way verification tracking strategy to track the objects in the video segment; At the edges of adjacent video segments, use IoU to connect the same objects in different segments, and finally obtain the 2D masks of all objects in the entire two-dimensional image sequence.

8. An electronic device, the electronic device includes a processor and a memory, and at least one instruction is stored in the memory, characterized in that, The at least one instruction is loaded and executed by the processor to implement the multi-view consistency-based zero-shot point cloud instance segmentation method according to any one of claims 1-6.

9. A computer-readable storage medium storing at least one instruction, characterized in that, The at least one instruction is loaded and executed by the processor to implement the multi-view consistency-based zero-shot point cloud instance segmentation method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Three-dimensional point cloud labeling method and system based on scene reconstruction

    CN118115994A

  • Instance segmentation method based on three-dimensional point cloud

    CN118351320A