Zero sample point cloud instance segmentation method and system based on multi-view consistency

By adopting a zero-sample method of multi-view consistency in three-dimensional point cloud instance segmentation, combined with SAM2 model and mask consistency weight correction, the problem of view angle inconsistency and long-distance projection error is solved, and the segmentation accuracy and performance are significantly improved.

CN120070901AActive Publication Date: 2025-05-30UNIV OF SCI & TECH BEIJING
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510546772.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-05-30
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

The existing three-dimensional point cloud instance segmentation method has problems of viewing angle inconsistency and long-distance projection error under multi-view conditions, resulting in a decrease in segmentation accuracy.

Method used

Using the zero-sample point cloud instance segmentation method based on multi-view consistency, by obtaining three-dimensional point clouds and two-dimensional image sequences, extracting keyframes and segmenting and tracking using SAM2 model, combining 3D to 2D projection and mask consistency weight correction, hyperpoint maps are constructed and multi-level iterative map clustering are performed to obtain the final 3D instance segmentation result.

Benefits of technology

It significantly enhances the segmentation accuracy at multi-view angles, effectively solves the long-distance projection error problem, and dynamically adjusts the affinity score threshold, prioritizes the processing of high-view angle consistency, and improves the performance of point cloud instance segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070901A_ABST
    Figure CN120070901A_ABST
Patent Text Reader

Abstract

The invention provides a zero sample point cloud instance segmentation method and system based on multi-view consistency. The method comprises the following steps: extracting a plurality of key frames from a two-dimensional image sequence corresponding to a three-dimensional point cloud; taking a 2D mask generated by segmenting the key frame as a prompt, and segmenting and tracking the two-dimensional image sequence to obtain 2D masks of all objects; performing super-segmentation on the three-dimensional point cloud to obtain a series of super-points, performing 3D-to-2D projection to obtain a 2D mask of each super-point on each image, representing edges between the super-points by affinity scores between the 2D masks, and constructing a super-point diagram; correcting the affinity score by using a mask consistency weight, and endowing a low-quality mask with a low weight to obtain a corrected super-point diagram; and through multi-level super-point iteration graph clustering, performing hierarchical combination on super-points with different view angle consistency levels to obtain a 3D instance segmentation result. According to the invention, instance segmentation can be carried out on the 3D point cloud.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of zero-shot point cloud instance segmentation, and particularly to a zero-shot point cloud instance segmentation method and system based on multi-view consistency. Background Art

[0002] The 3D point cloud instance segmentation task refers to dividing different points in three-dimensional point cloud data according to certain rules or objectives, so as to realize the recognition and separation of different objects, regions or categories. Point cloud instance segmentation has important applications in fields such as autonomous driving, robot navigation, and augmented reality. Compared with two-dimensional images, three-dimensional data (such as point cloud data, lidar scan data, etc.) provides richer spatial information, which helps to improve the accuracy and reliability of segmentation. However, at the same time, three-dimensional perception faces a series of new challenges and problems.

[0003] First of all, existing three-dimensional perception methods often rely on a large amount of high-quality labeled data, especially for the labeling of three-dimensional point clouds, which usually requires a large amount of manual intervention and has a high labeling cost. In some application scenarios, it is very difficult to obtain high-quality labeled data, which makes traditional supervised learning methods difficult to apply. To address this problem, researchers have begun to explore zero-shot learning methods, hoping to achieve the segmentation of new categories through existing models and data. In addition, researchers have also begun to explore transferring the powerful generalization ability of two-dimensional vision models to three-dimensional tasks, attempting to use the generalization ability of these models in different scenarios and categories to improve the effect of three-dimensional object detection. However, the application of zero-shot learning in point cloud instance segmentation still faces great challenges.

[0004] Combining two-dimensional image information with three-dimensional point cloud data has become an effective solution. Two-dimensional images have rich detail information and can provide higher resolution, which plays an important auxiliary role in fine-grained object segmentation. By fusing the two-dimensional image segmentation results with the three-dimensional point cloud data, the accuracy and robustness of point cloud instance segmentation can be improved.

[0005] An urgent problem to be solved is the perspective consistency problem. In 3D point cloud instance segmentation, existing methods usually introduce 2D image sequences to assist in completing point cloud instance segmentation. They complete image segmentation through multi-frame data and then project the segmentation results onto the point cloud. However, due to the appearance changes and occlusion phenomena of the target under different perspectives, the target segmentation results among multi-frame data are inconsistent. This perspective inconsistency directly affects the final segmentation accuracy. In addition, in 3D perception tasks, especially in the process of projecting from 2D images to 3D point clouds, as the object moves away from the camera, the error introduced by long-distance projection becomes a significant challenge. Existing research has not effectively solved this problem. Especially in the process of matching 2D images to 3D point clouds, objects at long distances and close distances are given the same weight during projection. This can cause 3D objects to be incorrectly matched with the 2D mask of another object, resulting in incorrect 3D segmentation. This problem worsens as the object moves away from the camera, ultimately leading to a decrease in segmentation accuracy. This challenge, especially in the process of processing complex scenes and large-scale point cloud data, seriously affects the accuracy of the point cloud instance segmentation task. Therefore, how to effectively solve the error problem caused by long-distance projection has also become a key difficult problem in current research. Summary of the Invention

[0006] To solve the above technical problems existing in the prior art, the present invention provides a zero-shot point cloud instance segmentation method and system based on multi-perspective consistency. The technical solutions are as follows:

[0007] On the one hand, a zero-shot point cloud instance segmentation method based on multi-perspective consistency is provided. The method includes:

[0008] S1. Obtain a 3D point cloud to be subjected to instance segmentation and a corresponding 2D image sequence;

[0009] S2. Extract multiple key frames from the 2D image sequence;

[0010] S3. Use the 2D masks generated by segmenting the key frames as prompts, and use the SAM2 model to segment and track the 2D image sequence to obtain the 2D masks of all objects in the 2D image sequence;

[0011] S4. Perform super-segmentation on the 3D point cloud to obtain a series of super points, perform 3D-to-2D projection, obtain the 2D masks of each super point on each image, represent the edges between super points by the affinity scores between the 2D masks corresponding to the super points, and construct a super point graph based on the super points and the edges between super points. The super points serve as the nodes of the super point graph, and the edges between super points serve as the edges between the nodes;

[0012] S5. Use the mask consistency weight to correct the affinity scores, assign low weights to low-quality masks, and obtain a corrected super point graph;

[0013] S6. Through multi-level superpoint iterative graph clustering, hierarchically merge superpoints with different levels of perspective consistency to obtain the final 3D instance segmentation result.

[0014] Optionally, the S2 specifically includes:

[0015] S21. Regard the two-dimensional image sequence as a video, and calculate the difference between two adjacent frames , and the calculation method is as follows:

[0016]

[0017] where, represents the pixel matrix of the current frame, represents the pixel matrix of the previous frame, and represent the height and width of the image respectively;

[0018] S22. Use a Hanning window to smooth the frame difference array, and the calculation formula for the window weight is as follows:

[0019]

[0020] where N is the length of the window, n is the index within the window, is the value of the window function at position n, and define the smoothed array as S. The value at the j-th position in S is expressed as:

[0021]

[0022] where, represents the value at the j - n position in . The index corresponding to the local maximum value in the array S represents the key frame, which indicates the frame where a new object appears.

[0023] Optionally, the S3 specifically includes:

[0024] S31. Use the image predictor of the SAM2 model to automatically segment each key frame and generate a 2D mask of the object, where the object includes new objects that appear in each key frame;

[0025] S32. Regard the image sequence between every two key frames as a video clip. The first frame and the last frame of the video clip are key frames respectively. Use the video predictor of the SAM2 model, and use the 2D masks generated by segmenting each key frame as prompts to track the object in the video clip using a two-way verification tracking strategy;

[0026] S33. At the edges of adjacent video segments, use IoU to connect the same objects in different segments, and finally obtain the 2D masks of all objects in the entire 2D image sequence.

[0027] Optionally, the bidirectional verification tracking strategy includes forward tracking and backward tracking. The targets missed during forward tracking will be identified during backward tracking, and the targets missed during backward tracking will be identified during forward tracking. The tracking results from the two directions complement each other, and the 2D masks from the two directions are merged.

[0028] Optionally, the S4 specifically includes:

[0029] S41. Use a graph cut algorithm to perform super-segmentation on the 3D point cloud and generate a series of 3D superpoints based on 3D geometric properties;

[0030] S42. Use a general pinhole camera matrix for 3D to 2D projection, and utilize the corresponding camera internal parameters and pose parameters to project the i-th 3D superpoint onto the k-th image and obtain its corresponding 2D mask:

[0031]

[0032] where is the projection operator;

[0033] S43. Calculate the normalized histogram based on the 2D mask and convert it into a vector, denoted as . This histogram vector represents the 2D instance features of the 3D superpoint. Obtain the affinity score by calculating the cosine similarity between the 2D instance features of different superpoints on the k-th image, as shown in the following formula:

[0034]

[0035] S44. Sum the affinity scores for each image to obtain the affinity score between all the 2D masks corresponding to the i-th superpoint and the j-th superpoint, which is used as the edge between the i-th superpoint and the j-th superpoint, as shown in the following formula:

[0036] .

[0037] Optionally, the S5 specifically includes:

[0038] S51. Calculate and obtain the visibility edge weight:

[0039] Project the points in the superpoint onto the image, filter out the points outside the image, and filter out the occluded points according to the depth of the points in the superpoint to obtain visible points, and define as the superpoint The ratio of the number of visible points on the k-th image to the total number of points in the superpoint. Define the visibility edge weight between the i-th superpoint and the j-th superpoint as:

[0040]

[0041] S52. Calculate and obtain the distance edge weight:

[0042] Define as the distance weighted average based on the k-th image, which is used to weight and average the 2D mask :

[0043]

[0044] where represents the distance from the visible points in the i-th superpoint to the k-th image

[0045] and represent the maximum and minimum distances from all visible points to the k-th image respectively. Define the distance edge weight between the i-th superpoint and the j-th superpoint as:

[0046]

[0047] S53. Calculate and obtain the purity edge weight:

[0048] The number of 2D masks corresponding to the superpoint is defined as its purity. Define as the maximum ratio of the superpoint projected onto different 2D masks in the k-th image:

[0049]

[0050] where represents the total number of visible points of the superpoint in the k-th image, represents the number of points belonging to the mask label value of 1, represents the number of points belonging to the mask label value of 2, represents the number of points belonging to the mask label value of n, where n is the total number of 2D masks. Define the purity edge weight between the i-th superpoint and the j-th superpoint as:

[0051]

[0052] S54. Merge the visibility edge weight, distance edge weight, and purity edge weight to form a mask consistency weight:

[0053]

[0054] S55. Obtain the final affinity score according to the mask consistency weight, and assign a low weight to low-quality masks. The low-quality masks include 2D masks with lower visibility, greater distance, and higher purity. The calculation formula is:

[0055] 。

[0056] Optionally, the multi-level superpoint iterative graph clustering specifically includes:

[0057] In each iterative clustering, set different thresholds according to the affinity score to merge superpoints. Prioritize processing superpoints with high view consistency. Superpoints with high view consistency and an affinity score greater than the threshold are merged first, and then the threshold is gradually decreased to merge superpoints with low view consistency.

[0058] On the other hand, a zero-shot point cloud instance segmentation system based on multi-view consistency is provided. The system includes:

[0059] An acquisition module for acquiring a three-dimensional point cloud to be subjected to instance segmentation and a corresponding two-dimensional image sequence;

[0060] A key frame extraction module for extracting multiple key frames from the two-dimensional image sequence;

[0061] A 2D mask acquisition module for using the 2D masks generated by segmenting the key frames as cues, segmenting and tracking the two-dimensional image sequence using the SAM2 model, and obtaining 2D masks of all objects in the two-dimensional image sequence;

[0062] A superpoint graph construction module for performing super-segmentation on the three-dimensional point cloud to obtain a series of superpoints, performing 3D to 2D projection, obtaining the 2D masks of each superpoint on each image, representing the edges between superpoints by the affinity scores between the 2D masks corresponding to the superpoints, and constructing a superpoint graph based on the superpoints and the edges between the superpoints. The superpoints are used as the nodes of the superpoint graph, and the edges between the superpoints are used as the edges between the nodes;

[0063] A superpoint graph correction module for correcting the affinity scores using the mask consistency weight, assigning a low weight to low-quality masks, and obtaining a corrected superpoint graph;

[0064] A 3D instance segmentation module for hierarchically merging superpoints with different view consistency levels through multi-level superpoint iterative graph clustering to obtain the final 3D instance segmentation result.

[0065] On the other hand, an electronic device is provided, which includes a processor and a memory. At least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the above-mentioned zero-shot point cloud instance segmentation method based on multi-view consistency.

[0066] On the other hand, a computer-readable storage medium is provided. At least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement the above-mentioned zero-shot point cloud instance segmentation method based on multi-view consistency.

[0067] The beneficial effects brought by the technical solution provided by the present invention at least include:

[0068] 1) The zero-shot point cloud instance segmentation method based on multi-view consistency of the present invention utilizes the complex memory mechanism of the SAM2 model and proposes a segmentation processing method based on key frames. The image sequence between every two key frames is regarded as an independent video segment. The key frames are automatically segmented by the image predictor of SAM2 to generate 2D masks (including 2D masks of new objects that appear in each key frame), and these are used as prompts. Even when an object is occluded or re-enters the scene (i.e., a new object appears), the video predictor is used to achieve continuous segmentation and tracking of all objects between video frames (existing methods use the objects in the first frame as prompts for tracking, and when new objects appear later, they cannot be tracked). And the target omission problem is solved through a two-way verification tracking strategy to maintain the consistency of segmentation and tracking. This mechanism significantly enhances the segmentation accuracy under multiple views.

[0069] 2) The mask consistency weight designed by the present invention not only effectively solves the long-distance projection error problem, but also assigns low weights to low-quality masks from multiple angles, thereby correcting the relationship (edges) between superpoints in the superpoint graph, so as to obtain a more accurate segmentation result when iteratively clustering and merging superpoints in the final step (the method of performing super-segmentation on the three-dimensional point cloud to obtain a series of superpoints (a superpoint is a set of multiple points with similar features) used in the present invention is an existing method, and the super-segmentation of this method is over-segmentation, that is, different parts of an instance may be segmented into different superpoints. For example, a table instance is over-segmented, and the table legs and the table top are segmented into different superpoints. Subsequently, it is necessary to perform combined segmentation on the over-segmented superpoints belonging to the same instance).

[0070] 3) The multi-level superpoint iterative graph clustering designed by the present invention dynamically adjusts the affinity score threshold according to the view consistency level during the iterative clustering process, preferentially processes superpoints with high view consistency, and significantly enhances the point cloud instance segmentation performance through the iterative graph clustering and hierarchical merging strategy. Description of the Drawings

[0071] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0072] Figure 1 is a flowchart of a zero-shot point cloud instance segmentation method based on multi-view consistency provided by an embodiment of the present invention;

[0073] Figure 2 is another flowchart of a zero-shot point cloud instance segmentation method based on multi-view consistency provided by an embodiment of the present invention;

[0074] Figure 3 is a flowchart of two-dimensional image sequence segmentation and tracking provided by an embodiment of the present invention;

[0075] Figure 4 is a block diagram of a zero-shot point cloud instance segmentation system based on multi-view consistency provided by an embodiment of the present invention;

[0076] Figure 5 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0077] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the drawings and specific embodiments.

[0078] An embodiment of the present invention provides a zero-shot point cloud instance segmentation method based on multi-view consistency. This method can be implemented by an electronic device, which can be a terminal or a server. Figure 1 and Figure 2 As shown in the flowchart of this method, the processing flow can include the following steps:

[0079] S1. Obtain a three-dimensional point cloud to be subjected to instance segmentation and a corresponding two-dimensional image sequence;

[0080] S2. Extract multiple key frames from the two-dimensional image sequence;

[0081] Optionally, S2 specifically includes:

[0082] S21. Regard the two-dimensional image sequence as a video, calculate the difference between two adjacent frames , and the calculation method is as follows:

[0083]

[0084] Among them, represents the pixel matrix of the current frame, represents the pixel matrix of the previous frame, and respectively represent the height and width of the image;

[0085] When the scene changes or a new object appears, the pixel difference between the previous frame and the current frame will increase significantly. Embodiments of the present invention use the pixel differences between these frames to identify key frames. In this way, the frames where new objects are located can be identified and marked, facilitating subsequent segmentation and tracking.

[0086] S22. Smooth the frame difference array using a Hanning window. The calculation formula for the window weight is as follows:

[0087]

[0088] where N is the length of the window, n is the index within the window, is the value of the window function at position n. Define the smoothed array as S. The value at the j-th position in S is expressed as:

[0089]

[0090] where, represents the value at the j - n position in . The index corresponding to the local maximum value in the array S represents the key frame, indicating the frame where a new object appears.

[0091] S3. Use the 2D mask generated by segmenting the key frame as a prompt, and use the SAM2 model to segment and track the two-dimensional image sequence to obtain the 2D masks of all objects in the two-dimensional image sequence;

[0092] Optionally, as Figure 3 shown, step S3 specifically includes:

[0093] S31. Use the image predictor of the SAM2 model to automatically segment each key frame and generate the 2D mask of the object, where the object includes the new objects that appear in each key frame;

[0094] S32. Regard the image sequence between every two key frames as a video segment. The first frame and the last frame of the video segment are key frames respectively. Use the video predictor of the SAM2 model, and use the 2D masks generated by segmenting each key frame as prompts to track the object in the video segment using a two-way verification tracking strategy;

[0095] S33. At the edges of adjacent video segments, use IoU to connect the same objects in different segments, and finally obtain the 2D masks of all objects in the entire two-dimensional image sequence.

[0096] Existing methods use the objects in the first frame as cues for tracking. When new objects appear later, they cannot be tracked. However, in the embodiments of the present invention, the 2D masks generated by segmenting each key frame are used as cues, so that tracking of all objects (including new objects that appear in each key frame) can be achieved.

[0097] Optionally, the two-way verification tracking strategy includes forward tracking and backward tracking. The targets missed in the forward tracking process will be identified in the backward tracking, and the targets missed in the backward tracking will be identified in the forward tracking. The tracking results from the two directions complement each other, and the 2D masks from the two directions are merged.

[0098] Although the embodiments of the present invention can obtain new objects by segmenting key frames, the key frames are not always the exact frames where new objects appear (there may be incorrect segmentation). If new objects appear before or after the key frames, they may be missed or omitted. Therefore, the embodiments of the present invention design a two-way verification tracking strategy to solve this problem.

[0099] S4. Perform super-segmentation on the three-dimensional point cloud to obtain a series of super points, perform 3D-to-2D projection, obtain the 2D masks of each super point on each image, represent the edges between super points by the affinity scores between the 2D masks corresponding to the super points, and construct a super point graph based on the super points and the edges between them. The super points are used as the nodes of the super point graph, and the edges between the super points are used as the edges between the nodes (the super point graph is represented as , where v represents the super point and E represents the edges between the super points);

[0100] Optionally, S4 specifically includes:

[0101] S41. Use a graph cut algorithm to perform super-segmentation on the three-dimensional point cloud and generate a series of 3D super points based on 3D geometric attributes;

[0102] The method used in the present invention to perform super-segmentation on the three-dimensional point cloud to obtain a series of super points (a super point is a set of multiple points with similar characteristics) is an existing method. The super-segmentation of this method is over-segmentation, that is, different parts of an instance may be segmented into different super points. For example, the instance of a table is over-segmented, and the table legs and the table top are segmented into different super points. Subsequently, it is necessary to merge the over-segmented super points belonging to the same instance to obtain an accurate instance segmentation result and complete the point cloud instance segmentation.

[0103] S42. Use a general pinhole camera matrix for 3D-to-2D projection and utilize the corresponding camera internal parameters and pose parameters Project the i-th 3D superpoint onto the k-th image and obtain its corresponding 2D mask:

[0104]

[0105] where is the projection operator;

[0106] S43. Based on the 2D mask calculate the normalized histogram and convert it into a vector, denoted as . This histogram vector represents the 2D instance features of the 3D superpoint. The affinity score is obtained by calculating the cosine similarity between the 2D instance features of different superpoints on the k-th image , as follows:

[0107]

[0108] S44. Sum the affinity scores for each image to obtain the affinity score between all 2D masks corresponding to the i-th superpoint and the j-th superpoint , which is used as the edge between the i-th superpoint and the j-th superpoint, as follows:

[0109] .

[0110] S5. Use the mask consistency weight to correct the affinity score, assigning a low weight to low-quality masks to obtain a corrected superpoint graph;

[0111] Optionally, S5 specifically includes:

[0112] S51. Calculate the visibility edge weight:

[0113] Project the points in the superpoint onto the image, filter out the points outside the image, and filter out the occluded points according to the depth of the points in the superpoint to obtain visible points. Define as the ratio of the number of visible points of the superpoint on the k-th image to the total number of points in the superpoint. Define the visibility edge weight between the i-th superpoint and the j-th superpoint as:

[0114]

[0115] S52. Calculate the distance edge weight:

[0116] Define as the distance weighted average based on the k-th image, which is used to weight the average of the 2D mask :

[0117]

[0118] Among them, represents the distance from the visible points in the i-th superpoint to the k-th image,

[0119] and respectively represent the maximum and minimum distances from all visible points to the k-th image. The edge weight of the distance between the i-th superpoint and the j-th superpoint is defined as:

[0120]

[0121] S53. Calculate the purity edge weight:

[0122] The number of 2D masks corresponding to the superpoint is defined as its purity. Define as the maximum ratio of the superpoint projected onto different 2D masks in the k-th image:

[0123]

[0124] Among them, represents the total number of visible points of the superpoint in the k-th image, represents the number of points belonging to the mask label value of 1, represents the number of points belonging to the mask label value of 2, represents the number of points belonging to the mask label value of n, where n is the total number of 2D masks. The purity edge weight between the i-th superpoint and the j-th superpoint is defined as:

[0125]

[0126] S54. Combine the visibility edge weight, distance edge weight, and purity edge weight to form a mask consistency weight:

[0127]

[0128] S55. According to the mask consistency weight, obtain the final affinity score, and assign a low weight to the low-quality masks. The low-quality masks include 2D masks with lower visibility, greater distance, and higher purity. The calculation formula is:

[0129] .

[0130] S6. Through multi-level superpoint iterative graph clustering, hierarchically merge the superpoints with different levels of perspective consistency to obtain the final 3D instance segmentation result.

[0131] Optionally, the multi-level superpoint iterative graph clustering specifically includes:

[0132] In each iterative clustering, different thresholds are set according to the affinity scores to merge the superpoints. Superpoints with high view consistency are processed preferentially. Superpoints with high view consistency and affinity scores greater than the threshold are merged first, and then the threshold is gradually decreased to merge superpoints with low view consistency.

[0133] The multi-level superpoint iterative graph clustering designed by the present invention dynamically adjusts the affinity score threshold according to the view consistency level during the iterative clustering process, preferentially processes superpoints with high view consistency, and significantly enhances the point cloud instance segmentation performance through the iterative graph clustering and hierarchical merging strategy.

[0134] As Figure 4 shown, an embodiment of the present invention further provides a zero-shot point cloud instance segmentation system based on multi-view consistency. The system includes:

[0135] An acquisition module 410, configured to acquire a three-dimensional point cloud to be subjected to instance segmentation and a corresponding two-dimensional image sequence;

[0136] A key frame extraction module 420, configured to extract a plurality of key frames from the two-dimensional image sequence;

[0137] A 2D mask obtaining module 430, configured to use the 2D mask generated by segmenting the key frames as a prompt, and segment and track the two-dimensional image sequence using the SAM2 model to obtain 2D masks of all objects in the two-dimensional image sequence;

[0138] A superpoint graph construction module 440, configured to perform super-segmentation on the three-dimensional point cloud to obtain a series of superpoints, perform 3D to 2D projection, obtain 2D masks of each superpoint on each image, represent the edges between superpoints by the affinity scores between the 2D masks corresponding to the superpoints, and construct a superpoint graph according to the superpoints and the edges between the superpoints. The superpoints are used as the nodes of the superpoint graph, and the edges between the superpoints are used as the edges between the nodes;

[0139] A superpoint graph correction module 450, configured to correct the affinity scores using the mask consistency weights, assign low weights to low-quality masks, and obtain a corrected superpoint graph;

[0140] A 3D instance segmentation module 460, configured to perform hierarchical merging on superpoints with different view consistency levels through multi-level superpoint iterative graph clustering to obtain a final 3D instance segmentation result.

[0141] The function structure of a zero-shot point cloud instance segmentation system based on multi-view consistency provided by an embodiment of the present invention corresponds to a zero-shot point cloud instance segmentation method provided by an embodiment of the present invention, and will not be elaborated here.

[0142] Figure 5 It is a schematic structural diagram of an electronic device 500 provided by an embodiment of the present invention. The electronic device 500 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 501 and one or more memories 502. Among them, at least one instruction is stored in the memory 502, and the at least one instruction is loaded and executed by the processor 501 to implement the steps of the above-mentioned zero-shot point cloud instance segmentation method based on multi-view consistency.

[0143] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions. The above instructions can be executed by a processor in a terminal to complete the above-mentioned zero-shot point cloud instance segmentation method based on multi-view consistency. For example, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0144] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware, or can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a magnetic disk, or an optical disc, etc.

[0145] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A zero-shot point cloud instance segmentation method based on multi-view consistency, characterized in that: The method comprises: S1, obtaining a three-dimensional point cloud and a corresponding two-dimensional image sequence to be segmented; S2, extracting a plurality of key frames from the two-dimensional image sequence; S3, using the 2D mask generated by segmenting the key frame as a hint, using the SAM2 model to segment and track the two-dimensional image sequence, and obtaining 2D masks of all objects in the two-dimensional image sequence; S4, super-segmenting the three-dimensional point cloud to obtain a series of super-points, performing 3D to 2D projection, obtaining a 2D mask of each super-point on each image, representing the edges between the super-points by affinity scores between the 2D masks corresponding to the super-points, constructing a super-point graph according to the super-points and the edges between the super-points, with the super-points serving as nodes of the super-point graph and the edges between the super-points serving as edges between the nodes; S5. Correcting the affinity score using the mask consistency weight, assigning a low weight to the low-quality mask, and obtaining a corrected superpoint map; S6. Through multi-level super-point iterative graph clustering, super-points with different levels of perspective consistency are hierarchically merged to obtain the final 3D instance segmentation result.

2. The method according to claim 1, characterized in that The S2 specifically includes: S21, the two-dimensional image sequence Treat it as a video and calculate the difference between two adjacent frames , the calculation method is as follows: ; in, Represents the pixel matrix of the current frame, represents the pixel matrix of the previous frame, and Respectively represent the height and width of the image; S22. Use the Hanning window to smooth the frame difference array. The calculation formula of the window weight is as follows: ; Where N is the length of the window, n is the index within the window, The value of the window function at position n, the smoothed array is defined as S, and the value of the jth position in S is expressed as: ; in, Indicated in The value at position jn in the array S, the index corresponding to the local maximum value, represents the key frame, indicating the frame where the new object appears.

3. The method according to claim 1, characterized in that The S3 specifically includes: S31, using the image predictor of the SAM2 model, automatically segmenting each key frame and generating a 2D mask of the object, wherein the object includes a new object appearing in each key frame; S32, treating the image sequence between every two key frames as a video segment, wherein the first frame and the last frame of the video segment are key frames respectively, using the video predictor of the SAM2 model, using the 2D mask generated by segmenting each key frame as a prompt, and using a two-way verification tracking strategy to track the object in the video segment; S33. At the edges of adjacent video clips, IoU is used to connect the same objects in different clips, and finally the 2D masks of all objects in the entire two-dimensional image sequence are obtained.

4. The method according to claim 3, characterized in that The two-way verification tracking strategy includes forward tracking and backward tracking. The target missed in the forward tracking process will be identified in the backward tracking, and the target missed in the backward tracking will be identified in the forward tracking. The tracking results from the two directions complement each other, and the 2D masks from the two directions are merged.

5. The method according to claim 1, characterized in that The S4 specifically includes: S41, super-segmenting the three-dimensional point cloud using a graph cutting algorithm, and generating a series of 3D super-points based on 3D geometric attributes; S42, use a general pinhole camera matrix for 3D to 2D projection, using the corresponding camera intrinsic parameters and attitude parameters The i-th 3D superpoint Project it onto the kth image and obtain its corresponding 2D mask: ; in, is the projection operator; S43, 2D mask based Calculate the normalized histogram and convert it into a vector represented as , this histogram vector represents the 2D instance features of the 3D superpoint, and the affinity score is obtained by calculating the cosine similarity between the 2D instance features of different superpoints on the kth image , as follows: ; S44. Sum the affinity scores of each image to obtain the affinity scores between the i-th superpoint and all 2D masks corresponding to the j-th superpoint. , as the edge between the i-th superpoint and the j-th superpoint, as follows: 。 6. The method according to claim 5, characterized in that The S5 specifically includes: S51. Calculate the visibility edge weight: Project the points in the superpoint onto the image, filter out the points outside the image, and filter out the occluded points according to the depth of the points in the superpoint to obtain the visible points. Define For super point The ratio of the number of visible points on the kth image to the total number of points in the superpoints, and the visibility edge weight between the i-th superpoint and the j-th superpoint is defined as: ; S52, calculate and obtain the distance edge weight: definition is the weighted average of the distance based on the kth image, used for weighted average 2D mask : ; in, represents the distance from the visible point in the i-th superpoint to the k-th image, and Respectively represent the maximum and minimum distances of all visible points to the k-th image, and define the distance edge weight between the i-th superpoint and the j-th superpoint as: ; S53. Calculate and obtain the purity edge weight: 2D mask corresponding to the superpoint The quantity is defined as its purity, and the definition For super point The maximum ratio of projections to different 2D masks in the kth image: ; in, Indicates the super point in the kth image The total number of visible points, Indicates the number of points whose mask label value is 1. Indicates the number of points with mask label value 2. represents the number of points with mask label value n, where n is the total number of 2D masks. The purity edge weight between the i-th superpoint and the j-th superpoint is defined as: ; S54, combining the visibility edge weight, distance edge weight, and purity edge weight to form a mask consistency weight: ; S55. According to the mask consistency weight, a final affinity score is obtained, and a low weight is assigned to a low-quality mask. The low-quality mask includes a 2D mask with lower visibility, longer distance, and greater purity. The calculation formula is: 。 7. The method according to claim 1, characterized in that The multi-level super-point iterative graph clustering specifically includes: In each iterative clustering, different thresholds are set according to the affinity scores to merge superpoints, giving priority to superpoints with high view consistency. Superpoints with high view consistency whose affinity scores are greater than the threshold are merged first, and then the threshold is gradually lowered to merge superpoints with low view consistency.

8. A zero-shot point cloud instance segmentation system based on multi-view consistency, characterized in that: The system comprises: An acquisition module, used to acquire a three-dimensional point cloud and a corresponding two-dimensional image sequence to be segmented; A key frame extraction module, used for extracting a plurality of key frames from the two-dimensional image sequence; A 2D mask acquisition module, used to segment and track the two-dimensional image sequence using the SAM2 model, using the 2D mask generated by segmenting the key frame as a prompt, to obtain 2D masks of all objects in the two-dimensional image sequence; A superpoint graph construction module is used to perform super-segmentation on the three-dimensional point cloud to obtain a series of superpoints, perform 3D to 2D projection, obtain a 2D mask of each superpoint on each image, represent the edges between superpoints by affinity scores between 2D masks corresponding to superpoints, and construct a superpoint graph according to the edges between superpoints, with superpoints serving as nodes of the superpoint graph and edges between superpoints serving as edges between nodes; A superpoint map correction module, used to correct the affinity score using mask consistency weights, assign low weights to low-quality masks, and obtain a corrected superpoint map; The 3D instance segmentation module is used to hierarchically merge superpoints with different levels of perspective consistency through multi-level superpoint iterative graph clustering to obtain the final 3D instance segmentation result.

9. An electronic device, comprising a processor and a memory, wherein at least one instruction is stored in the memory, wherein: The at least one instruction is loaded and executed by the processor to implement the zero-sample point cloud instance segmentation method based on multi-view consistency as described in any one of claims 1-7.

10. A computer-readable storage medium, wherein at least one instruction is stored in the storage medium, characterized in that: The at least one instruction is loaded and executed by the processor to implement the zero-sample point cloud instance segmentation method based on multi-view consistency as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Three-dimensional point cloud labeling method and system based on scene reconstruction

    CN118115994A

  • Instance segmentation method based on three-dimensional point cloud

    CN118351320A

  • Semantic SAM large model-based three-dimensional point cloud robustness component segmentation method

    CN118397282A

  • Open vocabulary 3D segmentation method based on three-dimensional Gaussian sputtering technology

    CN119445449A

  • Self-adaptive single object three-dimensional reconstruction and image point cloud synthesis method for automatic driving scene

    CN119516098A