3D segmentation method and system based on bidirectional fusion

By combining image and camera information on SAM3D for two-way fusion optimization, the problem of unclear target boundaries and unstable masks in outdoor scenes is solved, and stronger adaptability and efficiency is achieved, suitable for complex and sparse outdoor point cloud data segmentation.

CN120472164APending Publication Date: 2025-08-12ZHEJIANG UNIV

Patent Information

Application Number
CN202510562929.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing three-dimensional segmentation method based on SAM3D has problems such as unclear target boundaries and unstable mask fusion in outdoor scenes, making it difficult to adapt to complex and sparse outdoor point cloud data.

Method used

The 3D segmentation method based on bidirectional fusion is adopted, combined with image information and camera information, and the preliminary segmentation results are obtained through the three-dimensional image segmentation model, and the bottom-up two-way fusion and geometric information are performed. The weights of spatial position, feature similarity and semantic similarity are used for fusion optimization, and the final segmentation result is finally obtained.

Benefits of technology

It improves the consistency and applicability of target segmentation in outdoor scenes, enhances the model's perception ability and generalization performance, and can quickly adapt to different scenarios without additional training on image information extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472164A_ABST
    Figure CN120472164A_ABST
Patent Text Reader

Abstract

The invention discloses a 3D segmentation method and system based on bidirectional fusion, and the method comprises the steps: obtaining the image information and camera information of a fixed scene; then, the image information and the camera information pass through a three-dimensional image segmentation model, and a preliminary segmentation result of a three-dimensional space is obtained; and finally, based on the preliminary segmentation result, carrying out bidirectional fusion on two adjacent frames of point cloud images to obtain an optimized segmentation result, then obtaining a geometric information-based segmentation result of the scene, and carrying out bidirectional fusion on the point cloud image based on the geometric information segmentation result and the point cloud image based on the optimized segmentation result to obtain a final segmentation result. According to the method, the complexity of the outdoor scene and the difference of the sparse degrees of the point clouds are fully considered, the problem that segmentation masks of front and back frames of a long-distance object in an outdoor environment are inconsistent is solved, the segmentation consistency of an outdoor target is enhanced, the applicability to a large-range and complex scene is improved, and the method is suitable for large-scale and complex scenes. And the feasibility of the SAM3D-based three-dimensional segmentation method in practical engineering application is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of point cloud segmentation, and in particular relates to a 3D segmentation method and system based on bidirectional fusion. Background Art

[0002] As technologies such as autonomous driving, urban security, and augmented reality continue to be applied in outdoor scenarios, the demand for a more refined understanding and visualization of three-dimensional environments is also increasing. Currently, point cloud segmentation technology approaches are mainly divided into two categories: one is to directly build a segmentation model on point cloud data and use the spatial relationship of three-dimensional coordinates for feature learning, such as pointnet++, RandLaNet, OneFormer3D, and the technologies mentioned in patent CN116310349A; the other is to establish a cross-modal fusion mechanism of two-dimensional visual features and three-dimensional geometric features, fusing information extracted from RGB images with 3D scene information, such as UniSeg, TransFusion, and the technologies mentioned in patent CN114820369A. Although point cloud data has unique geometric representation capabilities, and modal fusion enables models to better understand scenes through information complementarity, two-dimensional visual models demonstrate significant advantages with the support of massive amounts of annotated data. The current large-scale image-based pre-trained model SAM (Segment Anything Model) in the computer vision field has accumulated a large amount of transferable semantic understanding capabilities, providing an important foundation for multimodal fusion methods, eliminating the need for separate training for the acquisition of 2D information. However, existing SAM-based three-dimensional segmentation methods, such as SAM3D and EmbodiedSAM, are mostly studied and verified in relatively closed and structured indoor scenes. When applied to outdoor environments, they are affected by various uncertain factors such as inconsistent scene characteristics and sparse point clouds, leading to problems such as unclear target boundaries and unstable mask fusion. Although the original method shows good accuracy and fine-grained segmentation effects on indoor data, its performance in outdoor scenes has certain defects. Summary of the Invention

[0003] In order to address the shortcomings of the existing technology, realize the purpose of applying SAM3D to outdoor and optimizing point cloud segmentation in combination with outdoor characteristics, the present invention adopts the following technical solutions:

[0004] The 3D segmentation method based on bidirectional fusion includes the following steps:

[0005] Step S1: Obtain image information and camera information of a fixed scene,

[0006] Step S2: The image information and camera information are passed through a 3D image segmentation model to obtain a preliminary segmentation result of the 3D space;

[0007] Step S3: Based on the preliminary segmentation results, a bottom-up bidirectional fusion OBM is performed on two adjacent frames of point cloud images to obtain an optimized category-independent segmentation result. Then, a normal-based graph cutting method is used to obtain a geometric segmentation result of the scene. The point cloud image based on the geometric information segmentation result and the point cloud image based on the category-independent segmentation result are bidirectionally fused OBM to obtain the final segmentation result. The bidirectional fusion OBM process is as follows:

[0008] Calculate point cloud image and point cloud images The corresponding mapping M between them, p1 and p2 represent the number of point clouds in the two point cloud images, and the corresponding mapping relationship is calculated by the nearest neighbor method; if (i, j)∈M, it means that point with dot There is a matching relationship between the two point cloud images. For each matching point pair, the fusion confidence weight W is defined c :

[0009]

[0010] Among them, W pose represents the spatial position weight, W feat represents the feature similarity weight, W sem Represents the semantic similarity weight; for any mask number m in the point cloud image X1, find all points Q1 belonging to the mask, and the number of points is recorded as Use the corresponding mapping M to find the corresponding point set Q2 in the point cloud image X2. For any area with mask number n in Q2 remember For the region Points, use Indicates the area with mask number n in X2 The number of points in the area, if it satisfies It is considered that the mask numbered m in the point cloud image X1 and the area numbered n in the point cloud image X2 are the same area, δ t Represents the attenuation factor, attenuation factor δ0 represents the preset parameter;

[0011] Based on the fusion confidence weight, the point numbered m in the point cloud image X2 is assigned a new number n to merge the mask; after completing the above processing from the point cloud image X1 to the point cloud image X2, a symmetric processing is performed, that is, the mask number in the point cloud image X2 is traversed, and the corresponding point and mask number are found in the point cloud image X1 to complete the two-way merging and obtain the optimized full-scene segmentation result.

[0012] Furthermore, in step S1, the image information includes depth information, and the camera information includes camera intrinsic and extrinsic parameters. In step S2, the two-dimensional mask is back-projected into three-dimensional space based on the extracted depth information and the camera intrinsic and extrinsic parameter information. The specific formula is as follows:

[0013] [x i ,y i , z i ] T =R -1 M -1 [u i , v i , 1] T -R -1 T

[0014] Where M represents the intrinsic camera calibration matrix, R and T represent the rotation matrix and translation matrix obtained from the external camera pose matrix, [u i , v i 1] and [x i ,y i , z i , 1] represent the two-dimensional homogeneous coordinates and the reconstructed three-dimensional homogeneous coordinates, u i Indicates the horizontal position parameter of the texture image, v i Indicates the position parameter of the texture image in the vertical direction, x i ,y i , z i They represent the position parameters of the three-dimensional image in the horizontal, vertical and depth directions respectively. The three-dimensional point cloud is extracted through a neural network for feature prediction to obtain the preliminary segmentation results of the three-dimensional space.

[0015] Furthermore, in step S3, for each matching point pair The spatial position weight is:

[0016]

[0017] Here, ||·||2 represents the l2 norm.

[0018] Furthermore, in step S3, for each matching point pair The feature similarity weight is:

[0019]

[0020] Among them, ||·||2 represents the l2 norm, represents the feature of point i in the first frame, Same thing.

[0021] Furthermore, in step S3, for each matching point pair The predicted semantic weight is:

[0022]

[0023] Furthermore, in step S3, for each matching point pair The confidence weight of the region is obtained by weighted averaging:

[0024]

[0025] Furthermore, the segmentation method is applied to outdoor vehicle driving, and road point cloud information is collected to obtain segmentation results. The ground is extracted separately according to the segmentation results and the masks are unified.

[0026] Furthermore, the image information in step S1 includes depth information, and the steps of obtaining the depth information through the point cloud are as follows:

[0027] Step S11: Convert the point cloud from the initial sensor coordinate system to a coordinate system aligned with the camera, including the following steps:

[0028] Step S111: converting the sensor coordinate system to the vehicle coordinate system;

[0029] Step S112: converting the vehicle coordinate system to the world coordinate system;

[0030] Step S113: converting the world coordinate system to the vehicle coordinate system corresponding to the camera's location at that moment;

[0031] Step S114: converting the vehicle coordinate system to the camera coordinate system;

[0032] Step S12: Projecting through the camera intrinsic parameter matrix to obtain the image coordinates and depth information corresponding to each point;

[0033] Step S13: Filter out unqualified points; due to visual and distance limitations, it is necessary to exclude points that are out of the field of view or too close or too far in depth, and only retain points that are within the boundary of the projected image and more than a certain distance in front of the camera;

[0034] Step S14: Generate a depth map on the image plane; initialize a two-dimensional array with the same size as the camera image, and fill the retained points into it according to their projection coordinates, and assign them the corresponding depths.

[0035] Furthermore, the image information in step S1 includes posture information, and the steps for obtaining the posture information are as follows:

[0036] Extract the rotation matrix and translation vector of the ego vehicle relative to the world coordinate system from the corresponding camera pose record, then extract the rotation matrix and translation vector of the camera relative to the ego vehicle coordinate system from the camera calibration information. Then combine these two transformations through linear algebra to obtain the overall transformation from the world coordinate system to the camera coordinate system, and save it in the Pose folder;

[0037] First, obtain the rotation matrix and displacement corresponding to the vehicle coordinate system to the world coordinate system, recorded as R geo→global With t ego→global Then obtain the rotation and translation corresponding to the camera coordinate system to the vehicle coordinate system, which is recorded as R cam→ego With t cam→ego , according to the coordinate transformation multiplication rule, we can get the rotation matrix and displacement vector corresponding to the world coordinate system to the camera coordinate system:

[0038] R world→camera =R ego→global ×R cam→ego

[0039] t world→camera =t ego→global -R ego→global ×t cam→ego

[0040] When the above two quantities are calculated, the complete posture information of the camera in the world coordinate system is obtained, and the posture is written into a fourth-order homogeneous transformation matrix. First, a unit matrix I is constructed. 4×4 , and then replace the first three rows and three columns with the rotation matrix R world→camera , and then replace the last column elements of the first three rows with the displacement vector t world→camera , and obtain the transformation matrix of the following form:

[0041]

[0042] Finally, a posture file is generated by writing each row of the 4×4 homogeneous matrix into a text file using a tab or space separator.

[0043] The 3D segmentation system based on bidirectional fusion includes an information extraction module, a three-dimensional image segmentation module and a bidirectional fusion module. The system adopts the 3D segmentation method based on bidirectional fusion to sequentially obtain and perform three-dimensional image segmentation based on image information of a fixed scene and camera information to obtain a preliminary segmentation result. The system then performs bidirectional fusion based on the preliminary segmentation result to obtain a final segmentation result.

[0044] The advantages and beneficial effects of the present invention are:

[0045] (1) The present invention integrates multimodal information, significantly improving the perception and generalization performance of the model and making it more adaptable.

[0046] (2) Since the SAM base model has powerful image segmentation capabilities, the present invention does not require additional training for the extraction of 2D information, and can thus be quickly applied to various scenarios. This method has an advantage in efficiency.

[0047] (3) The adaptive and flexible fusion conditions enable the present invention to be adjusted according to the characteristics of the scene, rather than using the same fixed threshold in all scenes. This is more flexible and adaptable than the SAM3D model.

[0048] (4) The present invention can be used in outdoor scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 It is a flow chart of a method according to an embodiment of the present invention.

[0050] Figure 2 This is a flow chart of obtaining fixed camera perspective depth information from point cloud in an embodiment of the present invention.

[0051] Figure 3 This is a flowchart of obtaining the pose Pose in an embodiment of the present invention.

[0052] Figure 4 3D point cloud feature extraction process according to an embodiment of the present invention.

[0053] Figure 5 This is a flow chart of bidirectional fusion OBM in an embodiment of the present invention. DETAILED DESCRIPTION

[0054] The following describes the specific embodiments of the present invention in detail with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to illustrate and explain the present invention and are not intended to limit the present invention.

[0055] The present invention is optimized based on the original SAM3D. SAM3D mainly optimizes the segmentation results based on the SAM base model using inter-frame information. When performing inter-frame fusion, the present invention considers the complexity of outdoor scenes and the different sparsity of point clouds, adopts adaptive and flexible fusion conditions rather than simple fixed thresholds to alleviate the problem of inconsistent segmentation masks between front and back frames for long-distance objects in outdoor environments. In addition, a U-Net network is added to extract 3D scene information to supplement modal information. To a certain extent, the segmentation consistency of outdoor targets is enhanced, the applicability to large-scale and complex scenes is improved, and the feasibility of the SAM3D-based three-dimensional segmentation method in practical engineering applications is improved.

[0056] like Figure 1 As shown, the present invention proposes a 3D segmentation method based on bidirectional fusion, comprising the following steps:

[0057] Step S1: Extract the depth, pose, color, and camera intrinsics of a fixed scene from the NuScenes dataset.

[0058] like Figure 2 As shown in the figure, when obtaining the depth information of a fixed scene, you can use a monocular camera to estimate the depth or extract the depth information from the point cloud. Monocular camera depth estimation can use models such as ZoeDepth. The steps to obtain depth information from the point cloud are as follows:

[0059] Step S11: Convert the point cloud from the initial sensor coordinate system to a coordinate system aligned with the camera, including the following steps:

[0060] Step S111: converting the sensor coordinate system to the vehicle coordinate system;

[0061] Step S112: converting the vehicle coordinate system to the world coordinate system;

[0062] Step S113: converting the world coordinate system to the vehicle coordinate system corresponding to the camera's location at that moment;

[0063] Step S114: Convert the vehicle coordinate system to the camera coordinate system.

[0064] Step S12: Projection is performed using the camera intrinsic parameter matrix. The points obtained after projection include the image coordinates and depth information corresponding to each point.

[0065] Step S13: Filter out unqualified points. Due to visual and distance limitations, points that are out of view or too close or too far away need to be excluded, and only points that are within the projected image boundary and at least a certain distance in front of the camera are retained.

[0066] Step S14: Generate a depth map on the image plane. Initialize a two-dimensional array depth_image with the same size as the camera image, and fill it with the retained points according to their projected coordinates, assigning them the corresponding depths.

[0067] like Figure 3 As shown in the figure, the specific steps for obtaining Pose information are: extracting the ego-vehicle's rotation matrix and translation vector relative to the world coordinate system from the camera's corresponding "ego_pose" record, and then extracting the camera's rotation matrix and translation vector relative to the ego-vehicle coordinate system from the camera's calibration information. These two transformations are then combined through linear algebra to obtain the overall transformation from the world coordinate system to the camera coordinate system, which is saved in the Pose folder.

[0068] Specifically, first obtain the rotation matrix and displacement corresponding to the "car coordinate system → world coordinate system" and record it as R ego→globalWith t ego→global Then obtain the rotation and translation corresponding to the "camera coordinate system → vehicle coordinate system" and record it as R cam→ego With t cam→ego , according to the coordinate transformation multiplication rule, we can get the rotation matrix and displacement vector corresponding to "world coordinate system → camera coordinate system":

[0069] R world→camera =R ego→global ×R cam→ego

[0070] t world→camera =t ego→global -R ego→global ×t cam→ego

[0071] When the above two quantities are calculated, the complete posture information of the camera in the world coordinate system is obtained. In order to facilitate subsequent use, the posture is written into a fourth-order homogeneous transformation matrix (i.e., a 4×4 matrix). Specifically, first construct a unit matrix I 4×4 , and then replace the first three rows and three columns with the rotation matrix R world→camera , and then replace the last column elements of the first three rows with the displacement vector t world→camera . In this way, the transformation matrix of the following form is obtained:

[0072]

[0073] Finally, the present invention generates a posture file named "{index}.txt" by writing each row of the 4×4 homogeneous matrix into a text file using a tab or space separator.

[0074] The steps for obtaining Color information are as follows: query the "sample_data" record corresponding to the camera sensor in the NuScenes database through the "camera_token", obtain the actual path of the camera image file on the local disk, read it, and save it in the Color folder in the format of {index}.jpg.

[0075] The steps to obtain Intrinsics information are as follows: Get the relevant parameters of the corresponding sensor from the calibrated_sensor.json provided by the NuScenes dataset.

[0076] Step S2: After obtaining the relevant data files, the SAM3D process can be applied to obtain the segmentation results:

[0077] Apply the SAM model to the RGB image to obtain the segmentation result. According to the extracted depth information and the camera internal and external parameters, the 2D mask is back-projected into the 3D space. The specific formula is as follows:

[0078] [x i ,y i , z i ] T =R -1 M -1 [u i , v i , 1] T -R -1 T

[0079] Where M is the intrinsic camera calibration matrix, R and T are the rotation and translation matrices obtained from the external camera pose matrix, [u i , v i , 1] and [x i ,y i , z i , 1] are the 2D homogeneous coordinates and the reconstructed 3D homogeneous coordinates, u i Indicates the horizontal position parameter of the texture image, v i Indicates the position parameter of the texture image in the vertical direction, x i ,y i , z i They represent the position parameters of the three-dimensional image in the horizontal, vertical and depth directions respectively, thereby obtaining the preliminary segmentation results of the 3D space.

[0080] like Figure 4 As shown in the figure, after obtaining the 3D point cloud by 2D back-projection, the 3D point cloud is input into the 3D U-Net, and the 3DUnet network is used to extract features, extract point cloud features, and pass the point cloud features through the MLP (multi-layer perceptron) prediction head to obtain the final semantic prediction results.

[0081] For the 3D scene feature extraction and semantic prediction model, the loss function is the point-by-point classification cross entropy loss L seg , specifically expressed as:

[0082]

[0083] Among them, C represents the number of categories, N represents the number of images, and y i,c represents the true category, y' i,c Indicates the predicted category.

[0084] Step S3: Perform a bottom-up OBM (Outdoor Bidirectional Merging) fusion process between image frames to obtain an optimized 3D point cloud category-independent segmentation result. Then, use a normal-based graph cutting method to obtain a scene segmentation mask based on geometric information. This segmentation result is fused with the obtained class-independent segmentation result through OBM. The ground is extracted separately from the final result and the mask is unified.

[0085] like Figure 5 As shown in the figure, OBM fusion is performed between frames. The specific description is as follows: Assume that the point clouds of two adjacent frames are and Where p1 and p2 represent the number of point clouds in X1 and X2 respectively; first calculate the corresponding mapping M between X1 and X2, and the corresponding mapping relationship is calculated by the nearest neighbor method; if (i, j)∈M, then it means that point and There is a matching relationship between the two frames. For each matching point pair, define the fusion confidence weight W c , the specific expression is:

[0086]

[0087] Among them, W pose is the spatial position weight, W feat is the feature similarity weight, W sem is the semantic similarity weight; for any mask number m in X1, find all the points belonging to the mask, denoted as Q1, and the number of points is Next, use the corresponding mapping M to find its corresponding point set Q2 in X2; for any region in Q2 with a mask number n remember For the region points, and use Indicates the area with mask number n in X2 The number of points in the area; if The mask numbered m in X1 and the region numbered n in X2 are considered to be the same region; is the attenuation factor, δ0 is taken as 0.5;

[0088] For each matching point pair The spatial position weight is:

[0089]

[0090] The feature similarity weight is:

[0091]

[0092] Among them, ‖·‖2 represents the l2 norm, represents the feature of point i in the first frame, Same thing.

[0093] The predicted semantic weight is:

[0094]

[0095] The confidence weight of the region is obtained by weighted averaging:

[0096]

[0097] Therefore, the point numbered m in X2 is assigned a new number n to merge the masks. After completing the above processing from X1 to X2, a symmetric processing is performed, that is, the mask numbers in X2 are traversed, and the corresponding points and mask numbers are found in X1 to achieve "bidirectional" merging and obtain the optimized full-scene segmentation result.

[0098] The result and the over-segmentation result obtained by applying geometric information are fused with OBM again. At this time, the parameter is δ0, and Set to 1.

[0099] The geometry-based segmentation mask can be obtained by graph-based segmentation, and the road surface extraction can be obtained by RANSAC.

[0100] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. 3D segmentation method based on bidirectional fusion, characterized by The steps include: Step S1: Obtain image information and camera information of a fixed scene, Step S2: The image information and camera information are passed through a 3D image segmentation model to obtain a preliminary segmentation result of the 3D space; Step S3: Based on the preliminary segmentation results, two adjacent frames of point cloud images are bidirectionally fused to obtain an optimized segmentation result. Then, a segmentation result based on geometric information of the scene is obtained. The point cloud image based on the geometric information segmentation result and the point cloud image of the optimized segmentation result are bidirectionally fused to obtain the final segmentation result. The bidirectional fusion process is as follows: Calculate point cloud image and point cloud images The corresponding mapping M between them, p1 and p2 represent the number of point clouds in the two point cloud images; if (i, j)∈M, it means point with dot There is a matching relationship between the two point cloud images. For each matching point pair, the fusion confidence weight W is defined c : Among them, W pose represents the spatial position weight, W feat represents the feature similarity weight, W sem Represents the semantic similarity weight; for any mask number m in the point cloud image X1, find all points Q1 belonging to the mask, and the number of points is recorded as Use the corresponding mapping M to find the corresponding point set Q2 in the point cloud image X2. For any area with a mask number n in Q2 remember For the region Points, use Indicates the area with mask number n in X2 The number of points in the area, if it satisfies It is considered that the mask numbered m in the point cloud image X1 and the area numbered n in the point cloud image X2 are the same area, δ t represents the attenuation factor; Based on the fusion confidence weight, the point numbered m in the point cloud image X2 is assigned a new number n to merge the mask, perform symmetric processing, and complete the bidirectional merging.

2. The 3D segmentation method based on bidirectional fusion according to claim 1, characterized in that: In step S1, the image information includes depth information, and the camera information includes camera intrinsic and extrinsic parameter information. In step S2, the two-dimensional mask is back-projected into three-dimensional space based on the extracted depth information and camera intrinsic and extrinsic parameter information. The specific formula is as follows: [x i ,y i ,z i ] T =R -1 M -1 [u i ,v i ,1] T -R -1 T Where M represents the intrinsic camera calibration matrix, R and T represent the rotation matrix and translation matrix obtained from the external camera pose matrix, [u i , v i , 1] and [x i ,y i ,z i , 1] represent the two-dimensional homogeneous coordinates and the reconstructed three-dimensional homogeneous coordinates, u i Represents the horizontal position parameter of the texture image, v i Indicates the position parameter of the texture image in the vertical direction, x i ,y i , z i They represent the position parameters of the three-dimensional image in the horizontal, vertical and depth directions respectively. The three-dimensional point cloud is extracted through a neural network for feature prediction to obtain the preliminary segmentation results of the three-dimensional space.

3. The 3D segmentation method based on bidirectional fusion according to claim 1, characterized in that: In step S3, for each matching point pair The spatial position weight is: Here, ||·||2 represents the l2 norm.

4. The 3D segmentation method based on bidirectional fusion according to claim 1, characterized in that: In step S3, for each matching point pair The feature similarity weight is: Among them, ||·||2 represents the l2 norm, represents the feature of point i in the first frame, Same thing.

5. The 3D segmentation method based on bidirectional fusion according to claim 1, characterized in that: In step S3, for each matching point pair The predicted semantic weight is:

6. The 3D segmentation method based on bidirectional fusion according to claim 1, characterized in that: In step S3, for each matching point pair The confidence weight of the region is obtained by weighted averaging:

7. The 3D segmentation method based on bidirectional fusion according to claim 1 is applied to outdoor vehicle driving, characterized in that: The road point cloud information is collected to obtain the segmentation results, and the ground is extracted separately according to the segmentation results and the masks are unified.

8. The 3D segmentation method based on bidirectional fusion according to claim 7, characterized in that: The image information in step S1 includes depth information, and the steps for obtaining the depth information through the point cloud are as follows: Step S11: Convert the point cloud from the initial sensor coordinate system to a coordinate system aligned with the camera, including the following steps: Step S111: converting the sensor coordinate system to the vehicle coordinate system; Step S112: converting the vehicle coordinate system to the world coordinate system; Step S113: converting the world coordinate system to the vehicle coordinate system corresponding to the camera's location at that moment; Step S114: converting the vehicle coordinate system to the camera coordinate system; Step S12: Projecting through the camera intrinsic parameter matrix to obtain the image coordinates and depth information corresponding to each point; Step S13: Filter out unsatisfactory points; retain points that are within the projected image boundary and more than a certain distance in front of the camera; Step S14: Generate a depth map on the image plane; initialize a two-dimensional array with the same size as the camera image, and fill the retained points into it according to their projection coordinates, and assign them the corresponding depths.

9. The 3D segmentation method based on bidirectional fusion according to claim 7, characterized in that: The image information in step S1 includes posture information, and the steps for obtaining the posture information are as follows: The rotation matrix and translation vector of the ego vehicle relative to the world coordinate system are extracted from the corresponding camera pose record. Then, the rotation matrix and translation vector of the camera relative to the ego vehicle coordinate system are extracted from the camera calibration information. Then, these two transformations are combined through linear algebra to obtain the overall transformation from the world coordinate system to the camera coordinate system.

10. A 3D segmentation system based on bidirectional fusion, comprising an information extraction module, a 3D image segmentation module, and a bidirectional fusion module, characterized in that: The 3D segmentation method based on bidirectional fusion according to any one of claims 1 to 9 is adopted to sequentially acquire and perform three-dimensional image segmentation based on image information of a fixed scene and camera information to obtain a preliminary segmentation result, and bidirectional fusion is performed based on the preliminary segmentation result to obtain a final segmentation result.

Citation Information

Patent Citations

  • Large-scale point cloud segmentation method and device based on deep learning, equipment and medium

    CN116310349A

Cited By

  • Bulk cargo cabin real-time fusion detection method based on laser point cloud

    CN120953342A

  • A Method and System for Decoupling 3D Scene Targets Based on Image Segmentation and 3D Reconstruction

    CN122676080A