Object segmentation method based on SAM and neural radiation field
Through the object segmentation method based on SAM and neural radiation field, bidirectional matching and iterative cycle optimization are used to solve the problems of unstable input and repeated operations in three-dimensional object segmentation, and efficient and stable three-dimensional object segmentation is achieved.
Patent Information
- Application Number
- CN202510349094.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-04
AI Technical Summary
In the existing three-dimensional object segmentation method, the prompt input is unstable and the segmentation of similar objects cannot be operated efficiently in batches, resulting in the problems of discontinuous and repeated operations in the segmentation results.
The object segmentation method based on SAM and neural radiation field is adopted, and the optimal prompt input is obtained through a two-way matching strategy, the point tracking algorithm is used to conduct the segmentation mask at the optimal perspective, and the segmentation results are optimized through iterative loops, and the neural radiation field is rendered.
Improve the stability and efficiency of segmentation results, ensure the selection of optimal prompt input, effectively utilize the spatial continuity of three-dimensional scenes, and reduce the discontinuity and repeated operations of segmentation results.
Smart Images

Figure CN120259340A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of object segmentation, and in particular to an object segmentation method based on SAM (Segment Anything Model) and neural radiance fields. Background Art
[0002] 3D object segmentation is one of the fundamental tasks in 3D vision and has wide applications in multiple fields such as medical images, 3D image editing, and autonomous driving. Given several pictures with known viewpoints, and cues representing the object to be segmented (which can be in the form of points, bounding boxes, images of the object to be segmented, or text descriptions, etc.), the 3D object segmentation technology can obtain the segmentation pictures of the target object from any viewpoint in the scene. The latest research has gradually focused on vision foundation models in order to improve the encoding and decoding capabilities of the models and their generalization capabilities on unseen data. Among them, SAM has achieved prompt-based segmentation tasks with its powerful zero-shot ability. The general encoding and decoding capabilities of SAM improve the quality of image feature extraction, and its strong versatility and robustness also improve the segmentation accuracy in complex scenes. At the same time, after receiving cues representing the object to be segmented in the form of points or bounding boxes, it can accurately segment the target object from various viewpoints, which strongly promotes unsupervised 3D object supervision and further reduces the requirements for data.
[0003] The 3D object segmentation technology combined with SAM still has limitations. The cue inputs of these methods are mainly points and bounding boxes, and there are two practical application problems with such inputs. First, there can be multiple choices for the cue points or bounding boxes representing the object range on the picture, and cue inputs with different description accuracies will affect the segmentation results, and the results are not stable. It should not be assumed that users will necessarily select the optimal cue input in actual applications. Second, for scenarios where multiple objects of the same type need to be segmented, these methods require re-specifying the cue input for each scenario, which is not conducive to batch and efficient operations. Therefore, it is possible to choose to use the text or picture describing the object to be segmented as the cue input, and let the model calculate the text image features and similarities and analyze the most suitable point or bounding box input cue by itself. This input method not only ensures stable results and will definitely select the optimal cue input, but also can use only one set of pictures or text for segmenting multiple objects of the same type in multiple scenarios, avoiding repeated operations and improving application efficiency. Currently, this method is still less studied in 3D image object segmentation. Summary of the Invention
[0004] Object of the Invention: The technical problem to be solved by the present invention is to provide an object segmentation method based on SAM and neural radiance fields in view of the deficiencies of the prior art.
[0005] The present invention discloses an object segmentation method based on SAM and neural radiance fields, comprising the following steps:
[0006] Step 1, collect multi-view scene images of the scene to be segmented, corresponding camera parameters, a target reference image representing the target object to be segmented, and an object segmentation mask;
[0007] Step 2, calculate the similarity between the target reference image and the scene images, and use a bidirectional matching strategy to obtain the object segmentation prompts in the form of points for each view;
[0008] Step 3, input the multi-view scene images and their corresponding object segmentation prompts into SAM to obtain the object segmentation masks for each view;
[0009] Step 4, evaluate the object segmentation masks for each view using the designed metrics, and select the view with the best segmentation quality;
[0010] Step 5, propagate the object segmentation prompts of the view with the best segmentation quality to other views using a point tracking algorithm, and use the propagated object segmentation prompts to re-segment each view using SAM to obtain the object segmentation masks segmented for each view;
[0011] Step 6, record the object segmentation masks segmented for each view in this round in Step 5 as the output of this round, and at the same time accumulate them into the object segmentation masks of the previous rounds. Return to Step 2 to loop through the subsequent steps. The object segmentation prompts selected in Step 2 for each round do not include the masks output in all previous rounds. Keep looping until the evaluation metric of the object segmentation mask is less than the preset threshold, and stop the iteration;
[0012] Step 7, merge the object segmentation masks output in all previous rounds for each view respectively as the final object segmentation masks for each view, and input them and the corresponding camera parameters for each view into the neural radiance field for training to perform rendering to obtain the object segmentation masks in a new view;
[0013] Step 8, perform target segmentation in a specific scene according to the obtained object segmentation masks.
[0014] The final result obtained is the specified view information or camera parameters by the user, and the model renders the segmentation mask or segmentation image of the scene image in the corresponding view as the output result. The segmentation mask and the segmentation image can be simply converted to each other, so the segmentation mask is used as the output.
[0015] Step 1 includes the following steps:
[0016] Step 1-1: Obtain the scene image data from the open-source dataset SPIn-NeRF. Each scene includes 20 to 40 scene images from different viewpoints and the parameters of the camera.
[0017] Step 1-2: For each scene, obtain the target reference picture ref_img of the target object to be segmented provided by the user and the object segmentation mask ref_mask of the object to be segmented in this picture ref_img i , where ref_img i is the reference picture of the i-th scene, and ref_mask i is the object segmentation mask representing the object to be segmented in the reference picture of the i-th scene; i
[0018] The provided reference picture involves other objects or backgrounds in addition to the target to be segmented. Therefore, the specific target to be segmented is specified in the form of a mask.
[0019] Step 1-3: For each scene, combine the multi-viewpoint scene images and camera parameters in Step 1-1 with the target reference picture and object segmentation mask of its corresponding scene in Step 1-2 to form the complete input data.
[0020] Step 1-3 includes the following steps:
[0021] Step 1-3-1: According to the scene number i, obtain the multi-viewpoint scene images T i : (target_img1, target_img2,..., target_img n ) of the corresponding scene from the scene image data in Step 1-1, where n is the number of viewpoints in this scene;
[0022] Step 1-3-2: According to the scene number i, obtain the multi-viewpoint camera parameters C i : (camera1, camera2,..., camera n ) of the corresponding scene from the camera parameter data in Step 1-1, where n is the number of viewpoints in this scene;
[0023] Step 1-3-3: From the multi-viewpoint RGB images T i of the corresponding scene in Step 1-3-1, the multi-viewpoint camera parameters C i of the corresponding scene in Step 1-3-2, the target reference picture ref_img of the target object to be segmented from Step 1-2 i and the object segmentation mask ref_mask of the object to be segmented in the target reference picture ref_img i i , combined into the final input data Input i :(T i ,C i ,ref_img i ,ref_img i ).
[0024] Step 2 includes the following steps:
[0025] Step 2-1: Use the pre-trained image feature extraction model DINOv2 to process the input data i Multi-view scene image T in i and the target reference image ref_img i , respectively get the multi-view scene image T i The feature map F of each view in i :(target_feature1,target_feature2,...,target_feature n ) and the target reference image feature map ref_feature i , where n is the number of viewing angles in the scene, the resolution of the feature map remains consistent with the original image, and there is a one-to-one mapping relationship with the pixels of the text image;
[0026] Step 2-2, F i All feature maps and target reference image feature maps ref_feature i All are expanded into vectors, and the expanded ref_feature i One by one and the expanded feature map F i Calculate the cosine similarity matrix of all feature graphs in and get the forward similarity Sim i :(sim1,sim2,...,sim n );
[0027] Step 2-3: For each view j, if merged_mask has not been initialized j If it passes, initialize merged_mask j , merged_mask j Represents the image segmentation results accumulated in the loop iteration at each perspective, and its size is the same as target_img j The same, the initial values are all 0, where j represents the view number, ranging from 1 to n;
[0028] Step 2-4: For the target reference image ref_img i Each object segmentation mask ref_mask in the object to be segmented iAmong the pixel points, find its corresponding feature point in the target reference picture feature map ref_feature i and, through the forward similarity sim j in each perspective, find the point set P with the highest cosine similarity of eigenvalues to these feature points j , and the points with the highest cosine similarity cannot have a non-zero value in merged_mask j . If the value in merged_mask j is non-zero, continue to find the point with the second-largest similarity, and so on until a point with a value of 0 in merged_mask j is found;
[0029] Step 2-5: Expand all the feature maps in the feature map F i and the target reference picture feature map ref_feature i into vectors. After expansion, for all features in F i calculate the cosine similarity matrix with the expanded object segmentation mask ref_mask Figure 1 to obtain the reverse similarity Rev_sim i : (rev_sim1, rev_sim2,..., rev_sim i ); n )
[0030] Step 2-6: For each perspective j, use the reverse similarity rev_sim j to find the point on the target reference picture ref_img j with the largest cosine similarity among all points in the point set P i . If the point with the largest similarity is not in the object segmentation mask ref_mask i , remove its corresponding point from the point set P j to obtain the prompt of the object to be segmented in the form of points for each perspective.
[0031] Step 3 includes the following steps:
[0032] Step 3-1: For each perspective j, use the point set P j as the forward point input of SAM in this perspective and combine it with the scene image target_img j to form the SAM input data prompt j : (P j , target_img j );
[0033] Step 3-2: For the input data prompt jInput them into SAM respectively to obtain the object segmentation mask coarse_mask from each perspective j 。
[0034] Step 4 includes the following steps:
[0035] Step 4-1: For each perspective j, calculate the purity j
[0036]
[0037] where Num(*) represents the number of points in point set *, Area(*) represents the area size of mask *, and P in_mask represents point set P j The part of the points in P j that are in coarse_mask j The object segmentation mask coarse_mask is obtained from Step 3-2. Purity measures how many points in point set P are contained in the segmented image per unit area, so it evaluates the accuracy of the segmentation result. If there are irrelevant parts in the segmentation result that do not belong to the segmentation result, or redundant parts that do not contain the points in point set P j then the purity will decrease. This indicator encourages the model to segment the picture as accurately and without redundancy as possible; j
[0038] Step 4-2: For each perspective j, calculate the coverage j
[0039]
[0040] where Num(*) represents the number of points in point set *, and P in_mask represents point set P j The part of the points in P j that are in coarse_mask j Coverage measures the proportion of point set P that is included in the segmentation result, so it evaluates the integrity of the segmentation result. If the model does not segment all the targets, or a large number of points in point set P j are not included in the segmentation result, then the coverage will decrease. This indicator encourages the model to segment the picture as completely and without omission as possible;
[0041] Step 4-3: For each perspective j, calculate the evaluation score score j
[0042] score j = purity j α * coverage j β
[0043] Among them, α and β are set hyperparameters, which are used to adjust the respective weights of purity and coverage in the evaluation score score. The larger α is, the more score focuses on purity. In this case, it is more inclined to give a higher score to the object segmentation mask that makes the segmentation result more accurate. However, in extreme cases, the model may only segment out a very small part of the target to obtain a higher purity. Therefore, β is introduced for balance. The larger β is, the more score focuses on coverage. In this case, it is more inclined to give a higher score to the object segmentation mask that segments out all targets as much as possible, balancing the previous extreme case. However, in this case, there is still an extreme where the model does not segment the image at all and outputs the entire image as the segmentation result to obtain a higher coverage. Therefore, α and β adjust and balance each other to make the score evaluate the segmentation result as accurately as possible;
[0044] Step 4-4: For scene i, compare the evaluation scores score of all perspectives therein, find the perspective with the highest score, record its perspective subscript as best, and the corresponding set of SAM forward point input points at this perspective is P best .
[0045] Step 5 includes the following steps:
[0046] Step 5-1: Transmit the point set P obtained in Step 4-4 best to other perspectives using the point tracking algorithm, and record the point set obtained in this way at perspective j as P` j ;
[0047] Step 5-2: For each perspective j, use the point set P` j as the forward point input of SAM at this perspective, and combine it with the scene image target_img j to form the transmitted SAM input data prompt` j : (P` j , target_img j );
[0048] Step 5-3: Input the prompt` j at each perspective into SAM respectively to obtain the object segmentation mask fine_mask j .
[0049] Step 6 includes the following steps:
[0050] Step 6-1: For each perspective j, accumulate the object segmentation mask fine_mask obtained in Step 5-3 j into the merged segmentation mask merged_mask j . Here, merged_mask j represents the accumulated value of fine_mask j in each round for perspective j, and its initial value is 0;
[0051] Step 6-2: Determine whether the score in this round is less than a preset threshold. If it is less than the threshold, stop the iteration, and use the merged segmentation mask merged_mask as the final segmentation mask for this perspective. If it is still greater than the threshold, return to Step 2 and continue to loop and execute.
[0052] Step 7 includes the following steps:
[0053] Step 7-1: According to the scene number i, obtain the multi-perspective segmentation mask M i :(merged_mask1, merged_mask2,..., merged_mask n ) for the corresponding scene from Step 6-2, and obtain the multi-perspective camera parameters C i for the corresponding scene from Step 1-3-2, and combine them into the input data NerfIn i :(M i , C i ) of the neural radiance field, where n is the number of perspectives in this scene;
[0054] Step 7-2: Train the neural radiance field through the input data NerfIn i of the neural radiance field;
[0055] Step 7-3: Use the trained neural radiance field to render the object segmentation mask in a new perspective.
[0056] The scenes described in Step 8 include medical image scenes and vehicle detection scenes. According to the obtained object segmentation mask, a series of downstream task applications can be carried out. For example, the segmentation target can be extracted from the 3D scene, allowing art creators to perform directional 3D editing on the target or realizing the removal of specific targets from the scene. It can also be used in medical images to achieve lesion detection and segmentation and other medical purposes, or in the field of autonomous driving to achieve the detection of pedestrians and vehicles, etc.
[0057] In the present invention, the SAM and the neural radiance field are both existing models.
[0058] Beneficial effects:
[0059] 1) Different from the previous method that uses points and bounding boxes as prompt inputs, this method chooses to use pictures describing the objects to be segmented as prompt inputs, ensuring stable results. It will definitely select the optimal prompt input, avoiding the situation where users cannot accurately select input points or input boxes. Moreover, it can use a single set of pictures for segmenting similar objects in multiple scenarios, avoiding repeated operations and facilitating users to achieve personalized batch processing.
[0060] 2) This method proposes a two-way matching method to efficiently extract high-quality SAM prompt inputs from pictures and obtain more accurate segmentation results with more accurate prompt inputs.
[0061] 3) Different from the previous methods in the same field that directly use SAM isolatedly on pictures of each perspective in a three-dimensional scene, ignoring the natural spatial continuity information of the three-dimensional scene pictures, this method proposes a multi-perspective conduction method to conduct the optimal SAM prompt inputs among different perspectives, effectively utilizing the spatial continuity of the three-dimensional scene, alleviating the problem that the segmentation target may be well segmented in some perspectives but difficult to segment in other perspectives, and mitigating the problem of discontinuous segmentation results among different perspectives.
[0062] 4) Different from the previous method that only performs segmentation once, this method adopts an iterative loop segmentation method to perform segmentation on each perspective multiple times, avoiding the situation where only a certain part or a single object with the most complete form and the highest similarity to the reference picture can be segmented each time, rather than segmenting all the target objects that meet the requirements.
[0063] The above operations propose a reasonable way to utilize three-dimensional spatial continuity information and iterative optimization, and at the same time select a stable and efficient input form, thereby further improving the effect of object segmentation based on SAM and neural radiance fields. Brief Description of the Drawings
[0064] Figure 1 It is a flowchart of the present invention.
[0065] Figure 2 It is a schematic diagram of an example of an input picture.
[0066] Figure 3 It is a schematic diagram of the object segmentation result of a three-dimensional image.
[0067] Figure 4 It is a system framework diagram of the method of the present invention.
[0068] Figure 5 It is a schematic diagram of the comparison of object segmentation between the method of the present invention and other methods. Detailed Embodiment
[0069] The method of the present invention is dedicated to solving the object segmentation method based on SAM and neural radiance fields. First, a bidirectional matching strategy is adopted to select input prompt points to be input to SAM on the target image. Then, the multi-view scene images and input prompt points are input to SAM to obtain a preliminary object segmentation mask. Next, the purity and coverage are used to evaluate the object segmentation quality in each view, and the view with the best segmentation quality is selected. Then, the input prompt points in the view with the best segmentation quality are propagated to other views using the point tracking algorithm PIPS to find the positions of these points in other views in 3D space. Then, SAM is used to segment each view again based on the propagated input prompt points, and the masks of this segmentation are accumulated into the cumulative masks recorded for each view. Then, this process is repeated in a loop until the score index of the segmented image is less than a preset threshold, and at the same time, points that have already been in the cumulative mask are not selected as input prompt points during the loop process. Then, the cumulative masks of each view and their respective camera parameters are input to the Nerf neural radiance field for training, and finally, the object segmentation mask in the new view is rendered.
[0070] Example:
[0071] The target task of this example is as Figure 2 and Figure 3 shown, Figure 2 The reference picture representing the object to be segmented and the source view of the target scene to be segmented. The leftmost column represents the reference picture of the object to be segmented, and the right three columns represent the source views of the target scene to be segmented. Figure 3 is the segmentation result in the new view inferred from Figure 2 The overall structural system of the entire method is as Figure 4 shown.
[0072] The following describes each step of the present invention according to the example.
[0073] Step (1), collect multi-view scene image data of the scene to be segmented, corresponding camera parameters, as well as a reference picture representing the target object to be segmented and an object segmentation mask, specifically including the following steps:
[0074] Step (1.1), obtain scene data from the open-source dataset SPIn-NeRF dataset. Each scene includes 20 to 40 RGB images at different views and the parameters of the camera.
[0075] Step (1.2), for each scene, obtain a reference picture ref_img provided by the user himself representing the target object to be segmented i and the object segmentation mask ref_mask representing the object to be segmented in the reference picture ref_img i , where ref_img i , where ref_imgi is the reference image for the i-th scene, ref_mask i is the object segmentation mask representing the object to be segmented in the reference image of the i-th scene;
[0076] Step (1.3), for each scene, combine the multi-view RGB images and camera parameters in step (1.1) with the reference image and object segmentation mask of its corresponding scene in step (1.2) to form the complete input data. This step specifically includes the following steps:
[0077] Step (1.3.1), according to the scene number i, obtain the multi-view RGB image T of the corresponding scene from the RGB image data in step (1.1) i :(target_img1,target_img2,...,target_img n ), where n is the number of viewpoints in this scene;
[0078] Step (1.3.2), according to the scene number i, obtain the multi-view camera parameters C of the corresponding scene from the camera parameter data in step (1.1) i :(camera1,camera2,...,camera n ), where n is the number of viewpoints in this scene;
[0079] Step (1.3.3), according to the scene number i, obtain the multi-view RGB image T of the corresponding scene from step (1.3.1) i , obtain the multi-view camera parameters C of the corresponding scene from step (1.3.2) i , obtain the reference image ref_img representing the target object to be segmented from step (1.2) i and the object segmentation mask ref_mask representing the object to be segmented in the reference image ref_img i , and combine them into the final input data Input i :(T i ,C i ,ref_img i ,ref_img i ,ref_img i ).
[0080] Step (2), calculate the similarity between the reference image of the target object and the multi-view scene images of the scene to be segmented, and use the bidirectional matching strategy to obtain the prompts of the object to be segmented in the form of points for each viewpoint. Step (2) specifically includes the following steps:
[0081] Step (2.1), use the pre-trained image feature extraction model DINOv2 to process Input iT in i and ref_img i to obtain the multi-view RGB images T i and the feature maps F i for each view in n :(target_feature1, target_feature2,..., target_feature n ) and the feature map ref_feature i , where n is the number of views in this scenario, the resolution of the feature maps remains the same as the original images, and there is a one-to-one mapping relationship with the pixel points of the text images;
[0082] Step (2.2), expand all the feature maps in F i and ref_feature i into vectors. After expansion, ref_feature i calculates the cosine similarity matrix with all the feature maps in the expanded F i to obtain the positive similarity Sim i :(sim1, sim2,..., sim n );
[0083] Step (2.3), for each view j, if merged_mask j has not been initialized yet, then initialize merged_mask j . merged_mask j represents the cumulative image segmentation results in each view during the loop iteration, and its size is the same as target_img j , with all initial values being 0, where j represents the view number, and the value range is from 1 to n;
[0084] Step (2.4), for each pixel point in ref_img i that is within ref_mask i , find its corresponding feature point in ref_feature i . In each view, find the set of points P j with the highest cosine similarity of eigenvalues to these feature points through sim j , and these points with the highest cosine similarity cannot have a non-zero value in merged_mask j . If the value in merged_mask j is non-zero, then continue to find the point with the second highest similarity, and so on until a point with a value of 0 in merged_mask j is found;
[0085] Step (2.5), expand all the feature maps and ref_feature in F i into vectors. After expansion, for all the features in F i and the expanded ref_mask i calculate the cosine similarity matrix to obtain the reverse similarity Rev_sim Figure 1 : (rev_sim1, rev_sim2,..., rev_sim i ); i n )
[0086] Step (2.6), for each perspective j, use rev_sim j to find the point in ref_img j with the maximum cosine similarity for all points in the point set P i . If the point with the maximum similarity is not in ref_mask i , then remove its corresponding point from the point set P j to obtain the prompt of the object to be segmented in the form of points for each perspective.
[0087] Step (3), input the multi-perspective scene image and its corresponding object segmentation prompt into SAM to obtain the object segmentation mask for each perspective. Step (3) specifically includes the following steps:
[0088] Step (3.1), for each perspective j, use the point set P j as the positive points input of SAM for this perspective, and combine it with target_img j to form the SAM input data prompt j : (P j , target_img j );
[0089] Step (3.2), input the prompts for each perspective j into SAM respectively to obtain the object segmentation mask coarse_mask j for each perspective.
[0090] Step (4), evaluate the object segmentation masks for each perspective using the designed metrics, and select the perspective with the best segmentation quality. Step (4) specifically includes the following steps:
[0091] Step (4.1), for each perspective j, calculate the purity j
[0092]
[0093] Among them, Num(*) represents the number of points in point set *, Area(*) represents the area size of mask *, and P in_mask represents point set P j in the part of points that are in coarse_mask j and coarse_mask j is obtained from step (3.2);
[0094] Step (4.2), for each view j, calculate the coverage rate coverage j
[0095]
[0096] Among them, Num(*) represents the number of points in point set *, and P in_mask represents point set P j in the part of points that are in coarse_mask j and forms a point set;
[0097] Step (4.3), for each view j, calculate the evaluation score score j
[0098] score j = purity j α * coverage j β
[0099] Among them, α and β are set hyperparameters, and finally α and β are set to 0.2 and 1.0 respectively through experiments;
[0100] Step (4.4), for scene i, compare the evaluation scores score of all views in it, find the view with the highest score, record its view subscript as best, and the corresponding SAM forward point input point set under this view is P best .
[0101] Step (5), propagate the prompt of the object to be segmented in the view with the best segmentation quality to other views using the PIPS model, and use the better-quality prompt obtained by propagation to re-segment each view using SAM to obtain the object segmentation masks segmented in each view. Step (5) specifically includes the following steps:
[0102] Step (5.1), use the point set P obtained in step (4.4) best , and conduct it to other views using the point tracking algorithm PIPS. Denote the point set obtained in this way in view j as P` j ;
[0103] Step (5.2), for each perspective j, use the point set P` j as the forward point input of SAM under this perspective, and combine it with target_img j to form the SAM input data prompt` j :(P` j , target_img j );
[0104] Step (5.3), input the prompt` j under each perspective into SAM respectively, and obtain the object segmentation mask fine_mask j under each perspective.
[0105] Step (6), record the object segmentation masks segmented under each perspective in Step 5 of this round as the output of this round, and at the same time accumulate them into the object segmentation masks of the previous rounds. Return to Step 2 to loop through the subsequent steps, but the prompts of the objects to be segmented selected in each Step 2 cannot be in the masks output in all previous rounds. Keep looping until the evaluation index of the object segmentation mask is less than the preset threshold, and then stop the iteration. Step (6) specifically includes the following steps:
[0106] Step (6.1), for each perspective j, accumulate the fine_mask j obtained in Step (5.3) into the merged segmentation mask merged_mask j , where the initial value of merged_mask j is 0, representing the accumulated value of fine_mask j under perspective j in each round;
[0107] Step (6.2), judge whether the score of this round is less than the preset threshold. If it is less than the threshold, stop the iteration and use merged_mask as the final segmentation mask for this perspective. If it is still greater than the threshold, return to Step 2 to continue looping and executing.
[0108] Step (7), merge the object segmentation masks output in all previous rounds under each perspective respectively as the final object segmentation masks for each perspective, input them and the corresponding camera parameters for each perspective into the neural radiance field training, and finally perform rendering to obtain the object segmentation masks in the new perspective. Step (7) specifically includes the following steps:
[0109] Step (7.1), according to the scene number i, obtain the multi-perspective segmentation masks N i for the corresponding scene from Step (6.2):(merged_mask1, merged_mask2,..., merged_mask n) Obtain the multi-view camera parameters C for the corresponding scenario from step (1.3.2). i Combine them into the input data NerfIn of the neural radiance field. i : (M i , C i ), where n is the number of viewpoints in this scenario;
[0110] Step (7.2): According to the scenario number i, obtain the input data NerfIn of the neural radiance field for the corresponding scenario from step (7.1). i Use it as the input to train the neural radiance field.
[0111] Step (7.3): Use the trained neural radiance field to render the object segmentation mask in a new viewpoint.
[0112] Step (8): According to the obtained object segmentation mask, perform target segmentation in a specific scenario;
[0113] Step (8.1): According to the obtained object segmentation mask, a series of downstream task applications can be carried out. For example, extract the segmented target from the 3D scene, enable art creators to perform targeted 3D editing on the target or remove specific targets from the scene. It can also be used in medical images to achieve lesion detection and segmentation and other medical purposes, or be applied in the field of autonomous driving to achieve the detection of pedestrians and vehicles, etc.
[0114] Result analysis:
[0115] The experimental environment parameters of the method of the present invention are as follows:
[0116] 1) The experimental platform parameters for the neural radiance field training and result testing process of the object segmentation method based on SAM and neural radiance field are Ubuntu 20.04.6 64-bit operating system, AMD Ryzen 9 5900X 12-Core, 64GB of memory, NVIDIA GeForce RTX 3090 24GB graphics card, using the Python programming language, the programming and development environment is Visual Studio Code, and third-party open-source libraries such as Pytorch and Numpy are used to implement.
[0117] The comparative experimental results (as shown in Table 1) of the method of the present invention and the methods in Document 7 (abbreviated as MVSeg) and Document 17 (abbreviated as SA3D) are analyzed as follows:
[0118] This method is an unsupervised method that does not require training and is directly tested on 10 scenes of the open-source dataset SPIn-NeRF dataset, numbered Orchids, Leaves, Fern, Room, Horns, Fortress, Fork, Pinecone, Truck, Lego. The comparison experiment results are shown in Tables 1 and 2. Among them, IoU and Acc are selected as metrics. IoU (Intersection over Union) is the intersection over union, which measures the overlapping degree between the predicted region and the true region, that is, the size of the intersection of the predicted segmentation mask and the true segmentation mask is divided by the size of the union. The larger the value, the better the segmentation quality. Acc (Accuracy) is the accuracy rate, which represents the number of correctly classified pixels in the segmented image divided by the total number of pixels in the image. Pixels predicted as the background outside the segmented object will also be counted as correctly classified as long as they are not actually part of the segmented object. Therefore, Acc is usually larger than IoU numerically, and the larger the value, the better the segmentation quality. Acc is suitable as an overall reference value to measure the classification accuracy of all pixels, but it may overestimate the situation with a large background area. IoU is a more direct evaluation criterion that can more accurately measure the matching of the target region and usually can obtain a more fair evaluation.
[0119] As Figure 5 shown, the comparison of the method of the present invention with MVSeg (Multispectral Video Semantic Segmentation) and SA3D (Segment Anything in 3D with NeRFs) (in order to show the segmentation results from different perspectives in the same scene, only the comparison of a single scene is shown). As shown in the comparison of the metrics in Tables 1 and 2 (Table 1 shows the comparison of the metrics of the method of the present invention with other methods on 10 scenes of the SPIn-NeRF dataset, and Table 2 shows the statistical comparison of the average metrics of the method of the present invention with other methods on 10 scenes of the SPIn-NeRF dataset), the method of the present invention is ahead of the MVSeg and SA3D methods, and exceeds the MVSeg and SA3D methods in both the average metrics and the vast majority of single-scene metrics, confirming that this experiment has good results for the 3D image object segmentation method.
[0120] Table 1 Comparison of the metrics of the method of the present invention with other methods on the SPIn-NeRF dataset
[0121]
[0122]
[0123] Table 2 Statistical comparison of the average metrics of the method of the present invention with other methods on the SPIn-NeRF dataset
[0124] Index MVSeg SA3D The method of the present invention IoU (%) 90.9 92.4 94.7 Acc (%) 98.9 98.9 99.2
[0125] In the self-comparison experiment, in the first group, the multi-view conduction operation was removed, and the prompt input points of the optimal view were no longer transmitted between each view. In the second group, the iterative loop segmentation operation was removed. Although the prompt input points of the optimal view were transmitted between each view, only one round of segmentation was performed, and the segmentation masks of each round were no longer loop-segmented and merged. The comparison with the final experimental result indicators is shown in Table 3, indicating that the method of multi-view conduction of prompt input points can effectively utilize the spatial continuity information of the three-dimensional image, and the method of iterative loop segmentation can greatly improve the segmentation quality, and can effectively reduce the situation where not all targets are segmented in the scenario with multiple segmentation targets.
[0126] Table 3 Comparison table of the final results of the method of the present invention and the method using inverse distance weighting
[0127]
[0128]
[0129] The present invention provides an idea and method for an object segmentation method based on SAM and neural radiance fields. There are many methods and ways to specifically implement this technical solution. The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by using the prior art.
Claims
1. A method for object segmentation based on SAM and neural radiance fields, characterized in that, It includes the following steps: Step 1: Collect multi-view scene images of the scene to be segmented, corresponding camera parameters, target reference images of the target object to be segmented, and object segmentation masks; Step 2: Calculate the similarity between the target reference image and the multi-view scene images, and use a bidirectional matching strategy to obtain the object prompts to be segmented in the form of points for each view; Step 3: Input the multi-view scene images and their corresponding object prompts to be segmented into SAM to obtain the object segmentation masks for each view; Step 4: Evaluate the object segmentation masks for each view using the designed metrics, and select the view with the best segmentation quality; Step 5: Propagate the object prompts to be segmented in the view with the best segmentation quality to other views using a point tracking algorithm, and use the propagated object prompts to be segmented to re-segment each view using SAM to obtain the object segmentation masks segmented for each view; Step 6: The object segmentation masks segmented for each view obtained in Step 5 are used as the output for this round, and are accumulated into the object segmentation masks of the previous rounds. Then, return to Step 2 to loop through the subsequent steps. The object prompts to be segmented selected in Step 2 for each round do not include the masks output in all previous rounds. Loop until the evaluation metric of the object segmentation mask is less than the preset threshold, and stop the iteration; Step 7: Merge the object segmentation masks output in all previous rounds for each view respectively as the final object segmentation masks for each view, and input them and the corresponding camera parameters for each view into a neural radiance field for rendering to obtain the object segmentation masks in the new view; Step 8: Perform target segmentation in the specific scene according to the obtained object segmentation masks.
2. The object segmentation method based on SAM and neural radiance fields according to claim 1, characterized in that, Step 1 includes the following steps: Step 1-1: Obtain scene image data, scene images from different views, and camera parameters from the existing dataset; Step 1-2: For each scenario, obtain the target reference image ref_img of the target object to be segmented provided by the user himself / herself i and the object segmentation mask ref_mask of the object to be segmented in this image ref_img i where ref_img i represents the reference image of the i-th scenario, and ref_mask i represents the object segmentation mask of the object to be segmented in the reference image of the i-th scenario; i Step 1-3: For each scene, combine the multi-view scene images and camera parameters in Step 1-1 with the target reference image and object segmentation mask of its corresponding scene in Step 1-2 to form the complete input data.
3. A method for object segmentation based on SAM and neural radiance fields according to claim 2, characterized in that Step 1-3 includes the following steps: Step 1-3-1: Obtain the multi-view scene images T of the corresponding scene from the scene image data in Step 1-1 according to the scene serial number i i :(target_img1, target_img2,..., target_img n ), where n is the number of viewpoints in this scene; Step 1-3-2: Obtain the multi-view camera parameters C for the corresponding scenario from the camera parameter data in Step 1-1 according to the scenario serial number i i :(camera1, camera2,..., camera n ), where n is the number of viewpoints in this scenario; Step 1-3-3: Combine the multi-view RGB images T i corresponding to the scene in Step 1-3-1, the multi-view camera parameters C i corresponding to the scene in Step 1-3-2, the target reference image ref_img of the target object to be segmented from Step 1-2 i and the object segmentation mask ref_mask of the object to be segmented in the target reference image ref_img i to form the final input data Input i : (T i , C i , ref_img i , ref_img i , ref_img i ).
4. A method for object segmentation based on SAM and neural radiance fields according to claim 3, characterized in that Step 2 includes the following steps: Step 2-1: Process the input data Input using a pre-trained image feature extraction model i for the multi-view scene image T i and the target reference picture ref_img i to obtain the feature maps F i for each view in the multi-view scene image T i : (target_feature1, target_feature2,..., target_feature n ) and the target reference picture feature map ref_feature i , where n is the number of views in this scene, the resolution of the feature map remains the same as the original image, and there is a one-to-one mapping relationship with the pixel points of the text image; Step 2-2: Take the feature map F i and all the feature maps in it and the target reference picture feature map ref_feature i and expand them all into vectors. After expansion, ref_feature i is calculated with all the feature maps in the expanded feature map F i to obtain a cosine similarity matrix, and the positive similarity Sim i is obtained: (sim1, sim2,..., sim n ); Step 2-3: For each pixel in the object segmentation mask ref_mask of the object to be segmented in the target reference image ref_img i find its corresponding feature point in the target reference image feature map ref_feature i and find the point set P with the highest cosine similarity of eigenvalues to these feature points from various perspectives through the forward similarity sim i where j represents the perspective serial number, and the value range is from 1 to n; j j Step 2-4: Unfold all the feature maps in the feature map F i and the target reference picture feature map ref_feature i into vectors. For each feature map in the unfolded F i , calculate the cosine similarity matrix with the unfolded object segmentation mask ref_mask i to obtain the reverse similarity Rev_sim i : (rev_sim1, rev_sim2,..., rev_sim n ); Step 2-5. For each perspective j, use the reverse similarity rev_sim j to find the point in the point set P j with the maximum cosine similarity in the target reference image ref_img i . If the point with the maximum similarity is not in the object segmentation mask ref_mask i , then remove its corresponding point from the point set P j to obtain the object to be segmented prompt in the form of points for each perspective.
5. A method for object segmentation based on SAM and neural radiance fields according to claim 4, characterized in that, Step 3 includes the following steps: Step 3-1: For each perspective j, use the point set P j as the forward point input of SAM under this perspective and combine it with the scene image target_img j to form the SAM input data prompt j : (P j , target_img j ); Step 3-2: Input the SAM input data prompt from each perspective j and input them into SAM respectively to obtain the object segmentation mask coarse_mask from each perspective j .
6. A method for object segmentation based on SAM and neural radiance fields according to claim 5, characterized in that, Step 4 includes the following steps: Step 4-1: For each perspective j, calculate the purity j : Among them, Num(*) represents the number of points in point set *, Area(*) represents the area size of mask *, and P in_mask represents point set P j in the part of the points that are in coarse_mask j formed by the point set, and the object segmentation mask coarse_mask j is obtained from step 3-2; Step 4-2: For each perspective j, calculate the coverage j : where Num(*) represents the number of points in point set *, and P in_mask represents point set P j in the part of coarse_mask j composed of the points in it; Step 4-3: For each perspective j, calculate the evaluation score score j score j = purity j α * coverage j β Where α and β are set hyperparameters; Step 4-4: Compare the evaluation scores score of all perspectives in scenario i, find the perspective with the highest score, record its perspective subscript as best, and the corresponding set of SAM positive point input points under this perspective is P best 。 7. A method for object segmentation based on SAM and neural radiance fields according to claim 6, characterized in that, Step 5 includes the following steps: Step 5-1: Propagate the point set P obtained in Step 4-4 best to other perspectives using the point tracking algorithm, and denote the point set obtained in this way in perspective j as P` j ; Step 5-2. For each perspective j, use the point set P` j as the positive point input of SAM under this perspective, and combine it with the scene image target_img j to form the SAM input data prompt` j : (P` j , target_img j ); Step 5-3: Input the prompt j from each perspective into SAM respectively to obtain the object segmentation mask fine_mask j from each perspective.
8. A method for object segmentation based on SAM and neural radiance fields according to claim 7, characterized in that, Step 6 includes the following steps: Step 6-1: For each perspective j, add the object segmentation mask fine_mask obtained in step 5-3 j to the merged segmentation mask merged_mask j . Here, merged_mask j represents the cumulative value of fine_mask j in each round under perspective j, and its initial value is 0; Step 6-2: Determine whether the evaluation score score for this round is less than the preset threshold. If it is less than the threshold, stop the iteration, and use the merged segmentation mask merged_mask as the final segmentation mask for this view. If it is still greater than the threshold, return to Step 2 to continue looping; Step 6-3: If continuing the loop, modify Step 2-3 to: For each pixel point marked by ref_mask in ref_img i find its corresponding feature point in ref_featue i , and through sim i find the point set P j with the highest cosine similarity of eigenvalues to these feature points in each perspective, and these points with the highest cosine similarity cannot be in merged_mask j . If they are in merged_mask j , continue to find the point with the second-largest similarity, and so on until a point not in merged_mask j is found. j 9. A method for object segmentation based on SAM and neural radiance fields according to claim 8, characterized in that, Step 7 includes the following steps: Step 7-1: According to the scene serial number i, obtain the multi-view segmentation mask M of the corresponding scene from Step 6-2 i :(merged_mask1,merged_mask2,...,merged_mask n ), obtain the multi-view camera parameters C of the corresponding scene from Step 1-3-2 i , and combine them into the input data NerfIn of the neural radiance field i :(M i ,C i ), where n is the number of viewpoints in this scene; Step 7-2: Input data NrefIn through the neural radiance field i Train the neural radiance field; Step 7-3: Use the trained neural radiance field to render the object segmentation masks in the new view.
10. A method for object segmentation based on SAM and neural radiance fields according to claim 9, characterized in that, The scenes described in Step 8 include medical image scenes and vehicle detection scenes.