Three-dimensional object segmentation method and system based on neural radiation field

Through multi-view training and two-dimensional segmentation mask fine-tuning methods, the high complexity and noise problems in NeRF three-dimensional segmentation are solved, and efficient and accurate three-dimensional segmentation is achieved, which is suitable for robot perception, augmented reality, autonomous driving, medical imaging and precision agriculture and other fields.

CN120388026APending Publication Date: 2025-07-29THE INST OF BIOTECHNOLOGY OF THE CHINESE ACAD OF AGRI SCI
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510424627.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing three-dimensional segmentation method based on NeRF lacks color and shadow clues for geometric learning, and directly trains the binary mask, or needs to modify the model architecture to add segmentation branching and loss functions, resulting in high complexity of the model and difficult training, and the output results of the post-processing method often contain noise.

Method used

The basic neural radiation field model is trained through multi-view original RGB images, and fine-tuned using a two-dimensional segmentation mask to generate a three-dimensional segmentation neural radiation field model, and the segmentation results are obtained through three-dimensional point cloud clustering. The two-step fine-tuning strategy is adopted without modifying the NeRF model architecture, and the same neural network architecture and loss function are used.

Benefits of technology

High-precision three-dimensional segmentation is realized, the computing resource requirements are reduced, the segmentation accuracy and accuracy are improved, and the generated segmentation result interchange ratio is higher than that of existing methods. It is suitable for multiple fields, especially in high-precision object counting and positioning tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388026A_ABST
    Figure CN120388026A_ABST
Patent Text Reader

Abstract

The invention relates to a three-dimensional object segmentation method and system based on a neural radiation field, and the method comprises the steps: training a basic neural radiation field model through a multi-view original RGB image, and obtaining a pre-trained neural radiation field model; fine tuning is performed on the pre-trained neural radiation field model according to a two-dimensional segmentation mask corresponding to the original RGB image, a three-dimensional segmentation neural radiation field model is obtained, and the two-dimensional segmentation mask comprises a first pixel value representing a foreground object and a second pixel value representing a background; segmenting an image to be segmented through the three-dimensional segmentation neural radiation field model to generate segmented three-dimensional point clouds; and clustering the three-dimensional point cloud to obtain a three-dimensional segmentation result of the to-be-segmented image. According to the method, high-precision three-dimensional segmentation is realized through a two-stage fine tuning strategy, a neural radiation field model architecture does not need to be changed, a three-dimensional segmentation process is simplified, and segmentation precision and counting accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image segmentation, and in particular, to a three-dimensional object segmentation method and system based on neural radiance fields. Background Art

[0002] Neural Radiance Fields (NeRF) have successfully achieved a breakthrough in accurately reconstructing 3D scenes from 2D images by learning volumetric scene representations. However, compared with NeRF-based view synthesis models, the exploration of object-level segmentation in the reconstructed 3D space is still relatively limited. Segmenting objects in 3D point clouds is crucial in fields such as robot perception, augmented reality (AR), autonomous driving, medical imaging, and agriculture.

[0003] Existing NeRF-based 3D segmentation methods, on the one hand, usually fail to directly train NeRF on binary masks because binary masks lack the color and shading cues required for geometric learning. On the other hand, existing NeRF segmentation methods either require significant modifications to the model architecture, introducing additional segmentation branches and loss functions, increasing model complexity and training difficulty, or only rely on post-processing methods to back-project 2D masks into 3D space, but due to the lack of density field optimization, the output results often contain a large amount of noise.

[0004] Therefore, there is an urgent need in the art for a 3D segmentation method that can fully utilize the advantages of NeRF in 3D reconstruction and is both simple and efficient. Summary of the Invention

[0005] The present invention provides a three-dimensional object segmentation method and system based on neural radiance fields to solve the defects of the prior art.

[0006] The present invention provides a three-dimensional object segmentation method based on neural radiance fields, including:

[0007] S1: Training a basic neural radiance field model with multi-view original RGB images to obtain a pre-trained neural radiance field model;

[0008] S2: Fine-tuning the pre-trained neural radiance field model according to the two-dimensional segmentation mask corresponding to the original RGB image to obtain a three-dimensional segmentation neural radiance field model, where the two-dimensional segmentation mask includes a first pixel value representing the foreground object and a second pixel value representing the background;

[0009] S3: Segmenting the image to be segmented through the three-dimensional segmentation neural radiance field model to generate a segmented three-dimensional point cloud;

[0010] S4: Perform clustering on the three-dimensional point cloud to obtain a three-dimensional segmentation result of the image to be segmented.

[0011] According to a three-dimensional object segmentation method based on a neural radiance field provided by the present invention, step S1 further includes:

[0012] S11: Perform volume rendering on the original RGB image through a basic neural radiance field model to obtain a rendered RGB image;

[0013] S12: Update the parameters of the basic neural radiance field model by calculating a first mean squared error loss function between the rendered RGB image and the original RGB image to obtain a pre-trained neural radiance field model.

[0014] According to a three-dimensional object segmentation method based on a neural radiance field provided by the present invention, step S11 further includes:

[0015] S111: Construct rays from multiple pixel points to the camera origin respectively according to multi-view original RGB images and corresponding camera parameters;

[0016] S112: Sample multiple three-dimensional space points along the rays;

[0017] S113: Map the three-dimensional space points and corresponding view directions to color values and volume density values through a mapping function of the basic neural radiance field model;

[0018] S114: Perform volume rendering of the original RGB image based on the color values and the volume density values to obtain a rendered RGB image.

[0019] According to a three-dimensional object segmentation method based on a neural radiance field provided by the present invention, the expression of the first mean squared error loss function in step S12 is:

[0020]

[0021] where l RGB is the mean squared error loss between the rendered RGB image and the original RGB image, r is the ray index value emitted by the camera, R is the set of rays emitted by the camera, is the RGB prediction value based on the rendered RGB image, and C gt (r) is the RGB ground truth value of the original RGB image.

[0022] According to a three-dimensional object segmentation method based on a neural radiance field provided by the present invention, step S2 further includes:

[0023] S21: Format the original RGB image into a two-dimensional segmentation mask, where the first pixel value is represented as [1, 1, 1] and the second pixel value is represented as [0, 0, 0].

[0024] S22: Use the same neural network architecture and loss function as the pre-trained neural radiance field model. Replace the original RGB image with the two-dimensional segmentation mask as the supervision signal. Update the parameters of the pre-trained neural radiance field model by minimizing the second mean squared error loss function between the rendered predicted mask and the true two-dimensional segmentation mask to obtain a three-dimensional segmentation neural radiance field model.

[0025] For a three-dimensional object segmentation method based on neural radiance field provided by the present invention, the expression of the second mean squared error loss function in step S22 is:[[]]

[0026]

[0027] where l mask is to minimize the mean squared error loss between the rendered predicted mask and the true two-dimensional segmentation mask, r is the ray index value emitted by the camera, is the RGB prediction value of the RGB image rendered based on the two-dimensional segmentation mask, and M gt (r) is the binary mask target.

[0028] For a three-dimensional object segmentation method based on neural radiance field provided by the present invention, step S3 further includes:[[]]

[0029] S31: Input the image to be segmented into the three-dimensional segmentation neural radiance field model, and extract the volume density field from the three-dimensional segmentation neural radiance field model.

[0030] S32: Sample sampling points in three-dimensional space and select high-density sampling points according to the volume density values.

[0031] S33: Use the high-density sampling points as foreground object points, and form a segmented three-dimensional point cloud by aggregation.

[0032] For a three-dimensional object segmentation method based on neural radiance field provided by the present invention, step S4 further includes:[[]]

[0033] S41: Remove outliers from the three-dimensional point cloud through an outlier removal algorithm to obtain a preprocessed three-dimensional point cloud.

[0034] S42: Downsample the preprocessed three-dimensional point cloud using voxel downsampling to obtain a downsampled three-dimensional point cloud.

[0035] S43: Apply the DBSCAN density clustering algorithm to the downsampled three-dimensional point cloud for clustering to obtain an initial clustering result.

[0036] S44: Merge small clusters and split large clusters in the initial clustering result to obtain the three-dimensional segmentation result of the target object.

[0037] According to a three-dimensional object segmentation method based on neural radiance fields provided by the present invention, step S44 further includes:

[0038] S441: Based on the clustering volume and the number of clustering points in the initial clustering result, merge the small clusters in the initial clustering result to obtain a merged clustering set;

[0039] S442: Calculate the clustering variance and the number of clustering points in the merged clustering set, and determine the number of sub-clusters of the large cluster;

[0040] S443: According to the number of sub-clusters, use the hierarchical clustering algorithm to split the large clusters in the merged clustering set to obtain the three-dimensional segmentation result of the target object.

[0041] The present invention also provides a three-dimensional object segmentation system based on neural radiance fields, including:

[0042] The first training module: used to train the basic neural radiance field model through multi-view original RGB images to obtain a pre-trained neural radiance field model;

[0043] The second training module: used to fine-tune the pre-trained neural radiance field model according to the two-dimensional segmentation mask corresponding to the original RGB image to obtain a three-dimensional segmentation neural radiance field model;

[0044] The segmentation module, which is configured with the three-dimensional segmentation neural radiance field model obtained by training the second training module, and is used for segmenting the image to be segmented to generate a segmented three-dimensional point cloud;

[0045] The clustering module: used to cluster the three-dimensional point cloud to obtain the three-dimensional segmentation result of the image to be segmented.

[0046] A 3D object segmentation method and system based on neural radiance fields provided by the present invention adopts a two-step fine-tuning strategy, without modifying the network architecture or loss function of the original neural radiance field (NeRF), maintaining the original structural integrity of the model. This zero-modification design enables the method to be seamlessly integrated into the existing NeRF framework, reducing the usage threshold for researchers and developers, and improving the universality and compatibility of the technology. Compared with complex models (such as FruitNeRF) that require joint training of RGB and segmentation tasks, the present invention only performs fine-tuning on the basis of the existing NeRF model, avoiding the computational overhead of additional segmentation heads, reducing the complexity of adjusting hyperparameters, significantly reducing the computational resource requirements, accelerating the training process, and enabling high-quality 3D segmentation to be achieved even under limited computational resources. In addition, the intersection over union (IoU) of the segmentation results generated by the method of the present invention is significantly higher than that of existing methods, the segmentation boundary is more accurate, and the point cloud noise is less. The high-quality segmentation directly improves the accuracy of downstream tasks (such as object counting and localization), making the system more reliable in practical applications. Secondly, the training process of the method of the present invention is more stable, especially when dealing with small target objects, avoiding the problem of mutual interference between RGB and segmentation tasks in joint training, making the model easier to converge to the ideal result. At the same time, this method has good generalization ability for different datasets and is applicable to various scenarios from synthetic fruit data to real crops. Thirdly, the present invention realizes semantic reshaping of the density field through the fine-tuning process. The density of the target object area is enhanced, while the density of the background area is suppressed, forming a clearer semantic field representation. The optimization of the density field not only improves the segmentation accuracy but also provides richer semantic information for 3D scene understanding.

[0047] Although the present invention is mainly designed for binary segmentation, the method of the present invention can be directly extended to the ability of multi-class segmentation tasks without changing the model architecture or loss function. Only by specifying different label values for different classes, the flexibility greatly expands the application scope of the method, making it applicable to more complex scene analysis requirements. At the same time, the 3D segmentation method of the present invention can be applied to multiple fields such as robot perception, augmented reality, autonomous driving, medical imaging, and precision agriculture, and is particularly suitable for application scenarios that require reconstructing and analyzing 3D objects from 2D images, especially showing significant advantages in tasks that require high-precision object counting, size measurement, or position localization. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0049] Figure 1 Schematic flowchart of a 3D object segmentation method based on neural radiance fields provided by an embodiment of the present invention;

[0050] Figure 2 Schematic structural diagram of a 3D object segmentation system based on neural radiance fields provided by an embodiment of the present invention;

[0051] Figure 3 Schematic comparison diagram of PSNR of RGB images rendered by FruitNeRF and InvNeRF-Seg on a target object provided by an embodiment of the present invention;

[0052] Figure 4 Schematic comparison diagram of IoU of mask prediction by FnvNeRF-Seg and FruitNeRF on a target object provided by an embodiment of the present invention;

[0053] Figure 5 Schematic comparison diagram of visualizations of apple segmentation point clouds of FruitNeRF, InvNeRF-Seg and SA3D provided by an embodiment of the present invention;

[0054] Figure 6 Schematic diagram of the double-pod segmentation point cloud of soybean plants generated using InvNeRF-Seg and SA3D provided by an embodiment of the present invention.

[0055] 100. First training module; 200. Second training module; 300. Segmentation module; 400. Clustering module. Detailed implementation manners

[0056] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention, and they should not be construed as limiting the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention. In the description of the present invention, it should be understood that the terms used are only for the purpose of description and cannot be construed as indicating or implying relative importance.

[0057] To better understand the present invention, the research background of the present invention will be first explained in detail below.

[0058] The current FruitNeRF based on NeRF integrates joint training through architectural modifications, enabling a single model to simultaneously predict RGB colors and segmentation masks, thereby completing the 3D reconstruction of RGB and segmented point clouds. In the specific implementation, a segmentation head and an auxiliary loss function are added to guide the learning of the segmentation mask. Taking the segmentation of fruit types as an example, although this method has successfully segmented different fruit types on synthetic and real datasets, it still faces several challenges.

[0059] In FruitNeRF, the geometric information of fruits is implicitly learned through smooth color gradients, while the sharp binary segmentation masks provide limited supervision for density field learning. The joint training process forces the network to balance these two conflicting signals, which may lead to poor performance in both tasks. When the target object is too small and the loss value of the segmentation prediction is much smaller than that of RGB reconstruction, the training process of generating clear segmentation masks will be significantly slowed down. The current version of FruitNeRF still relies on the old version of nerfstudio, making it difficult to integrate outdated components for the smooth training of the NeRF segmentation model. Therefore, it is difficult for this method to be easily adapted to its specific research field, and the potential of 3D segmentation will be severely limited.

[0060] However, when the present invention fine-tunes such arbitrarily modified RGB images on a pre-trained standard RGB NeRF model, an effective NeRF model is obtained.

[0061] To reduce the complexity of arbitrarily modifying RGB images, the present invention considers using a two-stage method to directly fine-tune on binary masks. Since the NeRF trained on RGB images has encoded accurate geometric information, the present invention assumes that fine-tuning the 2D mask without changing the model or loss function can make the density field focus on the foreground object. Finally, this simple fine-tuning strategy is proven to be extremely effective, achieving high-quality segmentation in both the rendered 2D masks and 3D point clouds.

[0062] The embodiments of the present invention will be described below with reference to the drawings.

[0063] As Figure 1 shown, the present invention provides a three-dimensional object segmentation method based on neural radiance fields, including:

[0064] S1: Training a basic neural radiance field model with multi-view original RGB images to obtain a pre-trained neural radiance field model.

[0065] Furthermore, the Neural Radiance Field (NeRF) is an implicit neural representation method for reconstructing 3D scenes from multi-view 2D images. In step S1, the multi-view RGB images refer to a set of images of the same scene taken from different positions and angles. The basic neural radiance field model is a multi-layer perceptron (MLP) network that can map 3D space coordinates and viewing directions to colors and volume densities. By training the basic neural radiance field model, a pre-trained neural radiance field model capable of representing the scene geometry and appearance is obtained in step S1.

[0066] Among them, step S1 further includes:

[0067] S11: Perform volume rendering on the original RGB image through the basic neural radiance field model to obtain a rendered RGB image.

[0068] In NeRF, volume rendering is the process of converting the 3D volume density and color values predicted by the neural network into a 2D image. In step S11, the basic neural radiance field model receives the input camera parameters and generates an RGB image that matches the original view through the volume rendering process, that is, the rendered RGB image.

[0069] Among them, step S11 further includes:

[0070] S111: Construct rays from multiple pixel points to the camera origin according to the multi-view original RGB images and the corresponding camera parameters.

[0071] Furthermore, a ray is a vector pointing from the camera origin (camera center) to a specific 3D point in the scene. For each pixel point in the image, a corresponding ray can be constructed according to the internal parameters (focal length, principal point coordinates) and external parameters (position, direction) of the camera.

[0072] S112: Sample multiple 3D space points along the ray.

[0073] Furthermore, ray sampling refers to the process of selecting multiple 3D space points on the ray emitted from the camera. The sampling points are distributed between the near plane and the far plane of the scene. In the present invention, a uniform sampling strategy is specifically adopted.

[0074] S113: Map the 3D space points and the corresponding viewing directions to color values and volume density values through the mapping function of the basic neural radiance field model.

[0075] Furthermore, the mapping function of the neural radiance field model is a multi-layer neural network that receives 3D space coordinates and viewing directions as inputs and outputs RGB color values and volume density values. The volume density value represents the degree of the existence of matter at the space point, while the color value represents the color of the point seen from that view.

[0076] Specifically, NeRF learns a continuous function f θ , which maps 3D spatial coordinates X ∈ ℝ 3 and viewing direction d ∈ ℝ 3 to emitted color c ∈ ℝ 3 and volume density σ ∈ ℝ ≥0 , and its training is based on multi-view 2D RGB images with known camera poses, and the expression is:

[0077] f θ (X, d) → (c, σ);

[0078] where θ represents the learnable parameters of the network.

[0079] S114: Perform volume rendering of the original RGB image based on the color value and the volume density value to obtain a rendered RGB image.

[0080] After obtaining the color value and the volume density value in step S113, in step S114, the obtained values are rendered to RGB values and evaluated with the input RGB values to guide the convergence of the NeRF model in the subsequent steps. Specifically, the expression of the rendered RGB image is:

[0081]

[0082] where is the RGB prediction value based on the rendered RGB image, a is the near plane distance, i.e., the starting sampling point of the ray in step S11, b is the far plane distance, t is the parametric distance along the ray, i.e., the ending sampling point of the ray, T(t) is the cumulative transmittance, σ(t) is the volume density at position t, c(t) is the color at position t, and σ(s) is the volume density at position s.

[0083] S12: Update the parameters of the basic neural radiance field model by calculating the first mean squared error loss function between the rendered RGB image and the original RGB image to obtain a pre-trained neural radiance field model.

[0084] Furthermore, the mean squared error (MSE) loss function is a common method to measure the difference between the predicted value and the true value. The calculation formula is the average of the sum of the squares of the differences between the rendered RGB values and the true RGB values of all pixel points, as shown below. By minimizing this loss function, the parameters of the basic neural radiance field model are continuously updated, making the rendered image closer and closer to the original image, and finally forming a pre-trained model that can accurately represent the 3D scene.

[0085] where the expression of the first mean squared error loss function in step S12 is:

[0086]

[0087] where l RGB is the mean squared error loss between the rendered RGB image and the original RGB image, r is the ray index value emitted by the camera, and R is the set of rays emitted by the camera. is the RGB prediction value based on the rendered RGB image, and C gt (r) is the RGB ground truth value of the original RGB image.

[0088] S₂: Fine-tune the pre-trained neural radiance field model according to the two-dimensional segmentation mask corresponding to the original RGB image to obtain a three-dimensional segmentation neural radiance field model, where the two-dimensional segmentation mask includes a first pixel value representing the foreground object and a second pixel value representing the background.

[0089] In step S₂ of the present invention, the pre-trained neural radiance field model is fine-tuned using a two-dimensional segmentation mask to obtain a neural radiance field model capable of performing three-dimensional segmentation tasks. The two-dimensional segmentation mask is a binary image that separates the target object (foreground) from the background in the image, where the foreground object is represented by the first pixel value and the background is represented by the second pixel value. The fine-tuning process keeps the architecture and loss function form of the original neural radiance field model unchanged, and only realizes the transformation of the model function by replacing the supervision signal (from RGB image to binary mask).

[0090] In the three-dimensional segmentation neural radiance field model, the density field is reshaped into a distribution consistent with the boundary of the target object, the density in the foreground object region is enhanced, and the density in the background region is suppressed. Therefore, based on the above principle, the present invention enables the model to directly separate the target object from the three-dimensional space through this semantic reconstruction of the density field, without the need for an additional segmentation network or post-processing steps.

[0091] where step S₂ further includes:

[0092] S₂₁: Format the original RGB image into a two-dimensional segmentation mask, where the first pixel value is represented as [1, 1, 1] and the second pixel value is represented as [0, 0, 0].

[0093] Specifically, in step S₂₁, the two-dimensional segmentation mask is a special image form used to distinguish the foreground target and the background region in the image. In the present invention, the two-dimensional segmentation mask adopts a binary form, that is, the pixel value of the foreground object (the target to be segmented) is set to [1, 1, 1], and the pixel value of the background region is set to [0, 0, 0]. The formatting process is the process of converting the semantic segmentation result into a normalized RGB format mask.

[0094] The semantic information in the original RGB image is usually extracted by manual annotation or automatic segmentation algorithms. For example, using advanced 2D segmentation models such as the Segment Anything Model (SAM) can generate accurate masks for objects. For segmentation results with multiple channels, they need to be converted into a three-channel form with values of 0 or 1 to be compatible with the RGB output format of the NeRF model.

[0095] The specific formatting process is as follows: First, for each original RGB image, generate or obtain the corresponding semantic segmentation result; then, set the foreground regions (such as apples) in the segmentation result to pixel values [1, 1, 1], and the background regions (such as tree trunks, leaves, sky) to [0, 0, 0]. Finally, use the formatted 2D segmentation mask as the supervision signal in the fine-tuning stage.

[0096] S22: Using the same neural network architecture and loss function as the pre-trained neural radiance field model, replace the original RGB image with the 2D segmentation mask as the supervision signal, and update the parameters of the pre-trained neural radiance field model by minimizing the second mean squared error loss function between the rendered predicted mask and the true 2D segmentation mask to obtain a 3D segmentation neural radiance field model.

[0097] Furthermore, the key innovation in the fine-tuning stage is to keep the architecture and loss function of the original neural radiance field model unchanged, only replacing the input supervision signal. The second mean squared error loss function calculates the pixel difference between the rendered predicted mask and the true 2D segmentation mask, and the specific formula is as follows.

[0098] The fine-tuning process uses the geometric information learned by the pre-trained model as a starting point, focusing on reshaping the distribution of the density field to be consistent with the foreground object boundaries. The volume rendering process still follows the volume rendering equation, but the output is interpreted as a binary mask instead of the true RGB color.

[0099] Specifically, in the rendering process, if the high-density regions passed through by each ray correspond to foreground objects (such as apples), their density is maintained or enhanced; if they correspond to background regions, their density is suppressed. The semantic reconstruction of the density field is achieved through backpropagation and gradient descent.

[0100] The fine-tuning process does not modify the neural network architecture, that is, it uses the same multi-layer perceptron (MLP) structure as in the pre-training stage of step S1, keeping the form of the input (spatial coordinates and viewing direction) and output (RGB color and volume density) unchanged. The only difference is that the training data changes from the original RGB image to the 2D segmentation mask, and the loss function calculates the difference between the rendered mask and the true mask, rather than the RGB color difference.

[0101] Among them, the expression of the second mean squared error loss function in step S22 is:

[0102]

[0103] where l mask is to minimize the mean square error loss between the predicted mask of the minimized rendering and the true two-dimensional segmentation mask, r is the ray index value emitted by the camera, is the RGB prediction value of the RGB image rendered based on the two-dimensional segmentation mask, M gt (r) is the binary mask target.

[0104] In step S2, the fine-tuned three-dimensional segmentation neural radiance field model not only inherits the geometric understanding ability of the pre-trained model but also incorporates the semantic information of the target object. This fine-tuning method without modifying the network architecture greatly simplifies the construction process of the three-dimensional segmentation model and at the same time ensures the high quality of the segmentation results. The fine-tuned density field presents obvious semantic distribution characteristics: the density value in the foreground target area is high and the boundary is clear, and the density value in the background area is close to zero, making the generation of three-dimensional point clouds and target extraction direct and efficient.

[0105] The finally fine-tuned three-dimensional segmentation neural radiance field model can accurately distinguish the apple and the background area, realizing an effective extension from the two-dimensional segmentation mask to three-dimensional segmentation. The key to the success of the fine-tuning strategy lies in using the geometric information learned by the pre-trained model, and only a small parameter adjustment is required to achieve the semantic reconstruction of the density field, greatly improving the efficiency and quality of three-dimensional segmentation.

[0106] S3: Segment the image to be segmented through the three-dimensional segmentation neural radiance field model to generate a segmented three-dimensional point cloud.

[0107] Step S3 is a process of applying the trained three-dimensional segmentation neural radiance field model to the actual segmentation task to generate segmented three-dimensional point cloud data. In step S3, the volume density field is the key information carrier, which records the degree of existence of objects in three-dimensional space and their distribution. By extracting the volume density field formed during the model training process and judging whether a spatial point belongs to the foreground target based on the high and low density values, a three-dimensional point cloud containing only the target object is finally generated. The above process converts the segmentation information of the two-dimensional image into the object segmentation in three-dimensional space, realizing the semantic extension from the plane to the three-dimensional space.

[0108] Among them, step S3 further includes:

[0109] S31: Input the image to be segmented into the three-dimensional segmentation neural radiance field model, and extract the volume density field from the three-dimensional segmentation neural radiance field model.

[0110] Specifically, the image to be segmented refers to the newly input image that needs to be segmented. The input image does not need to contain the annotated segmentation mask, but only needs to provide the RGB image and camera parameter information. When these images are input into the 3D segmentation neural radiance field model, the model uses the learned geometric and semantic information to calculate the density distribution in 3D space.

[0111] Specifically, the process of extracting the volume density field is as follows: First, determine the viewing angle according to the camera parameters (intrinsic matrix and extrinsic matrix) of the image to be segmented; then, emit rays from this viewing angle into the scene, and each ray corresponds to a pixel in the image; next, sample 3D space points along each ray; finally, input the 3D point coordinates and the corresponding viewing direction into the neural radiance field model to obtain the volume density value of each point.

[0112] S32: Sample to obtain sampling points in 3D space, and select high-density sampling points according to the volume density values.

[0113] Sampling in 3D space in step S32 refers to the process of selecting a series of discrete points in the scene space for evaluating the density distribution. The sampling strategy directly affects the accuracy and integrity of the segmented point cloud. High-density sampling points refer to the points in the sampling point set whose volume density values exceed a preset threshold, and these points are likely to be inside or on the surface of the foreground object.

[0114] The sampling process usually adopts two strategies: uniform grid sampling or adaptive sampling. Uniform grid sampling creates a 3D grid in the scene space at regular intervals, and the grid intersection points are used as sampling points; adaptive sampling adjusts the sampling density according to the complexity of the region, and increases the sampling points in the boundary or detail-rich regions.

[0115] The specific sampling and selection process includes: First, create a 3D sampling grid within the bounding box of the scene. For example, divide a 1m×1m×1m space region into a 100×100×100 grid to obtain one million uniformly distributed sampling points; then, input the 3D coordinates of each sampling point into the neural radiance field model to calculate its volume density value; next, set a density threshold (such as 0.5), and select the points with density values higher than the threshold as high-density sampling points; finally, these high-density sampling points are regarded as the points belonging to the foreground object.

[0116] S33: Use the high-density sampling points as foreground object points, and the set forms the segmented 3D point cloud.

[0117] In step S33, the foreground object points refer to the points determined to belong to the target object in three-dimensional space. The process of forming a segmented three-dimensional point cloud from the high-density sampling point set means organizing these individual points into a structured data set, where each point contains position coordinates and possibly additional attributes. The segmentation process converts the scattered high-density sampling points into a three-dimensional point cloud model for visualization and further analysis.

[0118] Specifically, first, the selected high-density sampling points are stored according to their spatial coordinates to form a point cloud data structure. Then, optionally, color information can be added to each point, which can be the corresponding color in the original RGB image or a unified identification color. Next, optional post-processing is performed on the point cloud, such as density uniformization, noise filtering, or hole filling. Finally, the processed point cloud data is saved in a standard format for subsequent use or visualization.

[0119] S4: Cluster the three-dimensional point cloud to obtain the three-dimensional segmentation result of the image to be segmented.

[0120] In the three-dimensional segmentation method based on neural radiance fields, the point cloud of the foreground object has been obtained in step S3. However, the point cloud contains multiple target objects and needs to be further divided into separate object instances. Therefore, in step S4, through a clustering analysis method, based on the spatial distribution characteristics of the point cloud, points that are close to each other are grouped together to identify independent three-dimensional objects.

[0121] Among them, step S4 further includes:

[0122] S41: Remove the outliers from the three-dimensional point cloud through an outlier removal algorithm to obtain a preprocessed three-dimensional point cloud.

[0123] Furthermore, the outlier removal algorithm is a basic operation in point cloud processing, which is used to remove those isolated points or small point clusters that do not belong to the main structure and are generated by noise or misjudgment. For outliers, they will interfere with subsequent clustering and lead to inaccurate clustering results.

[0124] The present invention uses the Statistical Outlier Removal (SOR) as the removal algorithm. By calculating the average distance from each point to its k nearest neighbor points and judging whether the point is an outlier based on the statistical characteristics of the overall distribution (such as the mean and standard deviation).

[0125] S42: Downsample the preprocessed three-dimensional point cloud using voxel downsampling to obtain a downsampled three-dimensional point cloud.

[0126] Furthermore, voxel downsampling is a point cloud downsampling method for spatial subdivision. Specifically, the three-dimensional space is divided into cubic grids of uniform size (referred to as voxels), and then all points within each voxel are replaced by the center point or average point of all points in that voxel, thereby reducing the number of points in the point cloud. While maintaining the overall shape of the point cloud, it significantly reduces the data volume and speeds up subsequent processing.

[0127] The core idea of voxel downsampling is spatial quantization, which discretizes the continuous three-dimensional space into voxel grids of fixed size. The choice of voxel size directly affects the degree of downsampling and the level of details retained: the smaller the voxel, the more details are retained, but the weaker the downsampling effect; the larger the voxel, the more obvious the simplification of the point cloud, but the greater the loss of details.

[0128] In the present invention, the specific process of voxel downsampling is as follows: First, determine the size of the voxel (such as 0.01 meters); then, divide the bounding box of the preprocessed three-dimensional point cloud into a series of voxels of equal size; next, for each voxel containing points, calculate the center point of all points within the voxel, that is, a simple average; finally, use these center points to form the downsampled three-dimensional point cloud.

[0129] S43: Apply the DBSCAN density clustering algorithm to cluster the downsampled three-dimensional point cloud to obtain an initial clustering result.

[0130] Furthermore, DBSCAN (Density-Based Spatial Clustering of Applications with Noise) is a density-based spatial clustering algorithm suitable for datasets with irregular shapes and possibly containing noise. Compared with traditional clustering algorithms such as k-means, DBSCAN does not require presetting the number of clusters. Instead, it automatically discovers the number and shape of clusters by defining the concept of "density", and finally can obtain an initial clustering result.

[0131] S44: Perform small cluster merging and large cluster splitting on the initial clustering result to obtain the three-dimensional segmentation result of the target object.

[0132] Furthermore, there may be two types of problems in the initial clustering result: one is over-segmentation (one object is divided into multiple clusters), and the other is under-segmentation (multiple objects are merged into one cluster). The small cluster merging and large cluster splitting processes aim to correct these problems so that the final clustering result corresponds to the actual objects one by one. Therefore, in step S44, by analyzing the geometric characteristics of the clusters (such as size, shape, number of points), the clustering result is adaptively adjusted to achieve accurate object-level segmentation.

[0133] Among them, step S44 further includes:

[0134] S441: Based on the cluster volumes and the number of cluster points in the initial clustering result, merge the small clusters in the initial clustering result to obtain a merged cluster set.

[0135] Furthermore, small cluster merging is a solution to the over-segmentation problem, aiming to merge those small clusters that are spatially close and feature-similar into a complete cluster. The criteria for determining whether to merge include the spatial distance between clusters, the size (volume or number of points) of the clusters, and the shape characteristics of the clusters.

[0136] In the present invention, the specific process of small cluster merging is as follows: First, calculate the volume and the number of points for each cluster in the initial clustering result; then, set a volume threshold (such as 0.0001 cubic meters) and a number-of-points threshold (such as 30 points), and mark the clusters with a volume less than the volume threshold or a number of points less than the number-of-points threshold as small clusters; next, for each small cluster, search for other clusters that are spatially adjacent to it, and calculate the closest point distance between them; if the distance is less than the specified threshold (such as 0.05 meters), then merge the small cluster into the closest cluster; finally, update the cluster information after merging to obtain a merged cluster set.

[0137] S442: Calculate the cluster variance and the number of cluster points in the merged cluster set, and determine the number of sub-clusters of the large clusters.

[0138] Cluster variance is a statistic that measures the degree of dispersion of the points within a cluster. It reflects the shape and internal structure of the cluster. A high variance indicates that the cluster contains multiple objects or has a complex shape. The said cluster variance is usually measured by calculating the average of the sum of the squares of the distances from the cluster points to the cluster center. The number of sub-clusters is the decision of how many smaller clusters a large cluster should be decomposed into, usually based on the size, shape, and expected object characteristics of the cluster.

[0139] In the present invention, the specific process of determining the number of sub-clusters is as follows: First, calculate the covariance matrix of the three-dimensional coordinates for each cluster in the merged cluster set; then, calculate the eigenvalues of the covariance matrix, and these eigenvalues represent the variances of the cluster in the principal directions; next, evaluate the magnitudes and ratios of these eigenvalues to judge the shape characteristics of the cluster; at the same time, consider the ratio of the number of points of the cluster to the number of points of a typical object; finally, based on these metrics and the pre-set empirical rules, determine the number of sub-clusters into which each cluster should be segmented.

[0140] S443: According to the number of sub-clusters, use the hierarchical clustering algorithm to segment the large clusters in the merged cluster set to obtain the three-dimensional segmentation result of the target object.

[0141] Further, the hierarchical clustering is a method of constructing a clustering tree from bottom to top or from top to bottom. The present invention adopts agglomerative clustering, starting from a single data point, gradually merging the most similar points or clusters until the specified number of clusters is reached.

[0142] In the present invention, the specific process of large cluster segmentation is as follows: First, for each large cluster to be segmented, its point set is extracted; then, the agglomerative hierarchical clustering algorithm is applied, and the number of clusters is set to the number of sub-clusters determined in the previous step; next, the hierarchical clustering algorithm gradually merges points based on the Euclidean distance or other similarity measures between points until the specified number of clusters is reached; finally, the segmented sub-clusters are added to the final result to replace the original large cluster.

[0143] As Figure 2 shown, the present invention also provides a three-dimensional object segmentation system based on a neural radiance field, including:

[0144] The first training module 100: used to train the basic neural radiance field model through multi-view original RGB images to obtain a pre-trained neural radiance field model;

[0145] The second training module 200: used to fine-tune the pre-trained neural radiance field model according to the two-dimensional segmentation mask corresponding to the original RGB image to obtain a three-dimensional segmentation neural radiance field model;

[0146] The segmentation module 300, the segmentation module 300 is configured with the three-dimensional segmentation neural radiance field model obtained by training the second training module, and is used for segmenting the image to be segmented to generate a segmented three-dimensional point cloud;

[0147] The clustering module 400: used to cluster the three-dimensional point cloud to obtain a three-dimensional segmentation result of the image to be segmented.

[0148] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0149] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute a three-dimensional object segmentation method based on neural radiance fields described in each embodiment or some parts of the embodiments.

[0150] The following describes a three-dimensional object segmentation method and system provided by the present invention in combination with specific embodiments.

[0151] The present invention uses the dataset of synthetic fruit tree apples and peaches in the publicly available FruitNeRF dataset as test samples. For the detailed information of data creation, reference can be made to FruitNeRF. Since the dataset contains segmentation masks, camera intrinsics, and poses, no additional processing is required.

[0152] For learning-based models, when the size of the target category is much smaller than the background, the model usually converges more difficultly. To further challenge the capabilities of the model, the present invention also self-collected a surround-shot video of soybean plant types and tried to segment individual pods from soybean plants for binary 3D segmentation and two pods for multi-class 3D segmentation.

[0153] Standard preprocessing includes using ffmpeg to extract frame images at a rate of 5 frames per second, and then using colmap for camera pose estimation. To generate the masks of the pods, the present invention uses the AnyLabeling annotation software installed with Segment Anything (mobile SAM) to automatically generate the corresponding pod masks.

[0154] The model trained based on a three-dimensional object segmentation method and system of the present invention is named InvNeRF-Seg. The model architecture of InvNeRF-Seg is consistent with the original Nerfacto of Nerfstudio. Hyperparameter tuning is performed to ensure that at least a high-quality RGB NeRF model can be exported, and the clustering method is also adjusted for the method used in FruitNeRF.

[0155] Application of NeRF in 3D segmentation: To achieve excellent segmentation performance at the lowest cost, the present invention adopts the Nerfacto model recommended by Nerfstudio and keeps all default configurations. To improve the learning efficiency, for the apple and peach datasets, only the far plane is adjusted from the default 1000 to 50.

[0156] For the soybean plant dataset, since the default configuration cannot reconstruct a high-quality 3D RGB point cloud, the present invention adjusts the following parameters to enhance the model capacity: hidden_dim, hidden_dim_color, hidden_dim_transient, appearance_embed_dim are both increased from 64 or 32 to 128 to support learning of more complex scene features; the far plane is reduced from 1000 to 2 to promote the effective convergence of the NeRF model. These adjustments for InvNeRF-Seg are also synchronously applied to FruitNeRF to ensure a fair comparison.

[0157] Comparison between the InvNeRF-Seg model of the present invention and SA3D: To provide a consistent comparison, the present invention projects the binary segmentation mask back into the 3D space by using the cumulative field density value. Different from the original SA3D that relies on volume rendering weights, the present invention extracts three-dimensional points by threshold processing the original field density. This method can achieve denser object reconstruction, can capture the internal volume signal of the target object more completely, and avoid generating only a surface "shell-like" point cloud, which is especially suitable for downstream tasks that require analyzing the internal structure of the object (such as counting fruits or soybean pods).

[0158] Clustering and counting: After the point cloud is generated, the fruits are first separated from the background to reduce computing power consumption, and then the same clustering process as FruitNeRF is adopted: outlier removal → voxel downsampling → DBSCAN clustering → merging small clusters.

[0159] In this stage, the present invention uses a new method to split large clusters: first, determine the number of sub-clusters according to the variance and number of points of each cluster point cloud; subsequently, split the cluster into 2 / 3 sub-clusters or keep it unchanged as needed; finally, use the Agglomerative Clustering in scikit-learn (1.6) to assign points to new clusters.

[0160] During the training process, the present invention evaluates the quality of the rendered 2D RGB images and segmentation masks using the peak signal-to-noise ratio (PSNR) and the intersection over union (IoU) respectively, and evaluates the quality of the generated 3D segmentation point cloud using the fruit counting accuracy.

[0161] The present invention uses SciPy (v1.15.2) to perform a two-tailed t-test to evaluate the statistical significance of the differences in PSNR and IoU scores between the comparison models.

[0162] The synthetic dataset of fruit trees evaluates that InvNeRF-Seg and FruitNeRF perform similarly in terms of the quality of the rendered RGB images on the apple and peach datasets, and there is no significant difference in the comparison of their PSNRs, as Figure 3 shown.Figure 3 In it, a represents the PSNR comparison between FruitNeRF and InvNeRF-Seg on the RGB rendered images of apples. Figure 3 In it, b represents the PSNR comparison between FruitNeRF and InvNeRF-Seg on the RGB rendered images of peaches, where ns indicates no statistical significance. The overall results are reasonable as they share the same architecture in the RGB 3D reconstruction task.

[0163] In contrast, the intersection over union (IoU) of the rendered masks of InvNeRF-Seg for apples and peaches is significantly higher than that of the FruitNeRF model (p-value < 0.001), highlighting its excellent accuracy in mask rendering, as Figure 4 shown, where *** indicates statistical significance, p < 0.001. The reason for the shown results is that FruitNeRF still relies on RGB image supervision, treats segmentation as an auxiliary task, and the RGB branch and the segmentation head share the same feature MLP network, which may limit the specialization ability of the segmentation head. Therefore, its segmentation may degenerate into a simple foreground and background separation, ultimately leading to a reduction in segmentation performance. Different from the joint learning method of FruitNeRF, InvNeRF-Seg directly fine-tunes the density field based on the real masks, enabling a more accurate alignment of the predicted density with the real object boundaries, thus obtaining a significantly improved segmentation quality.

[0164] For the final segmentation effect, as Figure 5 shown, Figure 5 In a, it is a schematic diagram of the apple segmentation point cloud generated by FruitNeRF; Figure 5 In b, it is a schematic diagram of the apple segmentation point cloud generated by InvNeRF-Seg; Figure 5 In c, it is a schematic diagram of the apple segmentation point cloud generated by SA3D; Figure 5 In d, it is a schematic diagram of the apple point cloud clustering result generated by FruitNeRF; Figure 5 In e, it is a schematic diagram of the apple point cloud clustering result generated by InvNeRF-Seg; Figure 5 In f, it is a schematic diagram of the apple point cloud clustering result generated by SA3D.

[0165] The counting accuracy of InvNeRF-Seg on the apple and peach datasets is significantly better than that of SA3D. In the apple counting task, InvNeRF-Seg performs equally with FruitNeRF, while in the peach counting, InvNeRF-Seg has a higher accuracy. The specific accuracy results are shown in Table 1.

[0166] Table 1 Apple and peach counting results of three models (FruitNeRF, InvNeRF-Seg, SA3D) based on the segmented point clouds

[0167] FruitNeRF (pred / gt) InvNeRF-Seg (pred / gt) SA3D (pred / gt) Apple 282 / 283 282 / 283 267 / 283 Peach 157 / 152 152 / 152 170 / 152

[0168] To gain a deep understanding of the internal behavior of InvNeRF-Seg, the present invention analyzed the changes in volume field density before and after fine-tuning. The analysis results showed that the density in the target object area continued to increase after fine-tuning, resulting in sharper and more locally concentrated density peaks along the sampled rays. In contrast, the non-target object and background areas exhibited density suppression or small fluctuations around zero, which minimized the cumulative weights and achieved effective background removal. The patterns in the analysis results confirmed that InvNeRF-Seg can achieve semantic reshaping of the radiation field without architectural modification.

[0169] Although InvNeRF-Seg is designed for binary segmentation, the present invention also tested its multi-class extension ability by assigning different label intensities to different classes in the input mask. The results are as Figure 6 shown. For the results of InvNeRF-Seg, Figure 6 a in Figure 6 is the original view, Figure 6 and c in Figure 6 is the local enlarged view; for the results of SA3D, Figure 6 b in Figure 6 is the original view, Figure 6 and d in Figure 6 is the local enlarged view. Although using the same mean squared error loss and architecture, the model still achieved feasible results for the three-dimensional segmentation of double pods, but the performance decreased slightly compared to binary segmentation, manifested as mis-segmented points across classes, as shown in Figure 6 a and Figure 6 c. The results of SA3D are also usable, but there is still obvious noise between the object and the background, as shown in Figure 6 b and Figure 6 d.

[0170] The experiments of the present invention on synthetic and real datasets show that InvNeRF-Seg using the two-step fine-tuning method (first training NeRF based on RGB images and then fine-tuning through the segmentation mask) is superior to SA3D and the joint training method FruitNeRF in terms of both segmentation accuracy and three-dimensional point cloud quality.

[0171] Specifically, compared with FruitNeRF, the segmentation accuracy (measured by IoU) of the fine-tuned model always remains at a higher level. The fine-tuning process can better maintain the quality of the three-dimensional point cloud, while the 3D segmentation point clouds generated by the joint training model and SA3D are noisier, resulting in a decrease in counting accuracy.

[0172] Compared with FruitNeRF, the training process of InvNeRF-Seg is more stable because joint training may be difficult to coordinate RGB and mask prediction. This problem is particularly obvious when the mask only occupies a very small area of the whole image (such as a single pod of a soybean plant). Sometimes FruitNeRF completely fails to predict the mask of a single pod. On the contrary, the fine-tuning method converges faster in the second-stage segmentation because it starts from a well-initialized geometric representation.

[0173] The color gradient of the RGB image determines the convergence of NeRF geometric learning, while the binary mask with hard boundaries (0-1 mutation) has little or even harmful impact on it. Joint training forces NeRF to learn geometry from limited information from scratch, while InvNeRF-Seg allows optimizing segmentation learning based on a defined 3D structure. In addition, InvNeRF-Seg avoids loss function conflicts: in FruitNeRF, the RGB loss (color mean square error) and the mask loss (binary cross-entropy) may compete with each other, making it difficult for the model to balance the two tasks and resulting in suboptimal performance. InvNeRF-Seg decouples these two learning stages, enabling NeRF to learn a strong shape representation before optimizing segmentation.

[0174] InvNeRF-Seg also improves the density field representation. The NeRF density field has been fully trained in the fine-tuning stage, so only minor adjustments are needed for segmentation fine-tuning. Compared with FruitNeRF and SA3D, its segmented 3D point cloud has sharper object boundaries and fewer artifacts. The field-level analysis of the present invention shows that InvNeRF-Seg not only improves the segmentation accuracy but also meaningfully reconstructs the NeRF field. The density in the masked object area is enhanced, while the background and non-target structures are continuously suppressed. This feature is particularly significant in sparse target datasets (for example, the model can effectively suppress background elements such as leaves and stems). Even in complex scenes, the segmentation quality remains high due to the reduced accumulation of background light weights.

[0175] In addition, joint training often requires fine hyperparameter tuning to balance the losses, while the two-step InvNeRF-Seg is more robust to hyperparameter selection and can adapt to new datasets without a large amount of tuning.

[0176] Although InvNeRF-Seg is mainly designed for binary segmentation, the present invention verifies its generalization ability in multi-class scenarios by assigning unique values to multiple classes and using the original MSE loss to fine-tune NeRF. The excellent results prove the inherent flexibility of this zero-modification framework for 3D segmentation. In contrast, extending FruitNeRF to multi-class segmentation requires modifying the network structure and loss function. At the same time, due to the lack of optimization in the post-processing process of SA3D, there is still noise in the target and background segmentation results.

[0177] The success of the two-step segmentation model is also of great significance for 3D perception. The present invention emphasizes that prior geometric learning is crucial for subsequent segmentation learning. Since many applications require accurate 3D object segmentation, the core idea of InvNeRF-Seg - enhancing object-level reconstruction by fine-tuning a mature NeRF model - provides a new direction for enhancing the scene understanding of interactive AI systems. This method may inspire the development of future learning-based 3D perception processes.

[0178] Generally, the present invention proposes a two-step zero-modification fine-tuning strategy based on NeRF to achieve efficient and accurate 3D segmentation. Compared with the state-of-the-art joint learning method FruitNeRF, InvNeRF-Seg not only trains more stably but also can generate higher-quality segmented 3D point clouds. Compared with SA3D based on post-processing, the 3D segmented point clouds generated by this method are purer, and the subsequent object-level clustering results are also more accurate. The internal analysis of the radiation field shows that the target area is strengthened while the background area is suppressed, highlighting the semantic optimization ability of the model. This method can also be naturally extended to multi-class segmentation.

[0179] It should be noted that the present invention only focuses on the segmentation process, assuming that accurate multi-view RGB images and corresponding 2D segmentation masks have been obtained. The present invention does not address the issue of mask generation itself, as it is beyond the scope of this work. However, with the increasing number of reliable 2D segmentation models, the simple and architecture-invariant strategy of this method builds a practical and effective bridge between 2D and 3D segmentation.

[0180] To address these limitations, the present invention provides a 3D object segmentation method and system based on neural radiance fields, proposing an invariant neural radiance field for segmentation (InvNeRF-Seg), which is a two-step zero-adjustment fine-tuning strategy for 3D segmentation based on NeRF. The present invention first trains a standard NeRF on RGB images and then uses 2D segmentation masks for fine-tuning - the entire process does not change the model architecture or loss function. This method directly generates higher-quality and clearer segmented point clouds from the optimized radiation field with minimal computational effort. Through the analysis of the density field distribution of the 3D scene, it is found that the model gradually enhances the semantic discrimination ability between the target object and the background through training: the density of the target area is enhanced while the background density is suppressed, ensuring clean and reliable segmentation results. The present invention verifies that InvNeRF-Seg outperforms SA3D and FruitNeRF in terms of performance on synthetic fruit datasets and real soybean datasets. This method effectively realizes the extension from 2D segmentation to high-quality 3D segmentation.

[0181] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A three-dimensional object segmentation method based on neural radiance fields, characterized in that, Including: S1: Training a basic neural radiance field model with multi-view original RGB images to obtain a pre-trained neural radiance field model; S2: Fine-tuning the pre-trained neural radiance field model according to the two-dimensional segmentation mask corresponding to the original RGB image to obtain a three-dimensional segmentation neural radiance field model, where the two-dimensional segmentation mask includes a first pixel value representing the foreground object and a second pixel value representing the background; S3: Segmenting the image to be segmented through the three-dimensional segmentation neural radiance field model to generate a segmented three-dimensional point cloud; S4: Clustering the three-dimensional point cloud to obtain a three-dimensional segmentation result of the image to be segmented.

2. The three-dimensional object segmentation method based on neural radiance fields according to claim 1, wherein Step S1 further includes: S11: Performing volume rendering on the original RGB image through the basic neural radiance field model to obtain a rendered RGB image; S12: Updating the parameters of the basic neural radiance field model by calculating the first mean square error loss function between the rendered RGB image and the original RGB image to obtain a pre-trained neural radiance field model.

3. A 3D object segmentation method based on neural radiance fields according to claim 2, wherein Step S11 further includes: S111: Respectively constructing rays from multiple pixel points to the camera origin according to the multi-view original RGB images and corresponding camera parameters; S112: Sampling multiple three-dimensional space points along the rays; S113: Mapping the three-dimensional space points and corresponding viewing directions to color values and volume density values through the mapping function of the basic neural radiance field model; S114: Performing volume rendering of the original RGB image based on the color values and the volume density values to obtain a rendered RGB image.

4. A 3D object segmentation method based on neural radiance fields according to claim 2, characterized in that The expression of the first mean square error loss function in step S12 is: Among them, l RGB is the mean squared error loss between the rendered RGB image and the original RGB image, r is the ray index value emitted by the camera, and R is the set of rays emitted by the camera. is the RGB prediction value based on the rendered RGB image, and C gt (r) is the RGB ground truth value of the original RGB image.

5. A three-dimensional object segmentation method based on neural radiance fields according to claim 1, characterized in that, Step S2 further includes: S21: Formatting the original RGB image into a two-dimensional segmentation mask, where the first pixel value is represented as [1, 1, 1] and the second pixel value is represented as [0, 0, 0]; S22: Using the same neural network architecture and loss function as the pre-trained neural radiance field model, using the two-dimensional segmentation mask to replace the original RGB image as the supervision signal, and updating the parameters of the pre-trained neural radiance field model by minimizing the second mean square error loss function between the rendered predicted mask and the true two-dimensional segmentation mask to obtain a three-dimensional segmentation neural radiance field model.

6. A three-dimensional object segmentation method based on a neural radiance field according to claim 5, characterized in that The expression of the second mean square error loss function in step S22 is: where, l mask is the mean squared error loss that minimizes the predicted mask rendered and the true two-dimensional segmentation mask, r is the ray index value emitted by the camera, is the RGB prediction value of the RGB image rendered based on the two-dimensional segmentation mask, M gt (r) is the binary mask target.

7. A three-dimensional object segmentation method based on a neural radiance field according to claim 1, characterized in that Step S3 further includes: S31: Inputting the image to be segmented into the three-dimensional segmentation neural radiance field model and extracting the volume density field from the three-dimensional segmentation neural radiance field model; S32: Sampling to obtain sampling points in three-dimensional space and selecting high-density sampling points according to the volume density values; S33: Using the high-density sampling points as foreground object points and aggregating them to form a segmented three-dimensional point cloud.

8. A three-dimensional object segmentation method based on a neural radiance field according to claim 1, characterized in that Step S4 further includes: S41: Removing outliers from the three-dimensional point cloud through an outlier removal algorithm to obtain a preprocessed three-dimensional point cloud; S42: Downsampling the preprocessed three-dimensional point cloud using voxel downsampling to obtain a downsampled three-dimensional point cloud; S43: Applying the DBSCAN density clustering algorithm to cluster the downsampled three-dimensional point cloud to obtain an initial clustering result; S44: Perform small cluster merging and large cluster splitting on the initial clustering result to obtain the three-dimensional segmentation result of the target object.

9. A three-dimensional object segmentation method based on neural radiance fields according to claim 8, characterized in that Step S44 further includes: S441: Based on the clustering volume and the number of clustering points in the initial clustering result, merge the small clusters in the initial clustering result to obtain a merged clustering set; S442: Calculate the clustering variance and the number of clustering points in the merged clustering set, and determine the number of sub-clusters of the large cluster; S443: According to the number of sub-clusters, use the hierarchical clustering algorithm to split the large cluster in the merged clustering set to obtain the three-dimensional segmentation result of the target object.

10. A three-dimensional object segmentation system based on neural radiance fields, characterized in that, Including: The first training module: used to train the basic neural radiance field model through multi-view original RGB images to obtain a pre-trained neural radiance field model; The second training module: used to fine-tune the pre-trained neural radiance field model according to the two-dimensional segmentation mask corresponding to the original RGB image to obtain a three-dimensional segmentation neural radiance field model; The segmentation module, the segmentation module is configured with the three-dimensional segmentation neural radiance field model obtained by training the second training module, and is used for segmenting the image to be segmented to generate a segmented three-dimensional point cloud; The clustering module: used to cluster the three-dimensional point cloud to obtain the three-dimensional segmentation result of the image to be segmented.

Citation Information

Cited By

  • Efficient decoupling and object removing method based on neural radiation field scene

    CN121937615A