Scene multi-object three-dimensional reconstruction representation method based on feature three planes
By extracting the mask area and design feature three-plane network model in RGB-D data frames, the accuracy and memory occupancy problems of existing SLAM methods during multi-object reconstruction are solved, and high-quality multi-object three-dimensional reconstruction and object separate reconstruction are realized.
Patent Information
- Application Number
- CN202510086780.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-06-13
AI Technical Summary
Existing NeRF-based SLAM methods have shortcomings in capturing geometric details and multi-object reconstruction, including low reconstruction accuracy and quality, high memory footprint and lack of support for individual reconstructions for each object.
By extracting the mask area in the RGB-D data frame, separating the background and object areas, and designing a featured three-plane network model, the object model is trained in parallel using vectorization operations, and multi-object reconstruction is performed in combination with implicit truncated symbol distance field (TSDF).
Achieve higher quality 3D reconstruction of multi-objects, reduces memory overhead, improves reconstruction speed and accuracy, and supports separate reconstruction of each object.
Smart Images

Figure CN120147508A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of deep learning and computer vision, and particularly relates to a method for three-dimensional reconstruction and representation of multi-objects in a scene based on feature tri-planes. Background Art
[0002] Dense Visual Simultaneous Localization and Mapping (SLAM) is a challenging task in the field of computer vision, which is defined as approximating the pose of a camera while constructing a 3D map of an unknown environment, and has wide applications in the fields of robotics, virtual reality, and autonomous driving. Traditional SLAM systems mainly focus on improving and enhancing the pose accuracy, because these methods cannot perform reasonable geometric estimation on unobserved regions. While based on the emerging neural rendering representation method Neural Radiance Field (NeRF), the way of using an MLP neural model to represent an implicit scene has achieved great breakthroughs in the fields of novel view synthesis, scene reconstruction, etc., providing reasonable but limited reconstruction accuracy for SLAM in reconstructing a global 3D map. With the emergence of NeRF, NeRF has gradually been used to infer the shape of geometric objects and reconstruct the 3D surface in SLAM, and based on this, NeRF-based SLAM methods have been developed.
[0003] Nowadays, there are many NeRF-based SLAM methods for scene reconstruction. iMAP represents the entire geometric scene with a single, huge multi-layer perceptron (MLP), and performs global updates through each new, potential partial scene. NICE-SLAM improves the existing problems of iMAP in reconstructing larger scenes by local storage on a voxel grid, and these methods have shown considerable effects in reconstruction accuracy.
[0004] However, these methods generally have the following deficiencies:
[0005] 1. There is a lack of ability to capture geometric details. Although the existing work has achieved dense reconstruction of objects, since NeRF is mainly used for view synthesis and learns the volume density field of objects during the training process, it is difficult to extract high-quality surfaces, which will result in over-smoothed scene reconstruction, so it is limited in the reconstruction accuracy and quality of scene reconstruction.
[0006] 2. Each MLP stores features on a voxel grid, which will lead to an increase in memory occupancy. Especially in the case of multi-object reconstruction, the memory occupancy will be very high.
[0007] 3. The existing work mainly focuses on the representation of a complete scene or a single object, and pays little attention to the separate reconstruction of each object in the scene, that is, while reconstructing the scene background, each object in the scene is reconstructed, and the reconstruction results of each object can be superimposed to form a complete scene representation. Summary of the Invention
[0008] In view of the deficiencies in the prior art, the present invention provides a method for three-dimensional reconstruction and representation of multi-objects in a scene based on a feature tri-plane.
[0009] The present invention inputs a set of consecutive RGB-D data frames, extracts the mask regions corresponding to each object in the scene, and divides the mask regions into a background region and an object region according to the item ID in the scene. By inputting consecutive RGB-D data frames and extracting the mask regions in the scene frame by frame, the corresponding neural network models can be initialized respectively according to the mask regions of the current frame items and the background. Since a similar design is adopted for all object models, the object models can be stacked and vectorized for training by using the vectorized operations provided in functorch, while the background model with a slightly larger network structure is trained separately in the training process. The model mainly includes a coordinate normalization module, a feature tri-plane module, and two MLP decoders. The color information and TSDF information of the sampling points are obtained through the model output. The result represented by TSDF can extract the 3D mesh by the marching cubes algorithm, where the distance to the nearest surface is represented by a positive sign in the free space, and the point is represented by a negative sign inside the surface.
[0010] The technical solutions adopted by the present invention to solve its technical problems include the following steps:
[0011] Step S1: Extract and associate the image masks of the items and the background in the RGB data set;
[0012] Step S2: Design a feature tri-plane network model;
[0013] Step S3: Initialize the object parallel network model and the background network model;
[0014] Step S4: Extract sampling points according to the mask region for in-region supervision;
[0015] Step S5: Obtain the output result of the network model according to the sampling point information;
[0016] Step S6: Design a loss function for the network model according to the output result of the network model;
[0017] Step S7: Train the object parallel network model and the background network model;
[0018] Step S8: Use the trained network model to perform combined scene grid extraction.
[0019] The beneficial effects of the present invention are as follows:
[0020] The present invention provides a method for three-dimensional reconstruction of multi-objects.
[0021] The present invention extracts the masks of various objects in the RGB-D image, and uses the mask information to extract and establish the network models of each object and the background, reconstructing multiple objects in the scene while reconstructing the scene background, and proposes a 3D reconstruction depth network learning framework for vectorized multi-object training.
[0022] The present invention proposes a multi-object reconstruction strategy based on an implicit truncated signed distance field, which is significantly faster than common rendering representations based on volume density and occupancy in terms of convergence speed, and can achieve a higher quality reconstruction effect.
[0023] The present invention uses the feature trilinear plane and the shallow decoder as the network backbone. For each sampling point in the continuous space, three feature vectors are obtained by projection and accumulated, and then converted into a truncated signed distance field and RGB values. Compared with the explicit representation method based on discrete voxel grids, it reduces a large amount of memory overhead. And compared with the fully implicit representation or the fully connected network that represents the scene as a continuous function, it improves the reconstruction speed without sacrificing the rendering quality. Brief Description of the Drawings
[0024] Figure 1 It is the flowchart of the multi-object three-dimensional reconstruction of the embodiment of the present invention;
[0025] Figure 2 It is the network structure diagram of the feature trilinear plane of the embodiment of the present invention;
[0026] Figure 3 It is the schematic diagram of the main step process of the embodiment of the present invention. Detailed Embodiment
[0027] To make the purpose, technical solution and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below in conjunction with the drawings of the present invention.
[0028] As Figure 1 shown, a method for representing the multi-object three-dimensional reconstruction of a scene based on a feature trilinear plane includes the following steps:
[0029] Step S1: Extract and associate the image masks of the items and the background in the RGB dataset. The specific details are as follows:
[0030] First, obtain the RGB dataset according to the actual application scenario. In this embodiment, the Replica dataset is used, and 8 synthetic scenes of the Replica dataset are used as the datasets for training and reconstruction quality evaluation. Replica is a synthetic dataset synthesized by the Blender software. Secondly, obtain the object mask according to the picture information provided in the RGB dataset and associate it. As a dataset synthesized by the Blender software, Replica is perfect in terms of depth and pose quality information, and the object mask is directly provided in the dataset. However, for other scene datasets, the object mask is usually not directly provided in the dataset. In this case, it is necessary to extract the object mask using the picture information. In previous practices, the extraction of the object mask was achieved using a pre-trained instance segmentation network. The user inputs a single picture to obtain a single segmentation map, and there is no association between the segmentation maps. However, in the actual application process, the same object in different frames will be in different positions, and the pre-trained instance segmentation network will assign different object IDs to it during the recognition process, which will cause it to be recognized as different objects in the subsequent training process, thus affecting the reconstruction quality of the object. Although the object ID can be changed manually, it takes a lot of time. Therefore, in order to ensure that the same object between the previous frame and the current active frame can be quickly associated as much as possible, we first synthesize each frame of the picture in the dataset into a video in chronological order, and then use the existing video object tracking network Track Anything to extract and continuously track each object in the video and generate a segmentation video. Secondly, the segmentation video is split according to the original number of images, and the associated object masks are obtained. Thus, each object can obtain a set of segmentation maps belonging to itself after the above processing. Similarly, the background in each frame of the picture is also processed in the same way to obtain a set of segmentation maps of the background. The segmentation map sets of the object and the background will help the initialization of the subsequent network model. Track Anything uses the method of Segment Anything (segment everything) interactive object segmentation, receives the user's scribble interaction as input, generates the corresponding mask, and uses the method of long-term video segmentation using a multi-store model (Xmem) to track the high-quality mask for a long time, so as to be able to generate the segmentation map of each frame of RGB image in the scene data.
[0031] Step S2: Design a feature triplane network model;
[0032] The feature triplane network model consists of a normalization module (input module), a feature triplane module (core module), and an output module. All subsequent references to the network model refer to the feature triplane network model.
[0033] The normalization module serves as the entry point of the feature tri-plane network model. It receives all input data (i.e., the sample points in S4 later) and normalizes their values to the range of [-1, 1], facilitating the fast training and inference of the network model.
[0034] Although the voxel grid structure based on explicit representation can complete computational operations quickly, it consumes a large amount of memory. While the network structure based on full implicit representation is efficient in terms of memory usage, it is slow in terms of computational speed and training speed. Based on these two aspects, the core part of the network model of the present invention (i.e., the feature tri-plane module) adopts a feature tri-plane structure, which combines local implicit representation and hybrid explicit-implicit representation, and simultaneously has the advantages of both. As Figure 2 shown, features are stored and optimized on the feature tri-planes. Since the surface texture information of an object in 3D reconstruction affects the reconstruction of the object's geometric information, different feature tri-planes are adopted in the design of the feature tri-planes, namely the appearance and geometric feature tri-planes, to represent the appearance and geometric features of the object respectively. The appearance and geometric feature tri-planes each include two scales, namely rough and fine. The rough feature tri-plane allows for efficient reconstruction of free space using fewer sample points and optimization iterations, and the fine feature tri-plane can reconstruct fine-grained appearance and geometric changes more meticulously. The geometric feature tri-plane includes rough geometric features and fine geometric feature tri-planes Similarly, the appearance feature tri-plane includes rough appearance feature tri-planes and fine appearance feature tri-planes The feature tri-planes obtain the feature values at the projected point positions by projecting the points in space onto three feature planes, bilinearly interpolating and summing the four nearest neighbors at the projected point positions on each feature plane, and then passing the feature values through the output module to obtain the final network output result. Both the object network model and the background network model adopt the feature tri-plane network model. For the core part of each object network model and background network model, the three feature planes of the two scales each use 32 channels. Therefore, the rough feature tri-plane and the fine feature tri-plane are concatenated in the channel dimension into a 64-channel feature, and this result will be fed into the output module.
[0035] The output module is implemented using two simple MLPs. The MLP is divided into a hidden layer and an output layer. The hidden layer has 32 channels, and a ReLu activation function is connected after the hidden layer. The output layer uses the Tanh function and the Sigmoid function to decode the truncated signed distance field value (TSDF value) and the rough appearance information respectively.
[0036] Step S3: Initialize the object parallel network model and the background network model. The specific details are as follows: Since splitting a large neural radiance field into several object networks and background networks can achieve more efficient training while obtaining more refined results, we need to initialize several object network models and background network models according to the set of self-segmentation maps of each object and the background segmentation map. The specific operations are as follows:
[0037] First, initialize the network models: In the actual application scenario of SLAM, images are accessed sequentially in chronological order. Therefore, the number of object network models to be actually initialized needs to be dynamically generated. When accessing the first image, first initialize the object network models according to the number of object masks in the image, and simultaneously establish the background network model; at the same time, each object network model and background network model maintain a key frame buffer. During the initialization process of the network models, the RGB images of the object mask and background mask regions are added to the key frame buffer as key frames to complete the initialization step of the key frame buffer. During the access of subsequent images, if an object network model has been established for the object mask in the accessed image, no further network model initialization is performed, and the RGB images of the mask part related to the object in the accessed image are added to their respective key frame buffers; for newly emerged object masks, new object network models are established and the initialization steps of the key frame buffer are performed. The above steps achieve the preliminary initialization and subsequent dynamic addition of the network models, that is, the initialization of the network models. Thus, the initialization of the background network model and the object network models is achieved.
[0038] Secondly, parallel processing is also required for the object network models to obtain the object parallel network model. The specific processing is as follows: Since all object network models adopt the same design, they can be parallelly trained under the parallelization functions provided by functorch. Before the formal start of training, all object network model decorders and optimizers also need to be stacked and sent into the functorch.combine_state_for_ensemble function (from the functorch library) to obtain the object parallel network model decoder_model and the parallel optimizer decoder_param. These two are the parallel network model and the parallel optimizer of all object models respectively, providing a basis for the parallel training of all object models. Since all object network models are not trained through loop statements, the resources of the GPU can be effectively utilized during the training process, thereby improving the training speed. Thus, the initialization of the object parallel network model is achieved.
[0039] Step S4: Extract sampling points according to the mask region for in-region supervision. The specific details are as follows:
[0040] The sampling method for different objects and the background is the same. Therefore, the sampling process for a single object will be described in this step.
[0041] After obtaining the mask of the object, extract the maximum boundary of the masked part and form a bounding box covering the object mask, i.e., the object object bounding box (x1, x2, y1, y2). Use the width and height of the bounding box to randomly select coordinates to obtain random coordinates. Perform ray sampling in the random coordinates so that object-level supervision is only applied to the pixels within the object object bounding box, and obtain a set of rays denoted as where N r represents the total number of rays in the ray set, which is taken as 1200 in the background network model of this embodiment and 120 in the object network model. For the pixels of the masked part within the bounding box, use the truncated signed distance field value (TSDF value), RGB value, and depth value generated by the sampling points to supervise these pixels. For other pixels within the object object bounding box, since they do not belong to the object, these pixels are set to empty during training. Since each object establishes its own corresponding object network model and key frame buffer through its own mask, and uses the key frame buffer of the object itself to extract image information for ray sampling, the training between different objects does not interfere with each other.
[0042] When processing the input of each frame, simulate the ray casting in NeRF, and use the current camera pose {R i |t i} to calculate the rays corresponding to the randomly selected pixels, and use them to render the truncated signed distance field, color information, and depth information of the rays. Since training only using RGB data will cause the training result to be more biased towards the texture of the object, and for the more important geometric shape part in 3D reconstruction, further optimization is required. Therefore, in order to further optimize the geometric information of the object, the sampling method of the present invention consists of the following three methods: hierarchical uniform sampling, depth information-guided sampling, and TSDF-guided sampling, sampling Nc, Ns, and Nt points along each ray respectively.
[0043] Hierarchical uniform sampling: Uniformly sample Nc points directly between the near boundary tn and the far boundary tf of the camera. Benefiting from the depth prior information, we use the maximum depth D surface of the object's own mask area to replace the far boundary tf. At this time, a total of Nc points are uniformly sampled. When there is incorrect depth data (noise information) at the sampling points, that is, the depth data is less than or equal to 0 or exceeds the set boundary, hierarchical uniform sampling will be used on the sampling ray instead of depth information-guided sampling and TSDF-guided sampling. At this time, a total of N all = Nc + Ns + Nt points are uniformly sampled. The sampling method is as follows:
[0044]
[0045] Among them, u represents a uniform distribution, and ti represents the i-th sampling point.
[0046] Depth information-guided sampling: The RGB-D dataset provides us with depth information generated based on the sensor, which supports the learning of a higher-quality volume density field. Under the guidance of the depth information, sampling is performed according to a normal distribution centered on the depth D surface A total of Ns points are sampled. In this embodiment, the variance is taken as d σ = 3 cm, which will be conducive to optimizing the object model towards more accurate geometric information. The specific sampling method is as follows:
[0047]
[0048] Among them, N represents a normal distribution.
[0049] If the current depth information has a large noise, the information of D surface is inaccurate. At this time, the far boundary tf is used to replace D surface , and the uniform sampling method is used to replace the depth-guided sampling.
[0050] TSDF-guided sampling: By measuring the distance from a point to the surface, TSDF can provide surface information more precisely. This guided sampling can effectively reduce the influence on the training of the object surface in the volume density field. Similar to the depth-guided sampling, TSDF sampling focuses on the vicinity of the object surface, thereby more effectively optimizing the geometric accuracy of the object. Specifically, as follows:
[0051]
[0052] Among them, T represents the truncation threshold in the truncated signed distance field.
[0053] Similarly, when the surface depth value is inaccurate, a uniform sampling strategy is introduced.
[0054] Finally, N all = Nc + Ns + Nt points are sampled respectively on each ray of all the ray sets, denoted as p n , where n ∈ [1, N all , and the set of all sampling points of the k-th object is denoted as
[0055] Step S5: Obtain the output result of the network model according to the sampling point information;
[0056] First, the sets of all sampling points of each object obtained in step S4 are superimposed to obtain the set of points of all objects after superposition, simply referred to as the superimposed point set, with a dimension of Send the obtained set of superimposed points into the object parallel network model.
[0057] In the first step, after passing through the normalization module, the set of superimposed points becomes a set of normalized points with a dimension of
[0058] In the second step, send into the core part of the network to obtain the rough geometry output and the fine geometry output Superimpose the two obtained features to get the final geometry output result, which is represented by the following formula:
[0059]
[0060] Similarly, the obtained appearance output result is:
[0061]
[0062] In the third step, after obtaining the geometry and appearance outputs, send and both through two shallow MLP decoders {M g , M a} (i.e., the output module) to obtain the output results and Among them, is the TSDF value, is the rough appearance information. The output of the MLP will be used for the extraction of depth values and color values. We refer to the SDF-based rendering method in StyleSDF to convert the TSDF value obtained for each ray into volume density. For each object, its conversion is specifically represented as follows:
[0063]
[0064] Among them, β represents a learnable parameter used to control the sharpness of the surface boundary. Subsequently, render the color and depth of each ray through the volume density, which is specifically represented as follows:
[0065]
[0066] Among them, z n represents the depth value of p n , where ω n represents the weight value, represents the rendered color value, represents the rendered depth value, p m represents the sampling point with subscript m under the current ray, p n Represent all the sampling points on this ray of light. The obtained depth and color will be used for the optimization of the network model and the extraction of the mesh after the network model training.
[0067] For obtaining the output result of the background network model, the steps are the same as those of the object parallel network model, that is, the obtained background sampling points are sent into the background network model to obtain the corresponding depth and color information.
[0068] Step S6: Design a loss function for the network model according to the output result of the network model. The specific details are as follows: For each object k, only sample the pixels within the 2D object bounding box where each object itself is located. The sampled ray of light is denoted as R k And only optimize the TSDF value, depth value, and color value of the pixels in the masked part within the 2D bounding box. The mask is denoted as M k , and the depth loss and color loss of each object are as follows:
[0069]
[0070] For the TSDF loss, it is divided into two parts: free space loss and loss for points close to the surface and within the truncation region. The free space loss of each object is expressed as:
[0071]
[0072] where is a set of points on the current object sampling ray r between the camera center and the surface truncation region measured by the depth sensor. This loss prompts the TSDF value to approach 1 in free space, so as to better optimize the structural features of the sampling points far from the object surface.
[0073] For the sample points close to the surface and within the truncation region, the loss is expressed as follows:
[0074]
[0075] where z(p) is the depth of point p on the plane relative to the camera, T is the truncation distance, D(r) is the true depth value measured by ray r, represents the set of points on ray r within the truncation distance. Considering that the points closer to the surface have stronger geometric constraints, the points close to the surface and within the truncation region are further divided. The part where |z(p)-D(r)|≤0.5T is defined as and the remaining points are defined as Therefore is divided into:
[0076]
[0077] Therefore, the TSDF loss of each object is expressed as follows:
[0078]
[0079] where λ fs and λ tr-n have greater weights than λ tr-f . In this embodiment: λ fs is set to 5, λ tr-n is set to 200, and λ tr-f is set to 10.
[0080] Above, the overall loss is defined as:
[0081]
[0082] where λ 1 , λ 2 , and λ 3 are hyperparameters for controlling the loss weights.
[0083] For the calculation of the background network model loss, the steps are the same as those of the object parallel network model, which can be regarded as a special case when there is only one object in the object parallel network model.
[0084] Step S7: Training of the object parallel network model and the background network model
[0085] The training of the network model will sequentially traverse each image in the dataset, and a round of training will be performed when each image is received.
[0086] First, in each round of the training process, each object and the background need to be iterated 20 times. Each round of iteration includes sampling of sample points, output of the network model results, calculation of the loss function, and backpropagation of the loss function and iterative optimization of the optimizer.
[0087] Second, in each round of training, first calculate the loss of the output result of the object parallel network model. Next, calculate the loss of the output result of the background network model. The result after adding the two losses will be used for backpropagation and optimization of the network model.
[0088] Finally, after each round of training, use the new value of the parallel optimizer decoder_param after this round of training to overwrite the old value of the optimizer optimisers of all object network models. Based on this, after each round of training of the object parallel network model is completed, the training results of each object network model in this round can be obtained synchronously.
[0089] Step S8: Use the trained network model to extract the combined scene grid. The specific details are as follows:
[0090] Since each network model has its own key-frame buffer, the depth information in the key-frame buffer of the network model can be combined with the camera intrinsics to convert the depth image of each frame into point cloud data, and the minimum oriented 3D bounding box of each object and the background can be determined. Next, random sample points are sampled for each 3D bounding box, and the depth and color information of the output are obtained by passing these sample points through their respective network models. Combining the obtained output results with the marching cube algorithm can obtain the grid of each object and the background in the scene. Each grid can be viewed and combined individually, and combining all the obtained grids gives the final result - the combined scene grid.
[0091] This embodiment also conducts a comparative experiment with other scene reconstruction methods in the prior art. The experimental results are measured and compared by sampling 200,000 points on the real grid and the network-reconstructed grid and calculating three metrics: Accuracy (cm), Completion (cm), and Completion Ration (<5cm%). Among them, Accuracy (cm) (accuracy) represents the average distance between the sampled points of the reconstructed grid and the nearest sampled points of the real grid; Completion (cm) (completion) represents the average distance between the sampled points of the real grid and the sampled points of the reconstructed grid;
[0092] Completion Ration (<5cm%) (reconstruction rate) represents the percentage of points with a completion less than 5 cm in the reconstructed grid. Accuracy reflects the accuracy of the reconstructed grid. The smaller the Accuracy, the closer the position of the point in the reconstructed grid is to the point in the real grid; Completion evaluates the integrity of the reconstructed grid. The smaller the Completion, the better the coverage of the point in the reconstructed grid and the fewer missing parts. Finally, the reconstruction rate is used to evaluate whether a large amount of information is missed during the entire reconstruction process and reflects the coverage effect of the reconstruction.
[0093] Therefore, these three metrics can objectively and comprehensively reflect the performance of the test model. The experimental results are shown in the following table (obtained by taking the average of 8 indoor scenes of Replica):
[0094]
[0095] According to the table data, this method has a more excellent performance in terms of completion and completion rate, and is second only to TSDF-Fusion in terms of accuracy, which can effectively improve the reconstruction result.
[0096] The above content is a further detailed description of the present invention in combination with specific / preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, they can also make several substitutions or modifications to these described embodiments, and these substitution or modification methods should all be regarded as belonging to the protection scope of the present invention.
[0097] The parts not detailed in the present invention belong to the well-known technologies in the art.
Claims
1. A method for 3D reconstruction of multiple objects in a scene based on characteristic three-planes, characterized in that: The steps include: Step S1: extract and associate the image masks of objects and backgrounds in the RGB dataset; Step S2: designing a characteristic three-plane network model; Step S3: Initialize the object parallel network model and the background network model; Step S4: extracting sampling points according to the mask area for intra-area supervision; Step S5: obtaining the network model output result according to the sampling point information; Step S6: designing a loss function for the network model according to the output result of the network model; Step S7: training the object parallel network model and the background network model; Step S8: Use the trained network model to perform combined scene grid extraction.
2. The method for 3D reconstruction of multiple objects in a scene based on characteristic three-planes according to claim 1, characterized in that: The specific operations of step S1 are as follows: first, obtain an RGB data set according to the actual application scenario; secondly, obtain the object mask according to the image information provided in the RGB data set and associate it; for data sets where the object mask is not directly provided, it is necessary to use the image information to extract the object mask; first, synthesize each frame of the data set into a video in chronological order, use the video target tracking network TrackAnything to extract and continuously track each object in the video and generate a segmented video, and then split the segmented video according to the number of original images to obtain the associated object mask, so that each object can obtain a segmentation map set belonging to the object after the above processing, and the background in each frame of the picture is also processed in the same way to obtain a background segmentation map set.
3. The method for 3D reconstruction of multiple objects in a scene based on characteristic three planes according to claim 1 or 2, characterized in that: The characteristic three-plane network model consists of a normalization module, a characteristic three-plane module and an output module; all subsequent network models are referred to as the characteristic three-plane network model; As the entry point of the feature three-plane network model, the normalization module receives all input data and normalizes its values to the range of [-1, 1], which facilitates the rapid training and reasoning of the network model. The characteristic three-plane module, i.e. the core part of the network model, adopts a characteristic three-plane structure, which combines local implicit expression and mixed explicit-implicit expression, and has the advantages of both. Different feature three planes are used in the design of feature three planes, namely, appearance and geometric feature three planes, which represent the appearance and geometric features of the object respectively; The three planes of appearance and geometric features include two scales, namely, coarse and fine; The geometric feature three planes include the rough geometric feature three planes and fine geometric features of the three planes Similarly, the appearance feature three-plane includes a rough appearance feature three-plane and fine appearance features three planes The feature three-plane is obtained by projecting points in space onto three feature planes, and bilinearly interpolating and summing the four nearest neighbors of the projection point position on each feature plane to obtain the feature value of the projection point position, and then using the feature value to pass through the output module to obtain the final network output result; both the object network model and the background network model use the feature three-plane network model, in which 32 channels are used for the three feature planes of the two scales in the core part of each object network model and the background network model. Therefore, the coarse feature three-plane and the fine feature three-plane are spliced into 64 channel features in the channel dimension, and this result will be sent to the output module; The output module is implemented using two MLPs, where the MLP is divided into a hidden layer and an output layer. The hidden layer has 32 channels and is followed by a ReLu activation function. The output layer uses a Tanh function and a Sigmoid function to decode the truncated signed distance field value and rough appearance information, respectively.
4. The method for 3D reconstruction of multiple objects in a scene based on characteristic three-planes according to claim 3, characterized in that: The specific method of step S3 is as follows: Initialize several object network models and background network models based on each object's own segmentation map set and background segmentation map set. The specific operations are as follows: First, initialize the network model: In the actual application scenario of SLAM, the images are accessed in chronological order. Therefore, the number of object network models to be initialized needs to be dynamically generated. When accessing the first image, the object network model is first initialized according to the number of object masks in the image, and the background network model is established synchronously. At the same time, each object network model and background network model maintains a key frame buffer. During the network model initialization process, the RGB images of the object mask and background mask areas are added to the key frame buffer as key frames to implement the initialization step of the key frame buffer. During the subsequent image access process, if the object mask in the accessed image has established an object network model, the network model will no longer be initialized, and the RGB image of the mask part of the object in the accessed image will be added to the respective key frame buffers. For newly appearing object masks, a new object network model is established and the key frame buffer is initialized. The above steps realize the initial initialization of the network model and the subsequent dynamic addition, that is, the initialization of the network model. This realizes the initialization of the background network model and the object network model. Secondly, the object network model needs to be parallelized to obtain the object parallel network model. The specific processing is as follows: Since all object network models use the same design, they can be trained in parallel under the parallelization function provided by functorch; before officially starting the training, all object network model decorders and optimizers optimisers need to be superimposed and sent to the functorch.combine_state_for_ensemble function from the functorch library to obtain the object parallel network model decoder_model and the parallel optimizer decoder_param, thereby realizing the initialization of the object parallel network model.
5. The method for 3D reconstruction of multiple objects in a scene based on characteristic three-planes according to claim 4, characterized in that: The specific operations of step S4 are as follows: After obtaining the mask of the object, the maximum boundary of the mask part is extracted, and a bounding box covering the object mask is formed, namely the object bounding box (x1, x2, y1, y2). The width and height of the bounding box are used to randomly select coordinates to obtain random coordinates; light sampling is performed in the random coordinates so that only the pixels within the object bounding box are subject to object-level supervision, and the obtained light set is recorded as Where N r Represents the total number of rays in the ray set; for the pixels in the mask part within the bounding box, the truncated signed distance field value, RGB value and depth value generated by the sampling point are used to supervise these pixels; for other pixels in the object bounding box, since they do not belong to the object, these pixels are left blank during training; since each object establishes its own corresponding object network model and key frame buffer through its own mask, and uses the object's own key frame buffer to extract image information for ray sampling, the training between different objects does not interfere with each other; When processing the input of each frame, we simulate the ray casting in NeRF and use the current camera pose {R i |t i }Calculate the light corresponding to the randomly selected pixel and use it to render the truncated signed distance field, color information and depth information of the light; the sampling method consists of the following three methods: layered uniform sampling, depth information guided sampling and TSDF guided sampling, sampling Nc, Ns and Nt points along each light respectively; Layered uniform sampling: Directly uniformly sample Nc points at the camera's near boundary tn and far boundary tf, using the maximum depth D of the object's own mask area surface Replace the far boundary tf, and uniform sampling will sample Nc points in total; when there is incorrect depth data at the sampling point, that is, the depth data is less than or equal to 0 or exceeds the set boundary, layered uniform sampling will be used on the sampling light to replace the depth information guided sampling and TSDF guided sampling. At this time, uniform sampling will sample N points in total all =Nc+Ns+Nt points; the sampling method is as follows: Among them, u represents uniform distribution, ti represents the i-th sampling point; Depth information guides sampling: Under the guidance of depth information, according to the depth D surface The normal distribution centered on is sampled, and a total of Ns points are sampled. The specific sampling method is as follows: Where N represents normal distribution; If the current depth information noise is large, D surface The information is inaccurate, so use the far boundary tf instead of D surface , and replace the depth-guided sampling with uniform sampling; TSDF guides sampling as follows: Where T represents the truncation threshold in the truncated signed distance field; Similarly, when the surface depth value is inaccurate, a uniform sampling strategy is introduced; finally, N all =Nc+Ns+Nt points, denoted as p n , where n∈[1,N all ], and the set of all sampling points of the kth object is recorded as 6. The method for 3D reconstruction of multiple objects in a scene based on characteristic three-planes according to claim 5, characterized in that: The specific method of step S5 is as follows: First, all the sampling points of each object obtained in step S4 are superimposed to obtain a superimposed point set of all objects, which is referred to as the superimposed point set. The dimension is The obtained superposition point set is sent to the object parallel network model; In the first step, the superposition point set is transformed into a dimension of 3] is a set of normalized points The second step is to Feed it into the core part of the network to get a rough geometry output and fine geometry output The two features are superimposed to obtain the final geometric output result, which is expressed as follows: Similarly, the appearance output is: The third step is to obtain the geometry and appearance output. and Both pass through two shallow MLP decoders {M g , M a ) to obtain the output result and in, is the TSDF value, The output of MLP will be used to extract depth and color values. The TSDF value obtained by each ray is converted into volume density. For each object, the conversion is specifically expressed as follows: Where β is represented as a learnable parameter that controls the sharpness of the surface boundary; the color and depth of each ray are then rendered by the volume density, as follows: where z n Represented as p n The depth value of n represents the weight value, Represents the color value of the rendering, Indicates the depth value of the rendering, p m Indicates the sampling point of the current light index m, p n Represents all sampling points on this ray; the depth obtained and color It will be used to optimize the network model and extract the mesh after the network model is trained; For obtaining the output results of the background network model, the steps are the same as those of the object parallel network model, that is, the obtained background sampling points are sent to the background network model to obtain the corresponding depth and color information.
7. The method for 3D reconstruction of multiple objects in a scene based on characteristic three planes according to claim 6, characterized in that: The specific method of step S6 is as follows: For each object k, only the pixels within the 2D object bounding box of each object are sampled, and the sampled ray is denoted as R k And only the TSDF value, depth value and color value of the mask part of the pixels within the 2D bounding box are optimized. The mask is represented by M k , the depth loss and color loss of each object are as follows: For TSDF loss, it is divided into two parts: free space loss and loss close to the surface and within the truncation region; The free space loss of each object is expressed as: in is a set of points on the current object sampling ray r between the camera center and the surface cutoff area measured by the depth sensor. This loss promotes the TSDF value The value in free space is close to 1, which better optimizes the structural features of sampling points far away from the surface of the object; For sample points close to the surface and within the cutoff region, the loss is expressed as follows: Where z(p) is the depth of point p on the plane relative to the camera, T is the cutoff distance, and D(r) is the true depth value measured by ray r. represents the set of points on the ray r within the cutoff distance; considering that points closer to the surface have stronger geometric constraints, the points close to the surface and within the cutoff area are further divided, and the part with |z(p)-D(r)|≤0.5T is defined as The remaining points are defined as therefore It is divided into: and Therefore, the TSDF loss for each object is expressed as follows: Among them, λ fs , tr-n Ratio tr-f Have greater weight; Above, the overall loss is defined as: Among them, λ1, λ2, λ3 are hyperparameters that control the loss weight; The calculation method of the background network model loss is the same as that of the object parallel network model.
8. The method for 3D reconstruction of multiple objects in a scene based on characteristic three-planes according to claim 7, characterized in that: The specific operations of step S7 are as follows: The training of the network model will traverse each image in the dataset in turn, and a round of training will be performed when each image is received; First, in each round of training, each object and background needs to be iterated 20 times. Each round of iteration includes the sampling of sample points, the output of network model results, the calculation of loss function, the back propagation of loss function and the iterative optimization of optimizer. Secondly, in each round of training, the loss of the output of the object parallel network model is calculated first, and then the loss of the output of the background network model is calculated. The result of adding the two losses will be used for back propagation and network model optimization; Finally, after each round of training, the old values of the optimizers optimisers of all object network models are overwritten with the new values of the parallel optimizer decoder_param after that round of training. Based on this, after each round of object parallel network model training is completed, the training results of each object network model in that round can be obtained synchronously.
9. The method for 3D reconstruction of multiple objects in a scene based on characteristic three planes according to claim 8, characterized in that: The specific operations of step S8 are as follows: Since each network model has its own key frame buffer, the depth information in the network model's own key frame buffer is combined with the camera's intrinsic parameters to convert the depth image of each frame into point cloud data, and determine the minimum directional 3D bounding box of each object and background. Next, random sample points are sampled for each 3D bounding box, and these sample points are output through their respective network models to obtain the depth and color information. The output results are combined with the marchingcube algorithm to obtain the grid of each object and background in the scene. Each grid can be viewed and combined separately, and all the grids are combined to obtain the final result: a combined scene grid.