A three-dimensional scene reconstruction method and apparatus

CN122841601APending Publication Date: 2026-09-29HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510391910.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

已有方案可能不同保证重建后的三维场景满足视觉真实与几何精准的要求,且可能存在重建不成功的情况

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122841601A_ABST
    Figure CN122841601A_ABST
Patent Text Reader

Abstract

This application provides a 3D scene reconstruction method and a 3D scene reconstruction platform, used to first perform a rough geometric reconstruction, and then further perform a 3D Gaussian reconstruction based on the rough reconstruction to obtain a more accurate 3D scene model. The method includes: firstly acquiring perceptual data, which includes data collected by a data acquisition device; then acquiring first reconstruction data, which is obtained by reconstructing geometric elements in the 3D scene based on the perceptual data, including multiple geometric elements constituting the 3D scene; then acquiring a semantic instance segmentation result based on the first reconstruction data and the perceptual data, which includes the semantics of at least one instance; and then acquiring second reconstruction data based on the perceptual data, the semantic instance segmentation result, and the first reconstruction data, which includes the data of the reconstructed 3D Gaussian model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision, and more particularly to a method and apparatus for three-dimensional scene reconstruction. Background Technology

[0002] 3D scene rendering is a widely used solution, particularly in applications such as autonomous driving, digital humans, and other vision-based applications. It is crucial for environmental awareness and device navigation. Furthermore, as devices become increasingly intelligent, the demand for realism and accuracy in 3D scene rendering continues to rise.

[0003] Data is collected from the real world to construct 3D scene models, providing reconstructed scene content for downstream tasks. This process demands high levels of visual realism and geometric accuracy in the reconstructed image. However, existing 3D reconstruction solutions, such as Simultaneous Localization and Mapping (SLAM) reconstruction, primarily focus on improving geometric accuracy, while differentiable reconstruction solutions aim to deliver visually realistic rendering from new perspectives. Existing solutions may not guarantee that the reconstructed 3D scene meets the requirements of visual realism and geometric accuracy, and reconstruction may fail in some cases. Therefore, how to reconstruct a more accurate 3D scene model has become a pressing issue. Summary of the Invention

[0004] This application provides a three-dimensional scene reconstruction method and apparatus, which first performs a rough geometric reconstruction, and then performs a three-dimensional Gaussian reconstruction based on the rough reconstruction to obtain a more accurate three-dimensional scene model.

[0005] In view of this, firstly, this application provides a three-dimensional scene reconstruction method. The method is applied to a three-dimensional reconstruction platform, which runs on an infrastructure including at least one computing node. The three-dimensional reconstruction platform and an acquisition device are communicatively connected. The acquisition device is used to acquire perceptual data of the three-dimensional scene to be reconstructed. The method includes: firstly acquiring perceptual data, which contains data of the three-dimensional scene to be reconstructed; subsequently acquiring first reconstruction data, which is obtained by reconstructing geometric elements in the three-dimensional scene to be reconstructed based on the perceptual data, including multiple geometric elements constituting the three-dimensional scene; subsequently acquiring a semantic instance segmentation result based on the first reconstruction data and the perceptual data, which includes at least one instance and the semantics of the at least one instance, each of the at least one instance being used to indicate an object in the three-dimensional scene to be reconstructed; subsequently acquiring second reconstruction data based on the perceptual data, the semantic instance segmentation result, and the first reconstruction data, which includes a three-dimensional model of the three-dimensional scene to be reconstructed, i.e., data obtained by three-dimensional reconstruction based on the input perceptual data, the semantic instance segmentation result, and the first reconstruction data.

[0006] In this embodiment, a rough geometric reconstruction is first performed, followed by further 3D reconstruction based on the reconstruction results. This allows the rough geometric reconstruction data to be used as a geometric prior for 3D reconstruction. Therefore, by using rough geometric reconstruction, the geometric accuracy of the 3D scene can be guaranteed. Combined with 3D reconstruction, the visual realism of the reconstructed scene can be improved, resulting in geometrically accurate and visually realistic reconstruction data.

[0007] Optionally, the aforementioned 3D reconstruction can specifically choose 3D Gaussian reconstruction or other 3D reconstruction schemes that can achieve more realistic visuals. The specific choice can be determined according to the actual application scenario, thereby obtaining a geometrically accurate and visually realistic 3D model.

[0008] In one possible implementation, the aforementioned perception data includes scene images and depth information. Obtaining the first reconstruction data may include: aligning the scene image with the depth information to obtain a depth value aligned with the scene image; obtaining point cloud map data based on the scene image and the depth value aligned with the scene image, for example, mapping the scene image into three-dimensional space using the depth value to obtain data composed of point clouds; subsequently dividing the point cloud map data into multiple local point cloud data, each local point cloud data representing at least one geometric element; and then determining a device path based on the multiple local point cloud data, the device path being the motion path of the acquisition device, including the pose changes of the acquisition device when acquiring perception data. The aforementioned first reconstruction data may include multiple local point cloud data and the device path. In this embodiment, when performing coarse geometric reconstruction, reconstruction can be performed at the voxel level. Compared to reconstructing the entire scene, reconstruction at the voxel level can improve the accuracy and stability of the output results.

[0009] In one possible implementation, the aforementioned determination of the device path based on multiple local point cloud data may include: dividing the points in the multiple local point cloud data into points of dynamic targets and points of static targets; and then determining the device path based on the points of the static targets. Typically, the targets included in the sensing data can be divided into dynamic targets and static targets. Since both dynamic targets and the acquisition device are usually in motion, static targets are more relevant for calculating the device path. Therefore, in this embodiment, using the points of static targets to determine the device path can more accurately determine the device path.

[0010] In one possible implementation, the aforementioned depth information includes a depth map. Obtaining point cloud map data based on a scene image and a depth value aligned with the scene image includes: aligning the scene image with the depth map to determine the depth value aligned with the scene image; extracting scene features from the scene image; and then projecting the scene features into a three-dimensional space based on the depth value to obtain point cloud map data. In this embodiment, when the depth information is a depth map, the features extracted from the scene image can be mapped into a three-dimensional space based on the depth map to construct point cloud map data.

[0011] In one possible implementation, the aforementioned depth information includes radar point cloud data, which includes point cloud data acquired by radar; the aforementioned acquisition of point cloud map data based on scene image and depth value aligned with scene image includes: downsampling the laser point cloud data to obtain downsampled point cloud data; and associating the downsampled point cloud data with scene image to obtain point cloud map data.

[0012] In this embodiment of the application, when the depth information is point cloud data collected by radar, the image can be directly associated with the point cloud to obtain point cloud map data that can represent a three-dimensional scene.

[0013] In one possible implementation, the aforementioned division of point cloud map data into multiple local point cloud data includes: segmenting the point cloud map data according to a preset size to obtain multiple grid point cloud data; and further segmenting the multiple grid point cloud data based on the flatness of the multiple grid point cloud data to obtain multiple local point cloud data. In this embodiment, after obtaining point cloud map data representing the overall 3D scene, the point cloud map can be further divided into multiple local point cloud data, and the point cloud can be segmented based on flatness to extract point cloud data that can represent local targets in the scene. This allows for subsequent scene reconstruction at a smaller granularity, improving the accuracy of the obtained reconstructed data.

[0014] In one possible implementation, the aforementioned method further includes: determining whether there is missing data in the perception data based on multiple local point cloud data, the missing data including data missing from the perception data and data from parts of the 3D scene; if it is determined that there is missing data in the perception data, generating online guidance information to instruct the user to collect the missing data. In this embodiment, when there is missing data in the perception data, such as when some areas of the scene are not scanned during the collection of perception data, resulting in possible image loss or loss of depth information, online guidance information can be generated to guide the user to scan the unscanned areas in order to collect the missing data.

[0015] In one possible implementation, the aforementioned method further includes: projecting multiple local point cloud data onto multiple voxels of a preset size according to the device path in the first reconstructed data to obtain multiple voxel point cloud data; extracting planar features from the multiple voxel point cloud data to obtain multiple first planar features; obtaining the correlation between each of the multiple first planar features and multiple observation poses, wherein the multiple observation poses are the poses of the virtual camera when observing each planar feature in three-dimensional space; subsequently adjusting the multiple local point cloud data and the device path according to the correlation and the multiple voxel point cloud data to obtain adjusted multiple local point cloud data and adjusted device path; and updating the first reconstructed data according to the adjusted multiple local point cloud data and adjusted device path to obtain third reconstructed data.

[0016] In this embodiment, the coarse reconstruction data can be optimized at the voxel level. The point cloud data output by the coarse reconstruction is combined with the device path and the first plane features for more refined optimization, thereby using more computing power to output more accurate reconstruction data.

[0017] In one possible implementation, the aforementioned instance segmentation based on perceptual data to obtain instance segmentation results may include: obtaining a two-dimensional instance segmentation result based on perceptual data, the two-dimensional instance segmentation result including at least one two-dimensional instance in two-dimensional space and the semantics of at least one two-dimensional instance; obtaining a three-dimensional instance segmentation result based on third reconstruction data, the three-dimensional instance segmentation result including at least one instance in three-dimensional space; aligning the two-dimensional instance segmentation result and the three-dimensional instance segmentation result to obtain an aligned three-dimensional instance segmentation result and a two-dimensional instance segmentation result, the instance segmentation result including the aligned three-dimensional instance segmentation result and the two-dimensional instance segmentation result. In this embodiment, two-dimensional (2D) instances and three-dimensional (3D) instances can be segmented separately, and then the 2D instances and 3D instances can be associated and aligned, thereby combining the 2D instances and 2D semantics to assign corresponding semantics to the 3D instances, and at the same time, the 3D instances can be used to correct the 2D instances and 2D semantics, thereby outputting a 3D instance point cloud with semantics and a 2D instance with spatial consistency and 2D semantics with the 3D instances through a complementary approach of 3D and 2D.

[0018] In one possible implementation, the aforementioned acquisition of 3D instance segmentation results based on third reconstructed data includes: extracting planar features from multiple local point cloud data included in the third reconstructed data, i.e., the adjusted multiple local point cloud data, to obtain at least one second planar feature. The second planar feature may include parameters such as the center point or normal vector of the fitted plane in the adjusted multiple local point cloud data; determining at least one 3D instance based on the at least one second planar feature, and the 3D instance segmentation result includes at least one 3D instance. In this embodiment, planar features can be extracted from local point clouds, thereby using the planar features to identify whether each point cloud belongs to the same instance, and then segmenting the point clouds belonging to the same instance to identify the 3D instance.

[0019] In one possible implementation, when the perceptual data includes scene images, obtaining a two-dimensional instance segmentation result based on the perceptual data includes: obtaining a two-dimensional instance segmentation result based on the scene image, for example, using a pre-trained segmentation model to segment the scene image; aligning the two-dimensional instance segmentation result with the three-dimensional instance segmentation result to obtain an aligned three-dimensional instance segmentation result and a two-dimensional instance segmentation result may include: obtaining a depth map corresponding to at least one two-dimensional instance based on at least one three-dimensional instance; correcting at least one two-dimensional instance based on the depth map to obtain a corrected at least one two-dimensional instance; associating at least one three-dimensional instance with at least one two-dimensional instance to obtain a graph structure, which represents the association relationship between the aligned at least one three-dimensional instance and at least one two-dimensional instance; updating the semantics of at least one two-dimensional instance based on the graph structure to obtain an updated semantics of at least one two-dimensional instance, wherein the two-dimensional instance segmentation result includes the corrected at least one two-dimensional instance and the updated at least one two-dimensional instance semantics, and the three-dimensional instance segmentation result includes at least one three-dimensional instance. In this embodiment, the depth map obtained by projecting the three-dimensional instance onto a two-dimensional plane can be used to correct the two-dimensional instance, and the two-dimensional instance can be associated with the three-dimensional instance through graph voting, and the semantics of the two-dimensional instance can be updated based on the association relationship to obtain a two-dimensional instance segmentation result with spatial consistency with the three-dimensional instance.

[0020] In one possible implementation, the aforementioned updating the semantics of at least one two-dimensional instance based on the graph structure to obtain the updated semantics of at least one two-dimensional instance includes: obtaining observation results of the same three-dimensional instance from multiple perspectives based on the graph structure; and updating the semantics of at least one two-dimensional instance based on the observation results to obtain the updated semantics of at least one two-dimensional instance. In this embodiment of the application, the semantics of two-dimensional instances can be updated by utilizing the behavior of the same instance under different observation perspectives, so that two-dimensional instances under different perspectives have the same semantics.

[0021] In one possible implementation, the aforementioned acquisition of a depth map corresponding to at least one two-dimensional instance based on at least one three-dimensional instance includes: dividing at least one three-dimensional instance into a foreground three-dimensional instance and a background three-dimensional instance based on at least one planar feature; and acquiring a depth map corresponding to at least one two-dimensional instance based on the foreground three-dimensional instance. In this embodiment, the three-dimensional instance can be divided into foreground three-dimensional instances and background three-dimensional instances based on planar features, and then the depth map projected onto the two-dimensional plane can be determined based on the foreground three-dimensional instance. Foreground reconstruction is crucial for three-dimensional scenes; therefore, updating the two-dimensional instance based on the more referential depth map projected onto the two-dimensional plane from the foreground three-dimensional instance can improve the overall scene reconstruction effect.

[0022] In one possible implementation, the aforementioned process of performing 3D Gaussian reconstruction based on perceptual data, semantic instance segmentation results, and first reconstruction data to obtain second reconstruction data includes: determining Gaussian parameters based on perceptual data, semantic segmentation results, and third reconstruction data. These Gaussian parameters include geometric attribute parameters, visual attribute parameters, and semantic attribute parameters. Geometric attribute parameters represent the geometric shape of the target in the 3D scene, visual attribute parameters represent parameters in the visual dimension of the 3D scene, and semantic attributes represent the semantics of the target in the 3D scene. The second reconstruction data is then obtained based on the Gaussian parameters. In this embodiment, 3D Gaussian reconstruction can be performed from geometric, visual, and semantic dimensions, thereby obtaining a 3D reconstruction model that performs better in all three dimensions.

[0023] In one possible implementation, the aforementioned determination of Gaussian parameters based on the perceived data and the third reconstructed data includes: constructing at least one of two-dimensional Gaussian constraints or three-dimensional Gaussian constraints based on the third reconstructed data, wherein the two-dimensional Gaussian constraints are used to constrain the distance between points in the three-dimensional scene and the fitting plane, and the three-dimensional Gaussian constraints are used to constrain the distribution of three-dimensional instances in the three-dimensional scene; and determining geometric attribute parameters based on the perceived data under the constraints of at least one of the two-dimensional Gaussian constraints or three-dimensional Gaussian constraints.

[0024] In the embodiments of this application, during the three-dimensional Gaussian reconstruction process, two-dimensional or three-dimensional Gaussian constraints are set for the reconstruction of geometric parameters, so that the reconstructed combined parameters are within the constraint range.

[0025] In one possible implementation, the aforementioned two-dimensional Gaussian constraint includes at least one of position constraint, orientation constraint, or shape constraint. Position constraint includes constraining the center position of the three-dimensional Gaussian sphere to be within the fitting plane; orientation constraint includes constraining the rotation direction of the three-dimensional Gaussian sphere to be a rotation about the normal vector of the fitting plane; and shape constraint includes constraining the shape of the Gaussian kernel to conform to the constrained shape. Specifically, in this application embodiment, the rotation direction of the three-dimensional Gaussian sphere and the shape of the Gaussian kernel can be constrained to ensure that the finally reconstructed three-dimensional Gaussian sphere meets the requirements.

[0026] In one possible implementation, when determining Gaussian parameters by combining instance segmentation results, perceptual data, and third-party reconstruction data, the Gaussian parameters include semantic attribute parameters, and the instance segmentation results include either two-dimensional or three-dimensional instance segmentation results. The aforementioned determination of Gaussian parameters by combining instance segmentation results, perceptual data, and third-party reconstruction data includes: encoding the instance segmentation results to obtain an encoded feature vector, which is used to represent instances in the three-dimensional scene; and decoding the encoded feature vector to obtain the semantic attribute parameters. In this embodiment, the instance segmentation results can be encoded and decoded to decode the semantic attribute parameters, thereby assigning semantics to each instance in the three-dimensional Gaussian model.

[0027] In one possible implementation, when the instance segmentation result includes both two-dimensional and three-dimensional instance segmentation results, encoding the instance segmentation result to obtain an encoded feature vector includes: encoding the semantic attributes and instance attributes corresponding to the two-dimensional and three-dimensional instance segmentation results to obtain the encoded feature vector; the semantic attribute parameters include the probability distribution of instances in two-dimensional space and the probability distribution of semantics. In this embodiment, two-dimensional instances and three-dimensional instances can be encoded separately to encode the probability distribution and semantic distribution of each instance.

[0028] In one possible implementation, the aforementioned method further includes: decoupling image signal processing (ISP) information from the scene image using a pre-trained bilateral mesh structure; and obtaining color parameters based on the ISP information, whereby visual attribute parameters include color parameters. In this embodiment, visual quality enhancement can also be performed based on exposure decoupling to improve the color performance of the 3D Gaussian model.

[0029] In one possible implementation, the aforementioned 3D Gaussian reconstruction process further includes: when acquiring multiple local point cloud data based on a scene image, performing 3D Gaussian reconstruction within the range corresponding to each local point cloud data to obtain second reconstructed data. In this embodiment, global constraints are also added, thereby ensuring that the constraints of each part in the overall 3D Gaussian model are reconstructed within a certain range.

[0030] In one possible implementation, the aforementioned second reconstruction data is used to plan motion paths for electronic devices.

[0031] In one possible implementation, the aforementioned method further includes rendering the second reconstruction data to obtain a rendered image, thereby displaying the reconstructed 3D scene to the user through the rendered image.

[0032] Secondly, this application provides a three-dimensional scene reconstruction platform. The device is applied to the three-dimensional reconstruction platform, which runs on an infrastructure including at least one computing node. The three-dimensional reconstruction platform and a data acquisition device are communicatively connected. The data acquisition device is used to acquire perceptual data of the three-dimensional scene to be reconstructed. The platform includes:

[0033] The input module is used to acquire sensor data;

[0034] The coarse reconstruction module is used to acquire the first reconstruction data, which is obtained by reconstructing the geometric elements in the 3D scene based on the perception data.

[0035] The semantic instance segmentation module is used to obtain semantic instance segmentation results based on the first reconstruction data and perception data. The semantic instance segmentation results include at least one instance and the semantics of the at least one instance. Each instance in the at least one instance is used to indicate an object in the 3D scene to be reconstructed.

[0036] The Gaussian reconstruction module is used to obtain second reconstruction data based on perceptual data, semantic instance segmentation results, and first reconstruction data. The second reconstruction data includes the data of the reconstructed 3D Gaussian model.

[0037] Each of the aforementioned modules can be deployed in at least one computing node within the 3D scene reconstruction platform. A computing node can be a computing device, such as a server, a chip, or something else, depending on the specific application scenario.

[0038] The effects achieved by the second aspect or any optional implementation thereof can be referred to the description of the first aspect or any optional implementation thereof, and will not be repeated hereafter.

[0039] In one possible implementation, the aforementioned perception data includes scene images and depth information. The coarse reconstruction module is specifically used for: aligning the scene images and depth information to obtain a depth value aligned with the scene images; acquiring point cloud map data based on the scene images and the depth values ​​aligned with the scene images; dividing the point cloud map data into multiple local point cloud data, where each local point cloud data represents at least one geometric element; and determining a device path based on the multiple local point cloud data, where the device path is the motion path of the acquisition device. The first reconstruction data includes the multiple local point cloud data and the device path.

[0040] In one possible implementation, the aforementioned apparatus further includes: a global optimization module, configured to: project multiple local point cloud data onto multiple voxels based on the device path in the first reconstructed data to obtain multiple voxel point cloud data; extract planar features from the multiple voxel point cloud data to obtain multiple first planar features; obtain the correlation between each of the multiple first planar features and multiple observation poses, wherein the multiple observation poses are the poses of the virtual camera when observing each planar feature in three-dimensional space; adjust the multiple local point cloud data and the device path based on the correlation and the multiple voxel point cloud data to obtain adjusted multiple local point cloud data and adjusted device path; and update the first reconstructed data with the adjusted multiple local point cloud data and adjusted device path to obtain third reconstructed data.

[0041] In one possible implementation, the aforementioned semantic instance segmentation module is specifically used for: obtaining a two-dimensional instance segmentation result based on perceptual data, the two-dimensional instance segmentation result including at least one two-dimensional instance in two-dimensional space and the semantics of at least one two-dimensional instance; obtaining a three-dimensional instance segmentation result based on third reconstruction data, the three-dimensional instance segmentation result including at least one instance in three-dimensional space; aligning the two-dimensional instance segmentation result with the three-dimensional instance segmentation result to obtain an aligned three-dimensional instance segmentation result and a two-dimensional instance segmentation result, the aforementioned semantic instance segmentation result including the aligned three-dimensional instance segmentation result and the two-dimensional instance segmentation result.

[0042] In one possible implementation, the aforementioned semantic instance segmentation module is specifically used to: extract planar features from multiple local point cloud data included in the third reconstructed data to obtain at least one second planar feature; determine at least one three-dimensional instance based on the at least one second planar feature, and the three-dimensional instance segmentation result includes at least one three-dimensional instance.

[0043] In one possible implementation, when the perceptual data includes a scene image, the aforementioned semantic instance segmentation module is specifically used for: obtaining a two-dimensional instance segmentation result based on the scene image; obtaining a depth map corresponding to at least one two-dimensional instance based on at least one three-dimensional instance; correcting at least one two-dimensional instance based on the depth map to obtain a corrected at least one two-dimensional instance; associating at least one three-dimensional instance with at least one two-dimensional instance to obtain a graph structure, the graph structure being used to represent the association relationship between the aligned at least one three-dimensional instance and at least one two-dimensional instance; updating the semantics of at least one two-dimensional instance based on the graph structure to obtain an updated semantics of at least one two-dimensional instance, wherein the two-dimensional instance segmentation result includes the corrected at least one two-dimensional instance and the updated at least one two-dimensional instance, and the three-dimensional instance segmentation result includes at least one three-dimensional instance.

[0044] In one possible implementation, the aforementioned semantic instance segmentation module is specifically used to: obtain observation results of the same three-dimensional instance from multiple perspectives based on the graph structure; update the semantics of at least one two-dimensional instance based on the observation results, and obtain the updated semantics of at least one two-dimensional instance.

[0045] In one possible implementation, the aforementioned semantic instance segmentation module is specifically used to: divide at least one three-dimensional instance into a foreground three-dimensional instance and a background three-dimensional instance based on at least one planar feature; and obtain a depth map corresponding to at least one two-dimensional instance based on the foreground three-dimensional instance.

[0046] In one possible implementation, the aforementioned Gaussian reconstruction module is specifically used to: determine Gaussian parameters based on perceptual data, semantic segmentation results, and third reconstruction data. The Gaussian parameters include geometric attribute parameters, visual attribute parameters, and semantic attribute parameters. The geometric attribute parameters are used to represent the geometric shape of the target in the three-dimensional scene, the visual attribute parameters are used to represent the parameters in the visual dimension of the three-dimensional scene, and the semantic attributes are used to represent the semantics of the target in the three-dimensional scene; and obtain second reconstruction data based on the Gaussian parameters.

[0047] In one possible implementation, the aforementioned Gaussian reconstruction module is specifically used to: construct at least one of two-dimensional Gaussian constraints or three-dimensional Gaussian constraints based on the third reconstruction data, wherein the two-dimensional Gaussian constraints are used to constrain the distance between points in the three-dimensional scene and the fitting plane, and the three-dimensional Gaussian constraints are used to constrain the distribution of three-dimensional instances in the three-dimensional scene; and determine geometric attribute parameters based on the perception data under the constraints of at least one of the two-dimensional Gaussian constraints or three-dimensional Gaussian constraints.

[0048] In one possible implementation, the aforementioned two-dimensional Gaussian constraint includes at least one of position constraint, attitude constraint, or shape constraint. The position constraint includes constraining the center position of the three-dimensional Gaussian sphere to be within the fitting plane. The attitude constraint includes constraining the rotation direction of the three-dimensional Gaussian sphere to be a rotation about the normal vector of the fitting plane. The shape constraint includes constraining the shape of the Gaussian kernel to conform to the constraint shape.

[0049] In one possible implementation, the aforementioned Gaussian reconstruction module is further configured to: decouple image signal processing (ISP) information in the scene image through a pre-trained bilateral grid structure; and obtain color parameters based on the ISP information, wherein the visual attribute parameters include color parameters.

[0050] Thirdly, embodiments of this application provide a computing device cluster, which includes one or more computing devices, each including a processor and a memory, wherein the processor and memory are interconnected via a circuit, and the processor calls program code in the memory to perform processing-related functions in the method shown in any of the first aspects above.

[0051] Fourthly, this application provides a three-dimensional reconstruction platform that runs on an infrastructure including at least one computing node. The three-dimensional reconstruction platform and an acquisition device are communicatively connected. The acquisition device is used to acquire perception data of the three-dimensional scene to be reconstructed. The computing node can be used to perform the steps as described in the first aspect or any optional implementation of the first aspect.

[0052] Fifthly, embodiments of this application provide a digital processing chip or chip, the chip including a processing unit and a communication interface, the processing unit obtaining program instructions through the communication interface, the program instructions being executed by the processing unit, the processing unit being used to perform processing-related functions as described in the first aspect or any optional embodiment of the first aspect based on data collected by the at least one sensor.

[0053] In a sixth aspect, embodiments of this application provide a computer-readable storage medium including instructions that, when executed on a computer, cause the computer to perform the method described in the first aspect or any optional implementation thereof.

[0054] In a seventh aspect, embodiments of this application provide a computer program product comprising a computer program / instructions, which, when executed by a processor, causes the processor to perform the method described in the first aspect or any optional implementation thereof. Attached Figure Description

[0055] Figure 1 A schematic diagram of the architecture of a three-dimensional reconstruction platform provided in an embodiment of this application;

[0056] Figure 2 A flowchart illustrating a three-dimensional reconstruction method provided in an embodiment of this application;

[0057] Figure 3 A flowchart illustrating another three-dimensional reconstruction method provided in an embodiment of this application;

[0058] Figure 4 A flowchart illustrating another three-dimensional reconstruction method provided in an embodiment of this application;

[0059] Figure 5 A flowchart illustrating another three-dimensional reconstruction method provided in an embodiment of this application;

[0060] Figure 6A flowchart illustrating another three-dimensional reconstruction method provided in an embodiment of this application;

[0061] Figure 7 A flowchart illustrating another three-dimensional reconstruction method provided in an embodiment of this application;

[0062] Figure 8 A flowchart illustrating another three-dimensional reconstruction method provided in an embodiment of this application;

[0063] Figure 9 A flowchart illustrating another three-dimensional reconstruction method provided in an embodiment of this application;

[0064] Figure 10 This is a schematic diagram of the structure of a three-dimensional reconstruction platform provided in an embodiment of this application;

[0065] Figure 11 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application;

[0066] Figure 12 This is a schematic diagram of the structure of a computing device cluster provided in an embodiment of this application;

[0067] Figure 13 This is a schematic diagram of another computing device cluster provided in an embodiment of this application. Detailed Implementation

[0068] The technical solutions of the embodiments of this application will now be described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0069] To facilitate understanding, some concepts or categories involved in the embodiments of this application will be explained below.

[0070] (1) Three-dimensional reconstruction

[0071] This refers to establishing a mathematical model of a three-dimensional object suitable for computer representation and processing. It forms the basis for processing, manipulating, and analyzing the properties of the object in a computer environment, essentially simulating a three-dimensional entity in a virtual three-dimensional space. Specifically, it involves determining the pose of each frame based on multiple images and the corresponding camera parameters, then mapping the pixel coordinates of each frame to three-dimensional space based on these poses, and finally reconstructing the three-dimensional model.

[0072] (2) Multilayer perceptron (MLP)

[0073] MLP (Multi-Level Processing) neural networks are a type of feedforward neural network (FFN). It is a fully connected (meaning each neuron is connected to all neurons in the previous layer) feedforward neural network model. For example, in classification problems, it transforms the raw scores for each class through multiple fully connected layers, and then uses a softmax function to obtain the predicted probability for each class. An MLP can be viewed as a directed graph composed of multiple node layers, each fully connected to the next. Besides the input node, each node is a neuron (or processing unit) with a non-linear activation function.

[0074] (3) transformer

[0075] A transformer architecture is a feature extraction network that includes both an encoder and a decoder. Of course, in some cases, a transformer architecture may not include an encoder but may include a decoder.

[0076] Encoder: Learns features, such as pixel features, within the global receptive field using self-attention.

[0077] Decoder: Learns the features of the desired module, such as the features of the output box, through self-attention and cross-attention.

[0078] For example, the structure of a Transformer layer in an existing scheme may include a multi-head attention network and a feedforward network module. Taking natural language processing as an example, the multi-head attention network obtains corresponding weight values ​​by calculating the relevance between words, thus obtaining context-related word representations, which is the core part of the Transformer structure. The feedforward network further transforms the obtained representations to obtain the final output of the Transformer layer. In addition to these two important components, residual layers (ADD) and linear normalization (Norm) are also stacked on these two components to optimize the output of the Transformer layer.

[0079] (4) Neural radiance field (NeRF)

[0080] NeRF is a method or model that uses neural networks to implicitly represent 3D scenes. A given scene can be learned using NeRF, and this scene is implicitly stored in the parameters of the NeRF neural network, i.e., the implicit representation of the scene. If a new perspective is needed, NeRF can be used to calculate the light and color values ​​at various locations within this scene, and after rendering, the new perspective can be output.

[0081] Typically, a neural field takes spatial coordinates or other dimensions such as time and camera pose as input, and simulates an objective function through a multilayer perceptron (MLP) network to generate an objective scalar (such as color, depth, etc.). The output of a neural radiation field (NeRF) is color and volume density. By sampling points along a camera ray, NeRF can predict the color and volume density for each point. Then, volume rendering projects the colors of all points along the ray onto the image plane to obtain the pixel color value corresponding to that ray. By sampling and volume rendering in parallel for each camera ray, an image from a new perspective can be generated.

[0082] (5) Simultaneous Localization and Mapping (SLAM):

[0083] This is a solution that enables robots or intelligent vehicles to navigate autonomously in unknown environments. It primarily addresses two key issues: first, determining the robot's position within the environment (i.e., localization); and second, constructing a map of the environment (i.e., mapping). The SLAM (Simultaneous Localization and Mapping) solution allows robots to achieve self-localization and environmental mapping using data collected by sensors, even without prior map information.

[0084] (6) Structure for Motion Inference (SFM):

[0085] It is a general term for algorithms that reconstruct three-dimensional structures from disordered images. Given a series of projected images of a three-dimensional structure taken from different viewpoints, SfM can reconstruct the corresponding 3D structure from these images.

[0086] (7) 3D Gaussian Splatting (3DGS)

[0087] Similar to NeRF, 3DGS is a scene representation and rendering scheme in the field of computer graphics. Unlike NeRF, which uses neural networks to implicitly represent scenes, 3DGS uses a set of Gaussian ellipsoids to explicitly represent 3D scenes. Each Gaussian ellipsoid has specific parameters, such as position, rotation, scale, opacity, and spherical harmonic coefficients representing view-dependent colors. While each Gaussian ellipsoid is discrete in space, it is continuous and differentiable within itself. This explicit representation uses relatively less memory and has lower computational complexity during training and rendering compared to NeRF. 3DGS information can be directly stored and manipulated, facilitating subsequent editing and processing, such as dynamic reconstruction, geometric editing, and physical simulation.

[0088] Specifically, 3DGS uses Gaussian functions to represent points or volumes in 3D space, achieving efficient and accurate representation of 3D scenes. Each point can be described by a Gaussian function, which defines the point's spatial location and its diffusion along various axes. In 3DGS, the probability of a point's existence is determined by the value of the Gaussian function, allowing for the modeling of blurred boundaries on surfaces, rather than strict geometric boundaries. 3DGS can represent 3D shapes simultaneously at multiple scales, thus capturing details at different levels, from microscopic to macroscopic.

[0089] (8) EmbodiedGS Reconstruction

[0090] In embodied intelligence applications, a scene reconstruction scheme using 3D Gaussian as the primary representation is employed. Given that embodied intelligence, in navigation and operational applications, requires both visual perception of the environment and interaction with the scene to alter the state of objects within it, several requirements are placed on the reconstruction model to enable rapid deployment of embodied models trained in simulators to perform embodied tasks within the simulator. These requirements include open-set semantics, collision models, embodied operability, and realism. In terms of the reconstruction scheme, these requirements translate to the need for semantically rich, geometrically accurate, and visually realistic characteristics.

[0091] (9) Depth camera (Color data with Depth processing, RGB-D)

[0092] By combining RGB three-channel color images with a depth map, pixel-level correspondence is achieved through registration, enabling the output of multimodal data on color and spatial distance. The depth map is typically acquired by an infrared sensor or other types of sensors, and the acquired depth map represents the distance between the object and the sensor. RGB-D cameras can perform depth sensing in various ways, such as through structured light, time-of-flight (ToF), or binocular imaging to output depth maps.

[0093] (10) Radar

[0094] Specifically, this can include lidar, millimeter-wave radar, or radio radar, which transmits wireless signals (such as radio waves, millimeter-wave signals, or laser signals) and receives the echoes reflected from the target, calculates the distance between the target and the radar, and outputs point cloud data. The points in the point cloud data can be used to represent targets in the environment, and the information of each point can include the distance between the target and the radar.

[0095] (11) LiDAR-Vision-IMU Odometry (LVIO)

[0096] A multi-sensor solution that integrates LiDAR, camera, and inertial measurement unit (IMU).

[0097] (12) Flatness

[0098] Flatness refers to the macroscopic deviation of the unevenness of all points on a planar surface relative to an ideal plane, reflecting the overall flatness of the surface. In the embodiments of this application, the ideal reference plane can be determined based on the least squares method, and the standard deviation of the perpendicular distance between each point on the surface and the reference plane can be calculated to quantify the overall flatness.

[0099] The 3D scene reconstruction method provided in this application can be applied to various scenarios that require scene rendering. For example, it can be applied to the construction of the vehicle's environment in intelligent driving scenarios, the construction of the environment of embodied devices, and the construction of virtual reality (VR) or augmented reality (AR) combined virtual and real scene scenarios.

[0100] Accordingly, the method provided in this application embodiment is applied to a 3D reconstruction platform. The 3D reconstruction platform runs on an infrastructure, which includes at least one computing node. The 3D reconstruction platform and an acquisition device are communicatively connected. The acquisition device is used to acquire perception data of the 3D scene to be reconstructed.

[0101] Specifically, the computing node can include user terminals, servers, or cloud service systems. Furthermore, in some scenarios, the computing node can also be a chip with computing capabilities, such as a graphics processing unit (GPU) or a neural-network processing unit (NPU).

[0102] For example, the method provided in this application embodiment can be deployed in various smart devices, such as in various vision task scenarios, including intelligent driving scenarios for robots, flight scenarios for drones, or driving scenarios for smart vehicles. It can utilize input data to perform scene modeling and output scene reconstruction data that can be used for tasks on smart devices.

[0103] For example, the method provided in this application embodiment can be deployed in a cloud service system. This system may include a client and a cloud server, and all or part of the steps in the method provided in this application embodiment can be deployed in the cloud server. Users can input data of the 3D scene to be reconstructed through the client, and the client transmits the data to the cloud server, which then performs the 3D scene reconstruction.

[0104] For example, the 3D reconstruction platform provided in this application can be deployed on a cloud server, with user terminals acting as acquisition devices to collect sensory data. Services are provided to users through cloud services. This 3D reconstruction platform can also be called a cloud service system or cloud service architecture, etc., depending on the specific application scenario.

[0105] For example, the architecture deployed on the 3D reconstruction platform used in the method provided in this application can be as follows: Figure 1 As shown. Figure 1 As shown, the architecture may include a cloud and acquisition devices 14. The cloud may include a cloud platform 12 and a data center 13. The 3D reconstruction platform 121 may deploy a 3D reconstruction platform 1211. The number of acquisition devices 14 may be one or more. Here, only one user terminal is used as an example for illustration and is not intended to be limited.

[0106] The cloud platform 12 may specifically include a 3D reconstruction platform 121, which may include a server cluster or a standalone computing device, or other devices with computing capabilities. Optionally, the 3D reconstruction platform 121 can work with other computing devices, such as data storage, routers, load balancers, etc. The 3D reconstruction platform 121 can use data from the data storage system or call program code in the data storage system to implement the method steps provided in the embodiments of this application.

[0107] Data center 13 can be used to store data for the 3D reconstruction platform 121 to query or write data, etc.

[0108] The acquisition device 14 establishes a connection with the 3D reconstruction platform 121. The acquisition device 14 is equipped with one or more sensors to collect sensory data about the environment where the user terminal is located. The number of acquisition devices 14 can be one or more. For example, in a scenario with multiple user terminals, multiple user terminals can upload collected data, such as images or point clouds, to a cloud server. The actions performed by the acquisition device in the method flow provided in this application embodiment can be performed by the acquisition device 14; in other words, one form of the acquisition device provided in this application embodiment can be the acquisition device 14.

[0109] The acquisition device 14 can send the acquired perception data to the 3D reconstruction platform 121. The 3D reconstruction platform 121 can use the received perception data to reconstruct a 3D scene using the method provided in this embodiment, and then send rendering data to the user terminal. After receiving the rendering data, the user terminal can display the rendering data on the screen.

[0110] In summary, the methods provided in this application embodiment can be applied to computing devices, that is, the aforementioned cloud 11 can be various computing devices, such as server clusters, cloud platforms, personal computers, smartphones or smart cars, etc.

[0111] In combination with the above Figure 1 As shown in the architecture, all or part of the steps of the method provided in this application can be executed by a computing device, which may specifically include a server cluster or a cloud platform, and can provide services to users through a client, or the computing device may also be other devices with computing capabilities.

[0112] In some existing 3D scene reconstructions, it is often impossible to simultaneously guarantee visual realism and geometric accuracy. Furthermore, the reconstructed scene often fails to accurately convey the specific meaning of each instance. Additionally, there may be issues such as reconstruction failures requiring multiple attempts, resulting in very low reconstruction efficiency.

[0113] For example, traditional SLAM reconstruction aims to minimize geometric errors from multi-view observations by adjusting pose and map point positions. Therefore, the advantage of SLAM reconstruction lies in its geometric accuracy. The point cloud obtained from SLAM reconstruction is then converted into a mesh model commonly used in simulators and textured. However, because meshes use fixed coloring for vertices or faces, their visual effect is not realistic enough from a roaming perspective.

[0114] For example, NeRF utilizes multilayer perceptrons (MLPs) to map the spatial coordinates of the input image and the viewing direction into color and volume density. Volume density is used to simulate the opacity of objects. By inputting the viewing direction, it learns images from different perspectives and then synthesizes images from unknown perspectives. This reconstruction method can simulate human perception, build world models through self-supervised learning, and is highly efficient in handling reflective scenes. It has been widely applied in fields such as metaverse, autonomous driving, game development, and 3D reconstruction. However, when dealing with large-scale scenes, its high computational cost and low efficiency in model training and rendering processes remain to be addressed.

[0115] Unlike NeRF's implicit representation, 3D Gaussian sphere modeling uses explicit representations of points in a scene using 3D Gaussian functions. Each 3D Gaussian function has specific geometric, color, and opacity properties. These Gaussian points, after proper positioning and scaling, effectively represent the overall shape and form of the scene, ultimately forming a continuous radiation field that encodes the light color and density information of each point within the 3D scene. This lays the foundation for efficiently rendering complex scenes and achieving superior visual quality. 3DGS, with its high-quality rendering effects, fast training and rendering speeds, and efficient scene representation, has gained widespread attention and application in real-time rendering, 3D reconstruction, augmented reality, 3D editing, and animation.

[0116] NeRF (or 3DGS) uses implicit (or explicit) representations for reconstruction. Their optimization goal is to minimize the photometric error between the rendered image and the training viewpoint real image by adjusting the MLP (or 3D Gaussian points). The advantage of differentiable reconstruction is visual realism. NeRF-based methods are essentially implicit neural network representations and do not have explicit geometric structures. Although 3DGS uses explicit representations, its trained model has Gaussian points that are spatially disorganized, making it impossible to extract accurate geometric structures from them.

[0117] Obtaining accurate geometric information from NeRF and 3DGS is one of the urgent problems to be solved. However, existing solutions are often only validated on public datasets, and their generalization ability is seriously insufficient in real-world large-scale scenes.

[0118] In real-world scenarios, NeRF-based solutions like Neuralangelo and 3DGS-based solutions like Gaussian Opacity Field both suffer from severe degradation in rendering quality and geometric accuracy. A comprehensive solution that simultaneously achieves visual realism and geometric precision still needs to address its generalization problem.

[0119] Therefore, this application provides a three-dimensional scene reconstruction method, which first performs a coarse reconstruction and then performs a three-dimensional Gaussian optimized reconstruction to obtain three-dimensional scene reconstruction data that simultaneously meets the requirements of visual realism and geometric accuracy.

[0120] The method flow provided in the embodiments of this application is described below.

[0121] See Figure 2 The following is a flowchart illustrating a three-dimensional scene reconstruction method provided in this application embodiment.

[0122] First, the method provided in this application embodiment can be applied to the aforementioned three-dimensional reconstruction platform, which runs in an infrastructure including at least one computing node. The three-dimensional reconstruction platform is communicatively connected to an acquisition device, which is used to acquire perception data of the three-dimensional scene to be reconstructed.

[0123] 201. Acquire sensory data.

[0124] The perception data includes data of the 3D scene to be reconstructed collected by at least one sensor in the acquisition device.

[0125] The perceived data can be divided into scene images and depth information. Scene images are images of the scene to be rendered, acquired by image sensors. Depth information can be used to represent the position of various targets in the scene to be rendered or their distance from the acquisition device.

[0126] Optionally, the depth information may specifically include depth maps or radar point cloud data. For example, the acquisition device may include a depth sensor, which can be used to acquire depth maps; if the acquisition device includes radar, point cloud data can be acquired using radar. Of course, if the acquisition device includes both a depth sensor and radar, depth information can be acquired using either one of the sensors, or both depth maps and point cloud data can be acquired simultaneously.

[0127] Optionally, after receiving the sensing data, the sensing data can be preprocessed, such as by denoising or desensitizing the data, to obtain more usable sensing data.

[0128] 202. Obtain the first reconstruction data.

[0129] The first reconstruction data is obtained by reconstructing geometric elements in a 3D scene based on perception data. That is, the first reconstruction data may include multiple geometric elements, which can represent one or more targets in the scene.

[0130] In combination with the above Figure 1In the 3D reconstruction platform to which the method provided in this application is applied, the computing node may specifically be a server, and the acquisition device may specifically include a user terminal. Optionally, step 202 may be executed by the client or by a cloud server.

[0131] For example, in one possible implementation, the reconstruction of geometric elements requires less computing power compared to 3D Gaussian reconstruction or other reconstruction algorithms. Therefore, the reconstruction of geometric elements can be performed by a client with lower computing power, thereby performing a coarse reconstruction on the client side to improve the geometric accuracy of the input data when performing subsequent 3D Gaussian reconstruction.

[0132] In another possible implementation, step 202 can also be performed by a cloud server. That is, the user can collect sensing data through a data acquisition device and transmit the sensing data to the cloud server for execution. This can be understood as the cloud server first performing a coarse geometric reconstruction and outputting more accurate coarse reconstruction data, i.e., the first reconstruction data.

[0133] Specifically, in one possible implementation, point cloud map data can be generated based on scene images and depth information included in the perception data. This point cloud map data can be used to represent various points in the 3D scene. For example, if the depth information includes a depth map, the information in the scene image is mapped to 3D space based on the depth map to obtain the point cloud map data; if the depth information includes radar point cloud data, the scene image and radar point cloud data are aligned to obtain the point cloud map data. The point cloud map is then divided into multiple local point cloud data sets, allowing subsequent inference at a local granularity, thus improving inference accuracy. Subsequently, a device path is determined based on the multiple local point cloud data sets. This device path is the movement path of the acquisition device, and the aforementioned first reconstruction data can include these multiple local point cloud data sets and the device path. Therefore, in this embodiment, the geometric reconstruction of the 3D scene can be achieved by dividing the point cloud map into multiple local parts to infer the device path.

[0134] In one possible implementation, the acquired sensing data typically includes dynamic and static targets. Dynamic targets are those in motion, usually with a velocity greater than 0, while static targets are those in a static state, usually with a velocity approximately equal to 0. When determining the device path, dynamic targets, due to their motion, are less relevant for path calculation, while static targets, being static, have their relative motion with the acquisition device dependent solely on the device's movement; therefore, static targets are more relevant for path calculation. Thus, in this embodiment, points in multiple local point cloud datasets can be divided into points representing dynamic and static targets, and the device path can be determined based on the points representing static targets, such as by determining the device path based on changes in the distance between the static target and the acquisition device.

[0135] Specifically, in one possible implementation, the aforementioned depth information may include a depth map. During the acquisition of the point cloud map, the scene image and the depth map can be aligned to determine the depth value corresponding to each pixel in the scene image. Scene features are extracted from the scene image, and the scene features are mapped into three-dimensional space based on the depth values ​​to obtain point cloud map data. Therefore, even in scenarios where point cloud data has not been acquired, such as when scene images and depth information are acquired using RGB-D, the scene image can still be mapped into three-dimensional space to obtain point cloud map data that can be used to represent a three-dimensional scene.

[0136] In one possible implementation, the aforementioned depth information includes radar point cloud data, which is the data collected by radar. When acquiring the point cloud map, the radar point cloud data can be downsampled to obtain downsampled point cloud data, thereby acquiring local data from the radar point cloud. The downsampled point cloud data is then correlated with the scene image to map the content of the two-dimensional image onto the three-dimensional point cloud, resulting in point cloud map data containing richer information.

[0137] Optionally, in one possible implementation, the point cloud map data can be segmented according to a preset size to obtain multiple grid point cloud data; the multiple grid point cloud data can then be further divided based on the flatness of each grid to obtain multiple local point cloud data. This allows the point cloud map to be divided into multiple voxels based on the flatness of each grid, i.e., the point cloud map data can be divided into multiple local data. One voxel can represent a local region; for example, point clouds belonging to the same plane can be divided into one voxel based on flatness, facilitating three-dimensional reconstruction in geometric dimensions.

[0138] In one possible implementation, if step 202 is performed by the client, it can also be determined whether there is missing data in the perception data based on multiple local point cloud data. Missing data includes data missing from the perception data and data from parts of the 3D scene. If missing data is determined, online guidance information is generated to instruct the user to collect the missing data. For example, if the client fails to collect information from a portion of the scene, guidance information can be generated for that portion, which may include the location or direction corresponding to the missing data. This reminds the user to use the client to collect the missing data to obtain more complete scene perception data.

[0139] In one possible implementation, multiple local point cloud data can be projected onto multiple voxels based on the device path in the first reconstructed data to obtain multiple voxel point cloud data; planar features are extracted from the multiple voxel point cloud data to obtain multiple first planar features; the correlation between each of the multiple first planar features and multiple observation poses is obtained, where the multiple observation poses are the poses of the virtual camera when observing each planar feature in three-dimensional space; then, the multiple local point cloud data and the device path are adjusted according to the correlation and the multiple voxel point cloud data to obtain adjusted multiple local point cloud data and adjusted device path; the first reconstructed data is updated according to the adjusted multiple local point cloud data and adjusted device path to obtain third reconstructed data. The first reconstructed data mentioned below can be replaced with the third reconstructed data, or the third reconstructed data can be used as the first reconstructed data for subsequent processing, which will not be elaborated further below.

[0140] 203. Obtain semantic instance segmentation results based on the first reconstructed data and the perception data.

[0141] Specifically, one or more instances and corresponding semantics can be segmented based on the first reconstructed data and the perceived data. Each of these one or more instances is used to indicate an object in the 3D scene to be reconstructed. Instances can be represented by masks or coordinates, etc. The following example illustrates how instances are represented by masks.

[0142] Optionally, a two-dimensional instance segmentation result can be obtained based on the perceptual data. The two-dimensional instance segmentation result includes at least one two-dimensional instance in the two-dimensional space and the semantics of at least one two-dimensional instance. A three-dimensional instance segmentation result can be obtained based on the first reconstruction data. The three-dimensional instance segmentation result includes at least one instance in the three-dimensional space. Then, the two-dimensional instance segmentation result and the three-dimensional instance segmentation result are aligned to obtain an aligned three-dimensional instance segmentation result and a two-dimensional instance segmentation result. The instance segmentation result includes the aligned three-dimensional instance segmentation result and the two-dimensional instance segmentation result.

[0143] In one possible implementation, a two-dimensional instance segmentation result can be obtained based on perceptual data. This result includes at least one two-dimensional instance in two-dimensional space and its semantics. For example, performing two-dimensional instance segmentation and semantic segmentation on a scene image included in the perceptual data yields one or more two-dimensional instances. A three-dimensional instance segmentation result can also be obtained based on first reconstructed data. This result includes at least one instance in three-dimensional space, which can also be referred to as a three-dimensional instance. The two-dimensional and three-dimensional instance segmentation results are then aligned to obtain an aligned three-dimensional and two-dimensional instance segmentation result. Therefore, in this embodiment, two-dimensional and three-dimensional instance segmentation can be performed separately, followed by alignment. This allows the three-dimensional instance to be used to correct the two-dimensional instance, improving the accuracy of the two-dimensional instance segmentation. Furthermore, the semantics of the two-dimensional instance can be used to add corresponding semantics to the three-dimensional instance, thus giving the instance in three-dimensional space corresponding semantics.

[0144] In one possible implementation, for a 3D instance, planar features of multiple local point cloud data included in the first reconstructed data can be extracted to obtain at least one second planar feature; subsequently, at least one 3D instance is determined based on the at least one second planar feature, and the 3D instance segmentation result includes at least one 3D instance. For example, planes belonging to the same instance can be determined based on the geometric relationships between planes, thereby obtaining multiple 3D instances.

[0145] In one possible implementation, the scene image can be segmented to obtain two-dimensional instance segmentation results. When aligning the two-dimensional segmentation results with the three-dimensional instance segmentation results, a depth map corresponding to at least one two-dimensional instance can be obtained based on at least one three-dimensional instance. This is equivalent to projecting the three-dimensional instance onto the plane corresponding to the two-dimensional instance to correct the two-dimensional instance and make it more accurate. Simultaneously, the three-dimensional and two-dimensional instances are associated to generate a graph structure representing the association between the aligned at least one three-dimensional instance and at least one two-dimensional instance. The semantics of the two-dimensional instance are then updated based on this association, thereby improving the accuracy of the two-dimensional instance's semantics. The aforementioned two-dimensional instance segmentation results include the corrected at least one two-dimensional instance and the updated semantics of at least one two-dimensional instance, and the three-dimensional instance segmentation results include at least one three-dimensional instance.

[0146] In one possible implementation, observation results of the same 3D instance from multiple perspectives can be obtained based on the graph structure; the semantics of at least one 2D instance can be updated based on the observation results to obtain the updated semantics of at least one 2D instance. For example, when performing 2D semantic segmentation, the same instance from different perspectives may be classified with different semantics. Through the embodiments of this application, the 2D semantics can be corrected by using the observation results of different perspectives on the same 3D instance, thereby improving the accuracy of the semantics corresponding to the 2D instance.

[0147] Optionally, at least one 3D instance can be divided into a foreground 3D instance and a background 3D instance based on at least one second plane feature; a depth map corresponding to at least one 2D instance is obtained based on the foreground 3D instance. Therefore, in the embodiments of this application, the 3D instance is divided into foreground and background, and the foreground 3D instance is used to determine the depth map corresponding to the 2D instance, so as to improve the accuracy of the obtained depth map corresponding to the 2D instance. It can be understood that the foreground instance in the 3D scene can be projected onto the plane corresponding to the 2D instance, and the 2D instance can be corrected using the depth map obtained after projection, so as to obtain a 2D instance that is more consistent with the 3D instance space.

[0148] 204. Obtain the second reconstruction data based on the perceptual data, semantic instance segmentation results, and the first reconstruction data.

[0149] After obtaining the semantic instance segmentation result and the first reconstruction data, a second reconstruction can be performed based on the perceptual data, the semantic instance segmentation result and the first reconstruction data. The second reconstruction can be carried out in three dimensions, and the second reconstruction data is output, which includes the data of the three-dimensional model of the three-dimensional scene to be reconstructed.

[0150] Therefore, in this embodiment, after performing a rough geometric model, semantic instance segmentation can be further performed, resulting in a segmentation result that includes both instances and semantics. Furthermore, during the subsequent reconstruction, the data from the rough model can be used as geometric guides to perform more refined modeling based on 3D reconstruction, resulting in a geometrically accurate and visually more realistic modeling result.

[0151] Optionally, the 3D reconstruction can specifically employ 3D Gaussian reconstruction or other reconstruction schemes. The following example uses 3D Gaussian reconstruction as an illustration; however, the 3D Gaussian reconstruction mentioned below can be replaced with other reconstruction schemes that can improve the visual reconstruction output, which will not be elaborated further.

[0152] In one possible implementation, the aforementioned process of performing 3D Gaussian reconstruction based on perceptual data, semantic instance segmentation results, and first reconstruction data to obtain second reconstruction data may specifically include: determining Gaussian parameters based on perceptual data, semantic segmentation results, and first reconstruction data. These Gaussian parameters include geometric attribute parameters, visual attribute parameters, and semantic attribute parameters. Geometric attribute parameters represent the geometric shape of the target in the 3D scene, visual attribute parameters represent parameters in the visual dimension of the 3D scene, and semantic attributes represent the semantics of the target in the 3D scene. Subsequently, the second reconstruction data is obtained based on these Gaussian parameters. In this embodiment, when performing 3D Gaussian reconstruction, reconstruction can be performed from geometric, visual, and semantic dimensions, thereby reconstructing scene reconstruction data that performs better in both geometric and visual dimensions. Furthermore, it may also include the semantics of instances, thus providing a more accurate representation of the 3D scene.

[0153] In one possible implementation, the aforementioned determination of Gaussian parameters based on the perceived data and the first reconstructed data may specifically include: constructing at least one of two-dimensional Gaussian constraints or three-dimensional Gaussian constraints based on the first reconstructed data, wherein the two-dimensional Gaussian constraints are used to constrain the distance between points in the three-dimensional scene and the fitting plane, and the three-dimensional Gaussian constraints are used to constrain the distribution of three-dimensional instances in the three-dimensional scene; and determining geometric attribute parameters based on the perceived data under the constraints of at least one of the two-dimensional Gaussian constraints or three-dimensional Gaussian constraints.

[0154] In this embodiment, for the Gaussian parameters of the geometric dimension, two-dimensional Gaussian constraints or three-dimensional Gaussian constraints can be constructed to constrain the geometric relationships in two or three dimensions, thereby constructing combined reconstruction data that conforms to the expected geometric distribution.

[0155] In one possible implementation, the aforementioned two-dimensional Gaussian constraint includes at least one of position constraint, orientation constraint, or shape constraint. Position constraint includes constraining the center position of the three-dimensional Gaussian sphere to lie within the fitting plane; orientation constraint includes constraining the rotation direction of the three-dimensional Gaussian sphere to be a rotation about the normal vector of the fitting plane; and shape constraint includes constraining the shape of the Gaussian kernel to conform to the constraint shape. Specifically, in this application, constraints can be applied from dimensions such as the instance's position, orientation, or shape to obtain instance reconstruction data with more accurate position, orientation, or shape.

[0156] In one possible implementation, the semantic instance segmentation result can be encoded to obtain an encoded feature vector. This encoded feature vector is used to represent instances in the 3D scene to be reconstructed. The encoded feature vector is then decoded to obtain semantic attribute parameters. In this embodiment, the semantic instance segmentation result can be encoded and decoded to reconstruct the semantics of the 3D scene from the semantically segmented instances, thus obtaining 3D scene reconstruction data including the semantics of each instance.

[0157] Optionally, the semantic attributes and instance attributes corresponding to the two-dimensional instance segmentation results and the three-dimensional instance segmentation results can be encoded to obtain encoded feature vectors. The semantic attribute parameters include the probability distribution of instances in two-dimensional space and the probability distribution of semantics, thereby reconstructing instances and their semantics into the three-dimensional scene.

[0158] In one possible implementation, the 3D Gaussian reconstruction process further includes: when acquiring multiple local point cloud data based on a scene image, performing 3D Gaussian reconstruction within the range corresponding to each local point cloud data to obtain second reconstructed data. In this embodiment, adding global constraints and constructing a 3D Gaussian sphere within a voxel range can further improve the output accuracy of the 3D Gaussian reconstruction.

[0159] In one possible implementation, the image signal processing (ISP) information in the scene image can be decoupled through a pre-trained bilateral grid structure; the color parameters can be obtained based on the ISP information, and the visual attribute parameters in the aforementioned Gaussian parameters can include the color parameters, thereby reconstructing scene reconstruction data with better color performance.

[0160] Furthermore, the second reconstruction data obtained can also be applied to downstream tasks. For example, the output second reconstruction data can be used to plan motion paths for electronic devices.

[0161] In one possible implementation, the second reconstructed data can also be rendered to obtain a rendered image. This allows the 3D scene reconstruction results to be displayed through the image, thus visualizing the 3D scene reconstruction results.

[0162] The foregoing has described the method flow provided by the embodiments of this application. The following will further describe the method flow provided by the embodiments of this application in more detail, in conjunction with specific application scenarios.

[0163] First, the method provided in the embodiments of this application can be divided into multiple stages or multiple modules, such as Figure 3 As shown, it can be specifically divided into acquisition device 31, coarse reconstruction stage 32, global optimization stage 33, semantic instance segmentation stage 34, Gaussian reconstruction stage 35 and downstream task stage 36, etc.

[0164] The acquisition device 31, also known as the acquisition terminal stage, can be equipped with one or more sensors to collect multimodal data, i.e., the aforementioned perception data. Typically, the collected multimodal data can be divided into scene images and depth information. Depth information can include depth maps or radar point cloud data. Depth maps can be acquired through depth sensors, binocular imaging systems, or other sensors, while radar point cloud data can be acquired through radar. For example, multimodal data can be acquired using an RGB-D camera or a combination of LiDAR and a camera. Common mapping equipment includes robots, backpack-mounted professional data acquisition devices, and handheld data acquisition devices. Users need to provide the intrinsic and extrinsic parameters of each sensor on the device, such as the camera's intrinsic parameters and the spatiotemporal extrinsic parameters between the LiDAR and the camera. Data synchronization between sensors is achieved using hardware-triggered timing.

[0165] Optionally, the acquisition device 31 may be equipped with a computing unit, a storage unit, or a communication unit, which can be used to deploy the subsequent coarse reconstruction stage 32 or to perform other data preprocessing functions.

[0166] The coarse reconstruction stage 32 can be used to reconstruct geometric elements in the 3D scene to be reconstructed based on multimodal data, outputting the reconstructed point cloud data and device paths, which is the aforementioned first reconstruction data. When the acquisition device collects multimodal data, it can simultaneously report the intrinsic and extrinsic parameters of each sensor. Coarse reconstruction can be performed based on the intrinsic and extrinsic parameters reported by the acquisition device. Specifically, the acquired point cloud and image data can be used to perform operations such as device path tracking, mesh occupancy estimation, point cloud stitching, and coloring, outputting the reconstructed point cloud data and device paths.

[0167] Optionally, during the coarse reconstruction stage 32, online guidance can be provided. For example, if geometric elements are missing in a certain area during reconstruction, prompts can be generated for the locations of the missing elements, indicating that data should be collected from the unscanned areas. This guides users to scan uncovered areas, effectively avoiding environmental structure gaps caused by incomplete data collection routes and improving collection efficiency.

[0168] The global optimization stage 33 takes the point cloud data and device paths output from the coarse reconstruction stage 32 as input and outputs higher-precision point cloud data and device paths. Based on the input initial point cloud dataset and initial device paths, it generates a high-precision point cloud map, further optimizes the device paths, colors the high-precision point cloud map, and extracts voxel geometry from the point cloud map data. The device paths and extracted voxel geometry, which are in the same coordinate system as the map, are then passed to the downstream stages.

[0169] The semantic instance segmentation stage 34 takes as input the higher-precision point cloud data and device path output from the global optimization stage 33, as well as images from the multimodal dataset, and outputs 2D instances, 3D instances, and their corresponding semantics. Specifically, 2D instance segmentation and 3D instance segmentation can be performed separately, and the 2D instances and semantics can be corrected by combining the 3D instances to obtain a 3D semantic instance point cloud map and spatially consistent 2D semantic instance segmented instances, which are then passed to the next stage.

[0170] The Gaussian reconstruction stage 35 takes as input the higher-precision point cloud data and device path output from the global optimization stage 33, and the semantic instance segmentation results output from the semantic instance segmentation stage 34. The output is the Gaussian reconstruction data after 3D Gaussian reconstruction, also known as the aforementioned second reconstruction data. Specifically, visual, geometric, and semantic instance information can be encoded into 3D Gaussian attributes to achieve 3D Gaussian sphere modeling.

[0171] Downstream task stage 36 can be used to execute downstream tasks, such as performing actions of the embodied device based on Gaussian reconstruction data, or embodied device simulation training.

[0172] Therefore, in this embodiment, the reconstruction of the 3D scene to be reconstructed can be divided into several parts. First, multimodal data collected by the acquisition device can be used to roughly reconstruct a color point cloud map of the surrounding environment and output the device path, providing accurate geometric priors for Gaussian mixture unified reconstruction. Second, semantic instance segmentation is performed based on the input scene image, point cloud map data, and device path, providing spatially consistent semantic instance prior information for 3D Gaussian reconstruction.

[0173] Furthermore, in conjunction with the aforementioned Figure 3 The architecture, for example, the complete architecture can be as follows Figure 4 As shown, the overall process can include the following steps.

[0174] First, user input is collected.

[0175] Users can collect raw data, i.e., the aforementioned sensory data, using robots, backpack-mounted professional data acquisition devices, and handheld data acquisition devices. Typically, the collected raw data includes at least hardware-synchronized RGB images and depth maps or point cloud data. A wider field of view can also be covered by combining multiple sensors. The intrinsic or extrinsic parameters of each sensor are recorded simultaneously. Furthermore, the acquisition device can be equipped with low-power computing and storage units to perform lightweight preprocessing of the raw data before storing it in a standardized format. Users can retrieve the collected raw data by copying or uploading it over a network.

[0176] A rough reconstruction was then carried out.

[0177] Coarse reconstruction can be performed either on the acquisition device or on the server. Specifically, based on the acquired scene images and depth information, coarse reconstruction of geometric elements can be performed using less computational power compared to 3D Gaussian reconstruction. This includes mapping the scene image into 3D space based on the depth map when the depth information is available, obtaining point cloud map data; and then dividing the point cloud map data into voxel-based spatial partitions to obtain an estimate of the voxel grid occupancy. Furthermore, based on the grid occupancy estimation results, online guidance can be provided during the data acquisition process, guiding users to scan uncovered areas, thereby significantly improving acquisition coverage and efficiency.

[0178] Then, a global optimization was performed.

[0179] Unlike the coarse reconstruction step, which has high real-time requirements, the global optimization step aims to achieve higher-precision reconstruction by leveraging cloud computing resources. Specifically, firstly, multiple point cloud data are stitched together using the initial input device path to obtain point cloud map data representing the complete scene. Next, planar features in the point cloud map, namely the aforementioned first planar features, are extracted using adaptive voxelization. Based on the co-view relationship of the same plane across multiple frames, a plane-based bundle adjustment (PBA) error is constructed. After nonlinear optimization, more accurate point cloud map data and device paths can be obtained.

[0180] Then semantic instance segmentation can be performed.

[0181] Semantic instance segmentation can take higher-precision point cloud data, device paths, and scene images as input, and output 3D semantic instance segmentation results and spatially consistent 2D semantic instance segmentation results. Typically, 2D semantic segmentation and instance segmentation may suffer from poor viewing angles and foreground / background occlusion, which may not guarantee the consistency of the same instance across multiple views after mapping to 3D space. Therefore, in this embodiment, planar clustering of the 3D point cloud is performed to obtain the separated foreground and background structures; 3D projection depth is introduced to improve the accuracy of 2D segmentation in complex environments such as foreground / background occlusion. Furthermore, a graph voting method is used to vote and associate the foreground and background results separately, thereby obtaining spatially consistent 2D semantic instance segmentation results and outputting the corresponding 3D segmentation results.

[0182] Then, a three-dimensional Gaussian reconstruction was performed.

[0183] Based on the higher-precision point cloud data, device path, 3D semantic instance segmentation results, and spatially consistent 2D semantic instance segmentation results output in the preceding steps, 3D Gaussian reconstruction is performed, and the reconstructed data is output. Specifically, the higher-precision point cloud data, device path, 3D semantic instance segmentation results, and spatially consistent 2D semantic instance segmentation results output in the preceding steps can be used as initialization parameters for the position, shape, color, and semantic instances of the 3D Gaussian sphere. The geometric prior information of the input precise point cloud can be used to supervise operations such as densification or deletion of the 3D Gaussian sphere. Specifically, a 2D Gaussian plane can be used to represent Gaussian regions in planar areas, and constraints such as pose, orientation, and shape can be introduced to control the training of 2DGS; a 3D Gaussian sphere can be used to represent Gaussian regions in non-planar areas, and normal consistency constraints or depth constraints can be introduced to accurately capture complex geometric information. Under the premise of ensuring accurate geometric supervision, the RGB image set and the semantic instance image set jointly use the 3D Gaussian rendering pipeline to complete the encoding training of 3D Gaussian color and semantic instances. Subsequently, the output 3D Gaussian model can be directly mapped to generate a 3D mesh model. The 3D Gaussian model and the mesh model with spatial mapping relationship together serve as the product of the reconstruction service, providing downstream applications with data derived from real-world reconstruction.

[0184] In downstream tasks, such as in embodied data generation and embodied simulation training applications, the 3D Gaussian and Mesh models output after the aforementioned reconstruction process serve as reconstruction models from the real world. They can directly act as real reconstruction assets, serving as seeds for data augmentation, thereby generating data for training large embodied perception models. They can also be imported into simulation platforms, allowing robots to conduct skill training in the reconstructed real-world environment, thus reducing the gap between simulation and reality during actual robot deployment.

[0185] The following section will provide a more detailed explanation of each step.

[0186] I. Rough Reconstruction

[0187] First, the preliminary reconstruction step can be deployed on either the data acquisition device or a server. The data acquisition device can specifically include...

[0188] For example, the coarse reconstruction process can be as follows: Figure 5 As shown.

[0189] The rough reconstruction process may include:

[0190] 501. Feature Extraction.

[0191] The data collected by the acquisition device may include scene images, or RGB images and depth information.

[0192] When depth information includes a depth map, visual features can be extracted from RGB images. Specifically, CNNs can be used to extract visual features, as can Oriented Fast and Rotated BRIEF (ORB) descriptor generation schemes based on keypoint detection, depending on the application scenario. Simultaneously, by querying the corresponding depth value in the depth map using pixel coordinates, and based on the depth value and the camera model corresponding to the sensor that acquired the RGB image, the visual features are mapped into space, reconstructing the corresponding 3D point cloud. This 3D point cloud can be considered as a 3D point cloud with ORB descriptors and colors.

[0193] When depth information includes radar point cloud data, such as for the LVIO scheme, the laser point cloud in each frame is first downsampled, also known as downsampling. For example, 1 / 3 of the points are taken by downsampling for subsequent association. To save computation, the image feature extraction can be performed without neural network. Instead, the RGB image is divided into multiple pixel blocks, and the error is calculated by pixel blocks, thus outputting the downsampled point cloud data.

[0194] 502. Local treatment.

[0195] To improve the output stability and accuracy of coarse reconstruction, point cloud data can be divided into local point cloud data with smaller granularity voxels.

[0196] Specifically, the point cloud data output in step 501 can be divided into multiple voxels of a pre-set size, such as 0.5m×0.5m×0.5m or other sizes. Then, based on the flatness of the point cloud within the voxel, further mesh cutting can be performed as needed. For example, if the average flatness of the plane within the voxel is greater than a preset value, the voxel can be further split to obtain more fine-grained voxels.

[0197] For example, for RGB-D cameras, the current feature point cloud frame can be transformed into the coordinate system of a local map based on a uniform motion model. The distance from a point to a surface and the descriptor distance are used to jointly establish association pairs for the point cloud features. For LIVO sensor combinations, the IMU's trajectory can be used to calculate the pose of the frame corresponding to the time in the radar point cloud data. Based on this pose, the point cloud data is transformed into a local map, and pixel blocks from points to surfaces are associated. Specifically, the distance from a point to a plane can be calculated, and points whose distance is not greater than a distance threshold are associated with surfaces. This distance threshold can be used to control the noise in the association process.

[0198] 503. Separation of static and dynamic states.

[0199] After obtaining local point cloud data containing feature association pairs, dynamic and static targets in the scene can be separated based on this local point cloud data.

[0200] Generally, static targets are more reliable for determining the device path of a data acquisition device. However, in real-world scenarios, the presence of dynamic objects in the environment leads to outliers in the association pairs. Directly substituting both outliers and inliers into subsequent algorithms, such as least squares optimization, will result in reduced accuracy and stability. Therefore, in this embodiment, dynamic and static targets are separated to allow for more accurate calculation of the device path of the data acquisition device using static targets.

[0201] For example, when dealing with dynamic-static separation, assuming that outliers are few, that is, there are few dynamic elements in the field of view, the random sample consensus (RANSAC) method is used to randomly sample all association pairs, then select the correct inliers, and dynamically label the corresponding voxels according to the voxels where the outliers are located, so as to achieve path determination.

[0202] 504. Path determined.

[0203] Based on the incoming values ​​separated in step 503 above, that is, the points of the static target, the positional change of the acquisition device can be determined by using the point-to-surface error. Based on the positional change of the acquisition device, the device path of the acquisition device can be determined.

[0204] For RGB-D cameras, which typically have a single measurement source, a pre-trained inliers point-to-area error model can be directly applied for registration of point cloud frames and local maps. Based on the registration results, feature association is performed again for registration. After multiple iterations of association and registration, pose optimization converges, thus outputting the path of the acquisition device.

[0205] For the LIVO sensor suite, which includes multiple sensors, an Error-State Iterative Kalman Filter (ESIKF) can be applied to fuse multi-source data, along with the IMU's high-frequency angular velocity and acceleration for trajectory estimation. Then, the aforementioned point-to-surface feature association is used to update the state of the inliers observed by laser. Subsequently, the camera's pose on the map is predicted based on pre-calibrated extrinsic parameters, and visible voxels are extracted from the local point cloud data based on this pose. Furthermore, based on the projection of the 3D points of the observed pixel blocks carried by the visible voxels onto the camera, a photometric error is established, the camera observation state is updated, and the device path of the acquisition equipment is output.

[0206] 505. Online guidance.

[0207] In the aforementioned coarse reconstruction process, a voxel map is used as the carrying structure of the local map. Based on the updated pose, ray-tracing can be applied to identify voxels within the current camera's field of view that are visible but lack laser points or have few laser points, i.e., voxels with missing data. When visualizing voxels, these voxels can be set as voxels to be scanned further. Based on the degree of aggregation of this type of voxel, scanning guidance can be provided, allowing the user to rescan these voxels and thus collect more complete data.

[0208] Therefore, in this embodiment, a rough geometric model is first performed to provide geometric priors for subsequent more refined modeling, thereby improving the accuracy of combining the three-dimensional Gaussian model obtained from the subsequent scene reconstruction.

[0209] II. Global Optimization

[0210] After the aforementioned coarse reconstruction, higher computing power can be used to globally optimize the coarse reconstruction data, so as to output higher precision point cloud data and the corresponding device paths of the acquisition equipment.

[0211] For example, the global optimization process can be as follows: Figure 6 As shown, it can be divided into point cloud data and device path optimization, and device path re-optimization. Specifically, it includes the following steps:

[0212] 601. Adaptive point cloud voxelization.

[0213] Specifically, during the global optimization phase, a voxel map of a specified size can be constructed. Based on the input device path, the point cloud data of each frame is transformed using the device path and then added to the voxel map. In other words, the aforementioned output local point cloud data is mapped to the voxel map based on the device path. For each point, voxel coordinates are calculated and added to the corresponding voxel.

[0214] Furthermore, to reduce computational complexity, local point cloud data can be mapped to a voxel map based on keyframes by extracting keyframes.

[0215] 602. Planar voxel extraction.

[0216] For point cloud data mapped to individual voxels, planar extraction can be performed for each voxel. Planar fitting calculations are then performed on the point cloud within each voxel. Based on the ratio between the minimum eigenvalue and the other two eigenvalues, it is determined whether the voxel meets the required planar feature, i.e., the aforementioned first planar feature. If so, the parameters of this planar feature are added to the list of voxels to be optimized, i.e., it is output as a voxel. If not, the voxel is further segmented at a lower resolution, and the determination of whether a planar feature meets the requirements is made within the smaller voxels, segmenting within the constraint of a preset maximum number of layers.

[0217] 603. Optimize point cloud data and device paths using PBA.

[0218] Typically, device paths can be represented by radar paths or camera paths, etc.

[0219] Specifically, based on the list of planar voxels, for each voxel, the timestamp of the device pose falling within that voxel is determined, establishing a correlation between each voxel and all observed poses. A nonlinear optimization problem is constructed based on the initial device path and the correlations between all planar voxels and all device paths. The minimum eigenvalue of the plane within the voxel is minimized by adjusting the device path. Finally, the optimized device path and point cloud map are obtained.

[0220] 604. Point cloud post-processing.

[0221] Point cloud maps store point clouds within each sub-region of space, and whether a point cloud is planar can be determined based on calculated planar fitting parameters. However, point clouds from radar scans or RGB-D cameras often contain varying degrees of noise, resulting in significant noise within planar voxels. For example, measurement noise may be particularly prominent at object edges. This measurement noise causes optimized point cloud maps to still suffer from problems such as excessively thick planar features and unclear object boundaries. Therefore, in this embodiment, voxel mapping can effectively solve this problem. For voxels with fewer than a preset threshold, noise can be removed by culling. For planar voxels, RANSAC planar fitting can be used to remove outliers, reducing the planar point cloud thickness to a specified parameter, such as 5mm.

[0222] Step 604 is optional. Whether or not to execute step 604 can be chosen based on the actual application scenario or computing power, and this application does not impose any restrictions on it.

[0223] 605. Further optimize equipment routing.

[0224] Specifically, based on the optimized device path and pre-calibrated extrinsic parameters, an initial device path in the optimized point cloud data can be obtained using linear interpolation. To further improve the accuracy of the obtained device path, it can be optimized again. Based on the initial device path obtained by linear interpolation, a many-to-one association can be established between the device poses of multiple frames and the commonly observed features on the same plane in the point cloud data. For each planar feature, a bundle adjustment based on photometric error is established using the many-to-one association. This enhances the projection consistency between image observation and laser point cloud, resulting in a better device path on a more accurate point cloud map. Based on the more accurate device path, high-precision coloring of the voxel point cloud map can be performed, providing better initial RGB values ​​for subsequent 3DGS training.

[0225] Therefore, in this embodiment, higher computing power can be used to globally optimize the coarsely reconstructed data, thereby obtaining more precise point cloud data and more accurate device paths.

[0226] For example, a comparison of the effects of coarse reconstruction and global optimization can be made as follows: Figure 7 As shown, compared to traditional methods, coarse reconstruction requires less computation and offers higher real-time performance, but its accuracy is insufficient. For example, it may suffer from significant point cloud noise and reconstruction errors, such as ghosting of walls in practical reconstruction scenarios. Global optimization, while computationally intensive and unable to achieve real-time reconstruction, offers higher accuracy, reducing point cloud noise and removing wall ghosting. Therefore, coarse reconstruction can be used to quickly reconstruct geometric elements, providing geometric priors for subsequent reconstructions and improving the accuracy of scene reconstruction.

[0227] III. Semantic Instance Segmentation

[0228] In the semantic instance segmentation step, semantic instance segmentation is performed based on the input scene images, point cloud data, and device paths to segment the objects included in the 3D scene to be reconstructed and their corresponding semantics. Furthermore, the semantic instance information from the point cloud map is used to correct potentially incorrect 2D semantic instance segmentations from certain viewpoints. In this step, on the one hand, the 3D semantic instance point cloud map can be used as a reconstruction product to accumulate scene and object assets; on the other hand, the semantic instance information on the 3D point cloud, along with its color and geometric voxel information, can be used as the initialization for a 3D Gaussian algorithm. The corrected 2D semantic instance segmentation image set is then used to supervise the training of the 3D Gaussian algorithm, resulting in semantic instance information optimized on the 3D Gaussian algorithm.

[0229] For example, the semantic instance segmentation process can be as follows: Figure 8 As shown.

[0230] 801. 2D semantic instance segmentation.

[0231] It can perform semantic instance segmentation on the input scene image and output 2D semantics and 2D instance mask.

[0232] Optionally, a pre-trained semantic instance segmentation model can be used for 2D semantic instance segmentation. For example, Grounded-SAM can be used to detect and segment images to obtain 2D semantics and instance masks. Grounded-SAM is composed of an open-set object detector (Grounding DINO) and a cue-enabled segmentation model (SAM), which can perform semantic instance segmentation of images captured in any scene based on text prompts, without the need to provide training samples in advance.

[0233] 802, 3D geometric clustering.

[0234] Specifically, geometric clustering can be performed on the input point cloud data to divide the point cloud into 3D instances.

[0235] For example, in the aforementioned global optimization step, the 3D space was divided using voxel structures. For the point cloud within each voxel, planar features (i.e., the aforementioned second planar features) were used for further hierarchical division, thus obtaining planar parameters such as the plane center point and plane normal vector. For instance, region growing can be used to extract planes from the reconstructed scene. The extracted planes already exist as instances (but their semantics are unknown), reducing computational load in subsequent association processes. Furthermore, this makes the association algorithm more robust to inaccurate edge masks output by Grounded-SAM.

[0236] 803, 3D-2D instance association.

[0237] Subsequently, the 2D semantics are associated with the mask and 3D instances, that is, the 3D instances belonging to the same instance are associated with the 2D semantics and the mask.

[0238] Specifically, semantic instance information of a 3D element can be obtained by voting on its semantic instance information through multiple observations of the same 3D element from multiple perspectives.

[0239] Typically, various problems can arise in 2D semantic and mask segmentation. For example, inconsistent semantic instance segmentation can occur due to changes in viewing perspective. In the right-hand view, Grounded-SAM correctly identifies the sofa, ground, and table masks; however, in the left-hand view, Grounded-SAM incorrectly identifies the sofa and part of the ground as a cart, leading to inconsistencies across multiple views. Furthermore, inaccurate semantic instance segmentation can be caused by foreground and background occlusion. For instance, Grounded-SAM incorrectly segments part of the background wall structure and the foreground trees as a single instance, introducing incorrect information into subsequent 3D-2D instance association. Therefore, this embodiment utilizes 3D instances to correct 2D instances.

[0240] In real-world scenarios, besides the inaccurate segmentation caused by Grounded-SAM, path errors in the point cloud map can also lead to inaccurate 3D-to-2D projection relationships, interfering with association voting. Therefore, to facilitate the establishment of 3D-to-2D projection relationships, NKSR can be used to convert the point cloud into a mesh, and the mesh's rendering pipeline can be used to establish a connection with the 2D image.

[0241] Specifically, 3D foreground-background separation can be performed first. This can include extracting areas with an area greater than a threshold (e.g., set to 5m) based on the results of 3D geometric clustering. 2 The plane is used as the background, and the remaining structure is used as the foreground.

[0242] Subsequently, the 3D projection depth map corrects the 2D instance mask: using the pose corresponding to each camera image, the foreground mesh is rendered to obtain the depth map corresponding to that camera image, and the depth map and the 2D mask are pixel-wise associated. For each mask, the depth can be used to cluster all pixels within the mask using region growing. Pixels with a depth difference not exceeding a threshold (e.g., 0.5m) are considered to be in the same class. Finally, the class with more pixel values ​​is selected to correct the mask.

[0243] Subsequently, 3D-2D instance association is performed based on graph voting. Foreground and background structures can be voted on separately to reduce interference between foreground and background due to association errors. For all unassociated mesh patches, their center points can be used as representatives of that mesh patch. For each frame of scene image, the corresponding pose can be used for rendering projection. If an instance mask contains the center points of two mesh patches, the instance association score of the edge formed by their center points is incremented by 1. By traversing all semantic segmentation image sets to vote on the association of patch center points, a graph structure can be constructed. Traversing this graph structure, center points with scores greater than a threshold (e.g., set to 3, indicating that they belong to the same instance in at least three viewpoints) are treated as instances, completing the instance association.

[0244] 804, 2D semantic instance correction.

[0245] In the aforementioned graph voting, instance masks were associated, but the semantic information of the same instance still differs across different perspectives. Assuming that the semantic information of the same instance is correct from most perspectives, the semantic information of the instance can be obtained by taking the mode. Furthermore, the semantic observations of the same instance from different perspectives can be calibrated using the 3D-to-2D projection relationship. If the semantic segmentation from a certain perspective does not match the semantics of the instance, the semantic segmentation from that perspective can be corrected to reflect the instance's semantics, thereby achieving spatial consistency of semantic instance information.

[0246] Then, 3D semantics and instances with spatial consistency can be output as well as 2D semantics and instances.

[0247] Therefore, in this embodiment, semantic instance segmentation of the point cloud map is performed based on the input scene image, point cloud data, and device path, providing spatially consistent semantic instance prior information for Gaussian mixture unified reconstruction. Addressing the common problems of imprecise masks and inconsistent multi-view observations in complex environments, such as those prevalent in 2D open-set semantic instance segmentation networks, a method is proposed to introduce 3D geometric information to automatically correct errors and improve spatial consistency. For example, 1. Depth prior information is introduced to improve the accuracy of 2D masks in complex environments with foreground and background occlusion; a 3D foreground / background separation mechanism is introduced to reduce the impact of 3D-2D incorrect projection relationships caused by path errors; and a graph-based voting mechanism is introduced to associate the foreground and background separately, obtaining a spatially consistent semantic instance segmentation mask and outputting the corresponding 3D segmentation results.

[0248] IV. 3D Gaussian Reconstruction

[0249] The input for 3D Gaussian reconstruction is the result of the global optimization step and the semantic instance segmentation step, and the output is a 3D Gaussian sphere model that has undergone 3D Gaussian reconstruction.

[0250] Specifically, a unified representation of the 3D scene to be reconstructed can be constructed based on the input data. This representation enables the reconstructed product to exhibit three major characteristics: visual realism, physical accuracy, and semantic richness. The unified representation of the 3D scene to be reconstructed uses a geometrically accurate Gaussian kernel as its carrier, and renderable texture features and semantic instance features are attached to it. This allows users to simultaneously extract high-quality 3D digital assets (such as mesh models, dense point clouds, etc.) and 2D digital assets (such as RGB images, semantic graphs, instance graphs, depth maps, etc.) from this representation.

[0251] In the process of performing 3D Gaussian reconstruction, such as Figure 9 As shown, reconstruction can be divided into multiple dimensions. For example, it can be divided into geometric, visual, and physical dimensions. Correspondingly, the reconstructed attributes can also be divided into geometric attribute parameters, visual attribute parameters, and semantic attribute parameters.

[0252] For example, the geometric attribute parameters may specifically include, but are not limited to, one or more of the following:

[0253] Location: also known as mean, represents the coordinates of the center of the Gaussian kernel in three-dimensional space;

[0254] Covariance: It can represent the shape distribution of the Gaussian kernel. The three column vectors in the covariance matrix represent the three principal axis directions of the Gaussian ellipsoid.

[0255] Scaling factor: Represents the size of each Gaussian kernel;

[0256] Ellipsoid property: Used to classify the planar properties of the Gaussian kernel. If the Gaussian kernel is located in a planar region, a 2D Gaussian kernel is used to represent the planar geometry, with its scaling factor in the normal direction approaching 0, resulting in an elliptical shape. If the Gaussian kernel is located in a non-planar region, a 3D Gaussian kernel is used to represent the solid geometry.

[0257] Visual attribute parameters may include, but are not limited to, one or more of the following:

[0258] Opacity: Represents the transparency of the Gaussian kernel. Higher opacity indicates that the Gaussian kernel is more likely to lie on the surface of an entity in the scene. During rendering, the maximum integral of the opacity values ​​of all Gaussian kernels along a ray is 1.

[0259] Spherical harmonic coefficient (SH) parameter: Encodes lighting information from different viewpoints and can reflect the changes in light and shadow within the 3D scene to be reconstructed.

[0260] Semantic attribute parameters may include:

[0261] Semantic encoding: used to encode the semantic attributes corresponding to the Gaussian kernel, representing the type of object (e.g., table, chair, ground, etc.). It is usually a set of high-dimensional feature vectors that can be decoded into semantic label values.

[0262] Instance encoding: Used to encode the instance attributes corresponding to the Gaussian kernel. For example, the corresponding attributes can be set as table1, table2, etc. It is usually a set of high-dimensional feature vectors that can be decoded into instance ID values ​​to identify the specific attributes of the instance.

[0263] Accordingly, 3D Gaussian reconstruction can be divided into geometric reconstruction, visual reconstruction, and semantic encoding, which will be introduced by example below.

[0264] (1) Geometric Reconstruction

[0265] The goal of geometric reconstruction is to optimize the geometric distribution of the Gaussian mixture representation so that it is highly consistent with the geometric features of the real scene.

[0266] Hybrid Gaussian representation uses Gaussian kernels with different geometric properties to represent the 3D scene to be reconstructed, employing 2D Gaussian to represent planar regions and 3D Gaussian to represent non-planar regions.

[0267] The geometric prior information of the scene comes from the high-quality voxel point cloud data output during the global optimization phase. This point cloud data provides the planar classification information (planar / non-planar) for each voxel mesh, as well as basic geometric attributes such as point cloud position, point cloud distribution, and normal vectors within the mesh. During the initialization phase of the Gaussian mixture representation, this information is used to construct the initial values ​​of the geometric attributes of the Gaussian kernel. Specifically, the planar classification information corresponds to the ellipsoidal attribute of the Gaussian kernel (including 2D / 3D Gaussians), which is usually a constant value and does not change during the scene optimization process. Additionally, the point cloud position information corresponds to the initial position value of the Gaussian kernel, and the point cloud distribution information and normal vectors correspond to the initial values ​​of the covariance and scaling factor of the Gaussian kernel. When initializing the covariance matrix, the three column vectors are arranged in ascending order of scaling factor, with the column vector with the smallest scaling factor corresponding to the normal vector direction of the voxel mesh.

[0268] During the training process of Gaussian mixture representation, constraints need to be imposed on the position, orientation, shape, and other dimensions of the Gaussian kernel. The geometric distribution of the Gaussian kernel is controlled by setting corresponding constraints.

[0269] For a 2D Gaussian in a plane, a planar voxel mesh is further used as geometric prior information to supervise its position, orientation, and shape, so that its distribution is completely consistent with the plane based on the voxel mesh fitting.

[0270] For example, it could include position constraints: constraining the position x of the center of the Gaussian ellipse. iThis constraint is limited to the current fitting plane by minimizing the distance from each Gaussian ellipse to the fitting plane (the planar distribution parameters of points within the voxel). This constraint can be expressed as:

[0271]

[0272] Attitude constraint: Constrain the rotation direction q of the Gaussian ellipse. rot This restricts the rotation to be around the normal direction, minimizing the Euler angles of the Gaussian ellipse rotating about the other two principal axes, as shown below:

[0273] q rot =q0+iq1+jq2+kq3

[0274]

[0275] Shape constraint: Minimize the scaling factor s0 of the Gaussian kernel along the normal direction, making the shape of the Gaussian kernel approximate a 2D ellipse. For example:

[0276]

[0277] For 3D Gaussian regions, the fine geometric structure is often difficult to capture directly due to the centimeter-level measurement errors commonly found in sensors such as radar and depth cameras. Therefore, the geometric constraints of non-planar regions need to be independent of those of planar regions to avoid introducing significant noise into the reconstruction results.

[0278] In the absence of detailed geometric priors, in addition to continuing to use the RGB constraints in the existing 3D Gaussian scheme to implicitly optimize non-planar regions, a self-supervised form of normal constraint is introduced. This constraint can include: given a training viewpoint, calculating the normal map estimated from the current pose distribution of the Gaussian mixture and the gradient map obtained by differentiating the depth map, and minimizing the residual between the two. When all Gaussian kernel pose distributions within the scene are consistent with the actual scene surface shape, the derivatives of the normal map and the depth map are equal.

[0279] Specifically, it includes:

[0280] The method of estimating 2D normal maps from 3D Gaussian remains consistent with RGB rasterization rendering:

[0281]

[0282] The depth map differential is calculated using the depth difference between the current pixel and its neighboring pixels.

[0283]

[0284] Based on the above expression, construct the normal consistency constraint:

[0285]

[0286] Furthermore, if the original data already provides dense depth prior information (such as a depth map obtained from a depth camera scan), then the following depth constraints can be set to directly supervise the distribution of the 3D Gaussian in non-planar regions:

[0287]

[0288] Furthermore, the 3D Gaussian pipeline also provides a control strategy for adaptive copying, splitting, merging, or deleting operations based on the current size, shape, or opacity of the Gaussian kernel. In the embodiments of this application, an adaptive global strategy for 3D Gaussian reconstruction is set, which may include:

[0289] Denseization strategy: Operations such as copying and splitting Gaussian spheres must be performed within a single voxel. The shape of the newly generated Gaussian kernel cannot exceed the boundary of the corresponding voxel mesh;

[0290] Deletion strategy: Deletion operation is performed on Gaussian kernels whose position and shape exceed a certain threshold of the voxel boundary.

[0291] Thus, a better three-dimensional Gaussian model can be obtained through reconstruction strategy.

[0292] (2) Visual reconstruction

[0293] For visual dimension reconstruction, it can be based on scene images or features extracted from scene images.

[0294] Typically, pre-trained encoders and decoders can be used for 3D Gaussian reconstruction. For example, the lighting and color changes that the viewpoint usually depends on can be encoded. By using weighted summation of low-order basis functions to achieve dynamic lighting simulation, the visual attribute parameters of the 3D Gaussian model, such as color under different viewpoints and lighting conditions, can be calculated.

[0295] Furthermore, when training Gaussian mixture representations, the input images need to possess high inter-frame consistency, meaning that the color and texture of the same region in a scene observed from different viewpoints should be highly consistent. However, due to the complex and varied lighting conditions in real-world scenes, and the automatic adjustment of exposure and other signals by the camera during acquisition, inter-frame consistency is often difficult to achieve. If original images with poor inter-frame consistency are used for training Gaussian mixtures, the algorithm's output will typically tend to fit a large number of false points in space due to excessive image noise, severely impacting the reconstruction quality.

[0296] Therefore, in this embodiment, a differentiable bilateral grid structure is set to decouple the image signal processing (ISP) information implicit in each image. This bilateral grid corresponds one-to-one with the input image and is optimized along with the entire scene during the Gaussian mixture representation training process.

[0297] During model training, given the pose information of the input image, the Gaussian mixture representation (GaMJ) renders an image based on that pose. After processing with a bilateral mesh, this image can be restored to have the same color and texture as the input image, which is then used to calculate the photometric loss function. In the inference phase after training, all bilateral mesh information can be removed. When the user renders an image from any new viewpoint in the scene, the result will be one with highly consistent texture. This approach significantly improves the visual quality when Gaussian mixture representation is applied to real-world scene reconstruction.

[0298] (3) Semantic reconstruction

[0299] Existing 3D Gaussian training and rendering pipelines only include the single attribute of visual texture features. However, for downstream tasks, such as in the application scenarios of embodied intelligent devices, embodied intelligent agents, in addition to visual perception requirements, also need spatial understanding of the 3D scene to be reconstructed, making downstream task execution possible. Therefore, in this embodiment, the product of 3D reconstruction possesses high-dimensional spatial understanding information, specifically including the semantic features of the scene (i.e., the category attributes of objects or regions, such as tables, chairs, etc.) and instance features (i.e., the unique ID label corresponding to each object or region, such as table 1, table 2, etc.). In this embodiment, for Gaussian mixture representation, a differentiable semantic / instance feature encoding scheme that can be stored in each Gaussian kernel is provided. This semantic / instance feature can be input into the original rasterization module of 3D Gaussian, rendered in the same way as color attributes, generating a semantic / instance feature image, and compared with the ground truth to calculate the semantic / instance residual, thereby supervising the training and optimization of the semantic / instance features.

[0300] Specifically, the semantic attributes attached to each Gaussian kernel can be encoded into a set of high-dimensional feature vectors s. i Its dimension is determined by the number of object types included in the scene. The semantic feature vectors activated by the softmax function (whose meaning is typically the probability value of each semantic label) are input into a Gaussian mixture rasterization module for rendering, resulting in a semantic probability distribution map in 2D camera space, as shown below:

[0301]

[0302] For instance branches, a similar approach is used to encode the instance attributes attached to each Gaussian kernel into a set of high-dimensional feature vectors l. iIts dimension is determined by the number of all objects included in the scene. Compared to semantic branches, since the dimension of instance feature vectors is usually much higher, directly performing rasterization rendering after softmax activation may cause considerable memory overhead. Therefore, this embodiment of the application can introduce a differentiable small MLP network h. i The instance feature vectors are decoded to obtain the probability vector corresponding to each instance ID. Then, rasterization rendering is performed to obtain the instance probability distribution map in the 2D camera space, as shown below:

[0303]

[0304] During the training phase, the semantic and instance branch loss functions share the same design form, for each training perspective's semantic / instance probability distribution I. S / L The semantic / instance 2D mask obtained using the semantic instance segmentation module. Supervised analysis is performed, and the cross-entropy loss is calculated using the following formula to obtain the semantic instance residuals. Then, the semantic instance feature vectors within the Gaussian kernel are optimized using gradient backpropagation, as shown below:

[0305]

[0306] Therefore, in this embodiment, the geometric coarse reconstruction result and the semantic instance segmentation result are used as inputs to construct a unified representation of the 3D scene to be reconstructed, so that the reconstruction product exhibits three major characteristics: visual realism, physical accuracy, and semantic richness. This unified representation of the 3D scene to be reconstructed uses a geometrically precisely distributed Gaussian kernel as a carrier, on which renderable texture features and semantic instance features are attached, enabling users to extract high-quality 2D / 3D digital assets simultaneously from this representation. Specifically, a 2D / 3D hybrid Gaussian representation is provided to achieve visual realism based on precise geometry, while introducing high-dimensional perceptual information such as semantic instances, supporting the generation of first-person perspective temporal data for simulating various robot forms and downstream applications such as mobile navigation and operation within the simulator. The quality of both visual rendering and geometric reconstruction surpasses the state-of-the-art (SOTA) methods in academia. In the visual part, the PSNR index is improved by 30%, and in the geometric part, the planar accuracy in large scenes is improved by approximately 20 times, such as reducing the plane thickness from 11.7cm to 5mm, significantly improving the accuracy of 3D scene reconstruction.

[0307] For example, the rendered image corresponding to the model reconstructed through the aforementioned process can be as follows: Figure 11 As shown in the rendering example, very good reconstruction results can be achieved from different dimensions.

[0308] Furthermore, the reconstructed 3D Gaussian model can be applied to scenarios such as VR / AR, games, autonomous driving, and surveying. For example, in a surveying scenario, a data acquisition device can be used to collect perceptual data in the scene to be reconstructed. Using the method provided in this application, scene reconstruction can be performed based on the perceptual data, outputting a 3D scene model with very high visual realism and geometric accuracy.

[0309] The foregoing has described the method flow provided in the embodiments of this application. The following describes the structure of the apparatus for executing the foregoing method flow.

[0310] See Figure 10 The present application provides a schematic diagram of the structure of a three-dimensional scene reconstruction platform, including:

[0311] Input module 1001 is used to acquire sensor data;

[0312] The coarse reconstruction module 1002 is used to acquire the first reconstruction data, which is obtained by reconstructing the geometric elements in the three-dimensional scene based on the perception data.

[0313] The semantic instance segmentation module 1003 is used to obtain semantic instance segmentation results based on the first reconstruction data and perception data. The semantic instance segmentation results include at least one instance and the semantics of the at least one instance. Each instance in the at least one instance is used to indicate an object in the three-dimensional scene to be reconstructed.

[0314] The Gaussian reconstruction module 1004 is used to obtain second reconstruction data based on the perceptual data, semantic instance segmentation results and first reconstruction data. The second reconstruction data includes the data of the reconstructed three-dimensional Gaussian model.

[0315] In one possible implementation, the aforementioned perception data includes scene images and depth information. The coarse reconstruction module 1001 is specifically used to: obtain point cloud map data based on the scene image and a depth value aligned with the scene image; divide the point cloud map data into multiple local point cloud data; determine the device path based on the multiple local point cloud data, where the device path is the motion path of the acquisition device, and the first reconstruction data includes multiple local point cloud data and the device path.

[0316] In one possible implementation, the aforementioned apparatus further includes: a global optimization module 1005, configured to: project multiple local point cloud data onto multiple voxels according to the device path in the first reconstructed data to obtain multiple voxel point cloud data; extract planar features from the multiple voxel point cloud data to obtain multiple first planar features; obtain the correlation between each of the multiple first planar features and multiple observation poses, wherein the multiple observation poses are the poses of the virtual camera when observing each planar feature in three-dimensional space; adjust the multiple local point cloud data and the device path according to the correlation and the multiple voxel point cloud data to obtain adjusted multiple local point cloud data and adjusted device path; update the first reconstructed data with the adjusted multiple local point cloud data and adjusted device path to obtain third reconstructed data.

[0317] In one possible implementation, the aforementioned semantic instance segmentation module is specifically used for 1003: obtaining a two-dimensional instance segmentation result based on perceptual data, the two-dimensional instance segmentation result including at least one two-dimensional instance in two-dimensional space and the semantics of at least one two-dimensional instance; obtaining a three-dimensional instance segmentation result based on third reconstruction data, the three-dimensional instance segmentation result including at least one instance in three-dimensional space; aligning the two-dimensional instance segmentation result and the three-dimensional instance segmentation result to obtain an aligned three-dimensional instance segmentation result and a two-dimensional instance segmentation result, the instance segmentation result including the aligned three-dimensional instance segmentation result and the two-dimensional instance segmentation result.

[0318] In one possible implementation, the aforementioned semantic instance segmentation module 1003 is specifically used to: extract at least one planar feature from multiple local point cloud data included in the third reconstruction data to obtain at least one second planar feature; determine at least one three-dimensional instance based on the at least one planar feature, and the three-dimensional instance segmentation result includes at least one three-dimensional instance.

[0319] In one possible implementation, when the perceived data includes scene images, the aforementioned semantic instance segmentation module 1003 is specifically configured to: obtain a two-dimensional instance segmentation result based on the scene image; obtain a depth map corresponding to at least one two-dimensional instance based on at least one three-dimensional instance; correct the at least one two-dimensional instance based on the depth map to obtain a corrected at least one two-dimensional instance; associate the at least one three-dimensional instance with the at least one two-dimensional instance to obtain a graph structure, the graph structure being used to represent the association relationship between the aligned at least one three-dimensional instance and the at least one two-dimensional instance; update the semantics of the at least one two-dimensional instance based on the graph structure to obtain an updated semantics of the at least one two-dimensional instance, wherein the two-dimensional instance segmentation result includes the corrected at least one two-dimensional instance and the updated at least one two-dimensional instance, and the three-dimensional instance segmentation result includes at least one three-dimensional instance.

[0320] In one possible implementation, the aforementioned semantic instance segmentation module 1003 is specifically used to: obtain the observation results of the same three-dimensional instance from multiple perspectives based on the graph structure; update the semantics of at least one two-dimensional instance based on the observation results, and obtain the updated semantics of at least one two-dimensional instance.

[0321] In one possible implementation, the aforementioned semantic instance segmentation module 1003 is specifically used to: divide at least one three-dimensional instance into a foreground three-dimensional instance and a background three-dimensional instance based on at least one planar feature; and obtain a depth map corresponding to at least one two-dimensional instance based on the foreground three-dimensional instance.

[0322] In one possible implementation, the aforementioned Gaussian reconstruction module 1004 is specifically used to: determine Gaussian parameters based on perceptual data, semantic segmentation results, and third reconstruction data. The Gaussian parameters include geometric attribute parameters, visual attribute parameters, and semantic attribute parameters. The geometric attribute parameters are used to represent the geometric shape of the target in the three-dimensional scene, the visual attribute parameters are used to represent the parameters in the visual dimension of the three-dimensional scene, and the semantic attributes are used to represent the semantics of the target in the three-dimensional scene; and obtain second reconstruction data based on the Gaussian parameters.

[0323] In one possible implementation, the aforementioned Gaussian reconstruction module 1004 is specifically used to: construct at least one of two-dimensional Gaussian constraints or three-dimensional Gaussian constraints based on the third reconstruction data, wherein the two-dimensional Gaussian constraints are used to constrain the distance between points in the three-dimensional scene and the fitting plane, and the three-dimensional Gaussian constraints are used to constrain the distribution of three-dimensional instances in the three-dimensional scene; and determine geometric attribute parameters based on the perception data under the constraints of at least one of the two-dimensional Gaussian constraints or three-dimensional Gaussian constraints.

[0324] In one possible implementation, the aforementioned two-dimensional Gaussian constraint includes at least one of position constraint, attitude constraint, or shape constraint. The position constraint includes constraining the center position of the three-dimensional Gaussian sphere to be within the fitting plane. The attitude constraint includes constraining the rotation direction of the three-dimensional Gaussian sphere to be a rotation about the normal vector of the fitting plane. The shape constraint includes constraining the shape of the Gaussian kernel to conform to the constraint shape.

[0325] In one possible implementation, the aforementioned Gaussian reconstruction module 1004 is further configured to: decouple ISP information in the scene image through a pre-trained bilateral grid structure; and obtain color parameters based on the ISP information, wherein the aforementioned visual attribute parameters include color parameters.

[0326] This application also provides a computing device 100. For example... Figure 11As shown, the computing device 100 includes a bus 102, a processor 104, a memory 106, and a communication interface 108. The processor 104, the memory 106, and the communication interface 108 communicate with each other via the bus 102. The computing device 100 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 100.

[0327] Bus 102 can be a Peripheral Component Interconnect Express (PCIe) bus, an Extended Industry Standard Architecture (EISA) bus, a Unified Bus (Ubus or UB), a Compute Express Link (CXL), a Cache Coherent Interconnect for Accelerators (CCIX), etc. The Unified Bus is also known as the Lingqu Bus. Buses can be divided into address buses, data buses, control buses, etc. For ease of representation, Figure 11 The bus is represented by only one line, but this does not mean that there is only one bus or one type of bus. Bus 104 may include a path for transmitting information between various components of computing device 100 (e.g., memory 106, processor 104, communication interface 108). The unified bus may also be referred to as the Lingqu bus.

[0328] The processor 104 may include any one or more of the following computing devices: central processing unit (CPU), graphics processing unit (GPU), microprocessor (MP) or digital signal processor (DSP), ASIC, FPGA, CPLD, NPU, SoC, offload card, accelerator card, etc.

[0329] Memory 106 may include volatile memory, such as random access memory (RAM). Processor 104 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD). Furthermore, memory 106 may also be implemented using storage class memory (SCM), phase change memory (PCM), or other types of storage media.

[0330] It is worth noting that the same type of storage medium can be configured in the same computing device to realize the function of memory 106, or two or more types of storage media can be configured to realize the function of memory 106. This application does not limit this.

[0331] The memory 106 stores executable program code, and the processor 104 executes the executable program code to implement the functions of the aforementioned computing device, thereby implementing the processing steps in the method provided in this application. That is, the memory 106 stores instructions for executing the method provided in this application.

[0332] Alternatively, the memory 106 may store executable code, which the processor 104 executes to implement the functions of the computing device, etc., thereby implementing the method provided in this application. That is, the memory 106 stores instructions for executing the method provided in this application.

[0333] The communication interface 103 uses transceiver modules such as, but not limited to, network interface cards and transceivers to enable communication between the computing device 100 and other devices or communication networks.

[0334] As one possible implementation, the computing device 100 may also include a chip system, which includes a processor and a power supply circuit. The power supply circuit supplies power to the processor, and the processor executes the operation steps corresponding to the method provided in this application. For simplicity, further details are omitted here. The processor can be implemented using a GPU, or it can be implemented using computing devices or AI chips such as a DPU, NPU, XPU, SoC, offload card, or accelerator card.

[0335] As one possible implementation, the computing device 100 may include various types of processors 104, meaning the computing device 100 is a heterogeneous device. For example, the computing device 100 may include a CPU and a GPU, and at least one of the processors 104 may execute the operation steps corresponding to the method provided in this application. For the sake of brevity, further details are omitted here.

[0336] This application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0337] like Figure 12 As shown, the computing device cluster includes at least one computing device 100. The memory 106 of one or more computing devices 100 in the computing device cluster may store the same instructions for performing the methods provided in this application.

[0338] In some possible implementations, the memory 106 of one or more computing devices 100 in the computing device cluster may also store partial instructions for executing the methods provided in this application. In other words, a combination of one or more computing devices 100 can jointly execute the instructions for executing the methods provided in this application.

[0339] It should be noted that the memory 106 in different computing devices 100 within the computing device cluster can store different instructions, each used to execute a portion of the document navigation system's functions. That is, the instructions stored in the memory 106 of different computing devices 100 can implement the functions of different processing steps executed by the aforementioned computing devices.

[0340] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 13 One possible implementation is shown. For example... Figure 13 As shown, two computing devices 100A and 100B are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device. In this possible implementation, the memory 106 in computing device 100A stores instructions for performing a portion of the computing device's functions. Simultaneously, the memory 106 in computing device 100B stores instructions for performing another portion of the computing device's functions, thus enabling computing devices 100A and 100B to jointly implement the aforementioned computing device functions.

[0341] It should be understood that Figure 13The functions of the computing device 100A shown can also be performed by multiple computing devices 100. Similarly, the functions of the computing device 100B can also be performed by multiple computing devices 100.

[0342] Figure 13 The connection method between the computing device clusters shown can be based on the fact that the method provided in this application requires a large amount of computing power, needs to achieve load balancing, or requires a large amount of data storage. Therefore, different modules are deployed in different computing devices. For example, the implementation of a part of the functions of the computing device is performed by computing device 100A, and the implementation of another part of the functions is performed by computing device 100B.

[0343] This application also provides another computing device cluster. The connection relationships between the computing devices in this computing device cluster can be similarly referred to... Figure 12 and Figure 13 The connection method of the computing device cluster. The difference is that the memory 106 of one or more computing devices 100 in the computing device cluster can store the same instructions for executing the method provided in this application.

[0344] In some possible implementations, the memory 106 of one or more computing devices 100 in the computing device cluster may also store partial instructions for executing the methods provided in this application. In other words, a combination of one or more computing devices 100 can jointly execute the instructions for executing the methods provided in this application.

[0345] It should be noted that the memory 106 in different computing devices 100 within the computing device cluster can store different instructions for executing some functions of the document navigation system provided in this application. That is, the instructions stored in the memory 106 of different computing devices 100 can implement the functions described above for the computing devices.

[0346] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, etc., including several instructions to cause a data quantization device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0347] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.

[0348] This application also provides a computer-readable storage medium storing a program for executing the aforementioned method flow. When the program is run on a computer, it causes the computer to perform the aforementioned... Figures 2 to 9 All or part of the steps in the method described in the embodiments shown.

[0349] This application also provides a digital processing chip. This digital processing chip integrates circuitry for implementing the aforementioned processor or processor functions, and one or more interfaces. When the digital processing chip integrates a memory, it can perform the method steps of any one or more of the foregoing embodiments. When the digital processing chip does not integrate a memory, it can be connected to an external memory via a communication interface. The digital processing chip implements the method steps of any one or more of the foregoing embodiments based on the program code stored in the external memory.

[0350] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium capable of being stored by a computing device, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct a computing device to perform the method provided in this application.

[0351] This application also provides a computer program product comprising one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., a solid-state disk (SSD)).

[0352] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.

[0353] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. The term "and / or" in this application is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Additionally, the character " / " generally indicates that the preceding and following related objects are in an "or" relationship. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or modules is not necessarily limited to those steps or modules explicitly listed, but may include other steps or modules not explicitly listed or inherent to such processes, methods, products, or devices. The naming or numbering of steps in this application does not imply that the steps in the method flow must be executed in the time / logical order indicated by the naming or numbering. The execution order of the named or numbered process steps can be changed according to the technical purpose to be achieved, as long as the same or similar technical effect can be achieved. The division of modules in this application is a logical division. In actual applications, there may be other division methods. For example, multiple modules may be combined into or integrated into another system, or some features may be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the modules shown or discussed may be through some ports, and the indirect coupling or communication connection between modules may be electrical or other similar forms, which are not limited in this application. Furthermore, the modules or sub-modules described as separate components may or may not be physically separated, may or may not be physical modules, or may be distributed in multiple circuit modules. Some or all of the modules can be selected to achieve the purpose of the solution in this application according to actual needs.

Claims

1. A method for reconstructing a three-dimensional scene, characterized in that, The method is applied to a 3D reconstruction platform, which runs on an infrastructure including at least one computing node. The 3D reconstruction platform is communicatively connected to an acquisition device, which is used to acquire perceptual data of the 3D scene to be reconstructed. The method includes: Acquire the perceived data; Acquire first reconstruction data, which is obtained by reconstructing geometric elements in the three-dimensional scene to be reconstructed based on the perception data; Semantic instance segmentation results are obtained based on the first reconstructed data and the perceived data. The semantic instance segmentation results include at least one instance and the semantics of the at least one instance. Each instance in the at least one instance is used to indicate an object in the three-dimensional scene to be reconstructed. Based on the perceptual data, semantic instance segmentation results, and the first reconstruction data, a second reconstruction data is obtained, which includes a 3D model of the 3D scene to be reconstructed.

2. The method according to claim 1, characterized in that, The perceived data includes scene images and depth information, and the acquisition of the first reconstructed data includes: Align the scene image with the depth information to obtain a depth value aligned with the scene image; Point cloud map data is obtained based on the scene image and the depth value aligned with the scene image; The point cloud map data is divided into multiple local point cloud data, and the point cloud in each local point cloud data represents at least one geometric element. The device path is determined based on the multiple local point cloud data, and the device path is the movement path of the acquisition device. The first reconstructed data includes the multiple local point cloud data and the device path.

3. The method according to claim 2, characterized in that, The step of obtaining the first reconstructed data based on the plurality of local point cloud data and the device path includes: The multiple local point cloud data are projected onto multiple preset voxels according to the device path to obtain multiple voxel point cloud data; Extract planar features from the multiple voxel point cloud data to obtain multiple first planar features; Obtain the association between each of the plurality of first planar features and a plurality of observation poses, wherein the plurality of observation poses are the poses of the virtual camera when observing each planar feature in the three-dimensional space; Based on the aforementioned correlation and the aforementioned voxel point cloud data, the aforementioned local point cloud data and the aforementioned device path are adjusted to obtain the adjusted local point cloud data and the adjusted device path. Based on the adjusted local point cloud data and the adjusted device path, the first reconstruction data is updated to obtain the third reconstruction data.

4. The method according to claim 3, characterized in that, The step of obtaining semantic instance segmentation results based on the first reconstructed data and the perceived data includes: Two-dimensional instance segmentation results are obtained based on the perceived data, and the two-dimensional instance segmentation results include at least one two-dimensional instance in the two-dimensional space and the semantics of the at least one two-dimensional instance; A three-dimensional instance segmentation result is obtained based on the third reconstruction data, wherein the three-dimensional instance segmentation result includes at least one instance in three-dimensional space; Align the two-dimensional instance segmentation result with the three-dimensional instance segmentation result to obtain an aligned three-dimensional instance segmentation result and a two-dimensional instance segmentation result. The semantic instance segmentation result includes the aligned three-dimensional instance segmentation result and the two-dimensional instance segmentation result.

5. The method according to claim 4, characterized in that, The step of obtaining the 3D instance segmentation result based on the third reconstruction data includes: Extract planar features from the adjusted local point cloud data to obtain at least one second planar feature; At least one three-dimensional instance is determined based on the at least one second planar feature, and the three-dimensional instance segmentation result includes the at least one three-dimensional instance.

6. The method according to claim 5, characterized in that, When the perceived data includes scene images, obtaining the two-dimensional instance segmentation result based on the perceived data includes: The two-dimensional instance segmentation result is obtained based on the scene image; Aligning the two-dimensional instance segmentation result with the three-dimensional instance segmentation result to obtain the aligned three-dimensional instance segmentation result and the two-dimensional instance segmentation result includes: Obtain a depth map corresponding to the at least one two-dimensional instance based on the at least one three-dimensional instance; The at least one two-dimensional instance is corrected according to the depth map to obtain the at least one corrected two-dimensional instance; By associating the at least one three-dimensional instance with the at least one two-dimensional instance, a graph structure is obtained, which is used to represent the association relationship between the at least one three-dimensional instance and the at least one two-dimensional instance after alignment; The semantics of the at least one two-dimensional instance are updated according to the graph structure to obtain the updated semantics of the at least one two-dimensional instance. The two-dimensional instance segmentation result includes the semantics of the corrected at least one two-dimensional instance and the updated at least one two-dimensional instance.

7. The method according to claim 6, characterized in that, The step of updating the semantics of the at least one two-dimensional instance according to the graph structure to obtain the updated semantics of the at least one two-dimensional instance includes: Based on the graph structure, obtain the observation results of the same 3D instance from multiple perspectives; The semantics of the at least one two-dimensional instance are updated based on the observation results to obtain the updated semantics of the at least one two-dimensional instance.

8. The method according to claim 6 or 7, characterized in that, The step of obtaining a depth map corresponding to the at least one two-dimensional instance based on the at least one three-dimensional instance includes: Based on the at least one planar feature, the at least one three-dimensional instance is divided into a foreground three-dimensional instance and a background three-dimensional instance; Based on the foreground 3D instance, obtain a depth map corresponding to at least one 2D instance.

9. The method according to any one of claims 3-8, characterized in that, The step of reconstructing the second reconstructed data based on the perceived data, the semantic instance segmentation result, and the first reconstructed data includes: Gaussian parameters are determined based on the perceptual data, semantic segmentation results, and the third reconstruction data. The Gaussian parameters include geometric attribute parameters, visual attribute parameters, and semantic attribute parameters. The geometric attribute parameters are used to represent the geometric shape of the target in the three-dimensional scene, the visual attribute parameters are used to represent the parameters in the visual dimension of the three-dimensional scene, and the semantic attributes are used to represent the semantics of the target in the three-dimensional scene. The second reconstruction data is obtained based on the Gaussian parameters.

10. The method according to claim 9, characterized in that, The step of determining the Gaussian parameters based on the perceptual data, semantic segmentation results, and the third reconstructed data includes: Based on the third reconstruction data, at least one of two-dimensional Gaussian constraints or three-dimensional Gaussian constraints is constructed. The two-dimensional Gaussian constraints are used to constrain the distance between points in the three-dimensional scene and the fitting plane, and the three-dimensional Gaussian constraints are used to constrain the distribution of three-dimensional instances in the three-dimensional scene. The geometric attribute parameters are determined based on the perceived data under at least one of the two-dimensional Gaussian constraints or the three-dimensional Gaussian constraints.

11. The method according to claim 10, characterized in that, The two-dimensional Gaussian constraint includes at least one of position constraint, attitude constraint, or shape constraint. The position constraint includes constraining the center position of the three-dimensional Gaussian sphere to be within the fitting plane. The attitude constraint includes constraining the rotation direction of the three-dimensional Gaussian sphere to be a rotation about the normal vector of the fitting plane. The shape constraint includes constraining the shape of the Gaussian kernel to conform to the constraint shape.

12. The method according to any one of claims 9-11, characterized in that, The method further includes: By using a pre-trained bilateral grid structure, the image signal processing (ISP) information in the scene image is decoupled; The color parameters are obtained based on the ISP information, and the visual attribute parameters include the color parameters.

13. A three-dimensional scene reconstruction platform, characterized in that, The 3D scene reconstruction platform runs on an infrastructure that includes at least one computing node. The platform is communicatively connected to a data acquisition device, which collects perceptual data of the 3D scene to be reconstructed. The 3D scene reconstruction platform includes: The input module is used to acquire the perceived data; A coarse reconstruction module is used to acquire first reconstruction data, which is obtained by reconstructing geometric elements in the three-dimensional scene to be reconstructed based on the perception data. The semantic instance segmentation module is used to obtain a semantic instance segmentation result based on the first reconstructed data and the perceived data. The semantic instance segmentation result includes at least one instance and the semantics of the at least one instance. Each instance in the at least one instance is used to indicate an object in the three-dimensional scene to be reconstructed. The Gaussian reconstruction module is used to reconstruct the second reconstruction data based on the perceptual data, semantic instance segmentation results and the first reconstruction data, and the second reconstruction data includes the three-dimensional model of the three-dimensional scene to be reconstructed.

14. The three-dimensional scene reconstruction platform according to claim 13, characterized in that, The perceived data includes scene images and depth information. The coarse reconstruction module is specifically used for: Align the scene image with the depth information to obtain a depth value aligned with the scene image; Point cloud map data is obtained based on the scene image and the depth value aligned with the scene image; The point cloud map data is divided into multiple local point cloud data, and the point cloud in each local point cloud data represents at least one geometric element. The device path is determined based on the multiple local point cloud data, and the device path is the movement path of the acquisition device. The first reconstructed data includes the multiple local point cloud data and the device path.

15. The three-dimensional scene reconstruction platform according to claim 14, characterized in that, The 3D scene reconstruction platform also includes a global optimization module, used for: The multiple local point cloud data are projected onto multiple preset voxels according to the device path to obtain multiple voxel point cloud data; Extract planar features from the multiple voxel point cloud data to obtain multiple first planar features; Obtain the association between each of the plurality of first planar features and a plurality of observation poses, wherein the plurality of observation poses are the poses of the virtual camera when observing each planar feature in the three-dimensional space; Based on the aforementioned correlation and the aforementioned voxel point cloud data, the aforementioned local point cloud data and the aforementioned device path are adjusted to obtain the adjusted local point cloud data and the adjusted device path. Based on the adjusted local point cloud data and the adjusted device path, the first reconstruction data is updated to obtain the third reconstruction data.

16. The three-dimensional scene reconstruction platform according to claim 14 or 15, characterized in that, The semantic instance segmentation module is specifically used for: Two-dimensional instance segmentation results are obtained based on the perceived data, and the two-dimensional instance segmentation results include at least one two-dimensional instance in the two-dimensional space and the semantics of the at least one two-dimensional instance; A three-dimensional instance segmentation result is obtained based on the third reconstruction data, wherein the three-dimensional instance segmentation result includes at least one instance in three-dimensional space; Align the two-dimensional instance segmentation result with the three-dimensional instance segmentation result to obtain an aligned three-dimensional instance segmentation result and a two-dimensional instance segmentation result. The semantic instance segmentation result includes the aligned three-dimensional instance segmentation result and the two-dimensional instance segmentation result.

17. The three-dimensional scene reconstruction platform according to claim 16, characterized in that, The semantic instance segmentation module is specifically used for: Obtain the planar features of the adjusted local point cloud data to obtain at least one second planar feature; At least one three-dimensional instance is determined based on the at least one second planar feature, and the three-dimensional instance segmentation result includes the at least one three-dimensional instance.

18. The three-dimensional scene reconstruction platform according to claim 17, characterized in that, When the perceived data includes scene images, the semantic instance segmentation module is specifically used for: The two-dimensional instance segmentation result is obtained based on the scene image; Obtain a depth map corresponding to the at least one two-dimensional instance based on the at least one three-dimensional instance; The at least one two-dimensional instance is corrected according to the depth map to obtain the at least one corrected two-dimensional instance; By associating the at least one three-dimensional instance with the at least one two-dimensional instance, a graph structure is obtained, which is used to represent the association relationship between the at least one three-dimensional instance and the at least one two-dimensional instance after alignment; The semantics of the at least one two-dimensional instance are updated according to the graph structure to obtain the semantics of the at least one two-dimensional instance after the update. The two-dimensional instance segmentation result includes the semantics of the at least one two-dimensional instance after the correction and the at least one two-dimensional instance after the update. The three-dimensional instance segmentation result includes the at least one three-dimensional instance.

19. The three-dimensional scene reconstruction platform according to claim 18, characterized in that, The semantic instance segmentation module is specifically used for: Based on the graph structure, obtain the observation results of the same 3D instance from multiple perspectives; The semantics of the at least one two-dimensional instance are updated based on the observation results to obtain the updated semantics of the at least one two-dimensional instance.

20. The three-dimensional scene reconstruction platform according to claim 18 or 19, characterized in that, The semantic instance segmentation module is specifically used for: Based on the at least one planar feature, the at least one three-dimensional instance is divided into a foreground three-dimensional instance and a background three-dimensional instance; Based on the foreground 3D instance, obtain a depth map corresponding to at least one 2D instance.

21. The three-dimensional scene reconstruction platform according to any one of claims 13-20, characterized in that, The Gaussian reconstruction module is specifically used for: Gaussian parameters are determined based on the perceptual data, semantic segmentation results, and the third reconstruction data. The Gaussian parameters include geometric attribute parameters, visual attribute parameters, and semantic attribute parameters. The geometric attribute parameters are used to represent the geometric shape of the target in the three-dimensional scene, the visual attribute parameters are used to represent the parameters in the visual dimension of the three-dimensional scene, and the semantic attributes are used to represent the semantics of the target in the three-dimensional scene. The second reconstruction data is obtained based on the Gaussian parameters.

22. The three-dimensional scene reconstruction platform according to claim 21, characterized in that, The Gaussian reconstruction module is specifically used for: Based on the third reconstruction data, at least one of two-dimensional Gaussian constraints or three-dimensional Gaussian constraints is constructed. The two-dimensional Gaussian constraints are used to constrain the distance between points in the three-dimensional scene and the fitting plane, and the three-dimensional Gaussian constraints are used to constrain the distribution of three-dimensional instances in the three-dimensional scene. The geometric attribute parameters are determined based on the perceived data under at least one of the two-dimensional Gaussian constraints or the three-dimensional Gaussian constraints.

23. The three-dimensional scene reconstruction platform according to claim 22, characterized in that, The two-dimensional Gaussian constraint includes at least one of position constraint, attitude constraint, or shape constraint. The position constraint includes constraining the center position of the three-dimensional Gaussian sphere to be within the fitting plane. The attitude constraint includes constraining the rotation direction of the three-dimensional Gaussian sphere to be a rotation about the normal vector of the fitting plane. The shape constraint includes constraining the shape of the Gaussian kernel to conform to the constraint shape.

24. The three-dimensional scene reconstruction platform according to any one of claims 21-23, characterized in that, The Gaussian reconstruction module is also used for: By using a pre-trained bilateral grid structure, the image signal processing (ISP) information in the scene image is decoupled; The color parameters are obtained based on the ISP information, and the visual attribute parameters include the color parameters.

25. A computing device cluster, characterized in that, The computing device cluster includes at least one computing device, the computing device including a memory and a processor; the memory stores code, the processor is configured to execute the code, and when the code is executed, the computing device performs the method as described in any one of claims 1-12.

26. A computer storage medium, characterized in that, The computer storage medium stores instructions that, when executed by the computer, cause the computer to perform the method according to any one of claims 1 to 12.

27. A computer program product, characterized in that, The computer program product stores instructions that, when executed by a computer, cause the computer to perform the method described in any one of claims 1 to 12.