Three-dimensional scene online reconstruction method, device, equipment and storage medium

By generating and fusing global Gaussian surface elements of sparse point clouds, and combining Poisson reconstruction algorithm, the problems of low efficiency and insufficient details of three-dimensional reconstruction in large-scale indoor scenes in the existing technology are solved, and a three-dimensional scene reconstruction with high precision and rich details are achieved, which is suitable for real-time and interactive applications.

CN119850849BActive Publication Date: 2025-06-24JIHUA LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510322096.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-24
Estimated Expiration
2045-03-19

AI Technical Summary

Technical Problem

Existing 3D reconstruction technology has problems of inefficiency and insufficient fine-grained appearance details in photo-level realism reconstruction and rendering of large indoor scenes, especially in real-time or interactive applications.

Method used

By obtaining multi-frame ordered scene RGB images, sparse point clouds under the camera coordinate system are generated and converted to the world coordinate system. The point cloud is initialized using local Gaussian surface elements, and the global Gaussian surface elements are generated through the fusion mechanism of the gated loop unit. Render global Gaussian surface elements from multiple perspectives, and combine Poisson reconstruction algorithm to generate high-quality scene reconstruction surfaces.

Benefits of technology

It realizes three-dimensional scene reconstruction with high precision and rich details, improves the accuracy of reconstruction and the coherence and integrity of the scene, and is suitable for real-time and interactive applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119850849B_ABST
    Figure CN119850849B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of three-dimensional reconstruction, and discloses an online three-dimensional scene reconstruction method, apparatus, device and storage medium. The method is used to realize the online reconstruction of a three-dimensional scene. The method includes: acquiring and processing multiple frames of ordered scene RGB images to obtain sparse point clouds of each frame of scene RGB image in the camera coordinate system, and converting them into sparse point clouds in the world coordinate system; initializing the sparse point clouds in the world coordinate system to obtain local Gaussian surface elements of each frame of scene RGB image; based on the fusion mechanism of the gated recurrent unit, reading the global features of the sparse point clouds of each frame of scene RGB image in the world coordinate system frame by frame starting from the first frame, and fusing the global features of the previous frame of scene RGB image with the local Gaussian surface element of the current frame of scene RGB image to generate the global Gaussian surface element of the current frame of scene RGB image; rendering the global Gaussian surface element of each frame of scene RGB image from multiple perspectives, and using the Poisson reconstruction algorithm to generate the scene reconstruction surface.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of three-dimensional reconstruction, and particularly to a method, device, equipment and storage medium for online reconstruction of a three-dimensional scene. Background Art

[0002] Photorealistic reconstruction and rendering of large indoor scenes are fundamental challenges in the fields of computer vision and graphics, and have wide applications in virtual reality, robotics, and immersive simulations. Traditional methods, such as visual simultaneous localization and mapping (SLAM) systems, can achieve real-time 3D reconstruction, but typically represent the scene using an entity surface with vertex colors or texture maps. These representations lack the fine-grained appearance details required for high-quality photorealistic rendering, limiting their effectiveness in scenarios that require a realistic visual experience.

[0003] Recent advances in the field of neural rendering, especially neural radiance fields (NeRF), have achieved remarkable success in novel view synthesis by learning implicit representations of scenes. However, these methods typically require extensive optimization for each scene and have slow rendering speeds, especially when scaled to large environments. Additionally, they usually require offline processing of all input images before optimization, which hinders the applicability of existing methods in real-time or interactive applications.

[0004] Therefore, the prior art still needs to be improved and developed. Summary of the Invention

[0005] The present invention provides a method, device, equipment and storage medium for online reconstruction of a three-dimensional scene, which is used to realize the online reconstruction of a three-dimensional scene.

[0006] In a first aspect of the present invention, a method for online reconstruction of a three-dimensional scene is provided. The method for online reconstruction of a three-dimensional scene includes: acquiring multiple frames of ordered scene RGB images, processing each frame of scene RGB image to obtain a sparse point cloud of each frame of scene RGB image in the camera coordinate system; converting the sparse point cloud in the camera coordinate system to the world coordinate system to obtain a sparse point cloud in the world coordinate system; initializing the sparse point cloud in the world coordinate system to obtain a local Gaussian surface element of each frame of scene RGB image; based on the fusion mechanism of a gated recurrent unit, sequentially reading the global features of the sparse point cloud of each frame of scene RGB image in the world coordinate system frame by frame starting from the first frame, and fusing the global features of the previous frame of scene RGB image with the local Gaussian surface element of the current frame of scene RGB image to generate a global Gaussian surface element of the current frame of scene RGB image; rendering the global Gaussian surface element of each frame of scene RGB image from multiple perspectives to obtain 2D images of each frame of scene RGB image from different perspectives; and generating a scene reconstruction surface using a Poisson reconstruction algorithm based on all the 2D images.

[0007] Preferably, the step of obtaining multiple ordered scene RGB images, processing each frame of the scene RGB image to obtain a sparse point cloud of each frame of the scene RGB image in the camera coordinate system includes: obtaining multiple ordered scene RGB images and the corresponding camera poses of each frame of the scene RGB image; preprocessing each frame of the scene RGB image; and based on the camera pose corresponding to each frame of the scene RGB image, using a depth estimation method to construct a sparse point cloud from the scene RGB image to obtain the sparse point cloud of each frame of the scene RGB image in the camera coordinate system.

[0008] Preferably, the step of converting the sparse point cloud in the camera coordinate system to the world coordinate system to obtain the sparse point cloud in the world coordinate system includes: using the position vector and rotation matrix of the camera to construct a 4x4 homogeneous transformation matrix; multiplying the sparse point cloud in the camera coordinate system by the transformation matrix to obtain the new coordinates of the sparse point cloud in the world coordinate system; and using the new coordinates to update the sparse point cloud to obtain the sparse point cloud in the world coordinate system.

[0009] Preferably, the step of initializing the sparse point cloud in the world coordinate system to obtain the local Gaussian patches of each frame of the scene RGB image includes: constructing an initial Gaussian patch based on the Gaussian function, where the parameters of the initial Gaussian patch include the center coordinates, rotation matrix, scaling vector, opacity, and spherical harmonic encoding; randomly initializing the parameters of the initial Gaussian patch to obtain the initial parameters; using a feature extraction algorithm to detect the key points in each frame of the scene RGB image and calculate the parameters of the key points; using the parameters of the key points to update the parameters of the initial Gaussian patch, and generating the local Gaussian patches of each frame of the scene RGB image based on the updated initial parameters.

[0010] Preferably, based on the fusion mechanism of the gated recurrent unit, starting from the first frame, reading the global features of the sparse point cloud of each frame of the scene RGB image in the world coordinate system frame by frame, and fusing the global features of the previous frame of the scene RGB image with the local Gaussian patches of the current frame of the scene RGB image to generate the global Gaussian patches of the current frame of the scene RGB image includes: reading the global features of the sparse point cloud of the previous frame of the scene RGB image in the world coordinate system based on a preset GRU model; reading the local features of the current frame of the scene RGB image based on the local Gaussian patches of the current frame of the scene RGB image; fusing the global features of the sparse point cloud of the previous frame of the scene RGB image in the world coordinate system with the local features of the current frame of the scene RGB image to obtain the fused features; inputting the fused features into the preset GRU model to obtain the fused features of the current frame of the scene RGB image output by the preset GRU model; and combining the fused features of the current frame of the scene RGB image with the sparse point cloud of the current frame of the scene RGB image in the world coordinate system to generate the global Gaussian patches of the current frame of the scene RGB image.

[0011] Preferably, rendering the global Gaussian patches of each frame of the scene RGB image from multiple perspectives to obtain 2D images of each frame of the scene RGB image from different perspectives includes: equally dividing each frame of the scene RGB image into a plurality of image patches; performing depth sorting on the global Gaussian patches detected within each image patch; traversing the global Gaussian patches within each image patch in the order of depth sorting, and performing alpha blending calculation on the colors of each intersecting global Gaussian patch along the direction of the camera ray to obtain the pixel color of each image patch; rendering and integrating each frame of the scene RGB image processed in blocks based on the pixel color to obtain 2D images of each frame of the scene RGB image from different perspectives.

[0012] Preferably, generating a scene reconstruction surface using the Poisson reconstruction algorithm based on all the 2D images includes: processing all the 2D images to obtain a sparse point cloud corresponding to the 2D images and its normal vectors; performing Poisson surface reconstruction using the Poisson reconstruction algorithm based on the sparse point cloud corresponding to the 2D images and its normal vectors to obtain an isosurface; performing smoothing processing and denoising on the isosurface to obtain a scene reconstruction surface.

[0013] A three-dimensional scene online reconstruction device according to a second aspect of the present invention includes: an acquisition module, configured to acquire a plurality of ordered scene RGB images, process each frame of the scene RGB image to obtain a sparse point cloud of each frame of the scene RGB image in the camera coordinate system; a conversion module, configured to convert the sparse point cloud in the camera coordinate system to the world coordinate system to obtain a sparse point cloud in the world coordinate system; an initialization module, configured to initialize the sparse point cloud in the world coordinate system to obtain local Gaussian patches of each frame of the scene RGB image; a fusion module, configured to, based on the fusion mechanism of a gated recurrent unit, sequentially read the global features of the sparse point cloud of each frame of the scene RGB image in the world coordinate system starting from the first frame, and fuse the global features of the previous frame of the scene RGB image with the local Gaussian patches of the current frame of the scene RGB image to generate global Gaussian patches of the current frame of the scene RGB image; a rendering module, configured to render the global Gaussian patches of each frame of the scene RGB image from multiple perspectives to obtain 2D images of each frame of the scene RGB image from different perspectives; a reconstruction module, configured to generate a scene reconstruction surface using the Poisson reconstruction algorithm based on all the 2D images.

[0014] A three-dimensional scene online reconstruction device according to a third aspect of the present invention includes: a memory and at least one processor, wherein computer-readable instructions are stored in the memory, and the memory and the at least one processor are interconnected by a line; the at least one processor invokes the computer-readable instructions in the memory so that the three-dimensional scene online reconstruction device executes each step of the three-dimensional scene online reconstruction method as described above.

[0015] The fourth aspect of the present invention provides a computer-readable storage medium, in which computer-readable instructions are stored. When it runs on a computer, it causes the computer to execute each step of the three-dimensional scene online reconstruction method described above.

[0016] In the technical solution provided by the present invention, by acquiring multiple frames of ordered scene RGB images and corresponding camera poses, sparse point clouds of each frame of image in the camera coordinate system can be accurately generated and then converted to a unified world coordinate system. This process ensures the accuracy and consistency of the data. Then, the sparse point clouds are initialized using local Gaussian patches, and through a fusion mechanism based on gated recurrent units, the global features of the previous frame and the local Gaussian patches of the current frame are effectively fused, thereby generating a more complete and accurate global Gaussian patch. This fusion strategy not only improves the reconstruction accuracy but also enhances the coherence and integrity of the scene. In addition, rendering the global Gaussian patch from multiple perspectives can capture more details and features in the scene, providing rich data support for the subsequent Poisson reconstruction algorithm. Finally, using the Poisson reconstruction algorithm can generate a high-quality scene reconstruction surface, which not only retains the high precision and detailed information in the original scene but also has good smoothness and continuity, providing a solid foundation for scene analysis, visualization, and interaction. Description of the Drawings

[0017] Figure 1 is a flowchart of the three-dimensional scene online reconstruction method provided by the embodiment of the present invention;

[0018] Figure 2 is a schematic structural diagram of the three-dimensional scene online reconstruction device provided by the embodiment of the present invention;

[0019] Figure 3 is a schematic structural diagram of the three-dimensional scene online reconstruction device provided by the embodiment of the present invention. Detailed Embodiments

[0020] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and do not have to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order different from that illustrated or described herein. In addition, the terms "comprising" or "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0021] For ease of understanding, the specific process of the embodiments of the present invention will be described below. Please refer to Figure 1 , the first embodiment of a three-dimensional scene online reconstruction method in the embodiments of the present invention includes:

[0022] S101. Obtain multiple frames of ordered scene RGB images, process each frame of the scene RGB image, and obtain a sparse point cloud of each frame of the scene RGB image in the camera coordinate system;

[0023] S102. Convert the sparse point cloud in the camera coordinate system to the world coordinate system to obtain a sparse point cloud in the world coordinate system;

[0024] S103. Initialize the sparse point cloud in the world coordinate system to obtain a local Gaussian surface element of each frame of the scene RGB image;

[0025] S104. Based on the fusion mechanism of the gated recurrent unit, read the global features of the sparse point cloud of each frame of the scene RGB image in the world coordinate system frame by frame starting from the first frame, and fuse the global features of the previous frame of the scene RGB image with the local Gaussian surface element of the current frame of the scene RGB image to generate the global Gaussian surface element of the current frame of the scene RGB image;

[0026] S105. Render the global Gaussian surface element of each frame of the scene RGB image from multiple perspectives to obtain 2D images of each frame of the scene RGB image from different perspectives;

[0027] S106. Based on all the 2D images, use the Poisson reconstruction algorithm to generate the scene reconstruction surface.

[0028] It can be understood that the execution subject of the present invention can be a three-dimensional scene online reconstruction device, or a terminal or a server. Specifically, it is not limited here. The embodiments of the present invention will be described by taking the server as the execution subject as an example.

[0029] In this embodiment, in step S101, obtaining multiple frames of ordered scene RGB images and the corresponding camera poses of each frame of the scene RGB image, processing each frame of the scene RGB image, and obtaining a sparse point cloud of each frame of the scene RGB image in the camera coordinate system specifically includes: obtaining multiple frames of ordered scene RGB images and the corresponding camera poses of each frame of the scene RGB image; preprocessing each frame of the scene RGB image; based on the corresponding camera pose of each frame of the scene RGB image, using a depth estimation method to construct a sparse point cloud from the scene RGB image to obtain a sparse point cloud of each frame of the scene RGB image in the camera coordinate system.

[0030] In this embodiment, multiple ordered scene RGB images can be captured by a camera or a camera array, and the camera pose (including position and orientation) corresponding to each frame of the scene RGB image is recorded. Multiple ordered scene RGB images and the camera pose corresponding to the scene RGB images can also be obtained from the collected video. The camera pose corresponding to the scene RGB image is obtained through camera calibration and pose estimation techniques. Camera calibration is used to determine the internal parameters of the camera, such as focal length, distortion coefficient, etc.; pose estimation is used to clarify the position and orientation of the camera in the external world.

[0031] In this embodiment, preprocessing is performed on each frame of the scene RGB image, including denoising, color correction, feature extraction, etc.

[0032] In this embodiment, a depth estimation method is used to generate a sparse point cloud from the scene RGB image to obtain the sparse point cloud of each frame of the scene RGB image in the camera coordinate system.

[0033] Depth estimation refers to the process of inferring the three-dimensional position (i.e., depth) of objects in a scene from a two-dimensional image. Depth estimation methods include stereo vision, structured light, lightfield cameras, and deep learning techniques (such as convolutional neural network CNN), etc. When different depth estimation methods are combined with camera parameters, specific processing needs to be carried out according to their principles. For example, the stereo vision method obtains disparity information through a binocular camera, and combines the internal parameters and pose of the camera to calculate the three-dimensional coordinates of the objects in the scene; the deep learning technique uses a trained model to predict the depth value of each pixel in the image, and then converts it into three-dimensional coordinates in combination with the camera parameters.

[0034] After obtaining the depth information of each frame of the scene RGB image, a sparse point cloud can be generated by combining it with the internal and external parameters of the camera (i.e., the results of camera calibration and pose estimation). These point clouds are represented in the camera coordinate system, and each point contains its position information (x, y, z) in three-dimensional space.

[0035] In this embodiment, in step S102, the sparse point cloud in the camera coordinate system is converted to the world coordinate system to obtain the sparse point cloud in the world coordinate system, which specifically includes: using the position vector and rotation matrix of the camera to construct a 4x4 homogeneous transformation matrix; multiplying the sparse point cloud in the camera coordinate system by the transformation matrix to obtain the new coordinates of the sparse point cloud in the world coordinate system; using the new coordinates to update the sparse point cloud to obtain the sparse point cloud in the world coordinate system.

[0036] In this embodiment, the position vector describes the position of the camera in the world coordinate system, and the rotation matrix describes the orientation of the camera.

[0037] In this embodiment, the sparse point cloud is represented in homogeneous coordinate form (i.e., each point is a 4x1 vector with the last element being 1).

[0038] In this embodiment, in step S103, the sparse point cloud in the world coordinate system is initialized to obtain the local Gaussian patches of each frame of the scene RGB image, which specifically includes: constructing the initial Gaussian patches based on the Gaussian function, and the parameters of the initial Gaussian patches include the center coordinates, rotation matrix, scaling vector, opacity, and spherical harmonic encoding; randomly initializing the parameters of the initial Gaussian patches to obtain the initial parameters; using a feature extraction algorithm to detect the key points in each frame of the scene RGB image and calculate the parameters of the key points; using the parameters of the key points to update the parameters of the initial Gaussian patches, and generating the local Gaussian patches of each frame of the scene RGB image based on the updated initial parameters.

[0039] In this embodiment, the feature extraction algorithm is such as the SIFT, SURF, or ORB algorithm.

[0040] In this embodiment, the key points are usually the points in the image with significant features (such as corner points, edges, or texture changes), and they are crucial for the subsequent update of the patch parameters.

[0041] In this embodiment, the parameters of the key points include the center coordinates, rotation matrix, scaling vector, opacity, and spherical harmonic encoding.

[0042] The rotation matrix can be estimated by analyzing the gradients or directions of the pixels around the key points. The scaling vector can be estimated according to the size of the area around the key points or the scale of the feature points. The opacity may need to be estimated according to the brightness, color, or other visual features of the area around the key points. The spherical harmonic encoding is usually used to represent the lighting and surface reflection properties.

[0043] In this embodiment, in step S104, based on the fusion mechanism of the gated recurrent unit, the global features of the sparse point cloud of each frame of the scene RGB image in the world coordinate system are read frame by frame starting from the first frame, and the global features of the previous frame of the scene RGB image are fused with the local Gaussian patch of the current frame of the scene RGB image to generate the global Gaussian patch of the current frame of the scene RGB image. Specifically, it includes: reading the global features of the sparse point cloud of the previous frame of the scene RGB image in the world coordinate system based on the preset GRU model; reading the local features of the current frame of the scene RGB image based on the local Gaussian patch of the current frame of the scene RGB image; fusing the global features of the sparse point cloud of the previous frame of the scene RGB image in the world coordinate system with the local features of the current frame of the scene RGB image to obtain the fused features; inputting the fused features into the preset GRU model to obtain the fused features of the current frame of the scene RGB image output by the preset GRU model; combining the fused features of the current frame of the scene RGB image and the sparse point cloud of the current frame of the scene RGB image in the world coordinate system to generate the global Gaussian patch of the current frame of the scene RGB image.

[0044] It should be noted that the GRU model is a recurrent neural network (RNN) model for processing sequential data. The core structure of the GRU model includes two gating mechanisms: the update gate and the reset gate.

[0045] The update gate determines whether the new input should be retained and how much information from the previous time step should be passed to the current time step. The update gate controls the flow of information by calculating a value between 0 and 1. When the value of the update gate is close to 1, it means that most of the information from the previous time step should be retained; when the value is close to 0, it means that the information from the previous time step should be ignored and more attention should be paid to the current input.

[0046] The reset gate determines whether the hidden state of the previous time step should be ignored. The reset gate also controls the flow of information by calculating a value between 0 and 1. When the value of the reset gate is close to 1, it means that the hidden state of the previous time step should be retained; when the value is close to 0, it means that the hidden state of the previous time step should be ignored and more attention should be paid to the relationship between the current input and the input of the previous time step.

[0047] The GRU model uses the current input and the hidden state of the previous time step for linear transformation and obtains the values of the reset gate and the update gate through the sigmoid activation function.

[0048] The GRU model uses the reset gate to control the retention degree of the hidden state of the previous time step, concatenates it with the current input, then performs linear transformation and tanh activation to obtain the candidate hidden state.

[0049] The GRU model uses update gates to control the degree of retention of the hidden state and candidate hidden state at the previous time step, and obtains the hidden state at the current time step through weighted summation.

[0050] In this embodiment, the preset GRU model will be used to process time series data and store and transmit the global feature information of the previous frame. The global features of the sparse point cloud of the previous frame of the scene RGB image in the world coordinate system can be read from the hidden state of the preset GRU model. These global features may include information such as shape, distribution, density, etc. that describe the overall properties of the point cloud.

[0051] In this embodiment, the local Gaussian patch of the current frame scene RGB image is converted into a local feature vector for fusion with the global features of the previous frame.

[0052] In this embodiment, operations such as concatenation or weighted summation are performed on the global features of the sparse point cloud of the previous frame scene RGB image in the world coordinate system and the local features of the current frame scene RGB image to obtain the fused features. The purpose of this step is to combine the global information of the previous frame with the local information of the current frame to form a richer feature representation.

[0053] In this embodiment, the fused features are used as input and passed to the preset GRU model. The GRU model updates its hidden state and output state according to the input fused features. In this process, the update gate and reset gate of the GRU will control the flow and memory of information to generate the fused feature representation of the current frame scene RGB image.

[0054] In this embodiment, the fused features of the current frame are combined with the sparse point cloud of the current frame scene RGB image in the world coordinate system. The purpose of this step is to attach the fused feature information to the sparse point cloud to generate a global Gaussian patch with richer information. According to the combined sparse point cloud and fused features, the global Gaussian patch of the current frame scene RGB image is generated. The global Gaussian patch represents the attributes such as shape, texture, and brightness of each part in the current frame scene, has global consistency and local details, and is modeled in the form of a Gaussian distribution.

[0055] In this embodiment, taking the process of obtaining the global Gaussian patch of the scene RGB image of the t-th frame as an example, the calculation process can be described as:

[0056]

[0057] Among them, and respectively represent the global Gaussian patch and the local Gaussian patch of the scene RGB image of the -th frame, represents the Global Gaussian patches of the RGB image of the scene of the frame is the gated recurrent unit

[0058] In this embodiment, in step S105, rendering the global Gaussian patches of each frame of the RGB image of the scene from multiple perspectives to obtain 2D images of each frame of the RGB image of the scene from different perspectives specifically includes:

[0059] Equally divide each frame of the RGB image of the scene into a plurality of image patches; perform depth sorting on the detected global Gaussian patches within each image patch; traverse the global Gaussian patches within each image patch in the order of depth sorting, and perform alpha blending calculation on the colors of each intersecting global Gaussian patch along the direction of the camera ray to obtain the pixel color of each image patch; render and integrate each frame of the RGB image of the scene after block processing based on the pixel color to obtain 2D images of each frame of the RGB image of the scene from different perspectives.

[0060] It can be understood that the Gaussian distribution presented by the Gaussian patch is elliptical. Therefore, the Gaussian distribution of points in the Gaussian local coordinate system can be expressed as:

[0061]

[0062] In the formula, represents uv the Gaussian distribution of point in space, u and v respectively represent the axis and axis coordinates of the Gaussian local coordinate system.

[0063] During the rendering process, the problem of shape degradation of the Gaussian patch may occur when observed from a specific perspective. When the Gaussian patch degrades at a specific time, the gradient of backpropagation cannot be propagated to the global Gaussian patch through the rendering loss (because the Gaussian distribution is 0). Therefore, a variance is set for this situation so that the algorithm is differentiable throughout even in the face of this situation, that is, is extended to:

[0064]

[0065] In the formula, represents the pixel block, represents the global Gaussian distribution of the image patch represents the image patch the coordinates in the uv local coordinate system of the corresponding Gaussian patch, represents based on the Gaussian distribution value,

[0066] c is the projection of the center point of the Gaussian surface element. Intuitively, The lower bound of is determined by a fixed screen-space Gaussian low-pass filter with center c and radius . In actual use, set to ensure that enough pixels are used during rendering.

[0067] In this embodiment, in the order of depth sorting, traverse the global Gaussian surface elements within each image block, and perform alpha blending calculation on the colors of each intersecting global Gaussian surface element along the direction of the camera ray to obtain the pixel color of each image block. Define the pixel color after rendering of a certain image block as , then the pixel color rendering process for the pixel color of is expressed as:

[0068]

[0069] In the formula, represents the color of the th global Gaussian surface element in the image block along the depth increasing direction, represents the opacity of the th global Gaussian surface element in the image block along the depth increasing direction, represents the opacity of the th global Gaussian surface element in the image block along the depth increasing direction, represents the global Gaussian distribution of the image block in the th Gaussian local coordinate system, represents the global Gaussian distribution of the image block in the th Gaussian local coordinate system.

[0070] In this embodiment, during the process of depth sorting the detected global Gaussian surface elements within each image block, a depth loss function is used. The depth loss function is expressed as:

[0071]

[0072] In the formula, represents the camera rays with valid depth information in the pixel set , and respectively represent the rendering depth and the true depth of the image block .

[0073] ​In this embodiment, in the process of traversing the global Gaussian patches within each image patch in the order of depth sorting and performing alpha blending calculation on the colors of each intersecting global Gaussian patch along the direction of the camera ray to obtain the pixel colors of each image patch, a color loss function is used. The color loss function is expressed as:

[0074] In the formula, and respectively represent the rendered pixel color and the true pixel color, is a set of a series of pixel points.

[0075] In this embodiment, during the rendering process, in order to ensure that all global Gaussian patches are locally aligned with the actual surface, a normal consistency loss function is used to align the rendered normal map with the normal prior map. The normal consistency loss function is expressed as:

[0076]

[0077] In the formula, represents the index of the patch intersected by the ray, is the blending weight of the intersection point, is the normal towards the camera strip, is the normal estimated by the gradient of the depth map.

[0078] In this embodiment, during the rendering process, the camera pose of the input view is required. The camera rays are projected according to the camera pose, the intersection points of the camera rays and the global Gaussian patches are calculated, and then the Gaussian distribution at the intersection points is calculated based on the intersection points, so as to solve the problem of shape degradation of the elliptical patches under specific viewing angles.

[0079] In this embodiment, in step S106, based on all 2D images, a Poisson reconstruction algorithm is used to generate the scene reconstruction surface, which specifically includes: processing all 2D images to obtain the sparse point cloud corresponding to the 2D images and their normal vectors; based on the sparse point cloud corresponding to the 2D images and their normal vectors, using the Poisson reconstruction algorithm to perform Poisson surface reconstruction to obtain an isosurface; performing smoothing processing and denoising on the isosurface to obtain the scene reconstruction surface.

[0080] In this embodiment, the Poisson reconstruction algorithm generates a continuous three-dimensional surface by solving the Poisson equation. This process utilizes the normal vector information to maintain the details and smoothness of the surface. The output of the algorithm is an isosurface, which represents the reconstructed three-dimensional surface.

[0081] This embodiment provides a method for online 3D scene reconstruction. By obtaining multiple ordered scene RGB images and their corresponding camera poses, it can accurately generate the sparse point cloud of each frame of image in the camera coordinate system and then convert it to the unified world coordinate system. This process ensures the accuracy and consistency of the data. Then, the sparse point cloud is initialized using local Gaussian patches, and through a fusion mechanism based on gated recurrent units, the global features of the previous frame are effectively fused with the local Gaussian patches of the current frame, thus generating a more complete and accurate global Gaussian patch. This fusion strategy not only improves the reconstruction accuracy but also enhances the coherence and integrity of the scene. In addition, rendering the global Gaussian patch from multiple perspectives can capture more details and features in the scene, providing rich data support for the subsequent Poisson reconstruction algorithm. Finally, using the Poisson reconstruction algorithm can generate a high-quality scene reconstruction surface, which not only retains the high precision and detailed information in the original scene but also has good smoothness and continuity, providing a solid foundation for scene analysis, visualization, and interaction.

[0082] The above describes the method for online 3D scene reconstruction in the embodiments of the present invention. Next, the device in the embodiments of the present invention will be described. Please refer to Figure 2 The implementation manner of the 3D scene online reconstruction device in the embodiments of the present invention includes:

[0083] An acquisition module 201, configured to acquire multiple ordered scene RGB images, process each frame of the scene RGB image, and obtain the sparse point cloud of each frame of the scene RGB image in the camera coordinate system;

[0084] A conversion module 202, configured to convert the sparse point cloud in the camera coordinate system to the world coordinate system to obtain the sparse point cloud in the world coordinate system;

[0085] An initialization module 203, configured to initialize the sparse point cloud in the world coordinate system to obtain the local Gaussian patch of each frame of the scene RGB image;

[0086] A fusion module 204, configured to, based on the fusion mechanism of gated recurrent units, read the global features of the sparse point cloud of each frame of the scene RGB image in the world coordinate system frame by frame starting from the first frame, and fuse the global features of the previous frame of the scene RGB image with the local Gaussian patch of the current frame of the scene RGB image to generate the global Gaussian patch of the current frame of the scene RGB image;

[0087] A rendering module 205, configured to render the global Gaussian patch of each frame of the scene RGB image from multiple perspectives to obtain 2D images of each frame of the scene RGB image from different perspectives;

[0088] A reconstruction module 206 for generating a scene reconstruction surface based on all 2D images using a Poisson reconstruction algorithm.

[0089] In this embodiment, by acquiring multiple ordered RGB images of the scene and their corresponding camera poses, sparse point clouds of each frame in the camera coordinate system can be accurately generated and then transformed into a unified world coordinate system. This process ensures the accuracy and consistency of the data. Then, the sparse point clouds are initialized using local Gaussian patches, and through a fusion mechanism based on gated recurrent units, the global features of the previous frame and the local Gaussian patches of the current frame are effectively fused to generate a more complete and accurate global Gaussian patch. This fusion strategy not only improves the reconstruction accuracy but also enhances the coherence and integrity of the scene. In addition, rendering the global Gaussian patch from multiple perspectives can capture more details and features in the scene, providing rich data support for the subsequent Poisson reconstruction algorithm. Finally, using the Poisson reconstruction algorithm, a high-quality scene reconstruction surface can be generated. This surface not only retains the high precision and detailed information in the original scene but also has good smoothness and continuity, providing a solid foundation for scene analysis, visualization, and interaction.

[0090] Figure 2 The structure of the shown three-dimensional scene online reconstruction device does not limit the three-dimensional scene online reconstruction device and can implement the steps of the three-dimensional scene online reconstruction method provided by the above method embodiments.

[0091] Above Figure 2 The three-dimensional scene online reconstruction device in the embodiments of the present invention is described in detail from the perspective of modular functional entities. Next, the three-dimensional scene online reconstruction device in the embodiments of the present invention is described in detail from the perspective of hardware processing.

[0092] Figure 3 FIG. is a schematic structural diagram of a three-dimensional scene online reconstruction device provided by an embodiment of the present invention. The device 300 may vary greatly due to different configurations or performances and may include one or more processors (central processing units, CPUs) 310 (for example, one or more processors) and a memory 320, and one or more storage media 330 for storing application programs 333 or data 332 (for example, one or more mass storage devices). Among them, the memory 320 and the storage media 330 may be transient storage or persistent storage. The program stored in the storage media 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the device 300. Further, the processor 310 may be configured to communicate with the storage media 330 and execute a series of instruction operations in the storage media on the device 300.

[0093] Device 300 may also include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input / output interfaces 360, and / or one or more operating systems 331, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, and so on.

[0094] An embodiment of the present invention also provides a computer-readable storage medium. The computer-readable storage medium can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions are run on a computer, the computer is caused to execute the steps of the three-dimensional scene online reconstruction method.

[0095] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, or units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0096] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to cause a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.

[0097] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of various embodiments of the present invention.

Claims

1. A three-dimensional scene online reconstruction method, characterized in that: The three-dimensional scene online reconstruction method comprises: Acquire multiple frames of ordered scene RGB images, process each frame of the scene RGB image, and obtain a sparse point cloud of each frame of the scene RGB image in the camera coordinate system; The sparse point cloud in the camera coordinate system is converted to the world coordinate system to obtain the sparse point cloud in the world coordinate system; Initialize the sparse point cloud in the world coordinate system to obtain the local Gaussian surface element of each frame of the scene RGB image; Based on the fusion mechanism of the gated recurrent unit, the global features of the sparse point cloud of each frame of the scene RGB image in the world coordinate system are read frame by frame starting from the first frame, and the global features of the scene RGB image of the previous frame are fused with the local Gaussian surface element of the scene RGB image of the current frame to generate the global Gaussian surface element of the scene RGB image of the current frame; Render the global Gaussian surface element of each frame of the scene RGB image from multiple perspectives to obtain a 2D image of each frame of the scene RGB image at different perspectives; Based on all 2D images, a scene reconstruction surface is generated using the Poisson reconstruction algorithm; The sparse point cloud in the initialized world coordinate system is obtained to obtain a local Gaussian facet of each frame of the scene RGB image, including: constructing an initialized Gaussian facet based on a Gaussian function, wherein the parameters of the initialized Gaussian facet include a center coordinate, a rotation matrix, a scaling vector, opacity, and a spherical harmonic code; randomly initializing the parameters of the initialized Gaussian facet to obtain an initialization parameter; using a feature extraction algorithm to detect key points in each frame of the scene RGB image, and calculating the parameters of the key points; using the parameters of the key points to update the initialization parameters, and generating a local Gaussian facet of each frame of the scene RGB image based on the updated initialization parameters; The fusion mechanism based on the gated recurrent unit reads the global features of the sparse point cloud of each frame of the scene RGB image in the world coordinate system frame by frame starting from the first frame, and fuses the global features of the scene RGB image of the previous frame with the local Gaussian surface elements of the scene RGB image of the current frame to generate the global Gaussian surface elements of the scene RGB image of the current frame, including: reading the global features of the sparse point cloud of the scene RGB image of the previous frame in the world coordinate system based on a preset GRU model; reading the local features of the scene RGB image of the current frame based on the local Gaussian surface elements of the scene RGB image of the current frame; fusing the global features of the sparse point cloud of the scene RGB image of the previous frame in the world coordinate system with the local features of the scene RGB image of the current frame to obtain fused features; inputting the fused features into the preset GRU model to obtain the fused features of the current frame scene RGB image output by the preset GRU model; combining the fused features of the scene RGB image of the current frame and the sparse point cloud of the scene RGB image of the current frame in the world coordinate system to generate the global Gaussian surface elements of the scene RGB image of the current frame.

2. The three-dimensional scene online reconstruction method according to claim 1, characterized in that: The method of acquiring multiple frames of ordered scene RGB images and processing each frame of the scene RGB images to obtain a sparse point cloud of each frame of the scene RGB images in a camera coordinate system includes: Obtain multiple frames of ordered scene RGB images and the camera pose corresponding to each frame of the scene RGB image; Preprocess each frame of scene RGB image; Based on the camera pose corresponding to each frame of the scene RGB image, a sparse point cloud is constructed from the scene RGB image using a depth estimation method to obtain a sparse point cloud of each frame of the scene RGB image in the camera coordinate system.

3. The three-dimensional scene online reconstruction method according to claim 1, characterized in that: The step of converting the sparse point cloud in the camera coordinate system to the world coordinate system to obtain the sparse point cloud in the world coordinate system includes: Use the camera's position vector and rotation matrix to construct a 4x4 homogeneous transformation matrix; Multiply the sparse point cloud in the camera coordinate system by the transformation matrix to obtain the new coordinates of the sparse point cloud in the world coordinate system; The sparse point cloud is updated with the new coordinates to obtain a sparse point cloud in the world coordinate system.

4. The three-dimensional scene online reconstruction method according to claim 1, characterized in that: The global Gaussian facets of each frame of the scene RGB image are rendered from multiple perspectives to obtain 2D images of each frame of the scene RGB image at different perspectives, including: Divide each frame of scene RGB image equally into several image blocks; Deep sorting of the global Gaussian surfels detected within each image patch; In the order of depth sorting, traverse the global Gaussian surface elements in each image block, perform alpha blending calculation on the color of each intersecting global Gaussian surface element along the direction of the camera light, and obtain the pixel color of each image block; The scene RGB image of each frame after block processing is rendered and integrated based on pixel color to obtain a 2D image of each frame of the scene RGB image at different viewing angles.

5. The three-dimensional scene online reconstruction method according to claim 1, characterized in that: Based on all the 2D images, a Poisson reconstruction algorithm is used to generate a scene reconstruction surface, including: All 2D images are processed to obtain sparse point clouds and their normal vectors corresponding to the 2D images; Based on the sparse point cloud corresponding to the 2D image and its normal vector, the Poisson reconstruction algorithm is used to reconstruct the Poisson surface and obtain the isosurface; The isosurface is smoothed and denoised to obtain a scene reconstruction surface.

6. A three-dimensional scene online reconstruction device, characterized in that: include: An acquisition module is used to acquire multiple frames of ordered scene RGB images, process each frame of the scene RGB image, and obtain a sparse point cloud of each frame of the scene RGB image in the camera coordinate system; A conversion module is used to convert the sparse point cloud in the camera coordinate system to the world coordinate system to obtain the sparse point cloud in the world coordinate system; An initialization module is used to initialize the sparse point cloud in the world coordinate system to obtain the local Gaussian surface element of each frame of the scene RGB image, specifically including: constructing an initialized Gaussian surface element based on a Gaussian function, wherein the parameters of the initialized Gaussian surface element include the center coordinates, the rotation matrix, the scaling vector, the opacity and the spherical harmonic coding; randomly initializing the parameters of the initialized Gaussian surface element to obtain the initialization parameters; using a feature extraction algorithm to detect the key points in each frame of the scene RGB image, and calculating the parameters of the key points; using the parameters of the key points to update the initialization parameters, and generating the local Gaussian surface element of each frame of the scene RGB image based on the updated initialization parameters; A fusion module is used for a fusion mechanism based on a gated recurrent unit, starting from the first frame, to read the global features of the sparse point cloud of each frame of the scene RGB image in the world coordinate system frame by frame, and to fuse the global features of the scene RGB image of the previous frame with the local Gaussian surface element of the scene RGB image of the current frame to generate the global Gaussian surface element of the scene RGB image of the current frame, specifically including: the fusion mechanism based on the gated recurrent unit, starting from the first frame, to read the global features of the sparse point cloud of each frame of the scene RGB image in the world coordinate system frame by frame, and to fuse the global features of the scene RGB image of the previous frame with the local Gaussian surface element of the scene RGB image of the current frame to generate the global Gaussian surface element of the scene RGB image of the current frame, including In summary: based on the preset GRU model, the global features of the sparse point cloud of the previous frame scene RGB image in the world coordinate system are read; based on the local Gaussian surface element of the current frame scene RGB image, the local features of the current frame scene RGB image are read; the global features of the sparse point cloud of the previous frame scene RGB image in the world coordinate system are fused with the local features of the current frame scene RGB image to obtain the fused features; the fused features are input into the preset GRU model to obtain the fused features of the current frame scene RGB image output by the preset GRU model; the global Gaussian surface element of the current frame scene RGB image is generated by combining the fused features of the current frame scene RGB image and the sparse point cloud of the current frame scene RGB image in the world coordinate system; A rendering module is used to render the global Gaussian surface element of each frame of the scene RGB image from multiple perspectives to obtain a 2D image of each frame of the scene RGB image at different perspectives; The reconstruction module is used to generate a scene reconstruction surface based on all 2D images using a Poisson reconstruction algorithm.

7. A three-dimensional scene online reconstruction device, characterized in that: comprising a memory and at least one processor, wherein the memory has computer-readable instructions stored therein; The at least one processor calls the computer-readable instructions in the memory to execute the various steps of the three-dimensional scene online reconstruction method as described in any one of claims 1-5.

8. A computer-readable storage medium having computer-readable instructions stored thereon, characterized in that: When the computer-readable instructions are executed by a processor, the steps of the three-dimensional scene online reconstruction method as described in any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Method for reconstructing flame CT (computed tomography) images with adaptive section sizes

    CN103971388A

  • Techniques for large-scale three-dimensional scene reconstruction via camera clustering

    US20240104831A1