Dynamic scene reconstruction method and system and related equipment
By dividing the scene into voxel grids of multiple scales and rendering them using a Gaussian model, the problem of insufficient detail expression in scene reconstruction in the existing NeRF method is solved, and more refined dynamic scene reconstruction is achieved.
Patent Information
- Application Number
- CN202511181944.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-08-22
AI Technical Summary
Existing neural radiance field (NeRF)-based methods have difficulty in effectively expressing the detailed information of the scene during scene reconstruction, resulting in the reconstruction results being not precise enough.
By dividing the scene to be reconstructed into voxel grids of multiple scales, obtaining voxel features, and rendering the scene using Gaussian models and multi-dimensional Gaussian features, the feature construction methods of voxel grids and Gaussian models are combined to enhance the expression of details.
The detail expression capability of dynamic scene reconstruction is improved, which can better reflect the dynamic changes and detailed characteristics of the scene.
Smart Images

Figure CN120689555A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information processing technology, and in particular to a dynamic scene reconstruction method, system and related equipment. Background Art
[0002] Scene reconstruction generally involves converting real-world physical scenes into digital 3D models through computer vision, graphics, or sensor technology. Its applications span a wide range of fields, including film and television production, interior design, and autonomous driving. Currently, a variety of methods are used to achieve scene reconstruction, including traditional computer vision, deep learning, and sensor fusion. However, achieving a scene that is consistent with the real world is a pressing issue in scene reconstruction.
[0003] One such scene reconstruction method, based on Neural Radiance Fields (NeRF), primarily meshes the reconstructed scene, constructs mesh features, and then uses NeRF network architecture training to obtain volume density features and color appearance features, thereby rendering the scene consistent with the real image. However, the scenes generated by such scene reconstruction methods still have limited ability to express the actual scene and cannot better capture more details. Summary of the Invention
[0004] The embodiments of the present invention provide a dynamic scene reconstruction method, system and related equipment, which enhance the expression of details in the reconstruction of real dynamic scenes.
[0005] An embodiment of the present invention provides a dynamic scene reconstruction method, including: Obtaining point cloud information of multiple frames of the scene to be reconstructed; For multiple frames of point cloud information, the scene to be reconstructed is divided into voxel grids of multiple scales, and voxel features of each voxel grid are obtained; Determining Gaussian models corresponding to the respective voxel grids according to the voxel features, and obtaining multi-dimensional Gaussian features of the Gaussian models; The Gaussian model is rendered according to the multi-dimensional Gaussian features to reconstruct the scene to be reconstructed.
[0006] Another embodiment of the present invention provides a dynamic scene reconstruction system, including: A point cloud acquisition unit, used to acquire point cloud information of multiple frames of the scene to be reconstructed; A voxel grid unit is used to divide the scene to be reconstructed into voxel grids of multiple scales based on the point cloud information of multiple frames, and obtain voxel features of each voxel grid; A Gaussian model unit, configured to determine Gaussian models corresponding to the respective voxel grids according to the voxel features, and obtain multi-dimensional Gaussian features of the Gaussian models; A rendering unit is used to render the Gaussian model according to the multi-dimensional Gaussian features to reconstruct the scene to be reconstructed.
[0007] Another aspect of an embodiment of the present invention further provides a computer-readable storage medium, which stores a plurality of computer programs, and the computer programs are suitable for being loaded by a processor and executing the dynamic scene reconstruction method as described in one aspect of an embodiment of the present invention.
[0008] Another aspect of the present invention provides a terminal device, including a processor and a memory; The memory is used to store multiple computer programs, and the computer programs are used to be loaded by the processor and execute the dynamic scene reconstruction method as described in one aspect of an embodiment of the present invention; the processor is used to implement each computer program in the multiple computer programs.
[0009] It can be seen that in the method of this embodiment, after obtaining multiple frames of point cloud information, the dynamic scene reconstruction system can divide the scene to be reconstructed into voxel grids of multiple scales based on the point cloud information, and obtain the voxel features of each voxel grid. Then, based on the voxel features, the corresponding Gaussian model is determined, and the multi-dimensional Gaussian features of the Gaussian model are obtained. Then, the Gaussian model is rendered based on the multi-dimensional Gaussian features to reconstruct the scene to be reconstructed. The main method is to combine the voxel grid constructed based on the actual scene to be reconstructed with the Gaussian model to render the scene to be reconstructed. The two different feature construction methods are used to enhance the expression of details in the reconstruction of real dynamic scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0011] Figure 1 is a schematic diagram of a dynamic scene reconstruction method provided by an embodiment of the present invention; Figure 2 is a flow chart of a dynamic scene reconstruction method provided by an embodiment of the present invention; Figure 3a Schematic diagram of optimizing a Gaussian model based on a dual-branch MLP model in an embodiment of the present invention; Figure 3b is a schematic diagram of another method for optimizing a Gaussian model based on a dual-branch MLP model according to an embodiment of the present invention; Figure 4Schematic diagram of sampling step adjustment in an embodiment of the present invention; Figure 5 This is a flow chart of a dynamic scene reconstruction method provided by a specific application embodiment of the present invention; Figure 6 This is a schematic diagram of the logical structure of a dynamic scene reconstruction system provided by an embodiment of the present invention; Figure 7 This is a schematic diagram of the logical structure of a terminal device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0012] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0013] The terms "first," "second," "third," "fourth," and so forth (if any) in the description and claims of the present invention and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the present invention described herein can, for example, be implemented in an order other than that illustrated or described herein. In addition, the terms "including" and "having," as well as any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus that includes a series of steps or elements is not necessarily limited to those steps or elements expressly listed, but may include other steps or elements not expressly listed or inherent to such process, method, product, or apparatus.
[0014] The embodiment of the present invention provides a dynamic scene reconstruction method, which is mainly used to perform 3D modeling of actual dynamic scenes, so as to be applied to some specific fields, such as Figure 1 As shown, the dynamic scene reconstruction system can achieve scene reconstruction through the following methods, including: Obtaining point cloud information of multiple frames of the scene to be reconstructed; For multiple frames of point cloud information, the scene to be reconstructed is divided into voxel grids of multiple scales, and voxel features of each voxel grid are obtained; Determining Gaussian models corresponding to the respective voxel grids according to the voxel features, and obtaining multi-dimensional Gaussian features of the Gaussian models; The Gaussian model is rendered according to the multi-dimensional Gaussian features to reconstruct the scene to be reconstructed.
[0015] Thus, in this embodiment, the voxel grid constructed based on the actual scene to be reconstructed and the Gaussian model are combined to render the scene to be reconstructed. Two different feature construction methods are used to enhance the expression of details in the reconstruction of the real dynamic scene.
[0016] An embodiment of the present invention provides a dynamic scene reconstruction method, the flow chart is as follows Figure 2 Shown, including: Step 101: Obtain point cloud information of multiple frames of a scene to be reconstructed.
[0017] It can be understood that the dynamic scene reconstruction system can periodically scan a scene to be reconstructed multiple times through sensors at certain time intervals to obtain point cloud information at each moment. The point cloud information obtained at each moment is a frame of point cloud information, and there is a certain time sequence between different frames of point cloud information.
[0018] In this embodiment, the multi-frame point cloud information is a time-series point cloud, which not only includes three-dimensional spatial information, but also incorporates the time dimension. It can dynamically reflect the evolution of the spatial structure of the scene or target over time, and emphasizes "temporal continuity" and "dynamic variability". It is a key data form for understanding the dynamic three-dimensional world.
[0019] Furthermore, in this embodiment, after scanning each frame of point cloud information, point cloud density may be counted to describe the degree of aggregation of each point cloud, such as 50 points / m2.
[0020] Step 102 : for the point cloud information of multiple frames, the scene to be reconstructed is divided into voxel grids of multiple scales, and voxel features of each voxel grid are obtained.
[0021] Specifically, voxel grids of corresponding scales can be set for different areas of the scene to be reconstructed according to the point cloud density of different areas of the scene to be reconstructed, so as to obtain original voxel grids of multiple scales and acquire voxel features of the original voxel grids.
[0022] A voxel is the smallest discrete unit in three-dimensional space, analogous to a pixel in a two-dimensional image. A voxel has a fixed spatial size (e.g., 1m×1m×1m). Each voxel stores attribute information within that space (such as color, density, and the presence of objects), serving to "grid" the continuous three-dimensional space. In this embodiment, if a certain area of the scene to be reconstructed is high-density, such as a point cloud density greater than 50 points / ㎡, a 12.5cm voxel grid can be used. If a certain area of the scene to be reconstructed is medium-density, such as a point cloud density between 10-50 points / ㎡, a 25cm voxel grid can be used. If a certain area of the scene to be reconstructed is low-density, such as a point cloud density less than 10 points / ㎡, a 50cm voxel grid can be used. This allows the scene to be reconstructed into voxel grids of multiple scales, i.e., multiple sizes, each reflecting different resolutions. Generally speaking, smaller voxel grids are used in areas with higher point cloud density to highlight specific details, while larger voxel grids are used in areas with lower point cloud density.
[0023] The voxel features of each original voxel grid may include: point cloud color, density, dynamic displacement information, static voxel normal vector information, etc. In this embodiment, dynamic and static attributes are set for each voxel grid to distinguish whether the same voxel grid corresponding to multiple frames of point cloud information is dynamic or static.
[0024] Among them, dynamic displacement information refers to the movement distance of the voxel grid between multiple consecutive frames (such as 3 frames). Specifically, the iterative closest point (ICP) algorithm and other algorithms can be combined with ORB feature matching to align the point clouds of 3 consecutive frames to a unified coordinate system, and calculate the displacement Δs of the original voxel grid. This displacement is the dynamic displacement information. If the displacement is greater than a threshold, such as 0.5m, the dynamic and static properties of the original voxel grid can be set to "dynamic", otherwise the dynamic and static properties of the original voxel grid can be set to "static".
[0025] The static voxel normal vector information reflects the geometric direction of the local surface represented by each original voxel grid and can represent the curvature of the original voxel grid. Specifically: A neighboring point cloud of the original voxel grid K=20 can be extracted, and the plane normal vector can be fitted using the progressive sampling (PROSAC) + maximum likelihood estimation (MLESAC) combined algorithm. The PROSAC algorithm selects initial samples according to probability, giving priority to samples that are more likely to be inliers of the original voxel grid; the MLESAC algorithm optimizes the fitting results based on maximum likelihood estimation, sets the inlier distance threshold to 0.05m, and quickly and accurately obtains static voxel geometric features.
[0026] It should be noted that after obtaining the original voxel grid, the following step 103 can be directly performed on the original voxel grid, or, preferably, the original voxel grid can be optimized first, and then the following step 103 can be performed on the optimized voxel grid. Specifically, when optimizing the original voxel grid: First, multiple original voxel grids may be constructed into a voxel multilayer structure tree based on the voxel features of the original voxel grids, and the original voxel grids may be adjusted based on the voxel multilayer structure tree to obtain adjusted voxel grids. For the adjusted voxel grids, the following step 103 may be performed.
[0027] In this embodiment, an octree can be used to construct a voxel multi-layer structure tree of the original voxel grid, specifically a voxel octree, where the octree is a tree-like data structure that recursively divides the three-dimensional space into eight equal cubes. Its root node is the entire three-dimensional space, and each non-leaf node contains 8 child nodes, corresponding to 8 equally divided areas formed by bisecting the three-dimensional space along the x, y, and z axes (similar to the 8 small cubes in a "magic cube").
[0028] The constructed voxel octree is a 3D spatial data structure that combines the advantages of voxels and octrees, used for efficiently representing, storing, and processing discrete information in 3D space. It uses the hierarchical, recursive partitioning of the octree to adaptively subdivide 3D space, using voxel grids as the basic unit of spatial discretization. These voxel grids are organized within the octree's hierarchical structure, with high-level nodes representing large-scale voxel grids and low-level nodes (child nodes) representing subdivided, smaller voxel grids. This creates a "coarse-to-fine" 3D spatial representation, making it a core tool for processing large-scale 3D data, such as point clouds, 3D models, and dynamic scenes.
[0029] In this embodiment, when constructing a voxel multi-layer structure tree, such as a voxel octree: The minimum bounding box of multiple original voxel grids is used as the root node of the voxel multi-layer structure tree, covering the entire scene to be reconstructed; According to the size of the original voxel grid, each original voxel grid is mapped to a different level corresponding to the multi-layer structure tree to form an initial level. Each original voxel grid corresponds to an initial level node. The formed pyramid structure includes a root node and an initial level node. Stores the characteristics of nodes at each level in the voxel multi-layer structure tree.
[0030] For example, when constructing a voxel multilayer structure tree such as a voxel octree, the root node covers the entire scene to be reconstructed, and the level l=0; according to the size of the original voxel grid, the initial level nodes are formed. Specifically, the larger original voxel grid (such as 1.0m) corresponds to the first level node of the voxel octree (l=1), the medium-sized original voxel grid (such as 0.5m) corresponds to the second level node of the voxel octree (l=2), and the smaller original voxel grid (such as 0.25m) corresponds to the third level node of the voxel octree (l=3). The size of each level node d l =1.0×2 -l m (l is the level, 0→7 corresponds to 1m→0.0078m), thus forming a pyramid structure, including the root node, initial layer nodes, etc.
[0031] Furthermore, the voxel octree can be dynamically adjusted later, such as splitting or merging nodes. When the initial level nodes are split, child nodes can be formed. The pyramid structure includes a root node, an initial level node, child nodes, etc.
[0032] In this embodiment, when storing the characteristics of nodes in each level: For nodes corresponding to voxel grids with static attributes, the normal vector, point density, and color mean of the corresponding voxel grid can be stored; For the nodes of the voxel grid corresponding to dynamic attributes, features such as normal vector, density, color, displacement and motion confidence can be stored. The motion confidence of the voxel grid with dynamic attributes can be mainly calculated by the displacement variance of the voxel grid between three consecutive frames:
[0033] Preferably, in a specific embodiment, when adjusting the original voxel grid based on the voxel multi-layer structure tree: According to the density change and feature similarity of the original voxel grid, some initial level nodes in the voxel multilayer structure tree are split to generate child nodes of a lower level, and the pyramid structure also includes child nodes, or some initial level nodes in the voxel multilayer structure tree are merged to generate nodes of a higher level, so as to split or merge the original voxel grids related to some initial level nodes; The node features between the voxel multi-layer structure trees obtained from the point cloud information of adjacent frames are smoothed.
[0034] In one specific embodiment, when adjusting the voxel multi-layer structure tree: In order to accurately capture details, in the original voxel grid, the branch where the first original voxel grid with dynamic attributes is located can be increased by 1 to 2 levels of detail. For the dynamic target object in the scene to be reconstructed corresponding to the first original voxel grid, the movement speed and direction of the dynamic target object are determined based on the continuous multi-frame point cloud information, and the subsequent movement area of the dynamic target object is predicted based on the movement speed and direction. Therefore, when dividing the voxel grid based on the subsequent frame point cloud information, the corresponding movement area is divided into fine-grained voxel grids.
[0035] Furthermore, when the curvature of the original voxel grid of static attributes is greater than 0.05, or the motion confidence of the original voxel grid of dynamic attributes is greater than 0.9, the branch level of these original voxel grids can be set to l=7 to ensure accurate capture of the details of the target object in the scene to be reconstructed.
[0036] In another embodiment, the original voxel grid is adjusted based on a voxel multilayer structure tree (such as a voxel octree) to adapt to density changes and feature similarities of related original voxel grids, specifically: If the point cloud density in a node in the voxel octree exceeds a threshold, such as 50%, the node is split, that is, a node is split into eight child nodes, so that the voxel grid corresponding to a node is split into eight voxel grids (corresponding to the eight nodes that are split); If the feature similarity of adjacent nodes in the voxel octree is higher than 90%, the adjacent nodes are merged to obtain a merged node, so that the voxel grids corresponding to the adjacent nodes are merged into one voxel grid (corresponding to the merged node); When the feature differences between adjacent nodes are too large and the point cloud density within the node is less than the threshold, the original voxel grid remains unchanged.
[0037] Furthermore, for nodes corresponding to static voxel grids in the voxel octree, normal vectors can be recalculated every 10 frames. This is typically done incrementally, using the previous frame's results and local changes in the current frame to quickly update the node normals for the current frame. For nodes corresponding to dynamic voxel grids in the voxel octree, displacements and confidence levels must be updated in real time every frame to reflect the motion of the corresponding target object.
[0038] Furthermore, the node features between the voxel multi-layer structure trees obtained from the point cloud information of adjacent frames can be smoothed. Specifically, the weighted average method can be used to fuse the feature vectors of the nodes of the current frame and the corresponding nodes of the previous frame: new feature = confidence × current frame node feature + (1-confidence) × previous frame node feature.
[0039] The confidence level of the target object in the voxel grid with dynamic attributes is high, so the corresponding proportion of the node features in the current frame is large. For example, when a vehicle brakes suddenly, the latest displacement data is used first. The confidence level of the target object in the voxel grid with static attributes is low, so the historical node features can be retained, such as roadside buildings, which can reduce inter-frame jitter.
[0040] It should be noted that, in the above process of optimizing the original voxel grids, before constructing the voxel multi-layer structure tree, preferably, each original voxel grid may be preprocessed first, and then the voxel multi-layer structure tree may be constructed based on the preprocessed original voxel grids. The preprocessing may include but is not limited to the following methods: In order to further increase the details of the reconstructed scene, the upsampling method can be used to further subdivide the original voxel grid. The voxel grid is progressively upsampled to , the size of the original voxel grid is reduced from 8cm to 2cm.
[0041] In order to minimize the amount of subsequent processing data, the original voxel grid in the invalid area of the original voxel grid can be filtered first. Specifically, the voxel occupancy rate Gσ(X) of the original voxel grid can be calculated = the number of point clouds in the voxel grid / voxel volume. When Gσ(X)>0.5, the original voxel grid is the voxel grid of the "valid area". When Gσ(X) is less than or equal to 0.5, the original voxel grid is the voxel grid of the "invalid area". Filtering is performed and only the voxel grid in the valid area is subsequently processed, which improves the computational efficiency.
[0042] Step 103 : determining the Gaussian model corresponding to each voxel grid according to the voxel features, and obtaining the multi-dimensional Gaussian features of the Gaussian model.
[0043] It can be understood that in the field of scene reconstruction and rendering, three-dimensional Gaussian modeling, as a typical explicit representation method, constructs a set of Gaussian distributions in three-dimensional space (each Gaussian distribution contains parameters such as position, shape, color, and weight), and projects them onto the image plane for weighted fusion based on the "Gaussian Splatting" rendering method, thereby achieving real-time rendering of high-precision scenes.
[0044] In this embodiment, the Gaussian model may be formed in this step by specifically performing the following steps: The voxel features obtained in step 102 can be used to determine the original Gaussian models corresponding to the voxel grids, and obtain multiple dimensional Gaussian features of the original Gaussian models; Compress the spatial position features and color features in multiple-dimensional Gaussian features to form a compressed Gaussian model; Preferably, the two branches of the dual-branch machine learning model can also be used to respectively obtain the first optimized value of the spatial position feature and the second optimized value of the color feature of the compressed Gaussian model, and optimize the spatial position feature and the color feature according to the first optimized value and the second optimized value to obtain the optimized Gaussian model.
[0045] Specifically, when determining the original Gaussian model, generally, at least one Gaussian model is generated within a voxel grid. In this process, it is mainly necessary to determine the number and radius of the Gaussian models within the voxel grid: The number of Gaussian models corresponding to a voxel grid is determined based on its voxel occupancy. Specifically, the scene to be reconstructed can be cut horizontally or vertically, the point cloud projected onto a plane, and the voxel occupancy Gσ(X) of each voxel grid estimated. When Gσ(X)>0.5, the Gaussian density within the voxel grid is increased, for example, 20 Gaussian models are generated per cubic meter. When Gσ(X)≤0.5, the number of Gaussian models within the voxel grid is reduced, or adjacent Gaussian models are merged (for example, 5-10 per cubic meter) to avoid redundant calculations.
[0046] Preferably, the Gaussian radius can also be dynamically adjusted by combining the voxel features of the voxel grid (such as curvature and height difference) and the semantic information of the point cloud within the voxel grid. Specifically, the actual category of the target object within the voxel grid can be obtained through a pre-trained semantic segmentation model. For example, for areas with high-detail objects such as building corners and vegetation edges, as well as high curvature areas of the voxel grid (such as curvature > 0.05), a small-radius Gaussian model (such as 0.1-0.3m) is used to preserve details; for smooth areas such as road surfaces and walls, a large-radius Gaussian model (such as 0.5-1m) is used to reduce the amount of computation.
[0047] Once the original Gaussian model is determined, multiple dimensional Gaussian features of the original Gaussian model can be obtained, including geometric features, radiation features, and spatial features: Geometric features can be obtained by fitting normal vectors and curvatures of neighboring point clouds, and the local surface change rate can be obtained by calculating the changes in normal vectors and curvatures between adjacent points. Radiation features include color variance (characterizing texture richness) and density gradient (detail saliency) based on texture directional features extracted from local binary patterns (LBP). Spatial features are sine-cosine position encoding: , enhancing location awareness.
[0048] In an embodiment, when compressing spatial position features and color features in multiple-dimensional Gaussian features: For spatial position feature compression: For voxel grids with different curvatures, different spatial position feature compression methods are used. For example, when the curvature of the voxel grid is low (such as <0.03), the third-order tensor T composed of Gaussian position coordinates can be decomposed into 32 groups of basis vectors through CP: When the curvature of the voxel grid is high (K ≥ 0.03), Tucker decomposition is used for the Gaussian position coordinate tensor, that is, T = G × 1 Ux × 2 Uy × 3 Uz, where G is the core tensor that stores the interaction information between dimensions, Ux, Uy, and Uz are the factor matrices of the corresponding dimensions, and the decomposition rank is set to 48 to capture complex geometric details.
[0049] For the compression of color features: specifically, the YCbCr color space can be used, 4-layer MLP fine encoding is used for the Y channel (luminance), and 2-layer MLP coarse encoding is used for the CbCr channel (chrominance). Combined with the visual attention mechanism: Ccompress = MLPY(Y) ⊕ MLPcbCr(CbCr) × Att(Y), the chrominance compression strength is dynamically adjusted, that is, fine encoding is used for the sub-features of the luminance channel, and coarse encoding is used for the sub-features of the chrominance channel.
[0050] In an embodiment, when optimizing the compressed Gaussian model, a dual-branch machine learning model, such as a dual-branch multi-layer perceptron (MLP) model, can be mainly used to obtain the first optimization value of the spatial position feature and the second optimization value of the color feature. Furthermore, the two branches of the dual-branch machine learning model can not only obtain the first optimization value and the second optimization value, but also respectively obtain the third optimization value of the shape feature of the compressed Gaussian model and the fourth optimization value of the weight information. For example Figure 3a As shown: Branch 1 of the MLP model is mainly used to optimize the spatial position features of the compressed Gaussian model, such as position and shape: a hierarchical attention mechanism is used. In an 8-layer fully connected network (128 dimensions per layer), a hierarchical attention module is introduced in each layer to focus on the voxel features of important voxel grids (such as high curvature areas and dynamic object boundaries). The 32-dimensional voxel features of the voxel grid are input into branch 1, and the Gaussian position offset of the compressed Gaussian model corresponding to the voxel grid is output. and scaling factor , namely the first optimization value and the third optimization value mentioned above, are used to adjust the position and shape of the compressed Gaussian model.
[0051] Branch 2 of the MLP model is mainly used to optimize the color features of the compressed Gaussian model, such as color and weight: based on the above-mentioned two-layer fully connected network, combined with the conditional generative adversarial network (cGAN) and activation function, the voxel grid voxel features and the semantic information of the scene to be reconstructed in the voxel grid can be used as conditions and input into branch 2. The RGB color correction value and weight coefficient, namely the second optimization value and the fourth optimization value mentioned above, can be output. In branch 2, the activation function can use Swish.
[0052] In this way, a joint loss function can be formed by the first optimization value, the second optimization value, the third optimization value and the fourth optimization value to optimize the Gaussian model formed above to form an optimized Gaussian model.
[0053] Furthermore, in a specific embodiment, when optimizing the Gaussian model, the following steps are also performed: For the original voxel grid with dynamic attributes, the displacement information of the next frame is predicted based on the displacement information of the previous multiple frames; based on the displacement information of the next frame and the first optimization value, second optimization value, third optimization value and fourth optimization value obtained by the two branches respectively, when optimizing the Gaussian model, such as Figure 3b As shown, a joint loss function can be formed based on the displacement information of the next frame and the first optimization value, the second optimization value, the third optimization value and the fourth optimization value obtained by the two branches to optimize the Gaussian model.
[0054] Step 104 : Rendering a Gaussian model according to the multi-dimensional Gaussian features to reconstruct a scene to be reconstructed.
[0055] In a specific embodiment, when rendering the Gaussian model, it can be achieved through the following steps: Constructing a hierarchical data structure of corresponding levels based on the multi-scale Gaussian feature pyramid and the sparse Gaussian update mechanism, wherein the levels of the multi-scale Gaussian feature pyramid correspond to the levels of the voxel multilayer structure tree formed by the voxel grid; Adjust the hierarchical data structure, such as the level of detail (LOD) hierarchical data structure, according to the viewing distance of the scene area to be reconstructed corresponding to the Gaussian model, so that different levels of LOD hierarchical data structure are used for different visual distances; Pruning the Gaussian model according to the global importance score of the Gaussian model, and adaptively adjusting the levels of the hierarchical data structure to form an adjusted hierarchical data structure; The sampling step size of ray tracing is determined according to the multi-dimensional Gaussian features, and the adjusted hierarchical data structure is rendered based on the determined sampling step size to form a scene to be reconstructed.
[0056] In this embodiment, when constructing a hierarchical data structure: Specifically, when rendering the Gaussian model, different rendering methods can be used based on Gaussian models of different resolutions. For example, for coarse-scale (such as 1m resolution) Gaussian models, the main purpose is to quickly render the entire reconstructed scene. Since the levels of the multi-scale Gaussian feature pyramid correspond to the levels of the voxel multi-layer structure tree formed above, the intersection detection of light and the Gaussian model can generally be accelerated based on the fast indexing structure of the voxel multi-layer structure tree (such as voxel octree) constructed above; and for fine-scale (such as 0.1m resolution) Gaussian models, the main purpose is to enhance local details, which can generally be sparsely processed through sparse representation learning algorithms to reduce the amount of data.
[0057] When rendering Gaussian models, the multi-dimensional Gaussian features of Gaussian models of different scales can be fused through trilinear interpolation with adaptive weights, where the weights can be dynamically adjusted according to the importance and similarity of Gaussian models of adjacent scales.
[0058] In this embodiment, when constructing a hierarchical data structure based on visual distance: The visual distance d of each area of the scene to be reconstructed refers to the straight-line distance from the rendering perspective to each area of the rendered scene to be reconstructed. When rendering a city street scene, the virtual camera is closer to nearby pedestrians and vehicles, so the visual distance is smaller, and farther away from distant buildings and the sky, the visual distance is larger. It can be seen that the visual distance can reflect the distant perspective or the near perspective of the corresponding area. Specifically, the relationship between the visual distance d and the level LOD of the hierarchical data structure can be expressed by the following function. In this way, the voxel-Gaussian model of a certain visual distance can form a hierarchical data structure of the corresponding level:
[0059] It can be seen that the area of the scene to be reconstructed with a farther visual distance has a smaller level of the hierarchical data structure, and can be rendered fuzzily with a lower resolution, while the area of the scene to be reconstructed with a closer visual distance has a larger level of the hierarchical data structure, and can be rendered accurately with a higher resolution.
[0060] In addition, if the voxel occupancy rate of the voxel grid with dynamic attributes in the scene to be reconstructed exceeds a certain value, such as 30%, or the user pays attention to the key area, the level of the hierarchical data structure is increased again, for example, by 1, and a Gaussian set of corresponding scale is selected from the Gaussian pyramid to ensure the rendering accuracy of the key area and the dynamic area.
[0061] Furthermore, when this embodiment dynamically adjusts the hierarchical data structure constructed above: The Gaussian model can be adjusted in combination with its global importance score GSj, thereby adaptively adjusting the level of the hierarchical data structure corresponding to the Gaussian model. This allows for dynamic adjustment of the hierarchical data structure constructed above. For example, the last 20% of the Gaussian models with the global importance scores can be pruned and the remaining Gaussian models can be distilled from the third order to the second order, thereby reducing the amount of rendering computation. The global importance score is a function related to the motion confidence of the voxel grid, as shown in the following formula:
[0062] in, is a kernel function that measures the correlation between the Gaussian model position and the features of the scene to be reconstructed. is the Gaussian density, is the motion confidence of the voxel grid (the higher the voxel occupancy of the voxel grid with dynamic properties, the higher the score).
[0063] Furthermore, in the embodiment, the sampling step size during ray tracing can be adjusted in real time, specifically including: Since ray tracing technology can be used for rendering at a high level of detail to provide more realistic visual effects, in this embodiment, after adaptively adjusting the hierarchical data structure constructed based on the Gaussian model, ray tracing can be used for hierarchical data structures that require high detail, using high-precision ray tracing to capture the detailed features of the Gaussian model through dense sampling; and the ray tracing process can be simplified for hierarchical data structures that require low detail.
[0064] In this embodiment, the sampling step size can be dynamically adjusted in real time during the ray tracing process, and can be adjusted in real time according to the multi-dimensional Gaussian features. Specifically, Figure 4 As shown in , the adjusted sampling step size is a function of the texture complexity of the target object corresponding to the Gaussian model, the geometric curvature of the Gaussian model, the corresponding motion confidence, and the Gaussian coverage. Δt represents the sampling step size during ray tracing, which is used to control the sampling interval in the scene during ray tracing. An appropriate sampling step size helps to improve rendering efficiency while ensuring rendering accuracy.
[0065] Represents the texture complexity of the target object corresponding to the Gaussian model. This term in the formula expresses the impact of texture complexity on the sampling step size. The more complex the texture, the smaller the sampling step size.
[0066] is the geometric curvature. This term in the formula expresses the influence of geometric curvature on the sampling step size. The larger the curvature, the smaller the formula and the smaller the step size. The limit is 0.3. The smaller the curvature, the larger the formula and the limit is 1.
[0067] It represents the motion confidence, which reflects the reliability of the voxel grid motion state. In the formula, this term expresses the influence of motion confidence on the sampling step size. The higher the motion confidence, the smaller the sampling step size.
[0068] Represents the radius of the Gaussian sphere, which measures the coverage of the Gaussian model in space. This term in the formula expresses the impact of the Gaussian sphere radius on the sampling step size. The larger the radius of the Gaussian sphere, the larger the sampling step size to avoid redundancy. The smaller the radius, the smaller the step size.
[0069] The overall formula mainly unifies the object material, geometric structure, motion confidence and representation granularity into a continuous control function, forming a method for adaptively controlling the light sampling step size of dynamic scenes. The coefficients in it can be adjusted according to actual conditions.
[0070] Furthermore, the sampling points obtained during ray sampling during ray tracing can be combined with the radiation characteristics of the Gaussian model and the RGB color correction value δ obtained when optimizing the Gaussian model above to eliminate the loss of high-frequency details of the corresponding voxel grid and improve rendering accuracy: Cfinal=Cvoxel+δ•αv, where Cvoxel is the original color of the voxel grid.
[0071] Furthermore, during the rendering process, real-time displacement prediction is performed to predict the displacement Δs^t+1 of the target object corresponding to each voxel grid of dynamic attributes in the next frame. This displacement is then offset in advance to eliminate artifacts. Specifically, displacement prediction can be performed using time series data processed by an LSTM module. For example, the displacement of the next frame can be predicted based on the displacements [Δst,…,Δst-4] of the previous frames (e.g., the first five frames).
[0072] It can be seen that in the method of this embodiment, after obtaining the point cloud information of any frame, the dynamic scene reconstruction system can divide the scene to be reconstructed into voxel grids of multiple scales based on the point cloud information, and obtain the voxel features of each voxel grid. Then, based on the voxel features, the corresponding Gaussian model is determined, and the multi-dimensional Gaussian features of the Gaussian model are obtained. Then, the Gaussian model is rendered based on the multi-dimensional Gaussian features to reconstruct the scene to be reconstructed. The main method is to combine the voxel grid and Gaussian model constructed based on the actual scene to be reconstructed to render the scene to be reconstructed. The two different feature construction methods are used to enhance the expression of details in the reconstruction of real dynamic scenes.
[0073] It should be noted that after the above steps 101 to 104, a frame of the scene to be reconstructed can be reconstructed. When multiple frames of the scene to be reconstructed are formed, when point cloud information of a new frame is subsequently obtained, it is not necessary to process all the point cloud information to update the subsequently reconstructed scene to be reconstructed. Instead, the following steps are performed to update the reconstructed scene to be reconstructed: The voxel grids with dynamic attributes in the voxel grids divided by the point cloud information of the subsequent frames are combined with the corresponding voxel grids of the previous multiple frames of the subsequent frame to calculate the relative displacement of the voxel grids with dynamic attributes between different frames; Determine feature information of the voxel grid of dynamic attributes according to the relative displacement and the corresponding voxel grids of adjacent frames; The above two-branch machine learning model is fine-tuned based on the feature information of the voxel grid with dynamic properties.
[0074] (1) Dynamically update the coordinates and features of the voxel grid to ensure the coherence of the target object's motion trajectory in the voxel grid with dynamic attributes. Specifically: For example, a sliding window is used to maintain the latest 20 frames of data, and the first 19 frames are reconstructed. For the point cloud data obtained in the 20th frame, after the voxel grid is divided, incremental ICP alignment can be started on the voxel grid with potential dynamic properties. The features of the voxel grid of the current frame are aligned with the features of the voxel grid of the previous 19 frames, aligned to a unified coordinate system, and the relative displacement of the voxel grid is calculated.
[0075] By combining the relative displacement (vold) of the voxel grid with potential dynamic properties and the information of the voxel grid of adjacent frames through a machine learning model (such as an MLP network), the voxel coordinates can be accurately updated: vnew = vold + MLP (vold, adjacent frame voxels), outputting the precise coordinates and feature vectors of the voxel grid with dynamic properties.
[0076] (2) Fine-tune the Gaussian model determined for subsequent frames.
[0077] After obtaining the feature information of the voxel grid with dynamic attributes based on the multi-frame point cloud information, during the training process of fine-tuning the dual-branch machine learning model, the parameter update threshold can be set to fine-tune only the dynamic feature layer parameters with larger changes in the dual-branch machine learning model and freeze the static feature layer.
[0078] Furthermore, the fine-tuning of the dual-branch machine learning model can be performed using asynchronous incremental training. When a new frame of data arrives, fine-tuning training begins immediately without affecting the dual-branch machine learning model's real-time optimization of the Gaussian model. This dynamic real-time adjustment of the dual-branch machine learning model allows for more precise optimization of the Gaussian model.
[0079] After training a dual-branch machine learning model, such as an MLP model, the output of a first optimization value, such as the aforementioned Δμ and ΔS, optimizes the position and size of the Gaussian model to accommodate the deformation of dynamic target objects in the reconstructed scene. A second optimization value can also be output to optimize the color features of the Gaussian model. Furthermore, the color features of the Gaussian model corresponding to the dynamic target object can be updated by interpolating between adjacent frames to reduce flicker.
[0080] The following is a specific application example to illustrate the dynamic scene reconstruction method of the embodiment of the present invention. Figure 5 Shown, including: Step 201 : obtaining point cloud information of multiple frames (eg, 2 or 3 frames) of a scene to be reconstructed, including point cloud density.
[0081] In step 202, voxel grids of corresponding scales are set for the scenes to be reconstructed in different regions according to the point cloud density of the scenes to be reconstructed, original voxel grids of multiple scales are obtained, and voxel features of the original voxel grids are acquired. In this embodiment, the voxel features may specifically include: point cloud color, density, dynamic displacement information, static voxel normal vector information, dynamic and static attributes of each voxel grid, etc.
[0082] Step 203: pre-process the original voxel grid to obtain a pre-processed original voxel grid. The following step 204 is performed on the pre-processed original voxel grid, specifically including: For example, on the one hand, the original voxel grid can be further subdivided by upsampling. On the other hand, in order to minimize the amount of subsequent processing data, the original voxel grid in the invalid area of the original voxel grid can be filtered first. For example, when the voxel occupancy rate Gσ(X) is less than or equal to 0.5, the original voxel grid is a voxel grid in the "invalid area". After filtering, only the voxel grid in the valid area is subsequently processed, which improves computational efficiency.
[0083] In step 204, the original voxel grid is optimized. For the optimized voxel grid, the following step 205 is executed. Specifically, when optimizing the original voxel grid: A1. Based on the voxel features of the original voxel grids, multiple original voxel grids are constructed into a voxel multilayer structure tree. Specifically: The minimum external bounding box of multiple original voxel grids is used as the root node of the voxel multi-layer structure tree, covering the entire scene to be reconstructed; according to the size of the original voxel grid, each original voxel grid is mapped to a different level corresponding to the multi-layer structure tree to form an initial level, and each original voxel grid corresponds to an initial level node. The formed pyramid structure includes a root node and an initial level node, etc.; the features of the nodes in each level of the voxel multi-layer structure tree are stored.
[0084] Among them, for the nodes of the voxel grid corresponding to static attributes, the normal vector, point density and color mean of the corresponding voxel grid can be stored; for the nodes of the voxel grid corresponding to dynamic attributes, features such as normal vector, density, color, displacement and motion confidence can be stored.
[0085] A2. Adjust the original voxel grid based on the voxel multi-layer structure tree to obtain an adjusted voxel grid. Specifically, for the adjusted voxel grid: According to the density change and feature similarity of the original voxel grid, some initial level nodes in the voxel multilayer structure tree are split to generate lower level child nodes, and the pyramid structure also includes child nodes; or some initial level nodes in the voxel multilayer structure tree are merged to generate higher level nodes. For example: If the point cloud density within a node in the voxel octree exceeds a threshold, such as 50%, the node is split, that is, a node is split into eight child nodes, so that the voxel grid corresponding to one node is split into eight voxel grids (corresponding to the eight child nodes). If the feature similarity of adjacent nodes in the voxel octree is higher than 90%, the adjacent nodes are merged to obtain a merged node, so that the voxel grids corresponding to the adjacent nodes are merged into one voxel grid (corresponding to the merged node). Furthermore, the node features of each node in the voxel multi-layer structure tree will be adjusted. For example, the node features between the voxel multi-layer structure trees obtained from the point cloud information of adjacent frames can be smoothed. Specifically, the weighted average method can be used to fuse the node feature vectors of the current frame and the corresponding node feature vectors of the previous frame: new feature = confidence × current frame node feature + (1-confidence) × previous frame node feature.
[0086] Step 205: Determine the Gaussian model corresponding to each voxel grid based on the voxel features, and obtain the multi-dimensional Gaussian features of the Gaussian model. This can be achieved by the following steps: B1. Determine the original Gaussian model corresponding to each voxel grid using the voxel features of the voxel grid after optimization in the above steps, and obtain multiple dimensional Gaussian features of the original Gaussian model. The multiple dimensional Gaussian features may include: geometric features, radiation features, and spatial features.
[0087] B2. Compress the spatial position features and color features in multiple-dimensional Gaussian features to form a compressed Gaussian model.
[0088] B3. Use the two branches of the dual-branch machine learning model to respectively obtain the first optimized value of the spatial position feature and the second optimized value of the color feature of the compressed Gaussian model, and optimize the spatial position feature and the color feature according to the first optimized value and the second optimized value to obtain the optimized Gaussian model.
[0089] Furthermore, the two branches of the dual-branch machine learning model can not only obtain the first optimization value and the second optimization value, but also respectively obtain the third optimization value of the shape features of the compressed Gaussian model and the fourth optimization value of the weight information. In this way, for the original voxel grid of dynamic attributes, the displacement information of the next frame can be predicted based on the displacement information of the previous multiple frames; based on the displacement information of the next frame, and the first optimization value, second optimization value, third optimization value and fourth optimization value obtained by the two branches respectively, a joint loss function is formed to optimize the Gaussian model.
[0090] Step 206: Render the Gaussian model based on the multi-dimensional Gaussian features to reconstruct the scene to be reconstructed. This may include: C1. Based on the multi-scale Gaussian feature pyramid and the sparse Gaussian update mechanism, a hierarchical data structure with corresponding levels is constructed, wherein the levels of the multi-scale Gaussian feature pyramid correspond to the levels of the voxel multi-layer structure tree formed by the above-mentioned voxel grid.
[0091] C2. Adjust the levels of the hierarchical data structure based on the viewing distance of the scene area to be reconstructed corresponding to the Gaussian model, using different levels of LOD hierarchical data structures for different visual distances; prune the Gaussian model based on its global importance score, and adaptively adjust the levels of the hierarchical data structure to form an adjusted hierarchical data structure; C3. Determine the sampling step size of ray tracing based on the multi-dimensional Gaussian features, and render the adjusted hierarchical data structure based on the determined sampling step size to form a frame of the scene to be reconstructed.
[0092] It should be noted that, through the above steps 201 to 206, a frame of the scene to be reconstructed can be reconstructed. By repeatedly executing 201 to 206, multiple frames of the scene to be reconstructed can be formed. After subsequently obtaining the point cloud information of a new frame, the following step 207 can also be executed.
[0093] Step 207, for the voxel grid of dynamic attributes in the voxel grid divided by the point cloud information of the subsequent frame, combined with the corresponding voxel grids of the previous multiple frames of the subsequent frame, calculate the relative displacement of the voxel grid of dynamic attributes between different frames; determine the feature information of the voxel grid of dynamic attributes based on the relative displacement and the corresponding voxel grids of adjacent frames; and fine-tune the above-mentioned dual-branch machine learning model based on the feature information of the voxel grid of dynamic attributes.
[0094] It can be seen that the following technical effects can be achieved through the above dynamic scene reconstruction method: 1. Divide the area of the scene to be reconstructed according to the point cloud density, generate a voxel grid of adaptive size, and construct a multi-scale voxel multi-layer structure tree, such as a voxel octree. It also distinguishes between dynamic and static voxel grids and stores motion features, such as dynamic displacement information; 2. Voxel-Gaussian mapping: The dynamically reconstructed scene is constructed as a multi-scale voxel radiation field. Each voxel grid contains radiation properties and geometric features. The multi-scale characteristics of the voxel grid are combined to dynamically adjust the Gaussian model, filling the gap in the lack of geometric features in Gaussian modeling. 3. Hybrid data compression: A curvature threshold (0.03) is introduced to switch compression modes to avoid detail loss in CP decomposition in high-curvature areas. Color parameter compression uses a 4-layer MLP network coding to address memory bottlenecks in large scenes. 4. Learn dynamic features through a two-branch machine learning model such as the MLP model and optimize the different properties of the Gaussian model; 5. Adaptively adjust the levels of the hierarchical data structure LOD based on visual distance and dynamic weight, and optimize the rendering process by combining pruning and distillation of the Gaussian model; 6. Light sampling strategies during rendering based on voxel feature weights, curvature, radiation radius (such as Gaussian coverage), and motion compensation; 7. Utilize the detail residuals output by the dual-branch machine learning model to modify the voxel radiation properties, breaking through the bottleneck of voxel grid representation of high-frequency details.
[0095] The embodiment of the present invention also provides a dynamic scene reconstruction system, the structural diagram of which is shown in FIG. Figure 6 Specifically, it may include: The point cloud acquisition unit 10 is used to acquire point cloud information of multiple frames of the scene to be reconstructed.
[0096] The voxel grid unit 11 is configured to divide the scene to be reconstructed into voxel grids of multiple scales based on the multiple frames of point cloud information acquired by the point cloud acquisition unit 10 , and acquire voxel features of each voxel grid.
[0097] The Gaussian model unit 12 is configured to determine the Gaussian models corresponding to the respective voxel grids according to the voxel features acquired by the voxel grid unit 11 , and to acquire multi-dimensional Gaussian features of the Gaussian models.
[0098] The rendering unit 13 is configured to render the Gaussian model according to the multi-dimensional Gaussian features acquired by the Gaussian model unit 12 to reconstruct the scene to be reconstructed.
[0099] Among them, if the point cloud information includes point cloud density, the voxel grid unit 11 is specifically used to set voxel grids of corresponding scales for different areas of the scene to be reconstructed according to the point cloud density of different areas of the scene to be reconstructed, obtain original voxel grids of multiple scales, and obtain voxel features of the original voxel grids; the voxel features include dynamic and static properties of the original voxel grids.
[0100] Furthermore, after obtaining the plurality of original voxel grids, the voxel grid unit 11 is further configured to construct the plurality of original voxel grids into a voxel multilayer structure tree based on the voxel features of the original voxel grids, and adjust the original voxel grids based on the voxel multilayer structure tree to obtain adjusted voxel grids; and perform the step of determining the Gaussian model corresponding to each voxel grid based on the voxel features on the adjusted voxel grids. Specifically, when constructing the voxel multilayer structure tree, the voxel grid unit 11 is configured to use the minimum external bounding box of the plurality of original voxel grids as the root node of the voxel multilayer structure tree, covering the entire scene to be reconstructed; map each original voxel grid to a different level corresponding to the multilayer structure tree based on the size of the original voxel grids to form an initial level, with each original voxel grid corresponding to an initial level node, and the formed pyramid structure includes the root node and the initial level node; and store the features of the nodes at each level in the voxel multilayer structure tree.
[0101] The voxel grid unit 11 is also used to adjust the original voxel grid, specifically to split certain initial level nodes in the voxel multi-layer structure tree according to the density change and feature similarity of the original voxel grid to generate lower-level child nodes, and the pyramid structure also includes the child nodes, or merge certain initial level nodes in the voxel multi-layer structure tree to generate higher-level nodes, so as to split or merge the original voxel grids related to the certain nodes; and to smooth the node features between the voxel multi-layer structure trees obtained from the adjacent frame point cloud information.
[0102] Furthermore, the above-mentioned Gaussian model unit 12 is specifically used to determine the original Gaussian model corresponding to each voxel grid according to the voxel features, and obtain multiple-dimensional Gaussian features of the original Gaussian model; compress the spatial position features and color features in the multiple-dimensional Gaussian features to form a compressed Gaussian model; the two branches of the dual-branch machine learning model respectively obtain the first optimized value of the spatial position feature and the second optimized value of the color feature of the compressed Gaussian model, and optimize the spatial position feature and color feature according to the first optimized value and the second optimized value to obtain the optimized Gaussian model.
[0103] Furthermore, the two branches also respectively obtain the third optimization value of the shape feature of the compressed Gaussian model and the fourth optimization value of the weight information. Then the Gaussian model unit 12 is also used to predict the displacement information of the next frame for the original voxel grid of dynamic attributes based on the displacement information of the previous multiple frames; and optimize the Gaussian model based on the displacement information of the next frame and the first optimization value, second optimization value, third optimization value and fourth optimization value respectively obtained by the two branches.
[0104] The rendering unit 13 is specifically used to construct a hierarchical data structure with corresponding levels based on a multi-scale Gaussian feature pyramid and a sparse Gaussian update mechanism, wherein the levels of the multi-scale Gaussian feature pyramid correspond to the levels of the voxel multi-layer structure tree formed by the voxel grid; adjust the levels of the hierarchical data structure according to the visual distance of the scene area to be reconstructed corresponding to the Gaussian model; prune the Gaussian model according to the global importance score of the Gaussian model, and adaptively adjust the levels of the hierarchical data structure to form an adjusted hierarchical data structure; determine the sampling step size of ray tracing according to the multi-dimensional Gaussian feature, and render the adjusted hierarchical data structure based on the determined sampling step size to form the scene to be reconstructed.
[0105] Furthermore, the dynamic scene reconstruction system of this embodiment also includes: The incremental data unit 14 is used to calculate the relative displacement of the voxel grid of dynamic attributes in the voxel grid divided by the point cloud information of the subsequent frame, in combination with the corresponding voxel grids of the previous multiple frames of the subsequent frame; determine the feature information of the voxel grid of dynamic attributes based on the relative displacement and the corresponding voxel grids of adjacent frames; and fine-tune the dual-branch machine learning model based on the feature information of the voxel grid of dynamic attributes.
[0106] The system in this embodiment mainly combines the voxel grid constructed based on the actual scene to be reconstructed with the Gaussian model to render the scene to be reconstructed, and adopts two different feature construction methods to enhance the expression of details in the reconstruction of the real dynamic scene.
[0107] The embodiment of the present invention further provides a terminal device, the structural diagram of which is shown in FIG. Figure 7 As shown, the terminal device may vary significantly due to configuration or performance differences. It may include one or more central processing units (CPUs) 20 (e.g., one or more processors), memory 21, and one or more storage media 22 (e.g., one or more mass storage devices) storing application programs 221 or data 222. The memory 21 and storage media 22 may be either transient or persistent storage. The programs stored in the storage media 22 may include one or more modules (not shown), each of which may include a series of instructions for operating on the terminal device. Furthermore, the CPU 20 may be configured to communicate with the storage medium 22 to execute the series of instructions stored in the storage medium 22 on the terminal device.
[0108] Specifically, the application 221 stored in the storage medium 22 includes an application for dynamic scene reconstruction, and the application may include the point cloud acquisition unit 10, the voxel grid unit 11, the Gaussian model unit 12, the rendering unit 13, and the incremental data unit 14 in the above-mentioned dynamic scene reconstruction system, which are not described in detail here. Furthermore, the central processing unit 20 can be configured to communicate with the storage medium 22 and execute a series of operations corresponding to the application for dynamic scene reconstruction stored in the storage medium 22 on the terminal device.
[0109] The terminal device may also include one or more power supplies 23, one or more wired or wireless network interfaces 24, one or more input and output interfaces 25, and / or one or more operating systems 223, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0110] The steps performed by the dynamic scene reconstruction system in the above method embodiment can be based on the Figure 7 The structure of the terminal device shown.
[0111] Furthermore, another aspect of an embodiment of the present invention provides a computer-readable storage medium, which stores a plurality of computer programs, and the computer programs are suitable for being loaded by a processor and executing the dynamic scene reconstruction method performed by the above-mentioned dynamic scene reconstruction system.
[0112] Another aspect of the present invention provides a terminal device, including a processor and a memory; The memory is used to store multiple computer programs, and the computer programs are used to be loaded by the processor and execute the dynamic scene reconstruction method performed by the above-mentioned dynamic scene reconstruction system; the processor is used to implement each computer program in the multiple computer programs.
[0113] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc.
[0114] The above is a detailed introduction to a dynamic scene reconstruction method, system and related equipment provided by an embodiment of the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
Claims
1. A dynamic scene reconstruction method, characterized in that: include: Obtaining point cloud information of multiple frames of the scene to be reconstructed; For multiple frames of point cloud information, the scene to be reconstructed is divided into voxel grids of multiple scales, and voxel features of each voxel grid are obtained; Determining Gaussian models corresponding to the respective voxel grids according to the voxel features, and obtaining multi-dimensional Gaussian features of the Gaussian models; The Gaussian model is rendered according to the multi-dimensional Gaussian features to reconstruct the scene to be reconstructed.
2. The method according to claim 1, wherein The point cloud information includes point cloud density. The point cloud information of multiple frames is divided into voxel grids of multiple scales, and voxel features of each voxel grid are obtained, specifically including: According to the point cloud density of different regions of the scene to be reconstructed, voxel grids of corresponding scales are set for the different regions of the scene to be reconstructed, to obtain original voxel grids of multiple scales, and to obtain voxel features of the original voxel grids; The voxel features include dynamic and static properties of the original voxel grid.
3. The method according to claim 2, wherein After obtaining the plurality of original voxel grids, the method further includes: constructing a plurality of original voxel grids into a voxel multilayer structure tree according to voxel features of the original voxel grids, and adjusting the original voxel grids based on the voxel multilayer structure tree to obtain adjusted voxel grids; The step of determining the Gaussian models corresponding to the respective voxel grids according to the voxel features is performed on the adjusted voxel grids.
4. The method according to claim 3, wherein The step of constructing a plurality of original voxel grids into a voxel multilayer structure tree based on the voxel features of the original voxel grids specifically includes: Taking the minimum external bounding box of the plurality of original voxel grids as the root node of the voxel multi-layer structure tree, covering the entire scene to be reconstructed; Mapping each original voxel grid to a different level corresponding to a multi-layer structure tree according to the size of the original voxel grid to form an initial level, wherein each original voxel grid corresponds to an initial level node, and the formed pyramid structure includes the root node and the initial level node; The features of the nodes in each level of the voxel multi-layer structure tree are stored.
5. The method according to claim 4, wherein The adjusting the original voxel grid based on the voxel multi-layer structure tree specifically includes: Splitting certain initial-level nodes in the voxel multilayer structure tree to generate child nodes of a lower level according to density changes and feature similarities of the original voxel grids, wherein the pyramid structure also includes the child nodes, or merging certain initial-level nodes in the voxel multilayer structure tree to generate nodes of a higher level, so as to split or merge the original voxel grids related to the certain nodes; The node features between the voxel multi-layer structure trees obtained from the point cloud information of adjacent frames are smoothed.
6. The method according to any one of claims 1 to 5, characterized in that The Gaussian models corresponding to the respective voxel grids are determined according to the voxel features, and multi-dimensional Gaussian features of the Gaussian models are obtained: Determining original Gaussian models corresponding to the respective voxel grids according to the voxel features, and obtaining Gaussian features of multiple dimensions of the original Gaussian models; The spatial position features and color features in the multiple dimensional Gaussian features are compressed to form a compressed Gaussian model.
7. The method according to claim 6, wherein After forming the compressed Gaussian model, the method further includes: The two branches of the dual-branch machine learning model respectively obtain the first optimized value of the spatial position feature and the second optimized value of the color feature of the compressed Gaussian model, and optimize the spatial position feature and color feature according to the first optimized value and the second optimized value to obtain the optimized Gaussian model.
8. The method according to claim 7, wherein The two branches further respectively obtain a third optimized value of the shape feature of the compressed Gaussian model and a fourth optimized value of the weight information, and the method further includes: For the original voxel grid with dynamic attributes, the displacement information of the next frame is predicted based on the displacement information of the previous frames; The Gaussian model is optimized according to the displacement information of the next frame and the first optimization value, the second optimization value, the third optimization value and the fourth optimization value respectively obtained by the two branches.
9. The method according to any one of claims 1 to 5, characterized in that Rendering the Gaussian model according to the multi-dimensional Gaussian features specifically includes: Constructing a hierarchical data structure of corresponding levels based on a multi-scale Gaussian feature pyramid and a sparse Gaussian update mechanism, wherein the levels of the multi-scale Gaussian feature pyramid correspond to the levels of the voxel multi-layer structure tree formed by the voxel grid; Adjusting the level of the hierarchical data structure according to the visual distance of the to-be-reconstructed scene area corresponding to the Gaussian model; Pruning the Gaussian model according to the global importance score of the Gaussian model, and adaptively adjusting the levels of the hierarchical data structure to form an adjusted hierarchical data structure; A sampling step size for ray tracing is determined according to the multi-dimensional Gaussian feature, and the adjusted hierarchical data structure is rendered based on the determined sampling step size to form the scene to be reconstructed.
10. The method according to any one of claims 1 to 5, characterized in that The method further comprises: For a voxel grid with dynamic attributes in a voxel grid divided by point cloud information of a subsequent frame, the relative displacement of the voxel grid with dynamic attributes between different frames is calculated by combining corresponding voxel grids of a plurality of frames preceding the subsequent frame; determining feature information of the voxel grid of the dynamic attribute according to the relative displacement and the voxel grids corresponding to the adjacent frames; A two-branch machine learning model is fine-tuned based on the feature information of the voxel grid of the dynamic attributes.
11. A dynamic scene reconstruction system, characterized in that: include: A point cloud acquisition unit, used to acquire point cloud information of multiple frames of the scene to be reconstructed; A voxel grid unit is used to divide the scene to be reconstructed into voxel grids of multiple scales based on the point cloud information of multiple frames, and obtain voxel features of each voxel grid; A Gaussian model unit, configured to determine Gaussian models corresponding to the respective voxel grids according to the voxel features, and obtain multi-dimensional Gaussian features of the Gaussian models; A rendering unit is used to render the Gaussian model according to the multi-dimensional Gaussian features to reconstruct the scene to be reconstructed.
12. The dynamic scene reconstruction system according to claim 11, wherein: The voxel grid unit is specifically used to set voxel grids of corresponding scales for different areas of the scene to be reconstructed according to the point cloud density of different areas of the scene to be reconstructed, obtain original voxel grids of multiple scales, and obtain voxel features of the original voxel grids; the voxel features include dynamic and static properties of the original voxel grids; according to the voxel features of the original voxel grids, multiple original voxel grids are constructed into a voxel multi-layer structure tree, and the original voxel grids are adjusted based on the voxel multi-layer structure tree to obtain adjusted voxel grids; for the adjusted voxel grids, the Gaussian model unit is notified to determine the Gaussian models corresponding to each voxel grid according to the voxel features.
13. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a plurality of computer programs, and the computer programs are suitable for being loaded by a processor and executing the dynamic scene reconstruction method according to any one of claims 1 to 10.
14. A terminal device, characterized in that: including processor and memory; The memory is used to store multiple computer programs, and the computer programs are used to be loaded by the processor and execute the dynamic scene reconstruction method according to any one of claims 1 to 10; the processor is used to implement each computer program in the multiple computer programs.
Citation Information
Patent Citations
Point cloud matching method and system based on derivative-free optimization
CN118314180A
Structured Gaussian splashing method based on image and radar data
CN119991902A
Multi-modal data fusion method and device based on dynamic Gaussian modeling and vehicle
CN120259827A
Interactive Relighting of Dynamic Refractive Objects
US20100033482A1
Cited By
Dynamic scene incremental reconstruction and rendering method based on 3DGS
CN120976447A
Three-dimensional grid construction method, electronic equipment, scanning equipment and storage medium
CN121053309A