Dynamic scene reconstruction method and device, equipment and medium

Through the combination of the 3D Gaussian sputtering module and the dynamic component module, the static scene Gaussian changes are generated, which solves the problem that dynamic scene reconstruction is difficult to achieve real-time rendering in the prior art, and realizes the dynamic scene reconstruction effect with high fidelity and low resource consumption.

CN120014150APending Publication Date: 2025-05-16TONGJI UNIV

Patent Information

Application Number
CN202411824799.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-12
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing dynamic scene reconstruction technology is difficult to achieve real-time rendering on the basis of ensuring high-fidelity scene reconstruction, and the computing resources are consumed and the training efficiency is inefficient.

Method used

The 3D Gaussian sputtering module is used for preliminary scene reconstruction, combining the dynamic component module and the color component module to generate the change in the position and color attributes of the Gaussian in the static scene to realize real-time reconstruction and rendering of the dynamic scene.

Benefits of technology

It realizes reconstruction and real-time rendering of high-fidelity dynamic scenes, reduces the consumption of computing resources, improves training efficiency, and makes the scene structure clearer and the data easy to interpret.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014150A_ABST
    Figure CN120014150A_ABST
Patent Text Reader

Abstract

The invention relates to a dynamic scene reconstruction method and device, equipment and a medium, and the method comprises the steps: obtaining a picture sequence; the image sequences are input into a dynamic scene reconstruction model, a dynamic scene between the image sequences is generated, and the dynamic scene reconstruction model comprises a 3D Gaussian sputtering module used for conducting preliminary scene reconstruction on the image sequences to obtain static scene Gaussian; the dynamic component module is used for generating the variable quantity of the position of the static scene Gaussian according to the position attribute of the static scene Gaussian in combination with the time characteristic; and the color component module is used for generating the variable quantity of the Gaussian color attribute of the static scene by combining the Gaussian motion deformation characteristics according to the Gaussian color attribute of the static scene. According to the invention, a real-time rendering function can be realized on the basis of ensuring high-fidelity scene reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of dynamic scene reconstruction, and in particular to a dynamic scene reconstruction method, device, equipment and medium. Background Art

[0002] Dynamic scene reconstruction is a core task in computer vision and computational graphics. Its goal is to reconstruct a realistic 3D model containing scene geometry, material properties, and dynamic changes from multi-view observation data (such as images, videos, or depth information). Compared with static scene reconstruction, it needs to deal with object motion, surface deformation, and illumination changes related to the time dimension, while ensuring the temporal consistency of the results. This technology is widely used in filmmaking, virtual reality, medical analysis, and robot navigation, but faces challenges such as complex data acquisition, high computational cost, occlusion processing, and realistic expression of dynamic changes. Traditional multi-view geometry-based technologies (such as SLAM and SfM) are mainly applicable to static scenes and have limited characterization of the time dimension; while deep learning-based methods, such as spatiotemporal networks and dynamic neural rendering based on neural radiance field extension, can reconstruct scenes containing complex time-varying features to a certain extent by analyzing and modeling continuous image sequences. However, these methods usually rely on pre-annotated keyframes or high-density sampling data, have high requirements on computing resources, and have problems of low training efficiency and high resource consumption when dealing with large-scale dynamic scenes.

[0003] The existing open patent document CN118485796A discloses a virtual reality and augmented reality fusion method based on indoor three-dimensional reconstruction, which uses feature descriptors and fusion constraint information to reconstruct the scene in three dimensions. This method has a relatively poor reconstruction effect for dynamic scenes, and the description scale space, especially the time scale, of the feature points represented by the feature descriptors is more difficult to design the model, and the consistency of time and space needs to be considered. The existing open patent document CN118710807A discloses an underwater scene representation method based on neural radiation field, which uses a fully connected neural network to model scene information, and after inputting the position and perspective posture, the density and color of the scene objects are obtained, and then the color of the pixel is obtained by volume rendering, wherein multiple MLPs are used, including MLP-S, MLP-D and MLP-I, which require a lot of time in training and rendering. The existing open patent document CN114973355A also uses a multi-layer perceptron MLP network based on neural radiation field to obtain the feature vector of the face. Although this implicit representation can obtain continuous dynamic changes, after implicit encoding, there are also problems of long training time and insufficient accuracy. Summary of the invention

[0004] The technical problem to be solved by the present invention is to provide a dynamic scene reconstruction method, device, equipment and medium, which can realize real-time rendering function on the basis of ensuring high-fidelity scene reconstruction.

[0005] The technical solution adopted by the present invention to solve the technical problem is: to provide a dynamic scene reconstruction method, comprising the following steps:

[0006] Get the image sequence and the corresponding camera pose;

[0007] The image sequence and the corresponding camera pose are input into a dynamic scene reconstruction model to generate a dynamic scene between the image sequences, wherein the dynamic scene reconstruction model includes:

[0008] A 3D Gaussian sputtering module, used for performing preliminary scene reconstruction on the image sequence based on the camera pose to obtain a static scene Gaussian;

[0009] The dynamic component module is used to generate the change of the position of the static scene Gaussian according to the position attribute of the static scene Gaussian and the time feature;

[0010] The color component module is used to generate the change amount of the color attribute of the static scene Gaussian according to the color attribute of the static scene Gaussian combined with the motion deformation characteristics of the Gaussian.

[0011] The 3D Gaussian sputtering module adopts Normal=F normalize (F align (F sort (s t ),r t ,v t )) Calculate the normal line approximated by the shortest axis. Normal is the normal line approximated by the shortest axis. F sort () is the function of sorting to obtain the shortest axis in Gaussian, F align () is the function of adjusting the shortest axis to the outer surface, F normalize () is the normalization function, s t is the scaling factor at time t, r t is the quaternion at time t, v t is the viewing direction at time t.

[0012] When the 3D Gaussian sputtering module optimizes the normal approximated by the shortest axis, the normal approximated by the shortest axis is first passed to the rasterizer to obtain a rendered depth map; then, based on the rendered depth map, the corresponding reference normal map is calculated using gradient information; finally, the normal approximated by the shortest axis is optimized by minimizing the difference between the rendered depth map and the reference normal map.

[0013] The dynamic component module includes:

[0014] A position encoding unit is used to encode the position attribute of the static scene Gaussian at time t to obtain the position feature;

[0015] Extract the splicing unit, which is used to project the position features onto six planes so that each point can be projected onto a plane formed by a pair of coordinate axes. The feature vectors of the projected points on the paired planes are multiplied respectively and then connected into a feature vector as the feature of the final Gaussian position point in the four-dimensional space.

[0016] The deformation decoder unit of the multi-layer perceptron is used to decode the features of the points of the Gaussian position in the four-dimensional space to obtain the change in the position of the Gaussian in the static scene.

[0017] The color component module includes:

[0018] A color encoding unit is used to encode the color attributes of the static scene Gaussian at time t to obtain color features;

[0019] Deformable perceptron unit, used to predict the color change of objects based on color features and Gaussian motion deformation features;

[0020] The color decoder unit of the multi-layer perceptron is used to decode the color change of the object to obtain the change amount of the color attribute of the static scene Gaussian.

[0021] The dynamic scene reconstruction model also includes: a shape destruction module, which is used to adjust the shape of the deformed Gaussian when a preset number of iterations is reached during training.

[0022] The technical solution adopted by the present invention to solve the technical problem is: to provide a dynamic scene reconstruction device, comprising:

[0023] The acquisition module is used to obtain the image sequence and the corresponding camera pose;

[0024] A reconstruction module is used to input the picture sequence and the corresponding camera pose into a dynamic scene reconstruction model to generate a dynamic scene between the picture sequences, wherein the dynamic scene reconstruction model includes:

[0025] A 3D Gaussian sputtering module, used for performing preliminary scene reconstruction on the image sequence based on the camera pose to obtain a static scene Gaussian;

[0026] The dynamic component module is used to generate the change of the position of the static scene Gaussian according to the position attribute of the static scene Gaussian and the time feature;

[0027] The color component module is used to generate the change amount of the color attribute of the static scene Gaussian according to the color attribute of the static scene Gaussian combined with the motion deformation characteristics of the Gaussian.

[0028] The technical solution adopted by the present invention to solve its technical problem is: to provide an electronic device, including a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned dynamic scene reconstruction method when executing the computer program.

[0029] The technical solution adopted by the present invention to solve its technical problem is: providing a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned dynamic scene reconstruction method are implemented.

[0030] Beneficial Effects

[0031] Due to the adoption of the above-mentioned technical scheme, the present invention has the following advantages and positive effects compared with the prior art: the present invention is based on 3D Gaussian splashing and integrates motion deformation field and color deformation field, and can process dynamic scenes with material color changes. Different from the traditional implicit scene representation that relies on deep learning, the present invention adopts an explicit representation method to make the scene structure clearer, the data easy to interpret, and the reconstruction process more efficient, so that the reconstruction results of dynamic scenes can be observed and analyzed more easily. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 is a flow chart of a dynamic scene reconstruction method according to a first embodiment of the present invention;

[0033] Figure 2 is a schematic diagram of a dynamic scene reconstruction model in a first embodiment of the present invention;

[0034] Figure 3 is a schematic diagram of a dynamic component module and a color component module in a first embodiment of the present invention;

[0035] Figure 4 It is a schematic diagram of an application scheme of an embodiment of the present invention. DETAILED DESCRIPTION

[0036] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and are not intended to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms fall within the scope limited by the appended claims of the application equally.

[0037] The first embodiment of the present invention relates to a dynamic scene reconstruction method, such as Figure 1 As shown, the following steps are included:

[0038] Step 1: Get the image sequence and the corresponding camera pose.

[0039] Step 2: Input the image sequence and the corresponding camera pose into a dynamic scene reconstruction model to generate a dynamic scene between the image sequences.

[0040] like Figure 2 As shown, the dynamic scene reconstruction model in this embodiment includes:

[0041] A 3D Gaussian sputtering module, used for performing preliminary scene reconstruction on the image sequence based on the camera pose to obtain a static scene Gaussian;

[0042] The dynamic component module is used to generate the change of the position of the static scene Gaussian according to the position attribute of the static scene Gaussian and the time feature;

[0043] The color component module is used to generate the change amount of the color attribute of the static scene Gaussian according to the color attribute of the static scene Gaussian combined with the motion deformation characteristics of the Gaussian.

[0044] The 3D Gaussian sputtering module in this embodiment uses the original 3DGS to perform preliminary scene reconstruction on the input pictures and camera poses, with the goal of quickly constructing the rough geometry and basic properties of the scene. Since the input picture sequence may have large differences, especially in dynamic scenes, changes in object shape and properties will lead to poor reconstruction results of the original 3DGS method. Therefore, the optimization of the 3D Gaussian sputtering module does not pursue high-precision results.

[0045] 3D Gaussians is a new display scene representation that generates sparse point clouds from sfm and constructs a set of Gaussians. Each 3D Gaussian is defined by the center position x and the covariance matrix, where the covariance matrix can be decomposed into a scaling matrix S and a rotation matrix R. The covariance matrix is ​​expressed as: Σ = RSS T R T , which can use the scaling factor s∈S 3 and the quaternion r∈R 4 3DGS improves the geometry of 3D Gaussian by optimizing the scaling factor s, quaternion r, position x, pruning and densification. When rendering a new view, the view transformation matrix W and the affine approximation Jacobian matrix J of the projective transformation are used to obtain the covariance matrix in the camera space: Σ′=JWΣW T J T Finally, the color of each pixel can be obtained by Calculate, where C represents the color of the final pixel, c i represents the color of the i-th Gaussian, α i represents the opacity of the i-th Gaussian, T i represents the transmittance reaching the i-th Gaussian.

[0046] When fitting the surface of an object, the geometry of the Gaussian sphere is crucial, especially for planar objects in the scene. During the training process, the normals of the various Gaussians that make up the plane should maintain consistency in direction to keep the plane smooth. The 3D Gaussian sputtering module of this embodiment needs to introduce the concept of Gaussian normals to describe and constrain the directionality of the Gaussian. Since the shape of the Gaussian is constantly changing over time in dynamic scenes, the shortest axis of the deformed Gaussian is constantly changing. If the normal is stored as a property of the Gaussian, the storage overhead will increase significantly, and it will be difficult to adapt to the needs of Gaussian normals changing over time in dynamic scenes. Therefore, this embodiment calculates the normal approximated by the shortest axis in the following way:

[0047] Normal=F normalize (F align (F sort (s t ),r t ,v t ));

[0048] Among them, Normal is the normal line approximated by the shortest axis, F sort () is the function of sorting to obtain the shortest axis in Gaussian, F align () is the function of adjusting the shortest axis to the outer surface, F normalize () is the normalization function, s t is the scaling factor at time t, r t is the quaternion at time t, v t is the viewing direction at time t.

[0049] For the optimization of normals, the normal data is first passed to the rasterizer to obtain the rendered depth map n1; then based on the rendered depth map n1, the corresponding reference normal map n2 is calculated using the gradient information, and finally the difference between the rendered depth map n1 and the reference normal map n2 (i.e., L normal =||n1-n2|| 2 ) is used to optimize the normal. Since this method calculates the normal in real time, it has good adaptability to dynamic changes in the scene.

[0050] like Figure 3As shown, the dynamic component module of this embodiment includes: a position encoding unit, which is used to encode the position attribute of the static scene Gaussian at time t to obtain the position feature; an extraction and splicing unit, which is used to project the position feature onto six planes so that each point can be projected onto a plane formed by a pair of coordinate axes, and the feature vectors of the projected points on the paired planes are multiplied respectively, and then connected into a feature vector as the feature of the final Gaussian position in the four-dimensional space; a deformation decoder unit of the multi-layer perceptron is used to decode the feature of the Gaussian position in the four-dimensional space to obtain the change in the position of the static scene Gaussian.

[0051] After the 3D Gaussian sputtering module completes the rough Gaussian geometry reconstruction, the dynamic component module of this embodiment refers to and modifies the hexplane method to continue processing dynamic scenes. In dynamic scenes, it is necessary to process changes in the time dimension, and the Gaussian coordinates are expanded to 4 dimensions (x, y, z, t). For a Gaussian with a position of (x, y, z) at a certain time t, the hexplane (i.e., the extraction splicing unit) projects (x, y, z, t) onto 6 planes, and each point can be projected onto a plane consisting of a pair of coordinate axes (such as xy and zt or xz and yt). The feature vectors of the projected points on the paired planes are multiplied separately, and then connected into a feature vector as the feature of the point in the four-dimensional space of the Gaussian with the final position (x, y, z, t), which can be expressed as:

[0052] f h =hex(x,y,z,t)=(P x,y ·P z,t )+(P x,z ·P y,t )+(P x,t ·P y,z );

[0053] Among them, P x,y Represents the corresponding features of the point projected onto the xy plane, and the paired plane P z,t After element-wise multiplication of the features, all the results are concatenated.

[0054] In this implementation, before inputting the Gaussian position coordinates (x, y, z, t) into the extraction splicing unit, position encoding is required to improve the expression ability of position features, so as to better capture the complex changes in dynamic scenes. The position encoding process is as follows:

[0055] E p =(sin(2 k πp),cos(2 k πp)),k=0,1,…,L-1;

[0056] Among them, p is the position coordinate before encoding, Ep is the encoded position coordinate.

[0057] In order to further capture the changes in Gaussian geometric features in dynamic scenes, this embodiment also uses a deformation decoder unit of a multi-layer perceptron (MLP), which extracts the feature vector f generated by the splicing unit h As input, the amount of change in the Gaussian attribute is predicted.

[0058] In the deformation decoder unit of the multi-layer perceptron, a separate decoder is designed for each attribute change, and the feature vector f h Mapped to the corresponding change. For example, the change in Gaussian position D p =φ p (hex(x,y,z))=φ p (f h )=(Δx,Δy,Δz), the change in scaling D s =φ s (f h ) = Δs, the change in rotation D r =φ r (f h )=Δr.

[0059] In this implementation, part of the network structure used for dynamic scene reconstruction is designed as an independent dynamic component module. This componentized design not only improves the flexibility of the method, but also allows the complexity of scene reconstruction to be dynamically adjusted according to demand.

[0060] like Figure 3 As shown, the color component module in this embodiment includes: a color encoding unit, which is used to encode the color attributes of the static scene Gaussian at time t to obtain color features; a deformable perceptron unit, which is used to predict the color change of the object based on the color features and the motion deformation features of the Gaussian; a color decoder unit of a multi-layer perceptron, which is used to decode the color change of the object to obtain the change amount of the color attribute of the static scene Gaussian.

[0061] After being processed by the dynamic component module, the scene Gaussian can already represent a certain degree of motion of dynamic objects. The dynamic component module predicts the motion deformation of the Gaussian. Although it can be reconstructed by generating a new Gaussian, there are more color changes related to time and space and changes in unrelated objects, rather than significant geometric motion. The color component module of this embodiment mainly focuses on color-related changes and drastic changes in objects.

[0062] This embodiment uses a deformable perceptron unit to predict the color change of an object. It needs to input the position encoding of the spherical harmonic function coefficients and the characteristics of the Gaussian's spatiotemporal position at the same time. The color change of the Gaussian is related to time and its position. Then, the MLP is used to decode the features to obtain the change of shs. The process can be expressed as:

[0063]

[0064] Among them, Θ() represents the deformed MLP of shs, E shs is the encoded shs, φ shs () is the feature decoder of shs. At the same time, normal optimization is performed in this stage to further improve the geometry of Gaussian and reduce surface discontinuities. The above network structure is also combined into a color component module, in which the most important part is the deformable perceptron unit.

[0065] In the process of dynamic scene reconstruction, there is a motion propagation effect. If a Gaussian moves, it means that the nearby Gaussians that constitute the same object may also have similar movements, but the neural network may ignore the influence of the surrounding Gaussians when predicting. Without an appropriate propagation mechanism, the overall movement of the object will become unnatural or discontinuous. Another problem is that the movement of some Gaussians may be too violent, and when there are too few training samples, the neural network cannot predict the changes well. To solve this problem, this embodiment adds a shape destruction module to the dynamic scene reconstruction model, and adds some noise through the shape destruction module. The specific operation is as follows: During training, when a fixed number of iterations is reached, this embodiment adjusts the shape of the deformed Gaussian through the shape destruction module (such as reducing the two long axes by half), which results in more small Gaussians in subsequent training to help quickly adapt to the dynamic changes of the object. At the same time, it can temporarily increase the sparsity of the surrounding space, and there are more possibilities to enrich the details of the scene.

[0066] This implementation constructs a dynamic scene reconstruction model, which includes the design of dynamic component modules and color component modules, and realizes modular decomposition and specialized optimization according to the characteristics of different scenes. During the multi-stage training process, the dynamic component module focuses on capturing the motion patterns and geometric changes in the scene, while the color component module is responsible for processing subtle changes in material color. The advantage of this modular design is that it can adapt to a variety of complex scenes, such as dynamic objects with complex texture changes on the surface of the material or non-rigid scenes with fast movement.

[0067] The present invention is further described in detail below through a specific implementation example.

[0068] The device used in this embodiment is a PC computer, which needs to be equipped with a high-performance GPU for convenient calculation; it needs to have a related Python environment and Pytorch package; and a device that can take photos, such as a mobile phone.

[0069] Figure 4 For practical application, the specific steps of the dynamic scene reconstruction method provided in this embodiment are as follows:

[0070] Step 1: Obtain a set of images of dynamic scenes at a certain time.

[0071] First, use a mobile phone or other device to shoot a scene video containing dynamic objects. Usually, you need to move the phone or camera in a stable manner and include as many angles as possible. Then process the video through ffmpeg and other tools, segment the video, obtain scene images from various perspectives, and then filter and remove photos with motion blur when shooting. Then use colmap and other methods to obtain the initial point cloud and the camera pose corresponding to the image set. During the training process, the camera pose is put into the stack, and then two cameras are randomly popped out for rendering in each iteration.

[0072] Step 2: Parameter setting and training setup.

[0073] This step gives the key parameter settings in the model. The original 3DGS in the first stage is set to train 3k times to get a rough scene Gaussian. Then in the dynamic stage and color stage, a total of 20k times. At the same time, the Adam optimizer is used for parameter optimization. The learning rate of the deformation network is exponentially decayed during training, from 1.6e-4 to 1.6e-6. The number of layers of the deformation network is initially 8, and a residual connection is performed every four layers.

[0074] Step 3: The original 3DGS processes the image set to obtain a rough static scene Gaussian.

[0075] According to the technical solution, the sparse point cloud, image set and corresponding camera pose in step 1 are first used as input. The number of Gaussian balls is equal to the number of sparse point clouds, and the coordinates of each point in the point cloud are used as the mean position of the Gaussian. The size and rotation direction of each axis of the Gaussian can be randomly initialized.

[0076] At the beginning of training, the camera's internal and external parameters are taken each time, and the 3D Gaussian is projected into the camera's 2D plane space. Then the final rendering is divided into blocks, and the Gaussian is projected onto the pixel blocks, and the depth is sorted, and alpha blending is performed to obtain the pixel color. The rendering is then compared with the real image, and the properties of the Gaussian are optimized through photometric loss.

[0077] Step 4: Input the properties of the static scene Gaussian into the dynamic component to obtain the change in the properties.

[0078] The scene Gaussian has a position attribute xyz, which and time t are used as inputs for the second stage. Figure 2 In the code, position coding is generally used to encode the xyzt values ​​through sine and cosine trigonometric functions, and then connected as space-time features. Then xyzt is grouped and projected onto paired planes, and then the features are connected and dot-multiplied, and finally the space-time related features are output.

[0079] At this time, we can use some functions of the multi-head decoder to input the spatiotemporal features and the encoded unprocessed color features, and obtain the changes in position, scale and rotation through position MLP, scaling MLP and rotation MLP, and then add them to the original value of the Gaussian to obtain the Gaussian at the next moment.

[0080] Step 5: Input the Gaussian from stage 2 to handle subtle color changes.

[0081] At this point, the scene Gaussian has position attributes xyz and color attributes rgb. At the same time, the deformation network in stage 2 can already obtain the change in shape when the position information is input. After the color component encodes the input rgb, it passes through the deformable perceptron unit to obtain the change in color. Then it also passes through the multi-head decoder to obtain the change in the attribute.

[0082] It is not difficult to find that the present invention is based on 3D Gaussian splashing and integrates motion deformation field and color deformation field. It can process dynamic scenes with material color changes. Different from the traditional implicit scene representation that relies on deep learning, the present invention adopts an explicit representation method to make the scene structure clearer, the data easy to interpret, the reconstruction process more efficient, and the reconstruction results of dynamic scenes can be observed and analyzed more easily.

[0083] A second embodiment of the present invention relates to a dynamic scene reconstruction device, comprising:

[0084] The acquisition module is used to obtain the image sequence and the corresponding camera pose;

[0085] A reconstruction module is used to input the picture sequence and the corresponding camera pose into a dynamic scene reconstruction model to generate a dynamic scene between the picture sequences, wherein the dynamic scene reconstruction model includes:

[0086] A 3D Gaussian sputtering module, used for performing preliminary scene reconstruction on the image sequence based on the camera pose to obtain a static scene Gaussian;

[0087] The dynamic component module is used to generate the change of the position of the static scene Gaussian according to the position attribute of the static scene Gaussian and the time feature;

[0088] The color component module is used to generate the change amount of the color attribute of the static scene Gaussian according to the color attribute of the static scene Gaussian combined with the motion deformation characteristics of the Gaussian.

[0089] The 3D Gaussian sputtering module adopts Normal=F normalize (F align (F sort (s t ),r t ,v t )) Calculate the normal line approximated by the shortest axis. Normal is the normal line approximated by the shortest axis. F sort () is the function of sorting to obtain the shortest axis in Gaussian, F align () is the function of adjusting the shortest axis to the outer surface, F normalize () is the normalization function, s t is the scaling factor at time t, r t is the quaternion at time t, v t is the viewing direction at time t.

[0090] When the 3D Gaussian sputtering module optimizes the normal approximated by the shortest axis, the normal approximated by the shortest axis is first passed to the rasterizer to obtain a rendered depth map; then, based on the rendered depth map, the corresponding reference normal map is calculated using gradient information; finally, the normal approximated by the shortest axis is optimized by minimizing the difference between the rendered depth map and the reference normal map.

[0091] The dynamic component module includes:

[0092] A position encoding unit is used to encode the position attribute of the static scene Gaussian at time t to obtain the position feature;

[0093] Extract the splicing unit, which is used to project the position features onto six planes so that each point can be projected onto a plane formed by a pair of coordinate axes. The feature vectors of the projected points on the paired planes are multiplied respectively and then connected into a feature vector as the feature of the final Gaussian position point in the four-dimensional space.

[0094] The deformation decoder unit of the multi-layer perceptron is used to decode the features of the points of the Gaussian position in the four-dimensional space to obtain the change in the position of the Gaussian in the static scene.

[0095] The color component module includes:

[0096] A color encoding unit is used to encode the color attributes of the static scene Gaussian at time t to obtain color features;

[0097] Deformable perceptron unit, used to predict the color change of objects based on color features and Gaussian motion deformation features;

[0098] The color decoder unit of the multi-layer perceptron is used to decode the color change of the object to obtain the change amount of the color attribute of the static scene Gaussian.

[0099] The dynamic scene reconstruction model also includes: a shape destruction module, which is used to adjust the shape of the deformed Gaussian when a preset number of iterations is reached during training.

[0100] The third embodiment of the present invention relates to an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the dynamic scene reconstruction method of the first embodiment when executing the computer program.

[0101] A fourth embodiment of the present invention relates to a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the dynamic scene reconstruction method of the first embodiment.

[0102] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) that contain computer-usable program code.

[0103] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0104] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction method, which is implemented in the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0105] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0106] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

Claims

1. A dynamic scene reconstruction method, characterized in that: The following steps are involved: Get the image sequence and the corresponding camera pose; The image sequence and the corresponding camera pose are input into a dynamic scene reconstruction model to generate a dynamic scene between the image sequences, wherein the dynamic scene reconstruction model includes: A 3D Gaussian sputtering module, used for performing preliminary scene reconstruction on the image sequence based on the camera pose to obtain a static scene Gaussian; The dynamic component module is used to generate the change of the position of the static scene Gaussian according to the position attribute of the static scene Gaussian and the time feature; The color component module is used to generate the change amount of the color attribute of the static scene Gaussian according to the color attribute of the static scene Gaussian combined with the motion deformation characteristics of the Gaussian.

2. The dynamic scene reconstruction method according to claim 1, characterized in that: The 3D Gaussian sputtering module adopts Normal=F normalize (F align (F sort (s t ),r t ,v t )) Calculate the normal line approximated by the shortest axis. Normal is the normal line approximated by the shortest axis. F sort () is the function of sorting to obtain the shortest axis in Gaussian, F align () is the function of adjusting the shortest axis to the outer surface, F normalize () is the normalization function, s t is the scaling factor at time t, r t is the quaternion at time t, v t is the viewing direction at time t.

3. The dynamic scene reconstruction method according to claim 1, characterized in that: When the 3D Gaussian sputtering module optimizes the normal approximated by the shortest axis, the normal approximated by the shortest axis is first passed to the rasterizer to obtain a rendered depth map; then, based on the rendered depth map, the corresponding reference normal map is calculated using gradient information; finally, the normal approximated by the shortest axis is optimized by minimizing the difference between the rendered depth map and the reference normal map.

4. The dynamic scene reconstruction method according to claim 1, characterized in that: The dynamic component module includes: A position encoding unit is used to encode the position attribute of the static scene Gaussian at time t to obtain the position feature; Extract the splicing unit, which is used to project the position features onto six planes so that each point can be projected onto a plane formed by a pair of coordinate axes. Multiply the feature vectors of the projected points on the paired planes respectively and then connect them into a feature vector as the feature of the final Gaussian position point in the four-dimensional space. The deformation decoder unit of the multi-layer perceptron is used to decode the features of the points of the Gaussian position in the four-dimensional space to obtain the change in the position of the Gaussian in the static scene.

5. The dynamic scene reconstruction method according to claim 1, characterized in that: The color component module includes: A color encoding unit is used to encode the color attributes of the static scene Gaussian at time t to obtain color features; The deformable perceptron unit is used to predict the color change of the object according to the color feature and the motion deformation feature of the Gaussian. The color decoder unit of the multi-layer perceptron is used to decode the color change of the object to obtain the change amount of the color attribute of the static scene Gaussian.

6. The dynamic scene reconstruction method according to claim 1, characterized in that: The dynamic scene reconstruction model also includes: a shape destruction module, which is used to adjust the shape of the deformed Gaussian when a preset number of iterations is reached during training.

7. A dynamic scene reconstruction device, characterized in that: include: The acquisition module is used to obtain the image sequence and the corresponding camera pose; A reconstruction module is used to input the picture sequence and the corresponding camera pose into a dynamic scene reconstruction model to generate a dynamic scene between the picture sequences, wherein the dynamic scene reconstruction model includes: A 3D Gaussian sputtering module, used for performing preliminary scene reconstruction on the image sequence based on the camera pose to obtain a static scene Gaussian; The dynamic component module is used to generate the change of the position of the static scene Gaussian according to the position attribute of the static scene Gaussian and the time feature; The color component module is used to generate the change amount of the color attribute of the static scene Gaussian according to the color attribute of the static scene Gaussian combined with the motion deformation characteristics of the Gaussian.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the dynamic scene reconstruction method as described in any one of claims 1-6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the dynamic scene reconstruction method as claimed in any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Face mouth reconstruction method and device

    CN114973355A

  • Virtual reality and augmented reality fusion method based on indoor three-dimensional reconstruction

    CN118485796A

  • Underwater scene characterization method based on neural radiation field

    CN118710807A

Cited By

  • Track scene prediction method and device, equipment and medium

    CN120766086A

  • Image rendering method based on dynamic Gaussian splashing

    CN122115664A

  • Image rendering method based on dynamic gaussian splatting

    CN122115664B