Three-dimensional dynamic scene reconstruction method, device and storage medium
By adopting Gaussian sputtering technology in dynamic three-dimensional scene reconstruction, Gaussian sputtering points are constructed from multi-view video and dynamic separation, the problems of low efficiency and inability to achieve streaming in the existing technology are solved, and rapid, streaming and real-time dynamic scene reconstruction is achieved.
Patent Information
- Application Number
- CN202411290735.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-14
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-09-14
AI Technical Summary
The prior art is inefficient in dynamic three-dimensional scene reconstruction, cannot effectively separate dynamic and static primitives, and it is difficult to realize streaming and incremental updates.
Using a Gaussian sputtering method, Gaussian sputtering points are constructed from multi-view videos, and dynamic and static separation is achieved through pixel-level changes, supporting dynamic scene reconstruction with fast, streaming and incremental updates.
It significantly improves the speed and efficiency of three-dimensional dynamic scene reconstruction, supports real-time rendering and streaming processing, and is suitable for applications such as virtual reality and robot navigation.
Smart Images

Figure CN119295651B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and in particular to a method, device and storage medium for reconstructing a three-dimensional dynamic scene based on multi-view videos. Background Art
[0002] With the rapid development of virtual reality and robotics in recent years, more and more application scenarios require not only understanding of two-dimensional scenes and images, but also requiring algorithms or intelligent agents to understand, reconstruct and analyze geometric and texture information in the three-dimensional world. Among them, reconstructing dynamic three-dimensional scenes from multi-view videos is a very important and practical task. For example, in virtual reality, dynamic scenes need to be reproduced in the virtual world to enable users to have a realistic experience of the real world in the virtual environment; or when intelligent robots navigate in dynamic environments, they may need to directly model the visual changes of the dynamic environment.
[0003] At present, researchers have proposed 3D scene reconstruction methods based on multi-plane images, layered rendering, neural radiation fields and dynamic point clouds. However, these methods do not perform well in dynamic scene reconstruction, have low efficiency, and do not support streaming and incremental updates.
[0004] In fact, for a 3D dynamic scene, it often contains a large number of motionless backgrounds or static objects or environments. This part of the data is actually independent of time. Therefore, if the primitives representing dynamic and static can be separated, the amount of calculation required for reconstruction can be greatly reduced and the reconstruction process can be greatly accelerated. However, such an explicit separation is impossible for the implicit 3D representation based on previous methods, because these methods do not have explicit primitives to represent the physical quantities of 3D points in space, but use continuous functions to approximate the shape and appearance of the entire scene. Therefore, in order to solve these problems, researchers began to pay attention to dynamic 3D scene reconstruction technology based on Gaussian sputtering. Gaussian sputtering is an effective method for converting 3D point cloud data into voxel or surface representation. Its basic idea is to take each 3D point as the center of the Gaussian kernel and "project" or "diffuse" it into the surrounding voxel space according to the density and color information of the point to form a continuous voxel model. In the reconstruction of dynamic 3D scenes, Gaussian sputtering can well capture the details and boundaries of moving objects, and due to its discretization characteristics, it is easier to achieve explicit separation of dynamic and static parts than the continuous function approximation method. By matching and tracking point cloud data between different frames, dynamic parts can be efficiently identified and updated while retaining a stable representation of the static background, thereby significantly improving the reconstruction speed and efficiency. In addition, Gaussian sputtering also supports streaming processing and incremental updates, which means that it can adapt to changes in real-time scenes, receive and process newly acquired data in real time, and then update the reconstructed three-dimensional scene model in a timely manner. This is undoubtedly an important advantage for virtual reality applications and robot navigation in dynamic environments. Existing dynamic scene reconstruction techniques based on Gaussian sputtering do not make good use of Gaussian sputtering as a property of displaying three-dimensional representations, but simply use the methods in neural radiation fields in Gaussian sputtering. Therefore, it is difficult to achieve efficient streaming reconstruction and real-time rendering, and usually requires a relatively large memory to store Gaussian sputtering points. Summary of the invention
[0005] The present disclosure aims to solve one of the technical problems existing in the existing related technologies at least to a certain extent.
[0006] To this end, the embodiments of the present disclosure propose a three-dimensional dynamic scene reconstruction method, device and storage medium based on multi-view video. The present disclosure utilizes the display characteristics of three-dimensional Gaussian sputtering to construct Gaussian sputtering points from multi-view video, and realizes the dynamic and static separation of Gaussian points according to pixel-level changes, thereby realizing fast, streamable and incrementally updateable dynamic scene reconstruction.
[0007] In order to achieve the above objectives, the present disclosure adopts the following technical solutions:
[0008] A first aspect of the present disclosure provides a method for reconstructing a three-dimensional dynamic scene, comprising the following steps:
[0009] S1, use multiple cameras to obtain multi-view synchronized videos of dynamic scenes as training data;
[0010] S2, calculate the matching points between video images of different perspectives to estimate the internal and external parameters of each camera;
[0011] S3, constructing a sparse point cloud according to the depth of each matching point, and generating an initial set of Gaussian sputtering point sets {p0} according to the sparse point cloud;
[0012] S4. For the first frame images of all videos in the training data, use each Gaussian splattering point in the Gaussian splattering point set {p0} and calculate the color radiation value of the three-dimensional space point through the Gaussian distribution, obtain the rendered image according to the color radiation value, and perform the following iterative optimization on the Gaussian splattering point set {p0}: calculate the gradient value of the two-dimensional Gaussian point projected on the plane according to the loss function between the rendered image and the first frame image, and determine whether to perform Gaussian densification processing on the Gaussian splattering point set {p0} according to the relationship between the gradient value and the splitting threshold, and delete the Gaussian points whose volume is greater than the volume setting threshold, the Gaussian points that are invisible from all viewing angles, and the Gaussian points whose transparency is less than the transparency setting threshold in the Gaussian splattering point set {p0}, and obtain the Gaussian splattering point set {p} for the first frame images of all videos;
[0013] For the remaining frame images of all videos, the pixel change variance of each perspective video is calculated, and the change voxel field is optimized according to the calculated pixel change variance. The Gaussian sputtering point set {p} is divided into a static point set {S} and a dynamic point set {D} according to the value of the change voxel field; the static point set {S} is no longer updated at the time step corresponding to the remaining frames; for the dynamic point set {D}, the attributes of each dynamic point at the time step corresponding to the remaining frame images are updated according to its rigid body constraints with the Gaussian sputtering point set {p} and its rendering loss with the remaining frame images;
[0014] The dynamic Gaussian sputtering point set is composed of the Gaussian sputtering point set {p}, the static point set {S} and the final dynamic point set {D}.
[0015] S5. Dynamic Gaussian sputtering point set combined with camera internal and external parameters Rendering is performed according to the Gaussian sputtering rendering pipeline to obtain rendering images at different times from a new perspective, thereby achieving dynamic three-dimensional scene reconstruction.
[0016] In some embodiments, in step S2, the pixel matching points between the video images of different viewing angles are calculated by using a structure-from-motion method.
[0017] In some embodiments, in step S4, the i-th Gaussian sputtering point in the Gaussian sputtering point set {p0} is represented as {x i ,Ri ,s i ,c i ,o i}, x i ,R i ,s i ,o i ,o i are the position, rotation angle, scale, color and transparency of the i-th Gaussian sputtering point respectively. The calculation formula of the color radiation value of the three-dimensional space point is:
[0018]
[0019] Where c(x) represents the color radiation value of the three-dimensional space point x, and N represents the total number of Gaussian sputtering points in the Gaussian sputtering point set {p0};
[0020] For the first frame image of each perspective video, each Gaussian splattering point in the Gaussian splattering point set {p0} is projected onto the imaging plane of the corresponding perspective. The projected two-dimensional Gaussian splattering points are sequentially differentiably rendered according to the depth from small to large, and the rendered image I corresponding to the Gaussian splattering point set {p0} is obtained. The formula for differentiable rendering is as follows:
[0021]
[0022] Where c(p) is the color of point p in the rendered image I, α i (p) represents the transparency of point p in the rendered image I obtained by the i-th Gaussian sputtering point in the Gaussian sputtering point set {p0}, μ i is the projection position of the i-th Gaussian sputtering point in the Gaussian sputtering point set {p0} in the rendered image I, Cov i is the covariance matrix of the projection of the i-th Gaussian splatter point in the rendered image I.
[0023] In some embodiments, in step S4, the iterative optimization of the Gaussian sputtering point set {p0} specifically includes:
[0024] At each iteration of the optimization process, the gradient value of the two-dimensional Gaussian point projected onto the plane is calculated according to the loss function between the rendered image and the first frame image;
[0025] After every K1 iterations, the following operations are performed: the average gradient value of each Gaussian sputtering point in the Gaussian sputtering point set {p0} recorded in the K1 iterations is calculated, and for each Gaussian sputtering point whose average gradient value is greater than the splitting threshold, the point is split or cloned according to the projection size of the Gaussian sputtering point, wherein, for Gaussian sputtering points whose own volume is greater than the volume threshold, the point is split according to the probability density sampling of the Gaussian distribution, and for Gaussian sputtering points whose own volume is less than or equal to the volume threshold, a cloning operation is used to generate a Gaussian sputtering point that is exactly the same as the original Gaussian sputtering point, thereby realizing Gaussian densification of the Gaussian sputtering points; at the same time, the Gaussian point pruning method is used to remove Gaussian sputtering points whose own volume is greater than the volume threshold or whose projection area is greater than the area threshold, and Gaussian sputtering points that are invisible from all viewing angles are removed;
[0026] After every K2 iterations, K2>K1, all Gaussian splattering points with transparency less than the transparency threshold are removed, and then the transparency of all Gaussian splattering points is set to a fixed value less than the transparency threshold, thereby deleting the obstructed Gaussian splattering points;
[0027] When the upper limit of the number of iterations is reached, the Gaussian sputtering point set {p} for the first frame images of all videos is obtained.
[0028] In some embodiments, in step S4, the loss function between the rendered image and the first frame image is Loss(I,I g ), the formula is as follows:
[0029] Loss(I,I g )=L2(I,I g )+λ1·L D-SSIM (I,I g )+λ2·L LPIPS (I,I g )
[0030] Among them, I, I g They represent the rendered image and the real image obtained by the Gaussian sputtering point set {p0}; L2, L D-SSIM They represent the 2-norm loss function and the structural dissimilarity loss function, respectively, both of which are used to constrain the pixel-level similarity between images; L LPIPS represents the learnable perceptual image patch similarity loss function, which is used to reduce the perceptual distance between the rendered image and the real image; λ1 and λ2 are the loss function L D-SSIM and L LPIPS The weight of .
[0031] In some embodiments, in step S4, for the remaining frame images of all videos, the specific steps of dividing the Gaussian sputtering point set {p} into a static point set {S} and a dynamic point set {D} include:
[0032] For each continuous video of the multi-view video, the variance of the pixel points of each video over time is calculated, and the Gaussian blur operation is performed on the pixel point variance to ensure the continuity of the boundary. Each pixel point is divided into a dynamic pixel point and a static pixel point according to whether the pixel point variance is greater than the dynamic threshold.
[0033] The change voxel field V(x,y,z) is randomly initialized to indicate whether a certain voxel (x,y,z) in the space is dynamic. The change voxel field V(x,y,z) is optimized by differentiable rendering and the variance of the pixel points of each video over time is used as the supervision, that is:
[0034]
[0035] Among them, L V represents the optimized changing voxel field; is the expected operator; M(r) represents the dynamic or static category of the pixel corresponding to the light ray r, M(r)=1 indicates that the pixel is a dynamic pixel, and M(r)=0 indicates that the pixel is a static pixel; It means that the voxel field V(r(y i ))The estimated pixel change field, represents the first sampling points, Indicates the position of the sampling point on the ray r; s represents the sigmoid activation function; N r represents the set of Gaussian sputtering points that the light ray r passes through;
[0036] According to the nearest neighbor principle, the voxels corresponding to each Gaussian sputtering point in the Gaussian sputtering point set {p} are found, and whether each Gaussian sputtering point is a dynamic Gaussian sputtering point is determined according to whether the voxel is a dynamic voxel. That is, for any Gaussian sputtering point in the Gaussian sputtering point set {p}, if the voxel corresponding to the Gaussian sputtering point is a dynamic voxel, then the Gaussian sputtering point is a dynamic Gaussian sputtering point, otherwise the Gaussian sputtering point is a static Gaussian sputtering point, thereby dividing the Gaussian sputtering point set {p} into a static point set {S} and a dynamic point set {D}.
[0037] In some embodiments, in step S4, the rigid body constraint between the kth Gaussian sputtering point in the dynamic point set {D} and the lth Gaussian sputtering point in the Gaussian sputtering point set {p} is set to Its expression is:
[0038]
[0039] in, and They represent the position and rotation angle of the kth Gaussian splash point in the dynamic point set {D} at the tth frame time, respectively. and represent the position and rotation angle of the lth Gaussian sputtering point in the Gaussian sputtering point set {p} at the tth frame time; w k,l Represents the weight of the loss function, which is negatively correlated with the distance between the two Gaussian sputtering points, and is specifically expressed as and They respectively represent the positions of the kth and lth Gaussian sputtering points in the first frame image.
[0040] In some embodiments, in step S4, the rigid body constraint between the dynamic point set {D} and the Gaussian sputtering point set {p} is L rigid , which is expressed as follows:
[0041]
[0042] in, represents the total number of Gaussian splash points in the dynamic point set {D}, represents the kth Gaussian sputtering point selected from the Gaussian sputtering point set {p} A set of adjacent points.
[0043] A second aspect of the present disclosure provides a reconstruction device based on the reconstruction method described in any embodiment of the first aspect of the present disclosure, comprising:
[0044] A first module stores training data, wherein the training data is a multi-view synchronized video of a dynamic scene acquired by multiple cameras;
[0045] The second module is used to calculate the matching points between video images of different perspectives to estimate the internal and external parameters of each camera;
[0046] The third module is used to construct a sparse point cloud according to the depth of each matching point, and generate an initial set of Gaussian splash point sets {p0} according to the sparse point cloud;
[0047] The fourth module is used for the first frame images of all videos in the training data, using each Gaussian splattering point in the Gaussian splattering point set {p0} and calculating the color radiation value of the three-dimensional space point through the Gaussian distribution, obtaining the rendered image according to the color radiation value, and performing the following iterative optimization on the Gaussian splattering point set {p0}: calculating the gradient value of the two-dimensional Gaussian point projected onto the plane according to the loss function between the rendered image and the first frame image, judging whether to perform Gaussian densification processing on the Gaussian splattering point set {p0} according to the relationship between the gradient value and the splitting threshold, deleting the Gaussian points whose volume is greater than the volume setting threshold, the Gaussian points that are invisible from all viewing angles, and the Gaussian points whose transparency is less than the transparency setting threshold in the Gaussian splattering point set {p0}, and obtaining the Gaussian splattering point set {p} for the first frame images of all videos;
[0048] For the remaining frame images of all videos, the pixel change variance of each perspective video is calculated, and the change voxel field is optimized according to the calculated pixel change variance. The Gaussian sputtering point set {p} is divided into a static point set {S} and a dynamic point set {D} according to the value of the change voxel field; the static point set {S} is no longer updated at the time step corresponding to the remaining frames; for the dynamic point set {D}, the attributes of each dynamic point at the time step corresponding to the remaining frame images are updated according to its rigid body constraints with the Gaussian sputtering point set {p} and its rendering loss with the remaining frame images;
[0049] The dynamic Gaussian sputtering point set is composed of the Gaussian sputtering point set {p}, the static point set {S} and the final dynamic point set {D}.
[0050] The fifth module is used to combine the internal and external parameters of the camera to dynamically calculate the Gaussian sputtering point set. Rendering is performed according to the Gaussian sputtering rendering pipeline to obtain rendering images at different times from a new perspective, thereby achieving dynamic three-dimensional scene reconstruction.
[0051] A third aspect of the present disclosure provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the reconstruction method described in any embodiment of the first aspect of the present disclosure.
[0052] Compared with the prior art, the present invention has the following characteristics and beneficial effects:
[0053] 1. The present invention introduces a voxel change field in three-dimensional dynamic reconstruction, which can make good use of the prior information in multi-view videos to capture the dynamic situation of three-dimensional space and greatly reduce the number of Gaussian sputtering points required to express three-dimensional dynamic scenes.
[0054] 2. The present invention uses Gaussian sputtering points and volume rendering to represent three-dimensional scenes and render new perspective images, which can complete the reconstruction of dynamic scenes within minutes. Compared with the method based on neural radiation field, it is less likely to have phenomena such as "floating objects" and "ghosting", and is more adaptable to situations with fewer input perspectives.
[0055] 3. The present disclosure can be applied to various fields that require three-dimensional dynamic scene reconstruction. It can quickly and incrementally reconstruct three-dimensional dynamic scenes, and can realize streaming and real-time rendering on any end-side device that supports CUDA or OpenGL. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 It is an overall flow chart of a method for reconstructing a three-dimensional dynamic scene provided by an embodiment of the first aspect of the present disclosure.
[0057] Figure 2 It is a structural schematic diagram of an electronic device provided by an embodiment of the third aspect of the present disclosure. DETAILED DESCRIPTION
[0058] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0059] On the contrary, the present application covers any substitution, modification, equivalent method and scheme made on the essence and scope of the present application as defined by the claims. Further, in order to make the public have a better understanding of the present application, some specific details are described in detail in the detailed description of the present application below. Those skilled in the art can fully understand the present application without the description of these details.
[0060] See also Figure 1 The first aspect of the present disclosure provides a method for reconstructing a three-dimensional dynamic scene, comprising the following steps:
[0061] S1, use multiple cameras to obtain multi-view synchronized videos of dynamic scenes as training data;
[0062] S2, calculate the matching points between video images of different perspectives to estimate the internal and external parameters of each camera;
[0063] S3, constructing a sparse point cloud according to the depth of each matching point, and generating an initial set of Gaussian sputtering point sets {p0} according to the sparse point cloud;
[0064] S4. For the first frame images of all videos in the training data, use each Gaussian splattering point in the Gaussian splattering point set {p0} and calculate the color radiation value of any point in the three-dimensional space through Gaussian distribution, obtain a rendered image according to the color radiation value, and perform the following iterative optimization on the Gaussian splattering point set {p0}: calculate the gradient value of the two-dimensional Gaussian point projected onto the plane according to the loss function between the rendered image and the first frame image, and determine whether to perform Gaussian densification processing on the Gaussian splattering point set {p0} according to the relationship between the gradient value and the splitting threshold, and delete the Gaussian points whose volume is greater than the volume setting threshold, the Gaussian points that are invisible from all viewing angles, and the Gaussian points whose transparency is less than the transparency setting threshold in the Gaussian splattering point set {p0}, and obtain the Gaussian splattering point set {p} for the first frame images of all videos;
[0065] For the remaining frame images of all videos, the pixel change variance of each perspective video is calculated, and the change voxel field is optimized according to the calculated pixel change variance. The Gaussian sputtering point set {p} is divided into a static point set {S} and a dynamic point set {D} according to the value of the change voxel field; the static point set {S} is no longer updated at the time step corresponding to the remaining frames; for the dynamic point set {D}, the attributes of each dynamic point at the time step corresponding to the remaining frame images are updated according to its rigid body constraints with the Gaussian sputtering point set {p} and its rendering loss with the remaining frame images;
[0066] The dynamic Gaussian sputtering point set is composed of the Gaussian sputtering point set {p}, the static point set {S} and the final dynamic point set {D}.
[0067] S5. Dynamic Gaussian sputtering point set combined with camera internal and external parameters Rendering is performed according to the Gaussian sputtering rendering pipeline to obtain rendering images at different times from a new perspective, thereby achieving dynamic three-dimensional scene reconstruction.
[0068] In some embodiments, when obtaining training data using step S1, multiple cameras are used to capture multi-perspective synchronized videos of a dynamic scene at different locations in the scene with fixed postures, and the synchronization of the multi-perspective videos is ensured by a hardware shutter connection between the multiple cameras.
[0069] In some embodiments, in step S2, a Structure from Motion (SfM) method is used to calculate pixel matching points between video images of different viewing angles, and an existing three-dimensional vision method is used to estimate the intrinsic and extrinsic parameters of each camera.
[0070] In some embodiments, in step S3, the depth corresponding to each pixel matching point is calculated from the pixel matching points generated in the motion recovery structure process, thereby constructing a set of sparse point clouds, and the positions of Gaussian sputtering points are initialized using the sparse point clouds to obtain an initial set of Gaussian sputtering point sets {p0}.
[0071] In some embodiments, step S4 first uses the first frame image of each video in the training data to perform static training on the initial Gaussian sputtering point set {p0} to obtain the Gaussian sputtering point set {p} of the first frame image, and then uses the remaining frame images of each video in the training data to dynamically expand the obtained Gaussian sputtering point set {p} to obtain the expanded Gaussian sputtering point set {p}. Finally, the Gaussian sputtering point set {p} and the extended Gaussian sputtering point set Constructing a dynamic Gaussian sputtering point set Step S4 specifically includes the following steps:
[0072] S41, fitting the Gaussian sputtering points of the first frame
[0073] After initialization, we first train static Gaussian sputtering points on the first frame of the multi-view synchronized video; Gaussian sputtering points use a set of anisotropic Gaussian functions in space to express the radiation field in three-dimensional space. Specifically, each Gaussian sputtering point contains the position x i , rotation angle R i , scale i 、Color c i and transparency i Equal parameters, denoted as {x i ,R i ,s i ,c i ,o i}, the subscript i represents the serial number of the Gaussian sputtering point; the specific process of static training is:
[0074] S411, using each Gaussian splattering point in the Gaussian splattering point set {p0} obtained in step S3 and calculating the color radiation value of the spatial point x through Gaussian distribution, the calculation formula is as follows:
[0075]
[0076] Where c(x) represents the color radiation value of the three-dimensional space point x, {x i ,R i ,s i ,c i ,o i} is the Gaussian sputtering point allowed to be optimized, N represents the total number of Gaussian sputtering points in the Gaussian sputtering point set {p0}, and the superscript T represents the transpose;
[0077] S412, for the first frame image of each viewing angle video, project each Gaussian splattering point in the Gaussian splattering point set {p0} onto the imaging plane of the viewing angle, and perform differentiable rendering on each two-dimensional Gaussian splattering point after projection in order of depth from small to large, to obtain a rendered image I corresponding to the Gaussian splattering point set {p0}. The formula for differentiable rendering is as follows:
[0078]
[0079] Where c(p) is the color of point p in the rendered image I, α i (p) represents the transparency of point p in the rendered image I obtained by the i-th Gaussian sputtering point, μ i is the projection position of the i-th Gaussian splash point in the rendered image I, Cov iis the covariance matrix of the projection of the i-th Gaussian splatter point in the rendered image I. Such a rendering process can be performed according to the image blocks, and the rendering of each image block is completed by a CUDA (Compute Unified Device Architecture) process block, thereby fully utilizing the powerful parallel computing capability of the image processing unit;
[0080] S413, let the loss function of the rendered image and the first frame image be Loss(I,I g ), the formula is as follows:
[0081] Loss(I,I g )=L2(I,I g )+λ1·L D-SSIM (I,I g )+λ2·L LPIPS (I,I g )
[0082] Among them, I, I g They represent the rendered image and the real image obtained by the Gaussian sputtering point set {p0}; L2, L D-SSIM They represent the 2-norm loss function and the structural dissimilarity loss function respectively, which are used to constrain the pixel-level similarity between images; L LPIPS represents the learnable perceptual image patch similarity loss function, which is used to reduce the perceptual distance between the rendered image and the real image; λ1 and λ2 are the loss function L D-SSIM and L LPIPS The weights are, in this embodiment, λ1=1.0, λ2=2.0.
[0083] S414, Processing of Gaussian sputtering point set {p0} during iteration
[0084] S4141, in each iteration of the optimization process, according to the loss function Loss(I,I g ) Calculate the gradient value of the two-dimensional Gaussian point projected onto the plane. In order to add Gaussian sputtering points to the positions where there were no Gaussian sputtering points in the space, the Gaussian densification process is used to add Gaussian sputtering points to the vacancies in the space: after every K1=100 iterations, calculate the average value of the gradient of the two-dimensional projection of each Gaussian sputtering point recorded in these 100 iterations; for each Gaussian sputtering point, if the average value of the two-dimensional point gradient in the past 100 iterations is greater than the splitting threshold thresh 2D (In this embodiment, thresh 2DAccording to experience, it is set to 0.001), then the Gaussian sputtering point is split or cloned according to the projection size. For Gaussian sputtering points with relatively large volumes, they are split according to the probability density sampling of the Gaussian distribution. For Gaussian sputtering points with relatively small volumes, a cloning operation is used to generate a Gaussian sputtering point that is exactly the same as the original Gaussian sputtering point. In this embodiment, the volume threshold is set to 0.0004. Gaussian sputtering points with a volume greater than 0.0004 are split, and Gaussian sputtering points with a volume less than or equal to 0.0004 are cloned. For the gradient average value less than or equal to the splitting threshold thresh 2D For each Gaussian sputtering point, no Gaussian densification is performed on it.
[0085] S4142. During the optimization process, the Gaussian point pruning method is used to remove unnecessary Gaussian splattering points: after every 100 iterations, Gaussian points with too large volume (the volume threshold is set to 0.0004 in this embodiment) or too large area after projection (the area threshold is set to 0.2 in this embodiment) are deleted, and Gaussian splattering points that are not visible from all training perspectives are removed.
[0086] S4143. While steps S4141 and S4142 are being performed, after each K2=1000 optimization process, all Gaussian splattering points whose transparency is less than a transparency threshold (in this embodiment, the transparency threshold is set to 0.02) are removed, and then the transparency of all Gaussian splattering points is set to 0.01, and then subsequent optimization is performed, and this step is repeated continuously to delete obstructed Gaussian splattering points; when the upper limit of the number of iterations is reached, the Gaussian splattering point set {p} for the first frame image of all videos is obtained.
[0087] S42, dynamically expanding the Gaussian sputtering point set {p} obtained in step S41 to obtain an extended Gaussian sputtering point set The specific process is as follows:
[0088] S421, for each continuous video of the multi-view video, calculate the variance of the pixel points of each video over time, and perform a Gaussian blur operation on the variance of the pixel points to ensure the continuity of the boundary, and divide each pixel point into a dynamic pixel point and a static pixel point according to whether the variance of the pixel point change is greater than a dynamic threshold. In this embodiment, the dynamic threshold is set to 0.03 according to experience;
[0089] S422, randomly initialize the change voxel field V(x, y, z) to indicate whether a certain voxel (x, y, z) in the space is dynamic, and optimize the change voxel field by differentiable rendering of the change voxel field V(x, y, z) and using the pixel change variance of the multi-view video obtained in step S421 as supervision, that is:
[0090]
[0091] Among them, L V represents the optimized changing voxel field; is the expected operator; M(r) represents the dynamic or static category of the pixel corresponding to the light r, M(r)=1 means that the pixel is a dynamic pixel, and M(r)=0 means that the pixel is a static pixel. Represents the voxel field at the ray r The estimated pixel change field, represents the first sampling points, Indicates the position of the sampling point on the ray r; s represents the sigmoid activation function; N r represents the set of Gaussian sputtering points that the light ray r passes through;
[0092] S423. Find the voxels corresponding to each Gaussian splattering point in the Gaussian splattering point set {p} according to the nearest neighbor principle, and determine whether each Gaussian splattering point is a dynamic Gaussian splattering point according to whether the voxel is a dynamic voxel, that is, for any Gaussian splattering point, if the voxel corresponding to the Gaussian splattering point is a dynamic voxel, then the Gaussian splattering point is a dynamic Gaussian splattering point, otherwise the Gaussian splattering point is a static Gaussian splattering point, thereby dividing the Gaussian splattering point set {p} into a static point set {S} and a dynamic point set {D}. For each static Gaussian splattering point, it will not change in the subsequent training process. For each dynamic Gaussian splattering point, step S424 is used to update it using multi-view video.
[0093] S424, using the multi-view video data from the second frame to the last frame according to the loss function in step S413, combined with the following rigid body regularization constraints to update all attributes of each dynamic Gaussian point in the dynamic point set {D};
[0094] In the training process of subsequent frames, the rigid body regularization term is used to constrain the training process. The rigid body constraint between the kth Gaussian sputtering point in the dynamic point set {D} and the lth Gaussian sputtering point in the Gaussian sputtering point set {p} is set to It is expressed as follows:
[0095]
[0096] in, and They represent the position and rotation angle of the kth Gaussian sputtering point at the tth frame, respectively. and represent the position and rotation angle of the lth Gaussian sputtering point at the tth frame; w k,l represents the corresponding loss function weight, which is negatively correlated with the distance between the two Gaussian sputtering points. It can be specifically expressed as and They respectively represent the positions of the kth and lth Gaussian sputtering points in the first frame image.
[0097] In actual calculation, for each dynamic Gaussian sputtering point in the dynamic point set {D}, only 20 neighboring points of the dynamic Gaussian sputtering point are selected from the Gaussian sputtering point set {p} to perform body rigid constraints on the dynamic Gaussian sputtering point, so that the rigid body constraint L between the dynamic point set {D} and the Gaussian sputtering point set {p} rigid It is expressed as follows:
[0098]
[0099] in, represents the total number of Gaussian splash points in the dynamic point set {D}, represents the set consisting of 20 adjacent points of the kth Gaussian sputtering point selected from the Gaussian sputtering point set {p}.
[0100] S43, the Gaussian sputtering point set {p}, the static point set {S} and the final dynamic point set {D} form a dynamic Gaussian sputtering point set
[0101] In some embodiments, step S5 specifically includes the following steps:
[0102] S51, according to the dynamic Gaussian sputtering point set obtained in step S4 And the intrinsic and extrinsic parameters of the target perspective camera that needs to be rendered, and use the internal and external parameters of the camera to construct the camera's transformation matrix and projection matrix.
[0103] S52, use the transformation matrix to transform the dynamic Gaussian sputtering point set The position is mapped to the imaging plane of the target view camera, and the two-dimensional covariance matrix corresponding to the three-dimensional Gaussian sputtering point is obtained according to the linear assumption combined with the Jacobian transformation matrix.
[0104] S53. For all two-dimensional Gaussian splattering points projected in step S52, Gaussian splattering points are rendered in order of depth from small to large. The specific calculation method is the same as that described in step S412. The obtained image is transmitted to the front-end display to complete the image synthesis under the new perspective.
[0105] S54 . For different time steps, switch the Gaussian splattering points to be rendered according to the required time step, and then render according to steps S51 to S53 .
[0106] Furthermore, for computing devices equipped with NVIDIA image processing units, the hierarchical design in the CUDA programming model is used to make a thread block containing 256 threads responsible for rendering a 16×16 image block in the image, where each thread is responsible for rendering one pixel. The Gaussian splattering points required for thread calculations are pre-calculated, and the relevant data is transferred to the shared memory of the graphics computing unit to achieve real-time rendering.
[0107] Furthermore, for computing devices that can apply the Open Graphics Standard Library (OpenGL), the projection of three-dimensional Gaussian splatter points and the transformation of the covariance matrix are implemented in the primitive shader, the rendering process described in step S412 is implemented in the fragment shader, and the depth test in OpenGL is turned off, the blending mode is turned on, and the blending mode is selected as ONE_MINUS_DST_ALPHA to achieve real-time rendering in the OpenGL framework.
[0108] Furthermore, for computing devices that support web pages, a WebGL framework is used, and efficient rendering is achieved using the same primitive shader and fragment shader as the computing devices that can apply OpenGL.
[0109] A third aspect of the present disclosure provides a device for reconstructing a three-dimensional dynamic scene, including:
[0110] A first module stores training data, wherein the training data is a multi-view synchronized video of a dynamic scene acquired by multiple cameras;
[0111] The second module is used to calculate the matching points between video images of different perspectives to estimate the internal and external parameters of each camera;
[0112] The third module is used to construct a sparse point cloud according to the depth of each matching point, and generate an initial set of Gaussian splash point sets {p0} according to the sparse point cloud;
[0113] The fourth module is used for the first frame images of all videos in the training data, using each Gaussian splattering point in the Gaussian splattering point set {p0} and calculating the color radiation value of the three-dimensional space point through the Gaussian distribution, obtaining the rendered image according to the color radiation value, and performing the following iterative optimization on the Gaussian splattering point set {p0}: calculating the gradient value of the two-dimensional Gaussian point projected onto the plane according to the loss function between the rendered image and the first frame image, judging whether to perform Gaussian densification processing on the Gaussian splattering point set {p0} according to the relationship between the gradient value and the splitting threshold, deleting the Gaussian points whose volume is greater than the volume setting threshold, the Gaussian points that are invisible from all viewing angles, and the Gaussian points whose transparency is less than the transparency setting threshold in the Gaussian splattering point set {p0}, and obtaining the Gaussian splattering point set {p} for the first frame images of all videos;
[0114] For the remaining frame images of all videos, the pixel change variance of each perspective video is calculated, and the change voxel field is optimized according to the calculated pixel change variance. The Gaussian sputtering point set {p} is divided into a static point set {S} and a dynamic point set {D} according to the value of the change voxel field; the static point set {S} is no longer updated at the time step corresponding to the remaining frames; for the dynamic point set {D}, the attributes of each dynamic point at the time step corresponding to the remaining frame images are updated according to its rigid body constraints with the Gaussian sputtering point set {p} and its rendering loss with the remaining frame images;
[0115] The dynamic Gaussian sputtering point set is composed of the Gaussian sputtering point set {p}, the static point set {S} and the final dynamic point set {D}.
[0116] The fifth module is used to combine the internal and external parameters of the camera to dynamically calculate the Gaussian sputtering point set. Rendering is performed according to the Gaussian sputtering rendering pipeline to obtain rendering images at different times from a new perspective, thereby achieving dynamic three-dimensional scene reconstruction.
[0117] It should be noted that the above-mentioned explanation of the embodiment of a method for reconstructing a three-dimensional dynamic scene is also applicable to the device for reconstructing a three-dimensional dynamic scene in this embodiment, and will not be repeated here.
[0118] In order to implement the above embodiment, the embodiment of the present disclosure further proposes a computer-readable storage medium, on which a computer program is stored. The program is executed by a processor to execute the three-dimensional dynamic scene reconstruction method of the above embodiment.
[0119] Reference below Figure 2 , which shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiments of the present disclosure. It should be noted that the electronic device in the embodiments of the present disclosure may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, servers, etc. Figure 2 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0120] like Figure 2As shown, the electronic device may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 101, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 102 or a program loaded from a storage device 108 to a random access memory (RAM) 103. In the RAM 103, various programs and data required for the operation of the electronic device are also stored. The processing device 101, the ROM 102, and the RAM 103 are connected to each other via a bus 104. An input / output (I / O) interface 105 is also connected to the bus 104.
[0121] Typically, the following devices may be connected to the I / O interface 105: an input device 106 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, etc.; an output device 107 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 108 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 109. The communication device 109 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Figure 2 An electronic device having various devices is shown, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.
[0122] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the present embodiment includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 109, or installed from the storage device 108, or installed from the ROM 102. When the computer program is executed by the processing device 101, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.
[0123] It should be noted that the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0124] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0125] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes the three-dimensional dynamic scene reconstruction method.
[0126] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, python, and conventional procedural programming languages, such as "C-" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0127] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0128] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of this application, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0129] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code that includes one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.
[0130] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or otherwise processing in a suitable manner if necessary, and then stored in a computer memory.
[0131] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0132] A person of ordinary skill in the art may understand that all or part of the steps carried out in the above-mentioned embodiment method may be implemented by instructing the relevant hardware through a program, and the developed program may be stored in a computer-readable storage medium, which, when executed, includes one of the steps of the method embodiment or a combination thereof.
[0133] In addition, each functional unit in each embodiment of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0134] The storage medium mentioned above may be a read-only memory, a magnetic disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limiting the present application. A person of ordinary skill in the art may change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. A method for reconstructing a three-dimensional dynamic scene, characterized in that: The following steps are involved: S1, use multiple cameras to obtain multi-view synchronized videos of dynamic scenes as training data; S2, calculate the matching points between video images of different perspectives to estimate the internal and external parameters of each camera; S3, constructing a sparse point cloud according to the depth of each matching point, and generating an initial set of Gaussian sputtering point sets {p0} according to the sparse point cloud; S4. For the first frame images of all videos in the training data, use each Gaussian splattering point in the Gaussian splattering point set {p0} and calculate the color radiation value of the three-dimensional space point through the Gaussian distribution, obtain the rendered image according to the color radiation value, and perform the following iterative optimization on the Gaussian splattering point set {p0}: calculate the gradient value of the two-dimensional Gaussian point projected on the plane according to the loss function between the rendered image and the first frame image, and determine whether to perform Gaussian densification processing on the Gaussian splattering point set {p0} according to the relationship between the gradient value and the splitting threshold, and delete the Gaussian points whose volume is greater than the volume setting threshold, the Gaussian points that are invisible from all viewing angles, and the Gaussian points whose transparency is less than the transparency setting threshold in the Gaussian splattering point set {p0}, and obtain the Gaussian splattering point set {p} for the first frame images of all videos; For the remaining frame images of all videos, the pixel change variance of each perspective video is calculated, and the change voxel field is optimized according to the calculated pixel change variance. The Gaussian sputtering point set {p} is divided into a static point set {S} and a dynamic point set {D} according to the value of the change voxel field; the static point set {S} is no longer updated at the time step corresponding to the remaining frames; for the dynamic point set {D}, the attributes of each dynamic point at the time step corresponding to the remaining frame images are updated according to its rigid body constraints with the Gaussian sputtering point set {p} and its rendering loss with the remaining frame images; The dynamic Gaussian sputtering point set is composed of the Gaussian sputtering point set {p}, the static point set {S} and the final dynamic point set {D}. S5. Dynamic Gaussian sputtering point set combined with camera internal and external parameters Rendering is performed according to the Gaussian sputtering rendering pipeline to obtain rendering images at different times from a new perspective, thereby achieving dynamic three-dimensional scene reconstruction.
2. The reconstruction method according to claim 1, characterized in that: In step S2, the pixel matching points between the video images of different viewing angles are calculated by using a structure from motion method.
3. The reconstruction method according to claim 1, characterized in that: In step S4, the i-th Gaussian sputtering point in the Gaussian sputtering point set {p0} is represented as {x i ,R i ,s i ,c i ,o i }, x i ,R i ,s i ,c i ,o i are the position, rotation angle, scale, color and transparency of the i-th Gaussian sputtering point respectively. The calculation formula of the color radiation value of the three-dimensional space point is: Where c(x) represents the color radiation value of the three-dimensional space point x, and N represents the total number of Gaussian sputtering points in the Gaussian sputtering point set {p0}; For the first frame image of each perspective video, each Gaussian splattering point in the Gaussian splattering point set {p0} is projected onto the imaging plane of the corresponding perspective. The projected two-dimensional Gaussian splattering points are sequentially differentiably rendered according to the depth from small to large, and the rendered image I corresponding to the Gaussian splattering point set {p0} is obtained. The formula for differentiable rendering is as follows: Where c(p) is the color of point p in the rendered image I, α i (p) represents the transparency of point p in the rendered image I obtained by the i-th Gaussian sputtering point in the Gaussian sputtering point set {p0}, μ i is the projection position of the i-th Gaussian sputtering point in the Gaussian sputtering point set {p0} in the rendered image I, Cov i is the covariance matrix of the projection of the i-th Gaussian splatter point in the rendered image I.
4. The reconstruction method according to claim 1, characterized in that: In step S4, the iterative optimization of the Gaussian sputtering point set {p0} specifically includes: At each iteration of the optimization process, the gradient value of the two-dimensional Gaussian point projected onto the plane is calculated according to the loss function between the rendered image and the first frame image; After every K1 iterations, the following operations are performed: the average gradient value of each Gaussian sputtering point in the Gaussian sputtering point set {p0} recorded in the K1 iterations is calculated, and for each Gaussian sputtering point whose average gradient value is greater than the splitting threshold, the point is split or cloned according to the projection size of the Gaussian sputtering point, wherein, for Gaussian sputtering points whose own volume is greater than the volume threshold, the point is split according to the probability density sampling of the Gaussian distribution, and for Gaussian sputtering points whose own volume is less than or equal to the volume threshold, a cloning operation is used to generate a Gaussian sputtering point that is exactly the same as the original Gaussian sputtering point, thereby achieving Gaussian densification of the Gaussian sputtering points; at the same time, the Gaussian point pruning method is used to remove Gaussian sputtering points whose own volume is greater than the volume threshold or whose projection area is greater than the area threshold, and Gaussian sputtering points that are invisible from all viewing angles are removed; After every K2 iterations, K2>K1, all Gaussian splattering points with transparency less than the transparency threshold are removed, and then the transparency of all Gaussian splattering points is set to a fixed value less than the transparency threshold, thereby deleting the obstructed Gaussian splattering points; When the upper limit of the number of iterations is reached, the Gaussian sputtering point set {p} for the first frame images of all videos is obtained.
5. The reconstruction method according to claim 1, characterized in that: In step S4, the loss function between the rendered image and the first frame image is assumed to be Loss(I,I g ), the formula is as follows: Loss(I,I g )=L2(I,I g )+λ1·L D-SSIM (I,I g )+λ2·L LPIPS (I,I g ) Among them, I, I g They represent the rendered image and the real image obtained by the Gaussian sputtering point set {p0}; L2, L D-SSIM They represent the 2-norm loss function and the structural dissimilarity loss function, respectively, both of which are used to constrain the pixel-level similarity between images; L LPIPS represents the learnable perceptual image patch similarity loss function, which is used to reduce the perceptual distance between the rendered image and the real image; λ1 and λ2 are the loss function L D-SSIM and L LPIPS The weight of .
6. The reconstruction method according to claim 1, characterized in that: In step S4, for the remaining frame images of all videos, the specific steps of dividing the Gaussian sputtering point set {p} into a static point set {S} and a dynamic point set {D} include: For each continuous video of the multi-view video, the variance of the pixel points of each video over time is calculated, and the Gaussian blur operation is performed on the pixel point variance to ensure the continuity of the boundary. Each pixel point is divided into a dynamic pixel point and a static pixel point according to whether the pixel point variance is greater than the dynamic threshold. The change voxel field V(x,y,z) is randomly initialized to indicate whether a certain voxel (x,y,z) in the space is dynamic. The change voxel field V(x,y,z) is optimized by differentiable rendering and the variance of the pixel points of each video over time is used as the supervision, that is: Among them, L V represents the optimized changing voxel field; is the expected operator; M(r) represents the dynamic or static category of the pixel corresponding to the light ray r, M(r)=1 indicates that the pixel is a dynamic pixel, and M(r)=0 indicates that the pixel is a static pixel; It means that the voxel field V(r(y i ))The estimated pixel change field, represents the first sampling points, Indicates the position of the sampling point on the ray r; s represents the sigmoid activation function; N r represents the Gaussian sputtering point set that the light ray r passes through; According to the nearest neighbor principle, the voxels corresponding to each Gaussian sputtering point in the Gaussian sputtering point set {p} are found, and whether each Gaussian sputtering point is a dynamic Gaussian sputtering point is determined according to whether the voxel is a dynamic voxel. That is, for any Gaussian sputtering point in the Gaussian sputtering point set {p}, if the voxel corresponding to the Gaussian sputtering point is a dynamic voxel, then the Gaussian sputtering point is a dynamic Gaussian sputtering point, otherwise the Gaussian sputtering point is a static Gaussian sputtering point, thereby dividing the Gaussian sputtering point set {p} into a static point set {S} and a dynamic point set {D}.
7. The reconstruction method according to claim 1, characterized in that: In step S4, the rigid body constraint between the kth Gaussian sputtering point in the dynamic point set {D} and the lth Gaussian sputtering point in the Gaussian sputtering point set {p} is Its expression is: in, and They represent the position and rotation angle of the kth Gaussian splash point in the dynamic point set {D} at the tth frame time, respectively. and represent the position and rotation angle of the lth Gaussian sputtering point in the Gaussian sputtering point set {p} at the tth frame time; w k,l represents the weight of the loss function, which is negatively correlated with the distance between the two Gaussian sputtering points, and is specifically expressed as and They respectively represent the positions of the kth and lth Gaussian sputtering points in the first frame image.
8. The reconstruction method according to claim 7, characterized in that: In step S4, the rigid body constraint between the dynamic point set {D} and the Gaussian sputtering point set {p} is set to L rigid , which is expressed as follows: in, represents the total number of Gaussian splash points in the dynamic point set {D}, represents the kth Gaussian sputtering point selected from the Gaussian sputtering point set {p} A set of adjacent points.
9. A reconstruction device based on the reconstruction method according to any one of claims 1 to 8, characterized in that: include: A first module stores training data, wherein the training data is a multi-view synchronized video of a dynamic scene acquired by multiple cameras; The second module is used to calculate the matching points between video images of different perspectives to estimate the internal and external parameters of each camera; The third module is used to construct a sparse point cloud according to the depth of each matching point, and generate an initial set of Gaussian splash point sets {p0} according to the sparse point cloud; The fourth module is used for the first frame images of all videos in the training data, using each Gaussian splattering point in the Gaussian splattering point set {p0} and calculating the color radiation value of the three-dimensional space point through the Gaussian distribution, obtaining the rendered image according to the color radiation value, and performing the following iterative optimization on the Gaussian splattering point set {p0}: calculating the gradient value of the two-dimensional Gaussian point projected onto the plane according to the loss function between the rendered image and the first frame image, judging whether to perform Gaussian densification processing on the Gaussian splattering point set {p0} according to the relationship between the gradient value and the splitting threshold, deleting the Gaussian points whose volume is greater than the volume setting threshold, the Gaussian points that are invisible from all viewing angles, and the Gaussian points whose transparency is less than the transparency setting threshold in the Gaussian splattering point set {p0}, and obtaining the Gaussian splattering point set {p} for the first frame images of all videos; For the remaining frame images of all videos, the pixel change variance of each perspective video is calculated, and the change voxel field is optimized according to the calculated pixel change variance. The Gaussian sputtering point set {p} is divided into a static point set {S} and a dynamic point set {D} according to the value of the change voxel field; the static point set {S} is no longer updated at the time step corresponding to the remaining frames; for the dynamic point set {D}, the attributes of each dynamic point at the time step corresponding to the remaining frame images are updated according to its rigid body constraints with the Gaussian sputtering point set {p} and its rendering loss with the remaining frame images; The dynamic Gaussian sputtering point set is composed of the Gaussian sputtering point set {p}, the static point set {S} and the final dynamic point set {D}. The fifth module is used to combine the internal and external parameters of the camera to dynamically calculate the Gaussian sputtering point set. Rendering is performed according to the Gaussian sputtering rendering pipeline to obtain rendering images at different times from a new perspective, thereby achieving dynamic three-dimensional scene reconstruction.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the reconstruction method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Gaussian scattering radiation field modeling method of dynamic threshold
CN117649479A
Plant image processing method, device and equipment based on three-dimensional phenotypic modeling
CN118447047A