Dynamic scene three-dimensional reconstruction method, cloud platform, system, equipment and storage medium
Through the combination of Gaussian splashing and residual features, the dynamic scene is reconstructed in three-dimensionally, which solves the shortcomings of traditional methods in dynamic scene processing and realizes efficient three-dimensional reconstruction and cloud rendering.
Patent Information
- Application Number
- CN202510136245.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-06-27
AI Technical Summary
Traditional three-dimensional reconstruction methods have problems such as visual artifacts, model size growth, computational load weight and insufficient detail recovery when dealing with dynamic scenarios. The existing 3DGS methods ignore time consistency during dynamic modeling, resulting in data redundancy and low transmission efficiency.
The static scenes of keyframes in dynamic scenes are trained in three-dimensional reconstruction of keyframes in dynamic scenes through Gaussian splashing method, and the Gaussian ball of keyframes is obtained, and the differences between non-keyframes and keyframes are described through residual features, realizing the three-dimensional reconstruction and rendering of dynamic scenes.
It improves the quality and efficiency of three-dimensional reconstruction of dynamic scenes, reduces model parameters and data redundancy, reduces flickering and inconsistency of 3D videos, and realizes efficient scheduling and resource management of cloud rendering platforms.
Smart Images

Figure CN120219607A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology. Specifically, it relates to a method for three-dimensional reconstruction of dynamic scenes, a cloud platform, a system, a device, and a storage medium. Background Art
[0002] With the rapid development of modern science and technology, three-dimensional reconstruction technology has become one of the key technologies in various fields, and its application demand continues to grow. Traditional three-dimensional reconstruction methods are mainly based on geometric reconstruction and multi-view geometry theory. However, these methods are highly dependent on the quality of input data and face problems such as heavy computational load and insufficient detail recovery when dealing with complex scenes. In recent years, the rise of deep learning technology has provided a new development direction for three-dimensional reconstruction technology. In particular, reconstruction methods based on neural networks have gradually become the mainstream of research. As an emerging deep learning-driven three-dimensional reconstruction method, Neural Radiance Fields (NeRF) realizes the mapping from light position and direction to color and opacity through implicit representation and uses volume rendering technology to generate images. Although the NeRF model structure is compact, it has limitations in terms of reconstruction quality and computational efficiency. Especially in large-scale scene reconstruction, NeRF's performance in detail reconstruction and neural network query speed is insufficient, resulting in a slow training process and increasing the computational burden of rendering.
[0003] To overcome these limitations, the 3D Gaussian Splatting (3DGS) method has been proposed. This method uses explicit Gaussian spheres to represent three-dimensional space and simulates geometric and color information in space through Gaussian distribution, significantly improving the rendering quality and speed. Compared with NeRF, 3D Gaussian Splatting has better performance when dealing with more complex and larger-scale scenes. Its explicit representation of Gaussian spheres can more efficiently handle complex structures and details in the scene and is considered a more suitable technical solution for three-dimensional reconstruction.
[0004] Although high-efficiency realistic static scene rendering has been achieved. However, for dynamic scenes, the independent frame-by-frame modeling 3DGS method ignores temporal consistency, resulting in visual artifacts and an increase in model size. Some methods model Gaussian attributes in time to represent dynamic scenes as a unified model, improving the quality but requiring all data to be loaded simultaneously, which limits practical applications in long-sequence streaming media. Other methods track Gaussian motion frame by frame, which is suitable for streaming media, but the large amount of data per frame hinders the transmission efficiency. This poses high requirements for the dynamic modeling method of 3DGS.
[0005] In addition, traditional 3DGS runs on the user's host device, which requires the user to provide a high-end GPU device for 3DGS rendering. Moreover, this solution only supports single-scene rendering at a time, greatly limiting the 3DGS rendering experience. To enable 3DGS rendering to benefit every user without the need for a high-end GPU device, a cloud rendering platform is required that allows multiple users to view the same 3DGS object from different perspectives. However, rendering separately for each user may result in high latency and degraded rendering performance. Therefore, an efficient scheduler is needed to allocate multiple rendering tasks to the most suitable graphics card to improve hardware utilization, rendering performance, and data consistency. However, existing work mainly focuses on task-based scheduling strategies, which introduce context switching overhead and result in low graphics card utilization. Summary of the Invention
[0006] To solve one of the above technical deficiencies, an embodiment of the present application provides a method, a cloud platform, a system, a device, and a storage medium for three-dimensional reconstruction of dynamic scenes.
[0007] According to the first aspect of the embodiments of the present application, a method for three-dimensional reconstruction of dynamic scenes is provided. The method includes:
[0008] Performing three-dimensional reconstruction training on the static scene of key frames in the dynamic scene by means of Gaussian splashing to obtain Gaussian spheres of key frame scenes;
[0009] Based on the Gaussian spheres of key frame scenes and taking the features of key frames as a reference, describing the differences between non-key frames and key frames in the dynamic scene through residual features;
[0010] Adding the residual features to the features of key frames to restore the complete scene information, and rendering the dynamic scene to complete the three-dimensional reconstruction of the dynamic scene.
[0011] In an optional embodiment of the present application, the step of performing three-dimensional reconstruction training on the static scene of key frames in the dynamic scene by means of Gaussian splashing to obtain Gaussian spheres of key frame scenes further includes:
[0012] Performing first-stage three-dimensional reconstruction training and second-stage three-dimensional reconstruction training successively with the same training loss. After completing the first-stage three-dimensional reconstruction training, Gaussian spheres with opacity rankings lower than the preset ranking threshold are screened out, and then the second-stage three-dimensional reconstruction training is performed. The number of Gaussian spheres remains unchanged during the second-stage three-dimensional reconstruction training to improve the fidelity of Gaussian spheres of key frame scenes and eliminate redundancy.
[0013] In an optional embodiment of the present application, the step of based on the Gaussian spheres of key frame scenes and taking the features of key frames as a reference, describing the differences between non-key frames and key frames in the dynamic scene through residual features further includes:
[0014] Segment the dynamic scene in the form of a Gaussian sphere group to eliminate the cumulative reconstruction error.
[0015] In an optional embodiment of the present application, the steps of adding the residual feature to the feature of the key frame to restore the complete scene information and rendering the dynamic scene to complete the three-dimensional reconstruction of the dynamic scene further include:
[0016] Save each channel of each feature of each non-key frame in the form of a grayscale image and splice them, and then save them in the form of a video after further compression to reduce data redundancy.
[0017] According to the second aspect of the embodiments of the present application, there is provided a three-dimensional reconstruction cloud platform for a dynamic scene, including:
[0018] A master node, configured to perform load balancing scheduling for the three-dimensional reconstruction task of the dynamic scene based on Gaussian splash;
[0019] A slave node, communicatively connected to the master node, and rendering the color and depth information of each individual three-dimensional reconstruction object of the dynamic scene by executing the three-dimensional reconstruction method of the dynamic scene according to any one of claims 1 to 4; the rendering result of the slave node is synchronously transmitted to the master node.
[0020] In an optional embodiment of the present application, the master node further includes:
[0021] Perform fusion on different dynamic scenes according to the depth information.
[0022] In an optional embodiment of the present application, the cloud platform performs peer-to-peer communication with the client through the WebRTC protocol, and transmits the camera control information and view matrix of the client through the data channel of the WebRTC protocol.
[0023] According to the third aspect of the embodiments of the present application, there is provided a three-dimensional reconstruction system for a dynamic scene, the system includes a key frame reconstruction module, a dynamic residual characterization module, and a dynamic scene reconstruction module; wherein,
[0024] The key frame reconstruction module is configured to perform three-dimensional reconstruction training on the static scene of the key frame in the dynamic scene in the Gaussian splash manner to obtain the key frame scene Gaussian sphere;
[0025] The dynamic residual characterization module is configured to, based on the key frame scene Gaussian sphere and with the feature of the key frame as a reference, describe the difference between the non-key frame and the key frame in the dynamic scene through the residual feature;
[0026] The dynamic scene reconstruction module is configured to add the residual feature to the feature of the key frame to restore the complete scene information and render the dynamic scene to complete the three-dimensional reconstruction of the dynamic scene.
[0027] According to a fourth aspect of the embodiments of the present application, there is provided a computer device, including: a memory;
[0028] a processor; and a computer program; wherein, the computer program is stored in the memory and is configured to be executed by the processor to implement the dynamic scene three-dimensional reconstruction method according to any one of the first aspects of the embodiments of the present application.
[0029] According to a fifth aspect of the embodiments of the present application, there is provided a computer-readable storage medium, on which a computer program is stored; the computer program is executed by a processor to implement the dynamic scene three-dimensional reconstruction method according to any one of the first aspects of the embodiments of the present application.
[0030] Adopting the dynamic scene three-dimensional reconstruction method provided in the embodiments of the present application has the following beneficial effects:
[0031] Through the reconstruction of static key frames 3DGS and the residual compensation of dynamic non-key frames 3DGS, the present application improves the reconstruction quality of 3D videos; based on the residual idea, the present application uses residual features to characterize the dynamic changes between frames, rather than constructing a complete scene for each frame, reducing the number of model parameters and minimizing the flickering and inconsistency of 3D videos. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0033] Figure 1 is a flowchart of the dynamic scene three-dimensional reconstruction method provided by the embodiments of the present application;
[0034] Figure 2 is a structural diagram of the dynamic scene three-dimensional reconstruction cloud platform provided by the embodiments of the present application;
[0035] Figure 3 is a structural diagram of the dynamic scene three-dimensional reconstruction system provided by the embodiments of the present application;
[0036] Figure 4 is a schematic structural diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] In order to make the technical solutions and advantages in the embodiments of the present application clearer and more understandable, the following further details the exemplary embodiments of the present application with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than an exhaustive list of all embodiments. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.
[0038] This application proposes a method for three-dimensional reconstruction of dynamic scenes. Figure 1 The flowchart of the method for three-dimensional reconstruction of dynamic scenes provided by the embodiments of this application is shown in Figure 1 :
[0039] S1: Perform three-dimensional reconstruction training on the static scene of the key frames in the dynamic scene in the Gaussian splashing manner to obtain the Gaussian sphere of the key frame scene.
[0040] In specific implementation, in order to reconstruct a more accurate and high-fidelity dynamic 3D scene, first train a high-quality static scene of the key frames as a reference for the dynamic scene. Specifically, first use a set of Gaussian basis elements G as an explicit representation similar to point clouds to represent the scene. Each Gaussian basis element has a set of optimizable parameters
[0041] {μ; R; f; s; α}
[0042] where μ is the central position, R is the rotation matrix, f represents the SH coefficient of the view-dependent color c, s is the scaling vector, and α is the transparency. For a point x located inside the Gaussian basis element, its spatial distribution is determined by :
[0043]
[0044] where Σ = Rss T R T . During rendering, the rendering color c of the pixel is calculated by alpha blending the overlapping Gaussians in depth order.
[0045] In some embodiments of this application, the first-stage three-dimensional reconstruction training and the second-stage three-dimensional reconstruction training are performed successively with the same training loss. After completing the first-stage three-dimensional reconstruction training, the Gaussian spheres with opacity rankings behind the preset ranking threshold are screened out, and then the second-stage three-dimensional reconstruction training is performed. The number of Gaussian spheres remains unchanged during the second-stage three-dimensional reconstruction training to improve the fidelity of the Gaussian sphere of the key frame scene and eliminate redundancy.
[0046] Specifically, in the first stage of static reconstruction, it is trained according to the following loss function:
[0047]
[0048] In the formula, is the photometric loss, is the set of training pixel rays, c g (r) and are the true and reconstructed colors of the ray r respectively; It is the D-SSIM evaluation index between the real picture and the rendering result during training, and λ2 is the weight parameter.
[0049] Preferably, in the first stage, train for 10,000 steps according to the above loss function.
[0050] Furthermore, in order to avoid waste of model size, after the reconstruction training in the first stage is completed, Gaussian spheres with opacity rankings lower than the preset ranking threshold are screened out. Preferably, 35% of the Gaussian spheres with the lowest opacity rankings are screened out. Then, the second stage of training is carried out, and training is carried out according to the same training loss. Preferably, train for 5,000 steps, and no addition or deletion of Gaussian spheres is carried out during this process. After the training is completed, a Gaussian sphere for the key frame scene with high fidelity and no redundancy is obtained.
[0051] S2: Based on the Gaussian sphere for the key frame scene, taking the features of the key frame as a reference, describe the differences between non-key frames and key frames in the dynamic scene through residual features.
[0052] In some embodiments of the present application, the dynamic scene is segmented in the form of a Gaussian sphere group to eliminate the cumulative reconstruction error.
[0053] In specific implementation, in order to handle continuous changes in the dynamic scene, especially scenes with long sequences, the embodiments of the present application use a Gaussian sphere group to segment the dynamic scene. Specifically, after having the Gaussian sphere for the key frame scene, due to the continuity between frames, some subsequent frames can reuse the information of the key frame to a certain extent. However, this approach may lead to continuous accumulation of reconstruction errors, thereby degrading the reconstruction quality. Therefore, the embodiments of the present application introduce the idea of a Gaussian sphere group. A key scene is reconstructed every time a certain number of frames of the dynamic scene are dynamically modeled. While fully utilizing the inter-frame continuity, data redundancy is avoided, and at the same time, the reconstruction quality is ensured. Preferably, in the embodiments of the present application, a key scene is reconstructed every time 20 frames of the dynamic scene are dynamically modeled.
[0054] Based on this, the embodiments of the present application achieve an improvement in the reconstruction quality of 3D videos through the reconstruction of static key frame 3DGS, the residual compensation of dynamic non-key frame 3DGS, and the idea of a Gaussian sphere group. Due to the unchanged number of Gaussians and the residual idea, the situation of 3D video flickering and inconsistency is reduced to the greatest extent.
[0055] In specific implementation, for each non-key frame, the embodiments of the present application perform residual modeling and learning on each feature of the Gaussian sphere while keeping the number of Gaussian spheres consistent. This approach represents the differences between the key frame and other frames, especially including newly emerging regions and moving parts. Through the residual features, the basic features of each frame can be corrected, thereby maintaining a high reconstruction quality. The features of the key frame are used as a reference, and subsequent frames describe their differences from the key frame through residual features.
[0056] S3: Add the residual features to the features of the key frame to restore the complete scene information, and render the dynamic scene to complete the 3D reconstruction of the dynamic scene.
[0057] In a specific implementation, during the reconstruction process, adding the residual features to the features of the key frame can quickly restore the complete scene information of the current frame. Taking the Gaussian sphere position x as an example, for a non-key frame x t = x k + Δx t , where x k represents the position information of the key frame, and Δx t represents the residual position information of the non-key frame. This method reduces the redundant information of long sequence frames and effectively improves the storage efficiency.
[0058] In some embodiments of the present application, each channel of each feature of each non-key frame is saved in the form of a grayscale image and stitched together, and then saved in the form of a video after further compression to reduce data redundancy.
[0059] In a specific implementation, after obtaining the Gaussian sphere features of the key frame and the non-key frame, in the embodiments of the present application, the entire Gaussian sphere group is further compressed to further reduce data redundancy and facilitate the subsequent construction of the cloud rendering system. Specifically, in this embodiment, for each channel of each feature of 20 frames of Gaussians, this embodiment first saves it as a png image in the form of a grayscale image to initially eliminate data redundancy. Then, the features of different frames are stitched together and further compressed using H.264 (a highly compressed digital video codec standard) and saved as an mp4 video. Through this compression method, this embodiment can greatly reduce the model data volume with almost no impact on the reconstruction quality, facilitating subsequent use.
[0060] Based on this, the embodiments of the present application use residual features to characterize the dynamic changes between frames instead of constructing a complete scene for each frame. This method reduces the number of model parameters, and through further compression after training is completed, the overall storage efficiency is greatly improved.
[0061] Further, please refer to Figure 2 , the embodiments of the present application provide a 3D reconstruction cloud platform for dynamic scenes, including:
[0062] The main node 100 is used for load balancing scheduling of the 3D reconstruction task of the dynamic scene based on Gaussian splashing;
[0063] The child node 200 is communicatively connected to the master node 100 and renders the color and depth information of each individual dynamic scene 3D reconstruction object by executing the dynamic scene 3D reconstruction method proposed in the embodiments of the present application. The rendering result of the child node 200 is synchronously transmitted to the master node 100.
[0064] In some embodiments of the present application, the master node 100 performs fusion on different dynamic scenes according to the depth information.
[0065] In some embodiments of the present application, the cloud platform performs peer-to-peer communication with the client through the WebRTC (Web Real-Time Communications) protocol, and transmits the camera control information and view matrix of the client through the data channel of the WebRTC protocol.
[0066] Traditional 3DGS runs on the user's host device, which requires the user to provide a high-end GPU device for 3DGS rendering. This solution only supports single-scene rendering each time, greatly limiting the 3DGS rendering experience. The embodiments of the present application propose a cloud platform. Specifically, the cloud platform consists of a master node 100 for deploying 3DGS task scheduling and child nodes 200 for providing high-performance computing, forming a 3DGS cloud platform. In addition, all 3DGS objects on the cloud platform can be freely combined to construct a 3DGS scene. The computing child node 200 provides the rendering of the color and depth information of each individual 3DGS object, and the rendering results of all child nodes 200 are synchronized to the master node 100. Then the master node 100 performs fusion between the scenes according to the depth information. Finally, through the streaming media strategy, the rendering results are streamed to the user's multi-display terminal.
[0067] Based on this, the embodiments of the present application construct a cloud rendering platform. By deploying high-performance computing resources in the cloud, the user's device only needs to have basic network connection and display capabilities, and the cloud is responsible for rendering calculations, reducing the dependence on local hardware. In addition, the cloud platform proposed in the embodiments of the present application has a distributed rendering architecture, which distributes the 3DGS rendering tasks to multiple computing child nodes 200. Each child node 200 independently renders different objects and synchronizes the results to the master node 100 in real time for scene synthesis. Depth information fusion is also performed. The master node 100 fuses the rendering results according to the depth information returned by each computing child node 200 to ensure the correct rendering of the overall scene.
[0068] Furthermore, for WebRTC peer-to-peer communication, embodiments of the present application use a signaling server to exchange media and network metadata to initiate a peer connection between the cloud and the client. Specifically, when the client sends a request to the signaling server to experience the cloud rendering service, the signaling server allocates a primary node 100 with service resources for the client in the cloud platform. A connection is established through a discovery and negotiation process. Then, the camera controller module of the client frequently samples the user's position and orientation, and converts the position and view direction into an affine transformation matrix. Embodiments of the present application use the data channel (RTCDataChannel) of WebRTC to directly transmit this matrix from the client to the primary node 100 every 20 milliseconds. Since there is no intermediate server, the number of hops for transmitting data using RTCDataChannel is less, which allows for lower transmission latency. With this matrix, the primary node 100 assigns the rendering task to the target GPU in the distributed parallel cluster computing system (RenderFarm), and combines the ray-based rendering frames for the client in the correct order. Finally, the rendering frames are sent to the encoder to generate encoded data packets.
[0069] Embodiments of the present application use an H.264 encoder implemented by NVENC (a dedicated hardware encoder based on NVIDIA) as the encoder of the primary node 100. By offloading the complete encoding task to NVENC, the graphics engine and the CPU can be freed up for other applications, such as rendering. To enable a typical end-to-end streaming scenario with low latency, embodiments of the present application minimize the encoding and decoding latency by selecting a GOP (Group of Pictures) with an IPPP (video coding frame structure composed of I frames and P frames) structure, no look-ahead, and the lowest possible VBV (Video Buffering Verifier) buffer to adapt to the given bitrate and available channel bandwidth. In the IPPP encoding structure (without B frames), all P frames are forward predicted without backward or bi-directional prediction, which will result in lower encoding and decoding latency, and thus is suitable for the situation in embodiments of the present application. Since the rendered frames are in RGB format, embodiments of the present application convert them to YUV420 format to achieve effective compression. The compressed bitstream is packed into RTP (Real-time Transport Protocol) packets, encrypted, and transmitted to the client through WebRTC with low latency, enabling the client to receive the required frames in real time after interaction. The received frames can be decoded and displayed immediately without buffering.
[0070] Furthermore, to provide viewing services for a large scene containing numerous objects to a single user, the simplest solution is to independently render each object and composite them into the final view. In this case, the master node 100 transmits the user's camera view to all 3DGS models (i.e., renderers) in the RenderFarm, enabling each renderer to independently generate frame images and depth information, and then collect and composite them into a large scene, and finally send the result back to the user. However, this solution encounters two obstacles: (1) The I / O bandwidth limitation of the master node 100 cannot handle simultaneous image transmissions, resulting in an unacceptable rendering response time for a single user when viewing multiple objects in a large scene; (2) The workloads of different GPUs in the RenderFarm are unbalanced because the rendering of each perspective in a large scene consumes differently and requires different amounts of rendering resources. Simply allocating uniform rendering resources to different objects will result in wasted GPU power for small-sized images and long-tailed latency for large-sized images.
[0071] In the embodiments of the present application, the RenderFarm is divided into two parts: HeavyFarm (the part for processing large-sized objects) and LightFarm (the part for processing small-sized objects). In HeavyFarm, tasks with higher pixel rate (pix_rate) or lower depth requirements will be allocated one or more dedicated graphics cards; the remaining tasks will run in LightFarm. Specifically, the embodiments of the present application introduce a load balancing scheduling strategy, dividing the objects into two groups: small-sized and large-sized. The master node 100 allocates large-sized objects to HeavyFarm so that each object occupies one graphics card for better performance, and allocates a group of small-sized objects to LightFarm to share GPU resources to improve utilization. Based on this, the embodiments of the present application have a 3DGS object combination function. Through the task scheduling and resource management of the cloud platform, users can freely combine different 3D objects to build personalized dynamic scenes.
[0072] The total time T for rendering a multi-object frame frame can be expressed as follows:
[0073]
[0074] where T i refers to the rendering time of the i-th 3DGS instance, D i is the result data size of the i-th NGP-NOLF instance related to the image resolution, BW is the bandwidth limitation of the master node 100, C is the time required to integrate all rendering results, and O is the set of rendering entities. The goal of the scheduler is to ensure a lower T while maintaining the predefined frame resolution (D i ).frame The scheduler should minimize the number of instances in O as much as possible and select an appropriate subset in the RenderFarm. The embodiments of this application use a pre - envisioned grid to estimate the number of pixels and depth on the final rendered screen to predict the corresponding number of nhits. With these metrics, the embodiments of this application can determine the necessity of the rendering task and the target RenderFarm. The embodiments of this application use EGL (an enterprise - oriented platform - independent high - level programming language) to generate a rough prior scene and perform off - screen rendering, which runs so fast that its impact on the rendering time can be ignored. The embodiments of this application adopt Infiniband (an "infinite bandwidth" technology) and CUDA - aware MPI (a technology that optimizes the CPU - GPU data transfer efficiency in GPU parallel computing) for cross - node communication, and the embodiments of this application can utilize the GPUDirect RDMA (a technology that allows direct data exchange between the GPU and third - party devices without involving the CPU) optimization on NVIDIA's advanced GPUs to accelerate the reading out of the rendering results, thereby reducing the cross - node overhead.
[0075] Based on this, the embodiments of this application have achieved the efficient characterization and compression of dynamic scenes based on 3D Gaussian splashing, significantly improving the reconstruction quality and compression efficiency of the model, reducing the restoration distortion, and at the same time constructing a dynamic 3D Gaussian cloud rendering platform, reducing the hardware requirements for users to watch 3D videos, enabling the dynamic scene reconstruction technology to be more efficiently applied to user devices.
[0076] It should be understood that although the steps in the flowchart are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the figure may include multiple sub - steps or multiple stages. These sub - steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub - steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub - steps or stages of other steps.
[0077] Please refer to Figure 3 , an embodiment of this application provides a three - dimensional reconstruction system for dynamic scenes, which includes a key - frame reconstruction module 10, a dynamic residual characterization module 20, and a dynamic scene reconstruction module 30; among them,
[0078] The key - frame reconstruction module 10 is used to perform three - dimensional reconstruction training on the static scene of the key - frame in the dynamic scene by means of Gaussian splashing to obtain the Gaussian spheres of the key - frame scene;
[0079] The dynamic residual representation module 20 is used to, based on the Gaussian sphere of the key-frame scene and with the features of the key frame as a reference, describe the differences between non-key frames and key frames in the dynamic scene through residual features.
[0080] The dynamic scene reconstruction module 30 is used to add the residual features to the features of the key frame to restore the complete scene information, and render the dynamic scene to complete the three-dimensional reconstruction of the dynamic scene.
[0081] For the specific limitations of the above dynamic scene three-dimensional reconstruction system, reference can be made to the limitations on the dynamic scene three-dimensional reconstruction method in the above text, which will not be elaborated here. Each module in the above dynamic scene three-dimensional reconstruction system can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0082] In one embodiment, a computer device is provided, and the internal structure diagram of the computer device can be as Figure 4 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements the above dynamic scene three-dimensional reconstruction method. It includes: a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the above dynamic scene three-dimensional reconstruction method.
[0083] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by the processor, it can implement the above dynamic scene three-dimensional reconstruction method.
[0084] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code. The solutions in the embodiments of the present application can be implemented in various computer languages. For example, C language, VHDL language, Verilog language, object-oriented programming language Java, and interpreted scripting language JavaScript, etc.
[0085] The present application is described with reference to the flowcharts and / or block diagrams of systems and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0086] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0087] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0088] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.
[0089] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these changes and modifications.
Claims
1. A method for 3D reconstruction of a dynamic scene, characterized in that: include: The key frame static scene in the dynamic scene is trained for 3D reconstruction by Gaussian splashing to obtain the Gaussian sphere of the key frame scene; According to the key frame scene Gaussian ball, taking the features of the key frame as a reference, describing the difference between the non-key frame and the key frame in the dynamic scene through residual features; The residual feature is added to the feature of the key frame to restore the complete scene information, and the dynamic scene is rendered to complete the three-dimensional reconstruction of the dynamic scene.
2. The dynamic scene 3D reconstruction method according to claim 1, characterized in that: The step of performing three-dimensional reconstruction training on the key frame static scene in the dynamic scene by Gaussian splashing to obtain the Gaussian sphere of the key frame scene also includes: The first stage 3D reconstruction training and the second stage 3D reconstruction training are performed successively with the same training loss. After completing the first stage 3D reconstruction training, Gaussian spheres with opacity ranking above a preset ranking threshold are screened out, and then the second stage 3D reconstruction training is performed. In the second stage 3D reconstruction training, the number of Gaussian spheres is kept unchanged to improve the fidelity of the Gaussian spheres of the key frame scene and eliminate redundancy.
3. The dynamic scene 3D reconstruction method according to claim 1, characterized in that: The step of describing the difference between the non-key frame and the key frame in the dynamic scene by residual features based on the key frame scene Gaussian ball and taking the features of the key frame as a reference also includes: The dynamic scene is segmented in the form of Gaussian sphere groups to eliminate the accumulated reconstruction error.
4. The dynamic scene 3D reconstruction method according to claim 1, characterized in that: The step of adding the residual feature to the feature of the key frame to restore the complete scene information and rendering the dynamic scene to complete the three-dimensional reconstruction of the dynamic scene also includes: Each channel of each feature of each non-key frame is saved and spliced in the form of a grayscale image, and is further compressed and saved in the form of a video to reduce data redundancy.
5. A dynamic scene 3D reconstruction cloud platform, characterized in that: include: The master node is used to perform load balancing scheduling of dynamic scene 3D reconstruction tasks based on Gaussian splashing; A child node is communicatively connected to the main node, and renders the color and depth information of each individual dynamic scene three-dimensional reconstruction object by executing the dynamic scene three-dimensional reconstruction method as described in any one of claims 1 to 4; the rendering result of the child node is synchronously transmitted to the main node.
6. The dynamic scene 3D reconstruction cloud platform according to claim 5, characterized in that: The master node also includes: Fusion is performed on different dynamic scenes according to the depth information.
7. The dynamic scene 3D reconstruction cloud platform according to claim 5, characterized in that: The cloud platform performs point-to-point communication with the client through the WebRTC protocol, and transmits the camera control information and viewing angle matrix of the client through the data channel of the WebRTC protocol.
8. A dynamic scene 3D reconstruction system, characterized in that: include: Key frame reconstruction module, dynamic residual representation module and dynamic scene reconstruction module; among them, A key frame reconstruction module is used to perform three-dimensional reconstruction training on the key frame static scene in the dynamic scene by Gaussian splashing to obtain the Gaussian sphere of the key frame scene; A dynamic residual characterization module, used to describe the difference between the non-key frame and the key frame in the dynamic scene through residual features based on the Gaussian ball of the key frame scene and the features of the key frame; The dynamic scene reconstruction module is used to add the residual features to the features of the key frame to restore the complete scene information, and render the dynamic scene to complete the three-dimensional reconstruction of the dynamic scene.
9. A computer device, characterized in that: include: Memory; processor; as well as Computer programs; The computer program is stored in the memory and is configured to be executed by the processor to implement the dynamic scene three-dimensional reconstruction method as claimed in any one of claims 1 to 4.
10. A computer-readable storage medium, characterized in that: A computer program is stored thereon; the computer program is executed by a processor to implement a dynamic scene three-dimensional reconstruction method as claimed in any one of claims 1 to 4.