Three-dimensional static background reconstruction method and device, equipment and storage medium

By combining the technology of 3DGS renderer and mask predictor, the problem of manually labeling dynamic object information in the existing technology is solved, and the method of automatically reconstructing static backgrounds is realized, which improves the reconstruction efficiency and accuracy.

CN120147532AActive Publication Date: 2025-06-13COWA TECHNOLOGY CO LTD +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510224592.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-13
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

When handling dynamic scenarios, existing three-dimensional scene reconstruction technology requires manual labeling of bounding boxes and trajectories of dynamic objects, which is time-consuming and labor-intensive and not suitable for automated processing.

Method used

Using a combination of 3DGS renderer and mask predictor, by obtaining the real image and camera parameters of the scene, initializing the 3D Gaussian parameters and inputting the renderer, outputting the rendered image, and then inputting the rendered image and the real image into the mask predictor, generating a mask image to distinguish between static background and dynamic objects, and finally training the 3D Gaussian parameters and mask predictor to reconstruct the static background.

Benefits of technology

It realizes that static background is automatically reconstructed from scenes containing dynamic objects without relying on manual annotation information, improving reconstruction efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147532A_ABST
    Figure CN120147532A_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional static background reconstruction method and device, equipment and a storage medium. The method comprises the following steps: acquiring a real image and camera parameters of a scene; inputting the initialized 3D Gaussian parameters and camera parameters into a 3DGS renderer; inputting a rendered image output by the 3DGS renderer and a real image into a mask predictor; training a 3D Gaussian parameter and a mask predictor; and inputting the trained 3D Gaussian parameters and camera parameters into a 3DGS renderer, and outputting a reconstructed rendering image of the static background. Based on the 3DGS renderer and the mask predictor, the static background is reconstructed from the scene data containing the dynamic object by using the 3DGS without depending on manual annotation information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of three-dimensional scene reconstruction, and particularly relates to a method, device, equipment and storage medium for reconstructing a three-dimensional static background. Background Art

[0002] Three-dimensional scene reconstruction refers to the process of reconstructing the original three-dimensional information based on multiple pictures taken from different perspectives in a scene. The reconstruction result can provide support for fields such as visual simulation, virtual reality, and game modeling. The 3DGS method has great advantages in terms of three-dimensional reconstruction fidelity and rendering efficiency, and has become a research hotspot. However, 3DGS assumes that the scene is static and is not applicable to dynamic scenes, that is, scenes containing dynamic objects. Dynamic objects cause artifacts in the scene reconstructed by 3DGS.

[0003] The existing solution is to reconstruct the entire dynamic scene by separately modeling the dynamic objects and the static background in the dynamic scene and then merging and rendering them, that is, reconstructing both the static background and the dynamic objects.

[0004] Specifically, through the scene point cloud and the rectangular contour frame (bounding box, bbox) of the dynamic object and its trajectory, the scene is divided into 1+N pieces of Gaussian point cloud: 1 piece representing the Gaussian point cloud of the dynamic background and N pieces representing the Gaussian point clouds of N dynamic objects. The parameters of these 1+N pieces of Gaussian point cloud are similar to those of 3DGS. The coordinates of the dynamic object in the static background are calculated according to its tracking trajectory, and these 1+N pieces of Gaussian point cloud are combined together and then rendered uniformly.

[0005] However, this technology requires the bbox of the dynamic object and its trajectory to complete the modeling strategy of separating the static and the dynamic. However, the information such as the bbox of the dynamic object and its trajectory depends on manual annotation, which is time-consuming and laborious. Summary of the Invention

[0006] In order to solve the above technical problems, the present invention proposes a method, device, equipment and storage medium for reconstructing a three-dimensional static background.

[0007] To achieve the above object, the technical solution of the present invention is as follows:

[0008] In a first aspect, the present invention discloses a method for reconstructing a three-dimensional static background, including:

[0009] Step S1: Obtain the real image of the scene and the camera parameters;

[0010] Step S2: Input the initialized 3D Gaussian parameters and the camera parameters into the 3DGS renderer, and the 3DGS renderer is used to output the reconstructed rendered image;

[0011] Step S3: Input the rendered image and the real image output by the 3DGS renderer into the mask predictor, which is used to output a mask image. The value of each pixel on the mask image can reflect the reconstruction effect of the rendered image and distinguish the static background from the dynamic objects;

[0012] Step S4: Train the 3D Gaussian parameters and the mask predictor;

[0013] Step S5: Input the trained 3D Gaussian parameters and camera parameters into the 3DGS renderer to output the rendered image of the reconstructed static background.

[0014] Based on the above technical solutions, the following improvements can be made:

[0015] As a preferred solution, in step S3, the value of each pixel on the mask image is between 0 and 1,

[0016] and the lower the value of the pixel, the more it represents that the pixel is occupied by the dynamic object;

[0017] The higher the value of the pixel, the more it represents that the pixel is occupied by the static background.

[0018] As a preferred solution, step S4 calculates the loss through the following steps:

[0019] Step A: Calculate the Hadamard products of the rendered image and the real image with the mask image respectively;

[0020]

[0021] where: I render and I real are the rendered image and the real image respectively;

[0022] M is the mask image;

[0023] ⊙ is the Hadamard product;

[0024] Step B: Based on the calculation results of step A, obtain the loss Loss through the following formula;

[0025]

[0026] where: is the L1 loss;

[0027] is the SSIM loss;

[0028] λ SSIM is the loss coefficient of SSIM.

[0029] As a preferred solution, during the training of the 3D Gaussian parameters and the mask predictor, when backpropagating, the gradient of the mask predictor does not propagate to the rendered image.

[0030] In a second aspect, the present invention also discloses a three-dimensional static background reconstruction device, including:

[0031] An acquisition module, configured to acquire a real image of a scene and camera parameters;

[0032] A rendering module, configured to input the initialized 3D Gaussian parameters and camera parameters into a 3DGS renderer, and the 3DGS renderer is configured to output a reconstructed rendered image;

[0033] A prediction module, configured to input the rendered image and the real image output by the 3DGS renderer into a mask predictor, and the mask predictor is configured to output a mask image, and the value of each pixel on the mask image can reflect the reconstruction effect of the rendered image and distinguish between a static background and a dynamic object;

[0034] A training module, configured to train the 3D Gaussian parameters and the mask predictor;

[0035] A reconstruction module, configured to input the trained 3D Gaussian parameters and camera parameters into the 3DGS renderer to output a rendered image of the reconstructed static background.

[0036] As a preferred solution, in the prediction module, the value of each pixel on the mask image is between 0 and 1,

[0037] and the lower the value of the pixel, the more it represents that the pixel is occupied by a dynamic object;

[0038] The higher the value of the pixel, the more it represents that the pixel is occupied by a static background.

[0039] As a preferred solution, the training module includes:

[0040] A first calculation unit, configured to calculate the Hadamard products of the rendered image and the real image with the mask image respectively;

[0041]

[0042] where: I render and I real are the rendered image and the real image respectively;

[0043] M is the mask image;

[0044] ⊙ is the Hadamard product;

[0045] A second calculation unit, configured to obtain a loss Loss through the following formula based on the calculation result of the first calculation unit;

[0046]

[0047] Wherein: is the L1 loss;

[0048] is the SSIM loss;

[0049] λ SSIM is the loss coefficient of SSIM.

[0050] As a preferred solution, during the training process of the 3D Gaussian parameters and the mask predictor, when backpropagating, the gradient of the mask predictor does not propagate to the rendered image.

[0051] In a third aspect, the present invention also discloses a computing device, including:

[0052] One or more processors;

[0053] A memory;

[0054] And one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include instructions for any of the above three-dimensional static background reconstruction methods.

[0055] In a fourth aspect, the present invention also discloses a storage medium, the storage medium stores one or more computer-readable programs, and the one or more programs include instructions, and the instructions are adapted to be loaded and executed by the memory to perform any of the above three-dimensional static background reconstruction methods.

[0056] The present invention discloses a three-dimensional static background reconstruction method, device, device and storage medium, having the following beneficial effects:

[0057] Based on the 3DGS renderer and the mask predictor, the present invention uses 3DGS to reconstruct the static background from the scene data containing dynamic objects without relying on manual annotation information. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0059] Figure 1 Is the flowchart of the three-dimensional static background reconstruction method provided by the embodiment of the present invention.

[0060] Figure 2Flow chart of the 3D static background reconstruction method provided by the embodiments of the present invention.

[0061] Figure 3 Block diagram of the 3D static background reconstruction device provided by the embodiments of the present invention.

[0062] Figure 4 Block diagram of the computing device provided by the embodiments of the present invention. Detailed implementation manners

[0063] The preferred implementation manners of the present invention will be described in detail below with reference to the accompanying drawings.

[0064] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0065] Using ordinal numbers such as "first", "second", "third", etc. to describe ordinary objects only represents different instances of similar objects, and does not intend to imply that the objects so described must have a given order in terms of time, space, sorting, or any other way.

[0066] In addition, the expression "including" elements is an "open" expression, and this "open" expression only means that there are corresponding components or steps, and should not be construed as excluding additional components or steps.

[0067] In order to achieve the purpose of the present invention, in some embodiments of the 3D static background reconstruction method, the 3D static background reconstruction method is based on 3DGS, as Figure 1-2 shown, including:

[0068] Step S101: Obtain the real image of the scene and the camera parameters;

[0069] Step S102: Input the initialized 3D Gaussian parameters and the camera parameters into the 3DGS renderer, and the 3DGS renderer is used to output the reconstructed rendered image;

[0070] Step S103: Input the rendered image output by the 3DGS renderer and the real image into the mask predictor, and the mask predictor is used to output the mask image, and the value of each pixel point on the mask image can reflect the reconstruction effect of the rendered image and distinguish the static background and the dynamic object;

[0071] Step S104: Train the 3D Gaussian parameters and the mask predictor;

[0072] Step S105: Input the trained 3D Gaussian parameters and camera parameters into the 3DGS renderer, and output the rendered image of the reconstructed static background.

[0073] Each of the above steps will be elaborated in detail below.

[0074] In step S101, the real image of the scene and camera parameters are obtained. The camera parameters include, but are not limited to, information such as the position, inclination angle, focal length, and optical center offset of the camera. These parameters determine the projection method from the 3D scene to the 2D image.

[0075] In step S102, the 3D Gaussian parameters include, but are not limited to, information such as the position, shape, opacity, and color of the Gaussian points. Multiple Gaussian points constitute the 3D representation of the entire scene.

[0076] The 3DGS renderer takes the 3D Gaussian parameters and camera parameters as inputs, and renders the rendered image of the 3D scene under the projection method determined by the camera parameters. Specifically, the 3DGS renderer first projects the 3D Gaussian onto the 2D plane according to the projection method, and then superimposes its color information based on the occlusion relationship and opacity of the Gaussian to calculate the RGB value of each pixel in the rendered image.

[0077] The 3DGS renderer is essentially a rendering algorithm that renders 3D Gaussian parameters into a 2D image in a manner determined by the camera parameters. The present invention uses the real image of the scene and the corresponding camera parameters as training data to train the 3D Gaussian parameters. It should be noted that the rendering process of the 3DGS renderer is well-known to those skilled in the art and will not be elaborated here.

[0078] Step S103 mainly involves a mask predictor. The mask predictor consists of a neural network, takes the rendered image and the real image as inputs, and outputs a mask image mask. The size of the mask image mask is the same as the size of the input images (rendered image and real image), specifically (H, W). The value of each pixel point on the mask image is between 0 and 1.

[0079] The lower the value of the pixel point (indicating that the 3D scene area corresponding to the position of the pixel point is reconstructed worse), the more it represents that the pixel point is occupied by a dynamic object;

[0080] The higher the value of the pixel point (indicating that the 3D scene area corresponding to the position of the pixel point is reconstructed better), the more it represents that the pixel point is occupied by the static background.

[0081] Corresponding to the mask image mask, the whiter the color, the closer the value is to 0.

[0082] For the stability during the training of the 3D Gaussian parameters, during backpropagation, the mask predictor does not propagate the gradient to the rendered image.

[0083] It should be noted that the neural network structure of the mask predictor can but is not limited to using U-Net, and the image output by this network can be interpolated to obtain the final mask image mask with the same size as the input image.

[0084] Step S104 calculates the loss through the following steps:

[0085] Step A: Calculate the Hadamard products of the rendered image and the ground truth image with the mask image respectively;

[0086]

[0087] where: I render and I real are the rendered image and the ground truth image respectively;

[0088] M is the mask image;

[0089] ⊙ is the Hadamard product;

[0090] Step B: Based on the calculation results of Step A, obtain the loss Loss through the following formula;

[0091]

[0092] where: is the L1 loss;

[0093] is the SSIM (Structural Similarity Index Measure, SSIM) loss;

[0094] λ SSIM is the loss coefficient of SSIM, which can be set to 0.2.

[0095] The present invention does not directly calculate the L1 loss and the SSIM loss of the rendered image and the ground truth image, but first calculates the Hadamard products of these two images and the mask image mask respectively, and then calculates the L1 loss and the SSIM loss using the results of the Hadamard products. Since the pixel values in the mask image mask are between 0 and 1, the smaller the value, the more likely the corresponding position in the image is a dynamic object. Using the results of the Hadamard product with the mask image mask to calculate the loss reduces the contribution of dynamic objects to the gradient, making the training focus more on the static background, thereby achieving a better reconstruction effect.

[0096] The above training of the mask predictor does not use additional labels and uses the same loss as the training of 3DGS.

[0097] In some other embodiments, the present invention also discloses a three-dimensional static background reconstruction device, as Figure 3 shown, including:

[0098] An acquisition module 201 for acquiring a real image of a scene and camera parameters;

[0099] A rendering module 202 for inputting initialized 3D Gaussian parameters and camera parameters into a 3DGS renderer, and the 3DGS renderer is used to output a reconstructed rendered image;

[0100] A prediction module 203 for inputting the rendered image and the real image output by the 3DGS renderer into a mask predictor, and the mask predictor is used to output a mask image, and the value of each pixel point on the mask image can reflect the reconstruction effect of the rendered image and distinguish between a static background and a dynamic object;

[0101] A training module 204 for training 3D Gaussian parameters and a mask predictor;

[0102] A reconstruction module 205 for inputting the trained 3D Gaussian parameters and camera parameters into the 3DGS renderer to output a rendered image of the reconstructed static background.

[0103] Further, in the prediction module, the value of each pixel point on the mask image is between 0 and 1,

[0104] and the lower the value of the pixel point, the more it represents that the pixel point is occupied by a dynamic object;

[0105] The higher the value of the pixel point, the more it represents that the pixel point is occupied by a static background.

[0106] Further, the training module includes:

[0107] A first calculation unit for respectively calculating the Hadamard products of the rendered image and the real image with the mask image;

[0108]

[0109] where: I render and I real are the rendered image and the real image respectively;

[0110] M is the mask image;

[0111] ⊙ is the Hadamard product;

[0112] A second calculation unit for obtaining a loss Loss through the following formula based on the calculation results of the first calculation unit;

[0113]

[0114] where: is the L1 loss;

[0115] is the SSIM loss;

[0116] λ SSIM is the loss coefficient of SSIM.

[0117] Furthermore, during the training process of the 3D Gaussian parameters and the mask predictor, when backpropagating, the gradient of the mask predictor does not propagate to the rendered image.

[0118] Furthermore, it should be noted that: when the three-dimensional static background reconstruction device provided in the above embodiment performs static background reconstruction, only the division of the above functional modules is used as an example for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the three-dimensional static background reconstruction device is divided into different functional modules to complete all or part of the functions described above.

[0119] In addition, the three-dimensional static background reconstruction device provided in the above embodiment and the embodiment of the three-dimensional static background reconstruction method belong to the same concept. For the specific implementation process, please refer to the method embodiment and will not be elaborated here.

[0120] In addition, in some other embodiments, as Figure 4 shown, the present invention also discloses a computing device, including:

[0121] one or more processors 301;

[0122] a memory 302;

[0123] and one or more programs, where one or more programs are stored in the memory 302 and are configured to be executed by one or more processors 301. One or more programs include the instructions of the three-dimensional static background reconstruction method disclosed in the above embodiment.

[0124] The processor 301 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 301 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 301 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 301 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 301 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0125] The memory 302 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 302 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 302 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 301 to implement the three-dimensional static background reconstruction method provided in the method embodiments of the present invention.

[0126] In addition, the computing device may optionally further include: a peripheral device interface and at least one peripheral device. The processor 301, the memory 302, and the peripheral device interface may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface through a bus, signal lines, or a circuit board. Schematically, the peripheral devices include but are not limited to: a radio frequency circuit, a touch display screen, an audio circuit, and a power supply, etc.

[0127] Of course, the computing device may also include fewer or more components, and this embodiment does not limit this.

[0128] In addition, in some other embodiments, the present invention also discloses a storage medium, and the storage medium stores one or more computer-readable programs. The one or more programs include instructions, and the instructions are adapted to be loaded and executed by the memory to perform the three-dimensional static background reconstruction method disclosed in the above embodiments.

[0129] The present invention discloses a three-dimensional static background reconstruction method, device, equipment and storage medium, which has the following beneficial effects:

[0130] Based on a 3DGS renderer and a mask predictor, the present invention uses 3DGS to reconstruct the static background from the scene data containing dynamic objects without relying on manually labeled information.

[0131] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification only illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.

Claims

1. A three-dimensional static background reconstruction method, characterized in that: include: Step S1: Obtain the real image and camera parameters of the scene; Step S2: inputting the initialized 3D Gaussian parameters and camera parameters into a 3DGS renderer, wherein the 3DGS renderer is used to output a reconstructed rendered image; Step S3: inputting the rendered image and the real image output by the 3DGS renderer into a mask predictor, wherein the mask predictor is used to output a mask image, wherein the value of each pixel on the mask image can reflect the reconstruction effect of the rendered image and distinguish static background from dynamic objects; Step S4: training 3D Gaussian parameters and mask predictor; Step S5: input the trained 3D Gaussian parameters and camera parameters into the 3DGS renderer, and output a rendered image of the reconstructed static background.

2. The three-dimensional static background reconstruction method according to claim 1, characterized in that: In step S3, the value of each pixel on the mask image is between 0 and 1. And the lower the value of a pixel, the more it represents that the pixel is occupied by a dynamic object; The higher the pixel value, the more likely it is that the pixel is occupied by a static background.

3. The three-dimensional static background reconstruction method according to claim 1, characterized in that: The step S4 calculates the loss by the following steps: Step A: Calculate the Hadamard product of the rendered image and the real image with the mask image respectively; Where: I render and I real They are rendered images and real images respectively; M is the mask image; ⊙ is Hadamard; Step B: Based on the calculation result of step A, the loss Loss is obtained by the following formula; in: is L1 loss; is the SSIM loss; λ SSIM is the loss coefficient of SSIM.

4. The three-dimensional static background reconstruction method according to claim 1, characterized in that: During the training of the 3D Gaussian parameters and the mask predictor, the gradients of the mask predictor are not propagated to the rendered image during back-propagation.

5. A three-dimensional static background reconstruction device, characterized in that: include: Acquisition module, used to obtain the real image and camera parameters of the scene; A rendering module, used for inputting the initialized 3D Gaussian parameters and camera parameters into a 3DGS renderer, wherein the 3DGS renderer is used for outputting a reconstructed rendering image; A prediction module, used for inputting the rendered image and the real image output by the 3DGS renderer into a mask predictor, wherein the mask predictor is used for outputting a mask image, wherein the value of each pixel on the mask image can reflect the reconstruction effect of the rendered image and distinguish static background from dynamic objects; Training module, used to train 3D Gaussian parameters and mask predictors; The reconstruction module is used to input the trained 3D Gaussian parameters and camera parameters into the 3DGS renderer and output the rendered image of the reconstructed static background.

6. The three-dimensional static background reconstruction device according to claim 5, characterized in that: In the prediction module, the value of each pixel on the mask image is between 0 and 1. And the lower the value of a pixel, the more it represents that the pixel is occupied by a dynamic object; The higher the pixel value, the more likely it is that the pixel is occupied by a static background.

7. The three-dimensional static background reconstruction device according to claim 5, characterized in that: The training modules include: A first calculation unit, used for respectively calculating the Hadamard product of the rendered image and the real image with the mask image; Where: I render and I real They are rendered images and real images respectively; M is the mask image; ⊙ is Hadamard; A second calculation unit is used to obtain a loss Loss through the following formula based on the calculation result of the first calculation unit; in: is L1 loss; is the SSIM loss; λ SSIM is the loss coefficient of SSIM.

8. The three-dimensional static background reconstruction device according to claim 5, characterized in that: During the training of the 3D Gaussian parameters and the mask predictor, the gradients of the mask predictor are not propagated to the rendered image during back-propagation.

9. A computing device, characterized in that include: one or more processors; Memory; And one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by one or more processors, and one or more of the programs include instructions for the three-dimensional static background reconstruction method described in any one of claims 1-4 above.

10. A storage medium, characterized in that The storage medium stores one or more computer-readable programs, and the one or more programs include instructions, and the instructions are suitable for being loaded by the memory and executing the three-dimensional static background reconstruction method described in any one of claims 1-4.

Citation Information

Patent Citations

  • Target reconstruction method based on differentiable SAR image renderer

    CN117437347A

  • Real-time high-quality dynamic human body rendering method based on multi-view image guidance

    CN118071908A

  • Traffic anomaly detection optimization strategy based on intelligent algorithm

    CN118262301A

  • Sparse visual angle three-dimensional reconstruction method based on depth prior information

    CN118657888A

  • Dynamic scene rendering method and system based on three-dimensional decomposition Hash coding

    CN118840471A