Gaussian splash reconstruction method based on depth camera and related equipment
Through the Gaussian splash reconstruction method based on the depth camera, combining the internal reference of the depth camera and camera pose, point cloud and image information, the Gaussian splash model is trained and image information is generated from multiple perspectives, which solves the reconstruction quality problem caused by scene changes and achieves more efficient and complex scene reconstruction.
Patent Information
- Application Number
- CN202510110393.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-23
AI Technical Summary
Existing Gaussian splash reconstruction methods cannot guarantee the reconstruction quality when scene changes, especially in complex or dynamic scenarios.
The Gaussian splatter reconstruction method based on the depth camera is adopted to initialize the point cloud of the target object by obtaining the internal parameters of the depth camera and the camera pose, and the Gaussian splatter model is trained in combination with RGB images and depth images. During the model training process, the generative model and monocular depth estimation algorithm are used to generate image information from multiple perspectives and added to the training set until the reconstruction is completed.
It improves the reconstruction effect of the target object from a sparse perspective, enhances the rendering and reconstruction capabilities of complex scenes, and ensures the stability of reconstruction quality.
Smart Images

Figure CN120032052A_ABST
Abstract
Description
Background Art
[0002] Three-dimensional (3D) scene reconstruction is an important research direction in the field of computer graphics and computer vision. Its goal is to recover the structure and appearance of a 3D scene from 2D image or video data. 3D scene reconstruction relies on a variety of technical means, such as stereo vision, structured light, laser scanning, etc. These methods may not work well in complex or dynamic scenes because it becomes particularly difficult to match corresponding points or handle light deformation. In addition, these methods may also have limitations when dealing with large areas or irregularly shaped objects.
[0003] Among the related technologies, Gaussian splashing, as a 3D scene reconstruction method that has emerged in recent years, has shown its strong application potential in many fields. Its core principle is to perform 3D reconstruction and new perspective synthesis based on Gaussian functions. By representing the scene as a combination of a set of Gaussian functions, it achieves efficient rendering of complex scenes. Although these rendering results are impressive, the same scene usually undergoes many changes over time, resulting in the inability to guarantee the quality of reconstruction.
[0004] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention
[0005] The present disclosure provides a Gaussian splash reconstruction method based on a depth camera and related equipment, which at least to a certain extent overcomes the problem that the reconstruction quality of Gaussian splash cannot be guaranteed due to scene changes in related technologies.
[0006] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by the practice of the present disclosure.
[0007] According to one aspect of the present disclosure, a Gaussian splash reconstruction method based on a depth camera is provided, including: obtaining the intrinsic parameters and camera pose of each depth camera; obtaining the RGB image and the corresponding depth image of the target object taken by each depth camera at the same time; initializing the point cloud of the target object according to the intrinsic parameters, the camera pose, the RGB image and the depth image; training a Gaussian splash model according to the point cloud, the RGB image and the depth image; in training the Gaussian splash model, generating the RGB images and the corresponding depth images of the target object under multiple perspectives according to the generation model and the monocular depth estimation algorithm, and adding them to the training set until the reconstruction of the Gaussian splash model is completed.
[0008] In some embodiments, obtaining the intrinsic parameters and camera pose of each depth camera includes: obtaining the intrinsic parameters and camera pose of each depth camera according to a standard checkerboard image calibration method.
[0009] In some embodiments, initializing the point cloud of the target object according to the intrinsic parameters, the camera pose, the RGB image and the depth image includes: mapping each pixel in the depth image to a three-dimensional space according to the intrinsic parameters and the camera pose to form the point cloud, wherein the RGB image is used to provide color information of the pixel points.
[0010] In some embodiments, the method also includes: creating a truncated signed distance function; traversing the point cloud, finding a corresponding volume element for each point in the truncated signed distance function volume; updating the signed distance value of the volume element based on the distance from the point to the center of the volume element, and determining the established truncated signed distance function, wherein the truncated signed distance function is used as the initial state of the Gaussian splash model.
[0011] In some embodiments, the training of the Gaussian splash model based on the point cloud, the RGB image and the depth image includes: using the point cloud as the initial state of the Gaussian splash model, using the RGB image and the depth image as training sets of the Gaussian splash model, and training the Gaussian splash model.
[0012] In some embodiments, generating RGB images and corresponding depth images of the target object under multiple perspectives based on the generation model and the monocular depth estimation algorithm includes: deflecting a depth camera by a preset angle to determine the pose matrix of the new perspective; rendering the Gaussian splatter model obtained by the current training from the perspective of the pose matrix to determine the rendered RGB image; using the rendered RGB image as input of the generation model to output the RGB image of the target object; using the RGB image of the target object as input of the monocular depth estimation algorithm to output the corresponding depth image.
[0013] According to another aspect of the present disclosure, a Gaussian splash reconstruction device based on a depth camera is also provided, including: an information acquisition module, used to obtain the intrinsic parameters and camera pose of each depth camera; an image acquisition module, used to obtain the RGB image and the corresponding depth image of the target object taken by each depth camera at the same time; an initialization module, used to initialize the point cloud of the target object according to the intrinsic parameters, the camera pose, the RGB image and the depth image; a training module, used to train a Gaussian splash model according to the point cloud, the RGB image and the depth image; a reconstruction module, used to generate RGB images and corresponding depth images of the target object under multiple perspectives according to a generation model and a monocular depth estimation algorithm in training the Gaussian splash model, and add them to a training set until the reconstruction of the Gaussian splash model is completed.
[0014] According to another aspect of the present disclosure, an electronic device is also provided, which includes: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute any one of the above-mentioned Gaussian splash reconstruction methods based on a depth camera by executing the executable instructions.
[0015] According to another aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the Gaussian splash reconstruction method based on a depth camera described above is implemented.
[0016] According to another aspect of the present disclosure, a computer program product is also provided, including a computer program, wherein when the computer program is executed by a processor, the computer program implements any one of the above-mentioned Gaussian splash reconstruction methods based on a depth camera.
[0017] The Gaussian splash reconstruction method based on the depth camera provided in the embodiments of the present disclosure includes: obtaining the intrinsic parameters and camera pose of each depth camera; obtaining the RGB image and the corresponding depth image of the target object taken by each depth camera at the same time; initializing the point cloud of the target object according to the intrinsic parameters, camera pose, RGB image and depth image; training the Gaussian splash model according to the point cloud, RGB image and depth image; in the training of the Gaussian splash model, generating the RGB images and corresponding depth images of the target object under multiple perspectives according to the generative model and the monocular depth estimation algorithm, and adding them to the training set until the reconstruction of the Gaussian splash model is completed. The present application can improve the reconstruction effect of the target object under sparse perspective by combining the generative model with the monocular depth estimation algorithm under sparse perspective, and continuously generating new perspective information that is increasingly far away from the existing perspective during the training process of the Gaussian splash model.
[0018] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification are used to explain the principles of the present disclosure. Obviously, the accompanying drawings described below are only some embodiments of the present disclosure, and for ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without creative work.
[0020] Figure 1 A schematic diagram showing a structure of a Gaussian splash reconstruction system based on a depth camera in an embodiment of the present disclosure;
[0021] Figure 2 A flowchart of a Gaussian splash reconstruction method based on a depth camera in an embodiment of the present disclosure is shown;
[0022] Figure 3 A flowchart showing a specific example of a Gaussian splash reconstruction method based on a depth camera in an embodiment of the present disclosure;
[0023] Figure 4 A flowchart showing another specific example of a Gaussian splash reconstruction method based on a depth camera in an embodiment of the present disclosure;
[0024] Figure 5 A flowchart showing another specific example of a Gaussian splash reconstruction method based on a depth camera in an embodiment of the present disclosure;
[0025] Figure 6 A flowchart showing another specific example of a Gaussian splash reconstruction method based on a depth camera in an embodiment of the present disclosure;
[0026] Figure 7 A schematic diagram of a Gaussian splash reconstruction device based on a depth camera in an embodiment of the present disclosure is shown;
[0027] Figure 8 A structural block diagram of a computer device according to an embodiment of the present disclosure is shown;
[0028] Fig. 9 A schematic diagram showing a computer-readable storage medium in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0029] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that the disclosure will be more comprehensive and complete and to fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0030] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the figures represent the same or similar parts, and their repeated description will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0031] For ease of understanding, before introducing the embodiments of the present disclosure, several terms involved in the embodiments of the present disclosure are first explained as follows:
[0032] TSDF: Truncated Signed Distance Function, truncated signed distance function
[0033] RGB: Red, Green, Blue, standard representation of color images.
[0034] The specific implementation of the embodiment of the present disclosure is described in detail below with reference to the accompanying drawings.
[0035] Figure 1 FIG. 2 shows an exemplary application system architecture diagram to which the Gaussian splash reconstruction method based on a depth camera in the embodiment of the present disclosure can be applied. Figure 1 As shown, the system architecture 100 may include an acquisition device 101 , a network 102 and a training device 103 .
[0036] The network 102 is a medium for providing a communication link between the acquisition device 101 and the training device 103, and may be a wired network or a wireless network.
[0037] Optionally, the wireless network or wired network described above uses standard communication technology and / or protocol. The network is usually the Internet, but it can also be any network, including but not limited to a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a dedicated network or any combination of a virtual private network). In some embodiments, the data exchanged through the network is represented by technologies and / or formats including Hyper Text Mark-up Language (HTML), Extensible Markup Language (XML), etc. In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPSec) can also be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above data communication technologies.
[0038] In some embodiments, the acquisition device 101 and the training device 103 may be located on different devices, and the acquisition device 101 may be a module on an electronic device with data acquisition capability (e.g., a depth camera). The training device 103 may be a module on an electronic device with data processing capability, such as a computer or a server.
[0039] In one example of the present disclosure, an acquisition device acquires the intrinsic parameters and camera pose of each depth camera; the acquisition device acquires the RGB image and the corresponding depth image of the target object taken by each depth camera at the same time. The training device initializes the point cloud of the target object according to the intrinsic parameters, camera pose, RGB image and depth image; the training device trains the Gaussian splash model according to the point cloud, RGB image and depth image; in training the Gaussian splash model, the training device generates the RGB images and corresponding depth images of the target object under multiple perspectives according to the generation model and the monocular depth estimation algorithm, and adds them to the training set until the Gaussian splash model reconstruction is completed.
[0040] Those skilled in the art will know that Figure 1 The number of acquisition devices, networks and training devices is only illustrative, and any number of terminal devices, networks and servers may be provided according to actual needs. This disclosure embodiment does not limit this.
[0041] Figure 2 A flowchart of a Gaussian splash reconstruction method based on a depth camera in an embodiment of the present disclosure is shown. Figure 2 As shown, the Gaussian splash reconstruction method based on the depth camera provided in the embodiment of the present disclosure includes the following steps:
[0042] S202, obtaining the intrinsic parameters and camera pose of each depth camera.
[0043] It should be noted that the above-mentioned depth camera can be a computer vision device that can obtain object distance information. Compared with traditional RGB cameras, depth cameras can obtain three-dimensional information of objects with higher accuracy. The above-mentioned depth camera is also called a three-dimensional camera or a stereo camera. The above-mentioned internal parameters can be parameters that describe the characteristics of the camera itself and the imaging process, wherein the parameters are the results of camera calibration and are used to map points in three-dimensional space to a two-dimensional image plane. The above-mentioned camera pose can be the position of the camera in space and the orientation of the camera. For example, the camera pose includes the transformation of the camera from a certain initial position (or reference coordinate system) to the current position. This transformation usually consists of a translation transformation and a rotation transformation. The translation transformation describes the movement distance and direction of the camera in space, while the rotation transformation describes the change in the orientation of the camera.
[0044] S204, obtaining an RGB image and a corresponding depth image of the target object captured by each depth camera at the same time.
[0045] It should be noted that the above RGB image is a color image, which consists of three channels: red, green, and blue. Each channel represents the brightness or intensity of a color. By combining the different brightness values of these three channels, various colors can be generated. RGB images can be used to describe the appearance color characteristics of an object and provide color information about the target object within the visible spectrum. The above depth image is a special type of image that represents the distance information from each point in the scene to the camera lens. Each pixel value in the depth image represents the depth or distance of the pixel in three-dimensional space, usually expressed in millimeters or other distance units. Unlike RGB images, depth images do not contain color information, but only provide geometric information about the distance of the object from the camera. Depth images play an important role in computer vision tasks such as three-dimensional reconstruction, object recognition, and pose estimation, because they allow algorithms to understand and analyze target objects in three-dimensional space. In a depth camera, RGB images are acquired together with depth images to provide complete visual information of the target object.
[0046] S206, initializing the point cloud of the target object according to the intrinsic parameters, camera pose, RGB image and depth image.
[0047] It should be noted that the above point cloud can be a collection of points with coordinate information in three-dimensional space. The above initialization can include: using the camera intrinsic parameters and the pixel values in the depth image, each pixel coordinate can be converted into a coordinate in three-dimensional space; applying the camera posture to transform the converted three-dimensional coordinates to obtain the coordinates of the target object in the global coordinate system; using the converted three-dimensional coordinates including color information as points in the point cloud to generate a point cloud.
[0048] S208, training a Gaussian splash model according to the point cloud, the RGB image, and the depth image.
[0049] It should be noted that the above training can be done through machine learning or optimization algorithms, so that the model can learn and understand the characteristics of the input data, so as to make accurate predictions or representations of new input data. The above Gaussian splatting is a rasterization technology for 3D Gaussian distribution description for real-time radiation field rendering, which can learn realistic scenes from small image samples and render these scenes in real time.
[0050] In a specific example, the point cloud is used as the initial state of the Gaussian splash model, and the RGB image and the depth image are used as the training set of the Gaussian splash model to train the Gaussian splash model.
[0051] S210, in training the Gaussian splash model, RGB images and corresponding depth images of the target object under multiple perspectives are generated according to the generation model and the monocular depth estimation algorithm, and added to the training set until the Gaussian splash model reconstruction is completed.
[0052] It should be noted that the above-mentioned generative model can be a machine learning model. During the training process, the generative model can generate RGB images of target objects from multiple perspectives. These RGB images can be used as part of the training set to help the Gaussian splatter model learn the ability to render target objects from different perspectives. The above-mentioned monocular depth estimation algorithm can be a technology that estimates the depth information of objects in the scene through images obtained by a single camera. It can perform depth inference for each pixel based on the visual information of the image to generate a depth map corresponding to the above-mentioned RGB image without the need for additional hardware support.
[0053] This application combines the generative model with the monocular depth estimation algorithm under sparse perspective, and continuously generates new perspective information that is increasingly farther away from the existing perspective during the training process of the Gaussian splash model, thereby improving the reconstruction effect of the target object under sparse perspective.
[0054] In one embodiment of the present disclosure, Figure 3 As shown, the Gaussian splash reconstruction method based on the depth camera provided in the embodiment of the present disclosure can obtain the intrinsic parameters and camera pose of each depth camera through the following steps, which can improve the measurement accuracy:
[0055] S302, obtaining the intrinsic parameters and camera pose of each depth camera according to a standard chessboard image calibration method.
[0056] In one embodiment of the present disclosure, Figure 4 As shown, the Gaussian splash reconstruction method based on the depth camera provided in the embodiment of the present disclosure can initialize the point cloud through the following steps, which can improve the accuracy of three-dimensional reconstruction:
[0057] S402, mapping each pixel in the depth image to a three-dimensional space according to the intrinsic parameters and the camera pose to form a point cloud, wherein the RGB image is used to provide color information of the pixel.
[0058] In one embodiment of the present disclosure, Figure 5 As shown, the Gaussian splash reconstruction method based on the depth camera provided in the embodiment of the present disclosure can determine the truncated signed distance function through the following steps, which can improve the effect of establishing the Gaussian splash model:
[0059] S502, creating a truncated signed distance function;
[0060] S504, traversing the point cloud, for each point, finding a corresponding volume element in the truncated signed distance function volume;
[0061] S506, updating the signed distance value of the volume element according to the distance from the point to the center of the volume element, and determining the established truncated signed distance function, wherein the truncated signed distance function is used as the initial state of the Gaussian splash model.
[0062] In one embodiment of the present disclosure, Figure 6 As shown, the Gaussian splash reconstruction method based on the depth camera provided in the embodiment of the present disclosure can train the Gaussian splash model through the following steps, which can improve the reconstruction effect of the target object under sparse viewing angle:
[0063] S602, a depth camera is deflected by a preset angle (e.g., 1°, 3°, 5°) to determine a pose matrix of a new viewing angle;
[0064] S604, rendering the currently trained Gaussian splash model from the perspective of the pose matrix to determine a rendered RGB image;
[0065] S606, using the rendered RGB image as input of the generation model, and outputting an RGB image of the target object;
[0066] S608: Use the RGB image of the target object as input to a monocular depth estimation algorithm, and output a corresponding depth image.
[0067] In a specific example, first, a small number of n fixed depth cameras are placed around the target, and the intrinsic camera pose W of the depth camera is obtained by using the chessboard pattern calibration method. ex . In the second step, use the depth camera in the first step to take a set of RGB images and corresponding depth maps of the target at the same time. In the third step, use the camera intrinsic parameters, camera pose, depth map and RGB image obtained in the first and second steps to initialize the truncated signed distance function and point cloud of the target. In the fourth step, use the point cloud in the third step as initialization and the RGB image and depth map in the second step as the training set to start the Gaussian splash model training. In the fifth step, during the training of the Gaussian splash model, after a certain number of training iterations (which can be set according to actual needs), use the generative model (such as the Stable Diffussion stable diffusion generative model) and the monocular depth estimation algorithm (such as the DINOv2 second-generation self-supervised visual transformer model) to generate RGB images and depth information (depth maps) under n new perspectives, and add them to the training set. Until the reconstruction is completed.
[0068] It should be noted that in the fifth step, the specific details of generating new perspective information include the following steps:
[0069] For the mth camera in the first step, when it is generated for the i-th time, the pose matrix of its new perspective is as follows (1):
[0070]
[0071] Where R is a rotation matrix with a smaller degree of deflection (e.g., 1°, 3°, 5°);
[0072] For the Gaussian splash model currently trained, i m Rendering is performed from the perspective of to obtain a rendered RGB image;
[0073] The above rendered RGB image is used as the input of the image generation model to obtain a generated RGB image;
[0074] The RGB image generated above is used as the input of the monocular depth estimation algorithm to obtain the depth image corresponding to the RGB image generated under the new perspective;
[0075] The obtained n generated RGB images and the corresponding n corresponding depth images are added to the training set to continue training.
[0076] Based on the same inventive concept, the present disclosure also provides a Gaussian splash reconstruction device based on a depth camera, as described in the following embodiments. Since the principle of solving the problem in the device embodiment is similar to that in the above method embodiment, the implementation of the device embodiment can refer to the implementation of the above method embodiment, and the repeated parts will not be repeated.
[0077] Figure 7 A schematic diagram of a Gaussian splash reconstruction device based on a depth camera in an embodiment of the present disclosure is shown. Figure 7 As shown, the device includes: an information acquisition module 71, an image acquisition module 72, an initialization module 73, a training module 74 and a reconstruction module 75.
[0078] The information acquisition module 71 is used to obtain the internal parameters and camera pose of each depth camera;
[0079] An image acquisition module 72 is used to acquire an RGB image and a corresponding depth image of a target object captured by each depth camera at the same time;
[0080] An initialization module 73 is used to initialize the point cloud of the target object according to the intrinsic parameters, the camera pose, the RGB image and the depth image;
[0081] A training module 74, for training a Gaussian splash model based on the point cloud, the RGB image and the depth image;
[0082] The reconstruction module 75 is used to generate RGB images and corresponding depth images of the target object under multiple perspectives according to the generation model and the monocular depth estimation algorithm in the training of the Gaussian splash model, and add them to the training set until the Gaussian splash model reconstruction is completed.
[0083] In one example of the present disclosure, the information acquisition module 71 is further used to obtain the intrinsic parameters and camera pose of each depth camera according to a standard checkerboard image calibration method.
[0084] In one example of the present disclosure, the initialization module 73 is also used to map each pixel in the depth image to a three-dimensional space according to the intrinsic parameters and the camera pose to form a point cloud, wherein the RGB image is used to provide color information of the pixel.
[0085] In one example of the present disclosure, the above-mentioned Gaussian splash reconstruction device based on the depth camera also includes a truncated signed distance function module, which is used to create a truncated signed distance function; traverse the point cloud, and find the corresponding volume element for each point in the truncated signed distance function volume; according to the distance from the point to the center of the volume element, update the signed distance value of the volume element, and determine the established truncated signed distance function, wherein the truncated signed distance function is used as the initial state of the Gaussian splash model.
[0086] In one example of the present disclosure, the training module 74 is also used to train the Gaussian splash model by using the point cloud as the initial state of the Gaussian splash model and using the RGB image and the depth image as the training set of the Gaussian splash model.
[0087] In one example of the present disclosure, the reconstruction module 75 is also used to deflect a depth camera by a preset angle to determine the pose matrix of the new perspective; render the Gaussian splash model obtained through current training from the perspective of the pose matrix to determine a rendered RGB image; use the rendered RGB image as input to the generation model to output an RGB image of the target object; use the RGB image of the target object as input to a monocular depth estimation algorithm to output a corresponding depth image.
[0088] It should be noted that the above information acquisition module 71, image acquisition module 72, initialization module 73, training module 74 and reconstruction module 75 correspond to S202 to S210 in the method embodiment, and the examples and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the contents disclosed in the above method embodiment. It should be noted that the above modules as part of the device can be executed in a computer system such as a set of computer executable instructions.
[0089] Those skilled in the art will appreciate that various aspects of the present disclosure may be implemented as systems, methods or program products. Therefore, various aspects of the present disclosure may be specifically implemented in the following forms, namely: complete hardware implementation, complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software, which may be collectively referred to herein as "circuits", "modules" or "systems".
[0090] Refer to the following Figure 8 The electronic device 800 according to this embodiment of the present disclosure is described. Figure 8 The electronic device 800 shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0091] like Figure 8 As shown, the electronic device 800 is in the form of a general computing device. The components of the electronic device 800 may include but are not limited to: at least one processing unit 810, at least one storage unit 820, and a bus 830 connecting different system components (including the storage unit 820 and the processing unit 810).
[0092] The storage unit stores program codes, which can be executed by the processing unit 810, so that the processing unit 810 executes the steps described in the above “exemplary method” section of this specification according to various exemplary embodiments of the present disclosure.
[0093] For example, the processing unit 810 can execute the following steps of the above-mentioned method embodiment: obtain the intrinsic parameters and camera pose of each depth camera; obtain the RGB image and the corresponding depth image of the target object taken by each depth camera at the same time; initialize the point cloud of the target object according to the intrinsic parameters, camera pose, RGB image and depth image; train the Gaussian splash model according to the point cloud, RGB image and depth image; in the training of the Gaussian splash model, generate RGB images and corresponding depth images of the target object under multiple perspectives according to the generation model and the monocular depth estimation algorithm, and add them to the training set until the Gaussian splash model reconstruction is completed.
[0094] For example, the processing unit 810 may execute the following steps of the above method embodiment: obtaining the intrinsic parameters and camera pose of each depth camera according to a standard checkerboard image calibration method.
[0095] For example, the processing unit 810 can perform the following steps of the above method embodiment: according to the intrinsic parameters and the camera pose, map each pixel in the depth image to the three-dimensional space to form a point cloud, wherein the RGB image is used to provide color information of the pixel.
[0096] For example, the processing unit 810 can execute the following steps of the above method embodiment: create a truncated signed distance function; traverse the point cloud, and find the corresponding volume element for each point in the truncated signed distance function volume; update the signed distance value of the volume element according to the distance from the point to the center of the volume element, and determine the established truncated signed distance function, wherein the truncated signed distance function is used as the initial state of the Gaussian splash model.
[0097] For example, the processing unit 810 may execute the following steps of the above method embodiment: using the point cloud as the initial state of the Gaussian splash model, using the RGB image and the depth image as the training set of the Gaussian splash model, and training the Gaussian splash model.
[0098] For example, the processing unit 810 can execute the following steps of the above-mentioned method embodiment: a depth camera is deflected by a preset angle to determine the pose matrix of the new perspective; the Gaussian splash model obtained by the current training is rendered from the perspective of the pose matrix to determine the rendered RGB image; the rendered RGB image is used as the input of the generation model to output the RGB image of the target object; the RGB image of the target object is used as the input of the monocular depth estimation algorithm to output the corresponding depth image.
[0099] The storage unit 820 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 8201 and / or a cache memory unit 8202 , and may further include a read-only memory unit (ROM) 8203 .
[0100] The storage unit 820 may also include a program / utility 8204 having a set (at least one) of program modules 8205, such program modules 8205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0101] Bus 830 may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.
[0102] The electronic device 800 may also communicate with one or more external devices 840 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 800, and / or communicate with any device that enables the electronic device 800 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed via an input / output (I / O) interface 850. Furthermore, the electronic device 800 may also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via a network adapter 860. As shown, the network adapter 860 communicates with other modules of the electronic device 800 via a bus 830. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 800, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0103] Through the description of the above implementation, it is easy for those skilled in the art to understand that the example implementation described here can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to the implementation of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the implementation of the present disclosure.
[0104] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer program product, which includes: a computer program, which implements the above-mentioned Gaussian splash reconstruction method based on the depth camera when executed by a processor.
[0105] In an exemplary embodiment of the present disclosure, a computer readable storage medium is also provided. Fig. 9 As shown, the computer-readable storage medium 900 may be a readable signal medium or a readable storage medium. A program product capable of implementing the above-mentioned method of the present disclosure is stored thereon. In some possible implementations, various aspects of the present disclosure may also be implemented in the form of a program product, which includes a program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to execute the steps according to various exemplary implementations of the present disclosure described in the above-mentioned "Exemplary Method" section of this specification.
[0106] For example, when the program product in the embodiment of the present disclosure is executed by a processor, the following steps are implemented: obtaining the intrinsic parameters and camera pose of each depth camera; obtaining the RGB image and the corresponding depth image of the target object taken by each depth camera at the same time; initializing the point cloud of the target object according to the intrinsic parameters, camera pose, RGB image and depth image; training the Gaussian splash model according to the point cloud, RGB image and depth image; in training the Gaussian splash model, generating RGB images and corresponding depth images of the target object under multiple perspectives according to the generation model and the monocular depth estimation algorithm, and adding them to the training set until the Gaussian splash model reconstruction is completed.
[0107] For example, when the program product in the embodiment of the present disclosure is executed by a processor, the following steps are implemented: according to the standard checkerboard image calibration method, the intrinsic parameters and camera pose of each depth camera are obtained.
[0108] For example, when the program product in the embodiment of the present disclosure is executed by a processor, the method implements the following steps: according to the intrinsic parameters and the camera pose, each pixel in the depth image is mapped to the three-dimensional space to form a point cloud, wherein the RGB image is used to provide the color information of the pixel.
[0109] For example, when the program product in the embodiment of the present disclosure is executed by a processor, the method implements the following steps: creating a truncated signed distance function; traversing the point cloud, and finding the corresponding volume element for each point in the truncated signed distance function volume; updating the signed distance value of the volume element based on the distance from the point to the center of the volume element, and determining the established truncated signed distance function, wherein the truncated signed distance function is used as the initial state of the Gaussian splash model.
[0110] For example, when the program product in the embodiment of the present disclosure is executed by a processor, the following steps are implemented: using the point cloud as the initial state of the Gaussian splash model, using the RGB image and the depth image as the training set of the Gaussian splash model, and training the Gaussian splash model.
[0111] For example, when the program product in the embodiment of the present disclosure is executed by a processor, the following steps are implemented: a depth camera is deflected by a preset angle to determine the pose matrix of the new perspective; the Gaussian splash model obtained through current training is rendered from the perspective of the pose matrix to determine a rendered RGB image; the rendered RGB image is used as input to the generation model to output an RGB image of the target object; the RGB image of the target object is used as input to a monocular depth estimation algorithm to output a corresponding depth image.
[0112] More specific examples of computer-readable storage media in the present disclosure may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0113] In the present disclosure, a computer readable storage medium may include a data signal propagated in baseband or as part of a carrier wave, wherein a readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A readable signal medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0114] Alternatively, the program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the foregoing.
[0115] In a specific implementation, the program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., using an Internet service provider to connect through the Internet).
[0116] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into multiple modules or units to be embodied.
[0117] In addition, although the steps of the method in the present disclosure are described in a specific order in the drawings, this does not require or imply that the steps must be performed in this specific order, or that all the steps shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps, etc.
[0118] Through the description of the above implementation, it is easy for those skilled in the art to understand that the example implementation described here can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to the implementation of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the implementation of the present disclosure.
[0119] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The description and examples are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the appended claims.
Claims
1. A Gaussian splash reconstruction method based on a depth camera, characterized in that: include: Get the intrinsic parameters and camera pose of each depth camera; Obtain an RGB image and a corresponding depth image of the target object taken by each depth camera at the same time; Initialize a point cloud of the target object according to the intrinsic parameters, the camera pose, the RGB image, and the depth image; Training a Gaussian splash model based on the point cloud, the RGB image, and the depth image; In training the Gaussian splash model, RGB images and corresponding depth images of the target object under multiple perspectives are generated according to the generation model and the monocular depth estimation algorithm, and added to the training set until the Gaussian splash model reconstruction is completed.
2. The Gaussian splash reconstruction method based on a depth camera according to claim 1, characterized in that: The step of obtaining the internal parameters and camera pose of each depth camera includes: According to the standard chessboard image calibration method, the intrinsic parameters and camera pose of each depth camera are obtained.
3. The Gaussian splash reconstruction method based on a depth camera according to claim 1, characterized in that: Initializing the point cloud of the target object according to the intrinsic parameter, the camera pose, the RGB image and the depth image includes: According to the intrinsic parameters and the camera pose, each pixel in the depth image is mapped to a three-dimensional space to form the point cloud, wherein the RGB image is used to provide color information of the pixel.
4. The Gaussian splash reconstruction method based on a depth camera according to claim 3, characterized in that: The method further comprises: Create a truncated signed distance function; Traversing the point cloud, for each point, finding a corresponding volume element in the truncated signed distance function volume; According to the distance from the point to the center of the volume element, the signed distance value of the volume element is updated, and the established truncated signed distance function is determined, wherein the truncated signed distance function is used as the initial state of the Gaussian splash model.
5. The Gaussian splash reconstruction method based on a depth camera according to any one of claims 1 to 4, characterized in that: The training of the Gaussian splash model according to the point cloud, the RGB image and the depth image comprises: The point cloud is used as the initial state of the Gaussian splash model, and the RGB image and the depth image are used as training sets of the Gaussian splash model to train the Gaussian splash model.
6. The Gaussian splash reconstruction method based on a depth camera according to claim 1, characterized in that: The step of generating the RGB image and the corresponding depth image of the target object under multiple viewing angles according to the generation model and the monocular depth estimation algorithm includes: A depth camera is deflected by a preset angle to determine the pose matrix of the new viewing angle; Rendering the currently trained Gaussian splash model from the perspective of the pose matrix to determine a rendered RGB image; Using the rendered RGB image as input of the generation model, and outputting an RGB image of the target object; The RGB image of the target object is used as the input of the monocular depth estimation algorithm, and the corresponding depth image is output.
7. A Gaussian splash reconstruction device based on a depth camera, characterized in that: include: Information acquisition module, used to obtain the intrinsic parameters and camera pose of each depth camera; An image acquisition module is used to acquire an RGB image and a corresponding depth image of the target object taken by each depth camera at the same time; An initialization module, used to initialize the point cloud of the target object according to the intrinsic parameters, the camera pose, the RGB image and the depth image; A training module, configured to train a Gaussian splash model according to the point cloud, the RGB image and the depth image; A reconstruction module is used to generate RGB images and corresponding depth images of the target object under multiple perspectives according to the generation model and the monocular depth estimation algorithm in training the Gaussian splash model, and add them to the training set until the Gaussian splash model reconstruction is completed.
8. An electronic device, characterized in that: include: processor; as well as A memory, configured to store executable instructions of the processor; The processor is configured to execute the Gaussian splash reconstruction method based on a depth camera as described in any one of claims 1 to 6 by executing the executable instructions.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the Gaussian splash reconstruction method based on a depth camera according to any one of claims 1 to 6 is implemented.
10. A computer program product comprising: A computer program or instruction, characterized in that when the computer program or instruction is executed by a processor, it implements the Gaussian splash reconstruction method based on a depth camera as described in any one of claims 1 to 6.
Citation Information
Cited By
Three-dimensional Gaussian visual positioning method for sparse visual angle
CN120339380A
A 3D Gaussian visual localization method for sparse view
CN120339380B
Reconstruction method and reconstruction system directly from color point cloud to three-dimensional Gaussian sputtering and storage medium
CN122289572A