Program, information processing device, and adaptive density control method
By projecting and segmenting 3D Gaussian in virtual space, and selecting based on the maximum projection size and opacity, the problem of insufficient detail reproduction in 3DGS is solved, achieving efficient and high-quality virtual image generation.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-27
- Publication Date
- 2026-04-08
AI Technical Summary
In existing 3D Gaussian Splatting (3DGS) technology, adaptive density control cannot effectively reproduce the details in the original captured image, resulting in insufficient reproduction of object details.
The density control of 3D Gaussian is optimized by projecting 3D Gaussian in virtual space and selecting and segmenting it according to its maximum projection size and opacity, while keeping the total volume constant.
It improves the spatial resolution and shape representation of virtual images, enables high-speed and highly defined virtual image generation, and reduces redundant processing time.
Smart Images

Figure 2026060022000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a technique for newly generating an image of an object as seen from an arbitrary viewpoint using images of the object taken from a plurality of viewpoints.
Background Art
[0002] Photogrammetry is a technique for estimating the three-dimensional shape of an object from images (also referred to as captured images) of the object taken from a plurality of viewpoints. For example, Patent Document 1 discloses an apparatus and method for generating a three-dimensional (3D) model of a subject from images of a plurality of cameras as a photogrammetry technique.
[0003] Furthermore, techniques for generating an image (also referred to as a virtual image) of an object (also referred to as a virtual object) in a virtual space as seen from an arbitrary viewpoint (also referred to as a virtual viewpoint) based on captured images of the object taken from a plurality of viewpoints have been studied. As such techniques, NeRF (Neural Radiance Fields) and NeuS (Neural Implicit Surface) are known, but there are problems such as a large load on machine learning and long time consumption.
[0004] On the other hand, 3DGS (3D Gaussian Splatting) has attracted attention as a technique for realizing high-speed and high-quality rendering. 3DGS is a technique for reproducing a 3D scene by arranging ellipsoids called 3D Gaussians in a 3D (three-dimensional) virtual space and optimizing the parameters of those 3D Gaussians to match the original captured images.
[0005] Unlike the implicit representation used by NeRF and other methods, 3DGS explicitly represents virtual objects in virtual space using a 3D Gaussian ellipsoid, enabling faster rendering. Furthermore, because 3DGS uses differentiable rasterization, processing can be accelerated compared to NeRF and other methods that use volume rendering, for example, by using a GPU (Graphics Processing Unit). For example, Non-Patent Documents 1-3 disclose various methods related to 3DGS. [Prior art documents] [Patent Documents]
[0006] [Patent Document 1] Japanese Patent Publication No. 2021-71749 [Non-patent literature]
[0007] [Non-Patent Document 1] https: / / arxiv.org / abs / 2308.04079 [Non-Patent Document 2] https: / / arxiv.org / abs / 2404.06109 [Non-Patent Document 3] https: / / ubc-vision.github.io / 3dgs-mcmc / [Overview of the project] [Problems that the invention aims to solve]
[0008] To recreate a 3D scene, a 3D Gaussian of a size appropriate to the details of the 3D scene is required. Obtaining such a 3D Gaussian has limitations if it is only optimized by comparing the size of the initial 3D Gaussian with the original captured image. Therefore, in addition to this optimization process, 3DGS duplicates or splits the 3D Gaussian placed during initialization according to predetermined conditions. The control that adjusts the density by splitting and duplicating the 3D Gaussian is generally called "adaptive density control."
[0009] However, in conventional techniques such as those described in Non-Patent Document 3, adaptive density control of 3D Gaussians involves replicating 3D Gaussians smaller than a certain size in a virtual space and splitting only those 3D Gaussians larger than a certain size. Therefore, conventional techniques may not be able to reproduce the details of objects that were represented in the original captured image.
[0010] One of the objectives of the present invention is to realize adaptive density control of 3D Gaussian, which is more suitable than the prior art for reproducing the details of objects represented in captured images. [Means for solving the problem]
[0011] In one embodiment, the present invention provides a program for a computer to function as: an acquisition means for acquiring a plurality of images of an object and camera information indicating the position, orientation, and shooting range of the camera that captured each of the plurality of images; a calculation means for calculating a projected Gaussian by projecting each of a plurality of 3D Gaussians arranged in a virtual space onto a screen corresponding to each of the cameras in the virtual space; an extraction means for extracting from the plurality of 3D Gaussians the one whose largest projected Gaussian projected onto each of the screens is greater than or equal to a threshold; and a splitting means for splitting a 3D Gaussian selected from the extracted 3D Gaussians with a probability according to opacity into a plurality such that the total volume is maintained. [Brief explanation of the drawing]
[0012] [Figure 1] A diagram showing an example of the configuration of an information processing device 1 according to an embodiment of the present invention. [Figure 2] A diagram showing an example of image DB121. [Figure 3] A diagram showing an example of camera information DB122. [Figure 4] A diagram showing the captured image and camera position. [Figure 5] A diagram showing the correspondence between feature points in each of the captured images. [Figure 6] A diagram showing three-dimensional feature points in a virtual space. [Figure 7] A diagram showing an example of Gaussian DB123. [Figure 8] A diagram for explaining a 3D Gaussian. [Figure 9] A diagram showing an example of the functional configuration of information processing apparatus 1. [Figure 10] A diagram showing an example of the configuration of adaptive density control means 115. [Figure 11] A flowchart showing an example of the operation flow of information processing apparatus 1. [Figure 12] / / 这里原文可能有误,推测为“ [Figure 12] ” A flowchart showing an example of the operation flow of splitting processing. [Figure 13] A flowchart showing an example of the operation flow of splitting processing in a modified example. [Figure 14] A diagram showing an example of the configuration of adaptive density control means 115 in a modified example. [Figure 15] A diagram showing an example of the functional configuration of information processing apparatus 1 in a modified example.
Embodiments of the Invention
[0013] <Embodiment> <Configuration of Information Processing Apparatus> / / 这里原文可能有误,推测为“ ” FIG. 1 is a diagram showing an example of the configuration of information processing apparatus 1 according to an embodiment of the present invention. Information processing apparatus 1 is an apparatus that generates a virtual image of a virtual object corresponding to an object as seen from a virtual viewpoint in a virtual space based on each of captured images of the object taken from a plurality of viewpoints, and is, for example, a computer. Information processing apparatus 1 includes a processor 11, a memory 12, and an interface 13. These components are communicably connected to each other, for example, via a bus.
[0014] The processor 11 controls each part of the information processing apparatus 1 by reading and executing a computer program (hereinafter simply referred to as a program) stored in the memory 12. The processor 11 is, for example, a CPU (Central Processing Unit) or a GPU. <000
[0015] Interface 13 is a communication circuit that connects the information processing device 1 to a communication line, external device, etc., via wired or wireless means, enabling communication.
[0016] Memory 12 is a storage means for storing the operating system, various programs, data, etc., that are loaded into the processor 11. Memory 12 has RAM (Random Access Memory) or ROM (Read Only Memory). Memory 12 may also have a solid-state drive, a hard disk drive, etc. Memory 12 also stores an image DB 121, a camera information DB 122, and a Gaussian DB 123.
[0017] The information processing device 1 may have an operation unit and a display unit. The operation unit is configured to receive operations and send signals corresponding to the content of those operations to the processor 11. This operation unit may include, for example, operation buttons, a keyboard, a touch panel, a mouse, or other controls for giving various instructions.
[0018] This display unit is configured to display images under the control of the processor 11. This display unit may have a display screen such as a liquid crystal display. Furthermore, the transparent touch panel of the aforementioned operation unit may be superimposed on this display screen. The information processing device 1 may be operated from an external device via the interface 13, or information may be presented to an external device.
[0019] <Image Database Structure> Figure 2 shows an example of the image DB 121. The image DB 121 is a database that stores images of an object taken from multiple viewpoints. The image DB 121 shown in Figure 2 stores image data associated with an image ID. The image ID is identification information that uniquely identifies each captured image. The image data is data that indicates the content of the captured image. These captured images stored in the image DB 121 are images of the same object, etc., taken from different viewpoints. The image DB 121 may also be grouped by the object (subject) that was photographed, or by scene. In other words, the image DB 121 only needs to store a series of captured images (groups of captured images) for each subject or scene so that the processor 11 can acquire them.
[0020] <Configuration of camera information database> Figure 3 shows an example of the camera information DB122. The camera information DB122 is a database that stores information such as the location of the camera that took each captured image (referred to as camera information), associating it with the image ID of that captured image. The camera information DB122 shown in Figure 3 stores the following items as associations: image ID, location, orientation, and shooting range.
[0021] In the camera information DB122, the "image ID" is identification information that uniquely identifies each captured image, and is the same information as in the image DB121 described above. In the camera information DB122, the "position" is information that indicates the position of the camera in the actual space (real space Sp) at the time the captured image identified by the corresponding image ID was taken, and is shown, for example, in three-dimensional coordinates in a Cartesian coordinate system.
[0022] For example, according to this camera information DB122, the captured image identified by image ID "I1" was taken from a camera located at the coordinates (x1, y1, z1).
[0023] In the camera information DB122, "orientation" indicates the direction in which the lens of the camera that captured the image is facing in real space Sp, and is represented, for example, by a three-dimensional vector. In the camera information DB122, "shooting range" is information such as parameters that determine the range captured by the camera that captured the image, such as focal length, angle of view, and the center position of the camera's image sensor.
[0024] The information processing device 1 shown in Figure 1 acquires images of an object taken from multiple viewpoints from a camera and stores them in the image DB 121. However, at that point, it does not store the position, orientation, shooting range, etc., of the camera that took those images. Therefore, the information processing device 1 calculates camera information, such as the position of the camera that took the images, from the captured images using, for example, SfM (Structure from Motion), and stores it in the camera information DB 122.
[0025] Figure 4 shows the captured image and camera position. In the following diagrams, the real space in which each component is arranged is represented as the xyz right-handed coordinate system, and the virtual space is represented as the XYZ right-handed coordinate system. In space, the direction along the x-axis is called the x-axis direction. Furthermore, within the x-axis direction, the direction in which the x component increases is called the +x direction, and the direction in which the x component decreases is called the -x direction. The y and z components, and the X, Y, and Z components are also defined according to the above definition as follows: y-axis direction, +y direction, -y direction, z-axis direction, +z direction, -z direction, X-axis direction, +X direction, -X direction, Y-axis direction, +Y direction, -Y direction, Z-axis direction, +Z direction, and -Z direction.
[0026] Figure 4 illustrates how an object J in real space Sp is photographed by cameras from viewpoints (x1, y1, z1) and (x2, y2, z2), resulting in captured images I1 and I2.
[0027] Figure 5 shows the correspondence between feature points in each of the captured images. The information processing device 1 extracts multiple feature points from each of multiple captured images taken from different viewpoints. The information processing device 1 then infers the correspondence between these feature points based on the arrangement of these feature points in the captured images, the surrounding shape, the distribution of grayscale values, the gradient, etc. For example, the information processing device 1 estimates that the feature points Pa1, Pb1, and Pc1 in captured image I1 correspond to the feature points Pa2, Pb2, and Pc2 in captured image I2, respectively.
[0028] Figure 6 shows three-dimensional feature points in a virtual space. As described above, the information processing device 1 identifies feature points in each captured image and estimates the correspondence between these feature points. This correspondence is the relationship that a single point of an object in real space has been photographed from multiple viewpoints. By estimating the correspondence between feature points, the information processing device 1 identifies a single point of the object that is the basis for those feature points. Then, the information processing device 1 calculates a point in the virtual space Sv corresponding to this single point as a "three-dimensional feature point".
[0029] For example, in the example shown in Figure 6, the three-dimensional feature points Pa, Pb, and Pc correspond to the feature points Pa1, Pb1, and Pc1, respectively, in the captured image I1 shown in Figure 5. Furthermore, the three-dimensional feature points Pa, Pb, and Pc correspond to the feature points Pa2, Pb2, and Pc2, respectively, in the captured image I2 shown in Figure 5. The information processing device 1 then calculates the camera position along the line connecting the calculated three-dimensional feature points and the feature points in the captured image, along with the camera orientation and shooting range.
[0030] <Gaussian DB Configuration> Figure 7 shows an example of Gaussian DB123. Gaussian DB123 is a database that stores 3D Gaussians. A 3D Gaussian is an ellipsoid used to reproduce virtual objects in a virtual space, and has various parameters such as position, orientation, shape, color, and opacity. The Gaussian DB123 shown in Figure 7 stores the Gaussian ID, center coordinates, rotation, scale, color, and opacity in an associated manner.
[0031] During the initialization phase, the 3D Gaussian is placed, for example, at the three-dimensional feature points mentioned above. As a result, the Gaussian DB123 stores the 3D Gaussians whose center coordinates coincide with the three-dimensional feature points. For example, as shown in Figure 7, the 3D Gaussian identified by Gaussian ID "E1" is placed at center coordinate "Pa". The center coordinate "Pa" is the three-dimensional feature point Pa corresponding to feature point Pa1 in captured image I1 and feature point Pa2 in captured image I2. The virtual object Jv shown in Figure 6 is constructed by the 3D Gaussians placed at the three-dimensional feature points in the virtual space Sv in this way.
[0032] Figure 8 is a diagram illustrating 3D Gaussian. Figure 8(a) shows an example of 3D Gaussian E. 3D Gaussian E is an ellipsoid in the virtual space Sv that has size (i.e., scale) and central coordinates, and rotation angles with respect to the X, Y, and Z axes, respectively. This 3D Gaussian also possesses color information that depends on the viewing direction, and opacity. These can be represented, for example, by a three-dimensional covariance matrix, harmonic functions, constants, etc.
[0033] Figure 8(b) shows how three 3D Gaussians E1, E2, and E3 are projected onto screen S. The information processing device 1 projects these multiple 3D Gaussians, which are placed in the virtual space Sv, onto screen S. Screen S is the surface in which virtual objects are reflected when any range is photographed from any viewpoint in the virtual space Sv. The 3D Gaussians E1, E2, and E3 are projected onto screen S by the information processing device 1, becoming projected Gaussians Pj1, Pj2, and Pj3, which are two-dimensional ellipses. These projected Gaussians Pj1, Pj2, and Pj3 may also be displayed overlapping in the virtual space Sv, depending on the order and arrangement of the 3D Gaussians E1, E2, and E3 as seen from screen S.
[0034] <Functional Configuration of Information Processing Devices> Figure 9 shows an example of the functional configuration of the information processing device 1. The processor 11 of the information processing device 1 functions as the image acquisition means 111, bundle adjustment means 112, initialization means 113, learning means 114, and adaptive density control means 115 shown in Figure 9 by reading and executing a program stored in the memory 12.
[0035] The image acquisition means 111 acquires captured images from the image DB 121 in the memory 12. The image acquisition means 111 then supplies the acquired captured images to the bundle adjustment means 112 and the learning means 114.
[0036] The bundle adjustment means 112 calculates camera information from the supplied captured images using the aforementioned SfM or the like.
[0037] The bundle adjustment means 112 shown in Figure 9 includes a feature point extraction means 1121, a feature point matching means 1122, a verification means 1123, and an estimation means 1124.
[0038] The feature point extraction means 1121 extracts the outlines of objects depicted in the captured image from the position gradient of the grayscale values of the captured image, and extracts feature points from those outlines.
[0039] The feature point matching means 1122 associates (also called matching) the feature points extracted from each of the multiple captured images and calculates the corresponding three-dimensional feature points in the virtual space.
[0040] The feature point extraction means 1121 and the feature point matching means 1122 are performed using feature point matching techniques such as SIFT (Scale-Invariant Feature Transform), SURF (Speeded-Up Robust Features), and ORB (Oriented FAST and Rotated BRIEF).
[0041] The verification means 1123 verifies the matched feature points using a method such as RANSAC (RANDOM SAmple Consensus). Based on the verification results, the verification means 1123 excludes feature points that were incorrectly matched due to the influence of outliers or the like.
[0042] The estimation means 1124 estimates camera information from the positional relationship between the calculated three-dimensional feature points in the virtual space and the corresponding feature points in the captured image, and stores this information in the camera information DB 122 of the memory 12.
[0043] As described above, the image acquisition means 111 acquires multiple captured images of an object. The bundle adjustment means 112 then acquires camera information indicating the position, orientation, and shooting range of the camera that captured each of the multiple captured images acquired by the image acquisition means 111. Therefore, the image acquisition means 111 and the bundle adjustment means 112 are examples of acquisition means that acquire multiple images of an object and camera information indicating the position, orientation, and shooting range of the camera that captured each of those multiple images, respectively.
[0044] The initialization means 113 is a means for placing an initial state of 3D Gaussian in a virtual space based on the captured image and camera information. For example, the initialization means 113 places the 3D Gaussian so that its center coordinates coincide with the three-dimensional feature points in the virtual space calculated by the bundle adjustment means 112. The initialization means 113 then determines the color and opacity of the 3D Gaussian using the gradation values of the feature points in the captured image that correspond to these three-dimensional feature points. The initialization means 113 also determines the rotation and scale of the 3D Gaussian in its initial state using the position gradient of the gradation values in the captured image, or pseudo-random numbers, etc. Once the initialization means 113 has determined the initial values of the 3D Gaussian, such as its center coordinates, rotation, scale, color, and opacity, it stores them in the Gaussian DB 123 of the memory 12.
[0045] The learning method 114 projects a 3D Gaussian onto a screen corresponding to the camera in a virtual space and performs machine learning by changing the parameters of the 3D Gaussian so that the resulting virtual image approaches the original captured image.
[0046] The learning means 114 shown in Figure 9 includes a calculation means 1141 and an error minimization means 1142. The calculation means 1141 reads all 3D Gaussians placed in the virtual space from the Gaussian DB 123 in memory 12, projects them onto screens in the virtual space corresponding to each camera that captured the captured image, and calculates the projected Gaussians for each.
[0047] Therefore, this calculation means 1141 is an example of a calculation means that calculates a projected Gaussian obtained by projecting each of a plurality of 3D Gaussians arranged in a virtual space onto a screen corresponding to each of the cameras that captured images in that virtual space.
[0048] The error minimization means 1142 compares the virtual image generated by the superposition of the calculated projected Gaussians with the captured image acquired by the image acquisition means 111, and changes the parameters of the 3D Gaussian so that the error (also called rendering error) between them is reduced. The calculation means 1141 projects the 3D Gaussian whose parameters have been changed by the error minimization means 1142 onto the aforementioned screen and calculates the projected Gaussian. The error minimization means 1142 continues this interaction with the calculation means 1141 until a predetermined condition is met. In this case, the predetermined condition is, for example, that the number of times the parameters of the 3D Gaussian have been changed exceeds a threshold.
[0049] The adaptive density control means 115 adjusts the density of the 3D Gaussians by first training the learning means 114 by changing the parameters of the 3D Gaussians until it satisfies the predetermined conditions described above, and then selecting and splitting those 3D Gaussians that satisfy the conditions.
[0050] Figure 10 shows an example of the configuration of the adaptive density control means 115. The adaptive density control means 115 shown in Figure 10 includes an extraction means 1151, a selection means 1152, a splitting means 1153, and an arrangement means 1154.
[0051] The extraction means 1151 extracts only those 3D Gaussians whose parameters have been learned by the learning means 114, and whose maximum projected Gaussian size projected onto each of the aforementioned screens is equal to or greater than a threshold. In other words, this extraction means 1151 is an example of an extraction means that extracts from multiple 3D Gaussians only those whose maximum projected Gaussian size projected onto each of the screens is equal to or greater than a threshold.
[0052] The "threshold" used to compare the maximum size of the projected Gaussian, as used here, is a reference value for size based on the resolution of the screen on which the projected Gaussian is projected, such as 1 pixel. In other words, the extraction means 1151 extracts only those projected Gaussians projected onto all screens, the largest of which is 1 pixel or larger. As a result, at least 3D Gaussians whose projected Gaussians are all less than 1 pixel in size are excluded from the extraction target of the extraction means 1151.
[0053] The selection means 1152 selects 3D Gaussians from among those extracted by the extraction means 1151 with a probability corresponding to the parameters of those 3D Gaussians. This probability is, for example, a probability corresponding to the opacity of the 3D Gaussians. Here, the selection means 1152 selects several 3D Gaussians from among those extracted by the extraction means 1151 with a probability proportional to the opacity of those 3D Gaussians. Therefore, the higher the opacity of a 3D Gaussian, the higher the probability with which the selection means 1152 will select that 3D Gaussian.
[0054] Note that "probability proportional to opacity" is one form of "probability according to opacity" and can be replaced by other forms. For example, the selection means 1152 may select a 3D Gaussian with a probability proportional to the logarithm of the opacity of the 3D Gaussian. Alternatively, for example, the selection means 1152 may select a 3D Gaussian with a probability proportional to a monotonically increasing function with the opacity of the 3D Gaussian as the independent variable.
[0055] The splitting means 1153 splits the 3D Gaussian selected by the selection means 1152. In this case, the splitting means 1153 splits the 3D Gaussian selected by the selection means 1152 into multiple parts such that its total volume is maintained.
[0056] In other words, the selection means 1152 and the splitting means 1153 are examples of splitting means that split a 3D Gaussian selected from the extracted 3D Gaussians with a probability corresponding to the opacity, so as to maintain the total volume.
[0057] The arrangement means 1154 arranges the 3D Gaussian, which has been split into multiple parts by the splitting means 1153, based on the position of the original 3D Gaussian before splitting. For example, the arrangement means 1154 arranges the split 3D Gaussian so that each part becomes a portion of the original 3D Gaussian. This arrangement means 1154 is an example of an arrangement means that arranges each of the split 3D Gaussian based on the position of the original 3D Gaussian.
[0058] The 3D Gaussian is adjusted in density by the adaptive density control means 115 and stored in the Gaussian DB 123 of memory 12. The learning means 114 reads the density-adjusted 3D Gaussians from the Gaussian DB 123 and repeatedly optimizes the parameters of these 3D Gaussians based on the difference between the virtual image and the captured image until predetermined conditions are met. This calculates a 3D Gaussian capable of generating a virtual image that closely resembles the captured image.
[0059] Figure 11 is a flowchart showing an example of the operation flow of the information processing device 1. The processor 11 of the information processing device 1 performs adaptive density control as shown in Figure 11.
[0060] Before performing adaptive density control, the processor 11 retrieves a series of captured images of the 3D scene from the image DB 121 in memory 12. The processor 11 then estimates the camera information used when these images were captured, associates it with the captured images, and places an initial state of 3D Gaussian in the virtual space. Furthermore, the processor 11 projects these 3D Gaussians onto a screen corresponding to the camera in the virtual space to generate a virtual image, for example, until a predetermined condition is met, and optimizes (learns) the parameters of the 3D Gaussian so that it approaches the captured image.
[0061] When the learning process satisfies the conditions described above, the processor 11 starts the adaptive density control operation shown in Figure 11. First, the processor 11 focuses on one of the undetermined 3D Gaussians (step S101). Then, the processor 11 projects the 3D Gaussian of interest onto all screens and compares each of the resulting projected Gaussians with a threshold (step S102).
[0062] The processor 11 identifies the largest projected Gaussian among those described above and determines whether its size (maximum size) is greater than or equal to a threshold (step S103). If it determines that the maximum size of the projected Gaussian is greater than or equal to the threshold (step S103; YES), the processor 11 extracts that 3D Gaussian of interest as the target of the splitting process described later (step S104) and proceeds to step S105.
[0063] On the other hand, if it is determined that the maximum size of the projected Gaussian is not greater than or equal to a threshold (step S103; NO), the processor 11 proceeds to step S105 without performing step S104 described above. In other words, in this case, the 3D Gaussian of interest is not extracted as a target for splitting.
[0064] Processor 11 determines that the 3D Gaussian that it was focusing on in steps S102 to S104 has been determined, and then determines whether there are any undetermined 3D Gaussians remaining (step S105). If it determines that there are any undetermined 3D Gaussians remaining (step S105; YES), processor 11 returns to step S101.
[0065] On the other hand, if it is determined that there are no undetermined 3D Gaussians remaining (step S105; NO), the processor 11 determines whether or not there are any 3D Gaussians extracted in step S104 (step S106).
[0066] If it is determined that no 3D Gaussians have been extracted (step S106; NO), the processor 11 terminates the process. On the other hand, if it is determined that there are 3D Gaussians that have been extracted (step S106; YES), the processor 11 performs a splitting process on those extracted 3D Gaussians (step S200).
[0067] Figure 12 is a flowchart illustrating an example of the operation flow of the splitting process. Processor 11 identifies the opacity of each 3D Gaussian extracted in step S104 and selects several 3D Gaussians with a probability proportional to their opacity (step S201). Processor 11 selects 3D Gaussians by sampling according to, for example, a multinomial distribution of the identified opacities.
[0068] Next, the processor 11 splits each of the selected 3D Gaussians into multiple parts while maintaining their total volume (step S202).
[0069] Furthermore, the processor 11 arranges the newly obtained 3D Gaussians based on the positions of the original 3D Gaussians before splitting (step S203).
[0070] In this way, the information processing device 1 performs adaptive density control, splitting the 3D Gaussian initialized according to the feature points of the captured image. This improves spatial resolution and shape representation, making it easier to accurately represent the details of the 3D scene.
[0071] In particular, this information processing device 1 selects the split 3D Gaussian based on the opacity of the 3D Gaussian optimized during the learning process, thereby precisely improving the spatial resolution of the parts with high opacity that have a high proportion of influence on the virtual image reconstructed (generated) on the screen. As a result, this information processing device 1 can achieve high-speed yet high-definition virtual image generation.
[0072] Furthermore, since the information processing device 1 does not perform 3D Gaussian replication, it can eliminate the processing time required for replication and improve rendering speed.
[0073] <Variation> The above describes the embodiment, but the contents of this embodiment can be modified as follows. Furthermore, the following modifications may be combined.
[0074] <1> In the embodiments described above, the processor 11 was a CPU or a GPU, but it may have other configurations. For example, the processor 11 may be an FPGA (Field Programmable Gate Array) or may include an FPGA. The processor 11 may also have an ASIC (Application Specific Integrated Circuit) or other programmable logic device. The information processing device 1 may also have multiple processors, multiple memories, and multiple interfaces. The information processing device 1 may be, for example, a mobile terminal such as a smartphone or a slate PC. The information processing device 1 may also be a virtual machine realized by the dynamic collaboration of multiple computing resources on the cloud via a communication line such as the internet.
[0075] <2> In the embodiment described above, the information processing device 1 selected a 3D Gaussian with a probability corresponding to the opacity and split it so that the total volume of the selected 3D Gaussian was maintained. However, the criteria for selecting the 3D Gaussian to be split are not limited to this.
[0076] Figure 13 is a flowchart showing an example of the operation flow of the splitting process in a modified example. In step S200 shown in Figure 13, step S201a is executed instead of step S201 shown in Figure 12. That is, the processor 11 selects a 3D Gaussian with a probability proportional to the opacity as well as the maximum size of the projected Gaussian (step S201a). As a result, the information processing device 1 becomes more likely to split the 3D Gaussian as the opacity increases, and also as the maximum size of the projected Gaussian increases. Therefore, the information processing device 1 can improve the spatial resolution of parts with high opacity that have a high proportion of influence on the virtual image reconstructed (generated) on the screen, as well as parts that have a high proportion of influence on the virtual image because they occupy a large area on the screen when projected.
[0077] In other words, the splitting means 1153 implemented by the processor 11 in this modified example is an example of a splitting means that splits a 3D Gaussian selected from the extracted 3D Gaussians with a probability corresponding to the opacity and the maximum size of the projected Gaussian, so as to maintain the total volume.
[0078] <3> In the embodiment described above, the information processing device 1 only splits the 3D Gaussian and does not perform replication; however, if certain conditions are met, one of the 3D Gaussians may be replicated.
[0079] Figure 14 shows an example of the configuration of the adaptive density control means 115 in a modified example. The adaptive density control means 115 shown in Figure 14 differs from that shown in Figure 10 in that it has a computation load determination means 1155 and a replication means 1156.
[0080] The computation load determination means 1155 is configured to monitor the computation load of the information processing device 1. This computation load determination means 1155 determines whether the load on the processor 11, or the computational resources controlled by the processor 11, is below a predetermined level.
[0081] If the computation load determination means 1155 determines that the computation load is below a predetermined level, it supplies the 3D Gaussian that was not extracted by the extraction means to the replication means 1156. The replication means 1156 replicates the 3D Gaussian supplied by the computation load determination means 1155 that meets the specified conditions.
[0082] For example, the replication means 1156 replicates a 3D Gaussian in virtual space if its size (three-dimensional size) is smaller than a certain threshold. This allows the information processing device 1 to improve the accuracy of rendering over a wide range of areas, which is difficult to achieve with splitting, and to enhance the general shape and low-frequency characteristics when reproducing an object as a virtual object. Furthermore, since the information processing device 1 only replicates the 3D Gaussian when sufficient computing resources are available, it is less likely to interfere with the splitting process of the 3D Gaussian in adaptive density control, and is less likely to reduce the rendering speed.
[0083] In this modified example, the computation load determination means 1155 and the replication means 1156 are examples of replication means that, when the computation load is below a predetermined level, replicates a plurality of 3D Gaussians that satisfy certain conditions.
[0084] <4> In the embodiment described above, the processor 11 of the information processing device 1 functioned as a bundle adjustment means 112 that calculated camera information at the time of capture from the captured images supplied by the image acquisition means 111. However, the camera information may be stored in the memory 12 in advance, associated with the captured images.
[0085] Figure 15 shows an example of the functional configuration of the information processing device 1 in a modified example. Camera information DB 122 stored in memory 12 has camera information pre-stored for each image ID that identifies an image in image DB 121. In this case, the processor 11 does not need to estimate camera information from the feature points of the captured image. As shown in Figure 15, the processor 11 simply needs to function as a camera information acquisition means 112a that reads the camera information associated with the image ID in camera information DB 122.
[0086] Image DB121 may be integrated with camera information DB122. For example, when a camera photographs an object, its own camera information may be stored as metadata for the captured image. Camera information may be generated when an object is photographed by a barometer, accelerometer, gyro sensor, GNSS, etc., attached to the camera itself. Also, if there is a moving device for moving the camera, camera information may be generated using a measuring instrument for the amount of movement provided on this moving device. Furthermore, camera information may be pre-associated with the position where the camera is attached, or with camera identification information, etc.
[0087] <5> Furthermore, in the embodiment described above, the information processing device 1 extracted only those 3D Gaussians whose maximum size projected onto each screen was greater than or equal to a threshold, from among the 3D Gaussians whose parameters had been learned, and selected several 3D Gaussians from among the extracted 3D Gaussians as targets for splitting, with a probability corresponding to the opacity. However, the order of extraction and selection may be reversed.
[0088] For example, the adaptive density control means 115 shown in Figure 15 is configured such that the extraction means 1151 is executed after the selection means 1152, followed by the splitting means 1153 and the placement means 1154. In this case, the selection means 1152 selects several 3D Gaussians from the learned 3D Gaussians with a probability corresponding to the opacity.
[0089] Then, the extraction means 1151 extracts 3D Gaussians from the 3D Gaussians selected by the selection means 1152 such that the maximum size of the projected Gaussian is greater than or equal to a threshold. The splitting means 1153 splits the 3D Gaussians extracted by the extraction means 1151 into multiple parts so as to maintain their total volume.
[0090] Therefore, the selection means 1152 shown in Figure 15 is an example of a selection means that selects a 3D Gaussian from among several 3D Gaussians with a probability corresponding to the opacity.
[0091] Furthermore, the extraction means 1151 and splitting means 1153 shown in Figure 15 are examples of splitting means that, from among 3D Gaussians selected according to opacity, etc., split a 3D Gaussian whose maximum size projected onto each screen exceeds a threshold, so as to maintain its total volume.
[0092] <6> In the embodiment described above, the program executed by the processor 11 is conceivable as a program that causes the computer to function as: an acquisition means that acquires a plurality of images of an object and camera information indicating the position, orientation, and shooting range of the camera that captured each of the plurality of images; a calculation means that calculates a projected Gaussian by projecting each of the plurality of 3D Gaussians placed in a virtual space onto a screen corresponding to each camera in that virtual space; an extraction means that extracts from the plurality of 3D Gaussians the one whose largest projected Gaussian projected onto each screen is greater than or equal to a threshold; and a splitting means that splits the 3D Gaussians selected from the extracted 3D Gaussians with a probability according to the opacity into multiple parts so as to maintain the total volume.
[0093] <7> The method of causing the processor 11 described above to execute the program described above can also be conceived as an adaptive density control method that causes the computer to control the density of 3D Gaussians. That is, the adaptive density control method of 3D Gaussians described above is an example of an adaptive density control method of 3D Gaussians characterized by including the steps of: the computer obtaining a plurality of images of an object and camera information indicating the position, orientation, and shooting range of the camera that took each of the plurality of images in association with each of them; the computer calculating a projected Gaussian by projecting each of the plurality of 3D Gaussians placed in a virtual space onto a screen corresponding to each camera in that virtual space; the computer extracting from the plurality of 3D Gaussians the one whose largest projected Gaussian projected onto each screen is greater than or equal to a threshold; and the computer splitting the 3D Gaussians selected from the extracted 3D Gaussians with a probability according to the opacity into a plurality so as to maintain the total volume. [Explanation of Symbols]
[0094] 1... Information processing device, 11... Processor, 111... Image acquisition means, 112... Bundle adjustment means, 1121... Feature point extraction means, 1122... Feature point matching means, 1123... Verification means, 1124... Estimation means, 112a... Camera information acquisition means, 113... Initialization means, 114... Learning means, 1141... Calculation means, 1142... Error minimization means, 115... Adaptive density control means, 1151... Extraction means, 1152... Selection means, 1153... Splitting means, 1154... Arrangement means, 1155... Computation load determination means, 1156... Replication means, 12... Memory, 121... Image DB, 122... Camera information DB, 123... Gaussian DB, 13... Interface.
Claims
1. Computers, An acquisition means for acquiring multiple images of an object and camera information indicating the position, orientation, and shooting range of the camera that took each of the multiple images, respectively, and associating them with each other. A calculation means for calculating a projected Gaussian obtained by projecting each of the multiple 3D Gaussians placed in a virtual space onto a screen corresponding to each of the cameras in the virtual space, An extraction means for extracting from the plurality of 3D Gaussians the one whose largest projected Gaussian projected onto each of the screens is greater than or equal to a threshold, A splitting method that selects a 3D Gaussian from the extracted 3D Gaussians with a probability corresponding to the opacity, and then splits it into multiple parts while maintaining the total volume. A program designed to function as such.
2. The splitting means splits a 3D Gaussian selected from the extracted 3D Gaussians with a probability corresponding to the opacity and the maximum size of the projected Gaussian, so as to maintain the total volume. The program according to claim 1.
3. The aforementioned computer, When the computational load is below a predetermined level, a replication means for replicating the one among the plurality of 3D Gaussians that satisfies the conditions, The program according to claim 1 for functioning as such.
4. The aforementioned computer, Arrangement means for positioning each of the split 3D Gaussian based on the position of the original 3D Gaussian, The program according to claim 1 for functioning as such.
5. Computers, An acquisition means for acquiring multiple images of an object and camera information indicating the position, orientation, and shooting range of the camera that took each of the multiple images, respectively, and associating them with each other. A calculation means for calculating a projected Gaussian obtained by projecting each of the multiple 3D Gaussians placed in a virtual space onto a screen corresponding to each of the cameras in the virtual space, A selection means for selecting a 3D Gaussian from among the aforementioned multiple 3D Gaussians with a probability corresponding to the opacity, A splitting means for splitting a 3D Gaussian from among the selected 3D Gaussians, such that the largest projected Gaussian projected onto each of the screens is greater than or equal to a threshold, while maintaining the total volume. A program designed to function as such.
6. An acquisition means for acquiring multiple images of an object and camera information indicating the position, orientation, and shooting range of the camera that took each of the multiple images, respectively, and associating them with each other. A calculation means for calculating a projected Gaussian obtained by projecting each of the multiple 3D Gaussians placed in a virtual space onto a screen corresponding to each of the cameras in the virtual space, An extraction means for extracting from the plurality of 3D Gaussians the one whose largest projected Gaussian projected onto each of the screens is greater than or equal to a threshold, A splitting method that splits a 3D Gaussian selected from the extracted 3D Gaussians with a probability corresponding to the opacity into multiple parts while maintaining the total volume, An information processing device having
7. The computer acquires multiple images of an object, and camera information indicating the position, orientation, and shooting range of the camera that took each of the multiple images, by associating them with each other. The steps include: a computer calculating a projected Gaussian by projecting each of the multiple 3D Gaussians placed in a virtual space onto a screen corresponding to each of the cameras in the virtual space; The computer extracts from the plurality of 3D Gaussians the one whose largest projected Gaussian projected onto each of the screens is greater than or equal to a threshold, The computer selects a 3D Gaussian from the extracted 3D Gaussians with a probability corresponding to the opacity, and then splits it into multiple parts while maintaining the total volume. A method for adaptive density control of a 3D Gaussian, characterized by including the following:
Citation Information
Patent Citations
Three dimensional model generation apparatus and method
JP2021071749A