Free-viewpoint video generation device, program, and free-viewpoint video generation method
The free-viewpoint video generation device efficiently reduces 3D Gaussians in dynamic scenes by adaptively generating regions and determining reduction amounts based on object movement, enhancing rendering speed and reducing memory usage.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2026-03-30
AI Technical Summary
Existing 3D Gaussian Splatting (3DGS) technologies struggle to maintain visual quality and reduce 3D Gaussian values efficiently in dynamic scenes, as they are designed for large-scale static scenes and do not account for object movement.
A free-viewpoint video generation device and method that adaptively generates regions based on 3D Gaussian information, determining reduction amounts for 3D Gaussians in each region, and selectively reduces them to generate free-viewpoint images, using learned information at predetermined time intervals.
Enables efficient removal of 3D Gaussians in dynamic scenes, reducing memory usage and accelerating rendering by adapting to object movement, while maintaining visual quality.
Smart Images

Figure 0007837489000001 
Figure 0007837489000002 
Figure 0007837489000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to a free viewpoint video generation device, a program, and a free viewpoint video generation method.
Background Art
[0002] In recent years, as a free viewpoint video generation technology in a 3D (Three Dimensions) space, the use of 3DGS (3D Gaussian Splatting) has attracted attention. 3DGS is a technology for representing a three-dimensional space using 3D Gaussians composed of color information depending on the center point, rotation, covariance, density, and viewpoint direction of a Gaussian distribution.
[0003] On the other hand, in 3DGS, the 3D Gaussian may increase according to the scaling or dynamic change of the scene. The increase in the 3D Gaussian may lead to an increase in the memory usage during rendering and a decrease in the rendering speed. Regarding this point, as a technique for efficiently rendering a large-scale static scene, the technique described in Non-Patent Document 1 has been proposed.
[0004] In the technique described in Non-Patent Document 1, first, regions are configured by dividing the three-dimensional space into block shapes. This is because in the scene configuration by 3DGS, there is no object boundary. Next, the reduction amount of the Gaussian within the region is determined from the region and the distance relationship between the camera during rendering. Specifically, in the region near the camera, the reduction number of the 3D Gaussian is made small, and in the region far from the camera, the reduction number of the 3D Gaussian is made large. This is because the region close to the camera needs to be displayed finely because it is largely displayed in the visual field, and the region far from the camera can be coarsely displayed without significantly deteriorating the visual quality because it is small in the visual field. By adopting such a method, the technique described in Non-Patent Document 1 can speed up rendering and reduce memory without significantly deteriorating the appearance during rendering.
Prior Art Documents
[0005] [Non-Patent Document 1] Liu, Yang, et al. “Citygaussian: Real-time high-quality large-scale scene rendering with gaussians.” European Conference on Computer Vision. Springer, Cham, 2025. [Overview of the project] [Problems that the invention aims to solve]
[0006] However, the technology described in Non-Patent Document 1, which targets large-scale static scenes, may not be able to efficiently maintain visual quality or reduce 3D Gaussian values in dynamic scenes. This is because the technology described in Non-Patent Document 1 constitutes a fixed region and cannot perform 3D Gaussian reduction that takes into account the movement of objects.
[0007] Therefore, one or more aspects of this disclosure aim to enable the removal of 3D Gaussian in a scene constructed by 3DGS in accordance with the movement of objects. [Means for solving the problem]
[0008] A free-viewpoint video generation device according to one aspect of this disclosure generates free-viewpoint video by 3DGS (3D Gaussian Splatting), and in a 3D scene model generated by learning 3D Gaussian information from video data, the learned information is used at predetermined time intervals. 、 3D Gaussian , to multiple dynamic objects A region generation unit that generates multiple regions including at least one region containing a 3D Gaussian classified by the aforementioned information, A reduction amount determination unit that determines the reduction amount for reducing the 3D Gaussian in each of the aforementioned multiple regions, In each of the aforementioned multiple regions, According to the aforementioned reduction amount,The system is characterized by comprising a reduction unit that reduces 3D Gaussian values, and a rendering unit that generates free-viewpoint images by rendering using the remaining 3D Gaussian values that were not reduced.
[0009] A program according to one aspect of this disclosure uses a computer to generate a 3D scene model by learning 3D Gaussian information from video data in order to generate free-viewpoint video using 3DGS (3D Gaussian Splatting), and the learned information is used at predetermined time intervals. 、 3D Gaussian , to multiple dynamic objects A region generation unit that generates multiple regions, including at least one region containing a 3D Gaussian classified by the aforementioned information, A reduction amount determination unit determines the reduction amount for reducing the 3D Gaussian in each of the aforementioned multiple regions. In each of the aforementioned multiple regions, According to the aforementioned reduction amount, This system is characterized by functioning as a reduction unit that reduces 3D Gaussian values, and a rendering unit that generates free-viewpoint images by rendering using the remaining 3D Gaussian values that were not reduced.
[0010] A free-viewpoint video generation method according to one aspect of this disclosure generates free-viewpoint video using 3DGS (3D Gaussian Splatting), and in a 3D scene model generated by learning 3D Gaussian information from video data, the learned information is used at predetermined time intervals. 、 3D Gaussian , to multiple dynamic objects By classifying, multiple regions are generated that include at least one region containing the 3D Gaussian classified by the aforementioned information. In each of the aforementioned multiple regions, the reduction amount for reducing the 3D Gaussian is determined. In each of the aforementioned multiple regions, According to the aforementioned reduction amount, This method is characterized by generating free-viewpoint video by reducing the 3D Gaussian and performing rendering using the remaining 3D Gaussian. [Effects of the Invention]
[0011] According to one or more aspects of this disclosure, 3D Gaussian can be removed from a scene composed of 3DGS in accordance with the movement of objects. [Brief explanation of the drawing]
[0012] [Figure 1] It is a block diagram schematically showing the configuration of the free viewpoint video generation device according to Embodiments 1 and 4. [Figure 2] It is a block diagram schematically showing the configuration of the 3D scene model generation unit. [Figure 3] (A) and (B) are schematic diagrams for explaining an example of the method for generating a region. [Figure 4] (A) and (B) are schematic diagrams for explaining the method for deleting 3D Gaussians. [Figure 5] It is a schematic diagram for explaining an example of the rendering method. [Figure 6] (A) and (B) are block diagrams showing an example of the hardware configuration. [Figure 7] It is a flowchart showing the operation during learning of the free viewpoint video generation device according to Embodiment 1. [Figure 8] It is a flowchart showing the operation during rendering of the free viewpoint video generation device according to Embodiment 1. [Figure 9] It is a block diagram schematically showing the configuration of the free viewpoint video generation device according to Embodiment 2. [Figure 10] It is a block diagram schematically showing the configuration of the free viewpoint video generation device according to Embodiment 3. [Figure 11] It is a flowchart showing the operation during learning of the free viewpoint video generation device according to Embodiment 3. [Figure 12] It is a flowchart showing a part of the processing in the region generation unit. [Figure 13] It is a flowchart showing the operation during rendering of the free viewpoint video generation device according to Embodiment 3. [Figure 14] (A) and (B) are schematic diagrams for explaining the first example for determining the reduction amount in Embodiment 4. [Figure 15] It is a schematic diagram for explaining the second example for determining the reduction amount in Embodiment 4. [Figure 16](A) and (B) are first schematic diagrams illustrating the process of generating regions in steps in Embodiment 5. [Figure 17] (A) and (B) are second schematic diagrams illustrating the process of generating regions in steps in Embodiment 5. [Modes for carrying out the invention]
[0013] Embodiment 1. Figure 1 is a block diagram schematically showing the configuration of the free-viewpoint video generation device 100 according to Embodiment 1. The free-viewpoint video generation device 100 first performs learning processing for 3D scene construction, and then performs rendering processing. The free-viewpoint video generation device 100 comprises a communication unit 110, a 3D scene model generation unit 111, a storage unit 115, a region generation unit 116, an input unit 120, a virtual camera configuration unit 121, a reduction amount determination unit 122, a reduction unit 123, a rendering unit 124, and a display unit 125.
[0014] The communication unit 110 communicates data with other devices. For example, the communication unit 110 receives video data captured by multiple cameras, which are multiple imaging devices not shown in the figure.
[0015] The 3D scene model generation unit 111 generates a 3D scene model by learning 3D Gaussian information from video data captured by multiple cameras in order to generate free-viewpoint video using 3DGS. The processing in the 3D scene model generation unit 111 can be implemented using known technologies such as 3DGS and 4D Gaussian Splatting.
[0016] Figure 2 is a block diagram that schematically shows the configuration of the 3D scene model generation unit 111. The 3D scene model generation unit 111 comprises a video processing unit 112, an initial point cloud generation unit 113, and a 3D scene construction unit 114.
[0017] The video processing unit 112 acquires video data captured by multiple cameras. In this example, the video processing unit 112 acquires video data from multiple cameras via the communication unit 110, but Embodiment 1 is not limited to this example. For example, if video data is already stored in the storage unit 115, the video processing unit 112 may retrieve the video data from the storage unit 115. Furthermore, if the video data is stored on another device (for example, a server), the video processing unit 112 may acquire the video data from the other device via the communication unit 110.
[0018] Next, the video processing unit 112 divides the video data from multiple cameras, in other words, the video data from multiple viewpoints, into frames and classifies them into groups of frames from the same time. A group of frames consists of multiple frames captured by multiple cameras at the same time, in other words, multiple frames obtained from multiple viewpoints at the same time. This is because multiple images taken at the same time are required to construct a 3D scene.
[0019] The initial point cloud generation unit 113 generates an initial point cloud from each frame in the frame group. The initial point cloud is represented as a three-dimensional point cloud. Known techniques such as SfM (Structure from Motion) can be used to construct the initial point cloud. Furthermore, the initial point cloud generation unit 113 can also use a 3D point cloud acquired by a laser scanner such as LiDAR (Light Detection And Ranging) as the initial point cloud, without using a frame group.
[0020] The 3D scene constructor 114 generates a 3D scene model by constructing a dynamic 3D scene from a group of frames and an initial point cloud. The 3D scene model is composed of a 3D Gaussian that contains the center point coordinates, rotation, covariance, density, and color information dependent on the viewpoint direction of the Gaussian distribution over time. By constructing the 3D scene model in this way, it is possible to represent a 3D scene that changes over time. Existing technologies such as 4D Gaussian Splatting can be utilized for these constructions.
[0021] For example, the 3D scene constructor 114 assigns a Gaussian distribution of a 3D Gaussian to the initial point cloud, visualizes the generated 3D Gaussian, and compares it with the original image identified from the frame set to determine which pixels are shifted and by how much. In this way, machine learning can be used to learn the 3D scene and construct a 3D scene model.
[0022] The memory unit 115 stores the 3D scene model generated in the manner described above.
[0023] The region generation unit 116 adaptively generates regions according to the 3D scene. For example, the region generation unit 116, in order to generate free-viewpoint images using 3DGS, learns 3D Gaussian information from video data to generate a 3D scene model. In this model, it classifies the 3D Gaussians at predetermined time intervals using the learned information, thereby identifying multiple regions that include at least one region containing a 3D Gaussian classified by that information. Here, the 3D Gaussian information consists of parameters such as the center point coordinates of the Gaussian distribution, rotation, covariance, density, and color information that depends on the viewpoint direction. The predetermined time is the time at which the frame group was captured.
[0024] Specifically, the region generation unit 116 receives 3D Gaussian information from the 3D scene model generation unit 111. The region generation unit 116 can calculate the amount of movement from the change in the center coordinates of the Gaussian between a predetermined time and the time before that time. Furthermore, if the amount of movement of the 3D Gaussian is calculated within the 3D scene model, such as in Deformable 3DGS, the region generation unit 116 can also utilize that amount of movement. Furthermore, the region generation unit 116 can also calculate the amount of rotation based on the change in rotation information between a predetermined time and the time before that time. Furthermore, the region generation unit 116 can calculate the density change based on the change in the 3D Gaussian density between a predetermined time and the time before that time. Furthermore, the region generation unit 116 can also include the above-mentioned amounts of movement, rotation, and density changes in the 3D Gaussian information.
[0025] The region generation unit 116 then generates regions by classifying 3D Gaussians with similar information into groups. Here, it is desirable for the region generation unit 116 to determine whether or not 3D Gaussians are similar by considering at least one of the parameters, displacement, rotation, and density change of the 3D Gaussians. This is because 3D Gaussian values contained within the same object can be expected to contain similar information.
[0026] For example, as shown in Figure 3(A), a region 130#1 is generated from 3D Gaussians 130A to 130D that have similar information at a certain time t, and a region 131#1 is generated from 3D Gaussians 131A to 131D that have similar information. If regions 130#1 and 131#1 correspond to objects, then, for example, as shown in Figure 3(B), at time t+1 following time t, the 3D Gaussians 130A to 130D contained in region 130 will have the same relative positional relationship as at time t, and the 3D Gaussians 131A to 131D contained in region 131 will also have the same relative positional relationship.
[0027] In this way, regions can be generated that correspond to objects within a 3D scene. These regions can be generated using well-known clustering methods, such as k-means.
[0028] The region generation unit 116 stores the label information of the regions (clusters) generated in this way as region information in the storage unit 115. The label information is information indicating the number assigned to each region. This number can be used as a region index during rendering, etc.
[0029] The input unit 120 accepts input of at least some of the camera information. The camera information includes external parameters of a virtual camera used when rendering a 3D scene using the 3D scene model stored in the memory unit 115, as well as internal parameters such as focal length.
[0030] The virtual camera configuration unit 121 receives external camera parameters as input via the input unit 120 and configures camera information for rendering. At this time, internal parameters can be estimated from the resolution of the rendered image, etc. Alternatively, the parameters of the camera that captured the video data used during training can also be used as internal parameters.
[0031] The reduction amount determination unit 122 determines the reduction amount to reduce the 3D Gaussian in each of the multiple regions generated by the region generation unit 116. For example, the reduction amount determination unit 122 determines the reduction amount such that it increases as the amount of movement of each of the multiple regions increases.
[0032] Here, the reduction amount determination unit 122 receives, in addition to camera information from the virtual camera configuration unit 121, the 3D scene model generated during training and region information as input. The reduction amount determination unit 122 identifies regions from the 3D scene model according to the region information and extracts 3D Gaussian information contained in the identified regions from the 3D scene model. Then, the reduction amount determination unit 122 determines how much to reduce the 3D Gaussian for each identified region based on the extracted information and outputs the result.
[0033] Specifically, the region information consists of region label information and the 3D Gaussian index that each region possesses. Therefore, the reduction amount determination unit 122 extracts information such as the amount of region movement from the 3D scene model based on the 3D Gaussian index for each target region, and determines the reduction amount from the extracted information.
[0034] For example, as shown in Figure 4(A), if the amount of movement of region 130 from time t to time t+1 is greater than the amount of movement of region 131, the reduction amount determination unit 122 makes the reduction amount of the 3D Gaussian of region 130 at time t+1 greater than the reduction amount of the 3D Gaussian of region 131 at time t+1, as shown in Figure 4(B).
[0035] The reduction unit 123 reduces the 3D Gaussian in each of the multiple regions generated by the region generation unit 116. For example, the reduction unit 123 reduces the 3D Gaussian from each of the multiple regions according to the reduction amount determined by the reduction amount determination unit 122. Specifically, the reduction unit 123 takes the reduction amount determined for each region and the 3D scene model as input and reduces the 3D Gaussian within the region from the 3D scene model according to a specific criterion.
[0036] The rendering unit 124 generates a free-viewpoint image by rendering using the remaining 3D Gaussian that has not been reduced. For example, the rendering unit 124 receives the 3D scene model after it has been reduced by the reduction unit 123 and performs rendering processing using that 3D scene model. Specifically, as shown in Figure 5, the rendering unit 124 can render a free-viewpoint image by using the remaining 3D Gaussian that has not been reduced and projecting it using a virtual camera position CP indicated by an external parameter.
[0037] The display unit 125 displays the image rendered by the rendering unit 124.
[0038] Figures 6(A) and (B) are block diagrams showing some hardware configuration examples of the free-viewpoint video generation device 100. Some or all of the 3D scene model generation unit 111, region generation unit 116, virtual camera configuration unit 121, reduction amount determination unit 122, reduction unit 123, and rendering unit 124 can be configured, for example, as shown in Figure 6(A), with a memory 10 and a processor 11 such as a CPU (Central Processing Unit) that executes the program stored in the memory 10. Such a program may be provided via a network or by being recorded on a recording medium. That is, such a program may be provided, for example, as a computer program product. As described above, the free-viewpoint video generation device 100 can be realized using a computer.
[0039] Furthermore, some or all of the 3D scene model generation unit 111, region generation unit 116, virtual camera configuration unit 121, reduction amount determination unit 122, reduction unit 123, and rendering unit 124 can also be composed of processing circuits 12 such as a single circuit, a composite circuit, a program-operated processor, a program-operated parallel processor, an ASIC (Application Specific Integrated Circuit), or an FPGA (Field Programmable Gate Array), as shown in Figure 6(B). As described above, the 3D scene model generation unit 111, region generation unit 116, virtual camera configuration unit 121, reduction amount determination unit 122, reduction unit 123, and rendering unit 124 can be realized by a processing circuit network.
[0040] The communication unit 110 can be implemented using a communication interface (I / F) such as a NIC (Network Interface Card) that performs communication. The input unit 120 can be implemented using an input interface such as a keyboard or mouse. The display unit 125 can be implemented using a display.
[0041] Figure 7 is a flowchart showing the operation of the free-viewpoint video generation device 100 according to Embodiment 1 during the learning phase. First, the video processing unit 112 acquires video data captured from multiple viewpoints and generates a group of frames by dividing it into multiple images with different viewpoints for each frame (S10).
[0042] Next, the initial point cloud generation unit 113 generates an initial point cloud from the frame group, which consists of divided multiple viewpoint images (S11). At this time, the initial point cloud is composed of a three-dimensional point cloud. As a generation method, the three-dimensional point cloud is generated from the relationship between similar areas between each image and the camera position. In this process, the initial point cloud generation unit 113 can utilize techniques such as SfM. Point clouds from laser scanners such as LiDAR can also be used. In this case, the point cloud data is associated with the video data before input by calibration, etc.
[0043] Next, the 3D scene constructor 114 generates a 3D scene model (S12) by constructing a dynamic 3D scene from the generated 3D point cloud and information from multiple viewpoint images. This 3D scene is constructed using a time-dependent 3D Gaussian model. Existing methods such as 4D Gaussian Splatting can be used for scene construction.
[0044] Then, the region generation unit 116 repeats the processes of the next steps S14 and S15 at predetermined intervals (S13, S16). In step S14, the region generation unit 116 extracts information from the 3D Gaussian that constitutes the 3D scene model over time, including the center point coordinates, rotation, covariance, density, color information, amount of movement due to changes in coordinate information, amount of rotation due to changes in rotation information, and amount of change in density. In step S15, the region generation unit 116 applies a clustering method to classify 3D Gaussians with similar information in the extracted information into groups and generates a region for each object. In clustering to generate a region of 3D Gaussians for each object, the region generation unit 116 uses a combination of at least one piece of information possessed by the 3D Gaussian, such as the center coordinates, color, density, rotation angle, covariance, and past movement information of the Gaussian over a predetermined period.
[0045] Figure 8 is a flowchart showing the rendering operation of the free-viewpoint video generation device 100 according to Embodiment 1. During rendering, camera information, frame number, region information, and a 3D scene model are required. Camera information consists of external and internal parameters of the virtual camera that represents the user's viewpoint, and is identified by the virtual camera component 121. External parameters can be input via the input unit 120. Frame numbers are information about the time when rendering is to be performed, and may be identified from the time of the video displayed by the free-viewpoint video generation device 100, or they can be input via the input unit 120. Region information consists of region label information and the 3D Gaussian index of each region, and is stored in the storage unit 115.
[0046] First, the reduction amount determination unit 122 identifies regions from the 3D scene model according to the region information. Then, the reduction amount determination unit 122 and the reduction unit 123 repeat the following steps S21 and S22 for each identified region (S20, S23).
[0047] In step S21, the reduction amount determination unit 122 determines the reduction amount of the 3D Gaussian contained within the region. In determining the reduction amount, the reduction amount determination unit 122 determines the reduction amount of the 3D Gaussian based on the movement information of the entire region estimated by the average of the movement amounts of the 3D Gaussian. Here, the reduction amount is determined by gradually increasing the reduction amount according to the magnitude of the movement amount of the entire region. In other words, the reduction amount determination unit 122 prepares stepwise thresholds for the degree of movement according to the scene in advance, and determines the amount of reduction by comparing these thresholds with the movement amount. For example, the reduction amount determination unit 122 considers a region to be a fast-moving object if the amount of movement between predetermined time intervals is large, and decides to reduce a large amount of 3D Gaussian. Furthermore, for displacements below a certain threshold, it is possible to set the reduction amount of the 3D Gaussian from that region to "0".
[0048] In step S22, the reduction unit 123 reduces the 3D Gaussian contained in that region by the determined reduction amount. Here, the reduction unit 123 selects a 3D Gaussian that has little impact on rendering and reduces the selected 3D Gaussian. The degree of impact during rendering can be determined using information such as the density and covariance of the 3D Gaussian, or how much of each 3D Gaussian was used during training or rendering.
[0049] The extent to which 3D Gaussian is used can be calculated by determining how many 3D Gaussian curves intersect the light rays traveling from the camera center to the image plane during rendering. This method can be based on existing techniques such as Light Gaussian.
[0050] Another approach is to prioritize reducing 3D Gaussians corresponding to pixels with strong edges in the rendered image. This is because, in dynamic objects, edge information is lost due to factors such as motion blur. One method to achieve this is to prioritize reducing 3D Gaussians with a small scale or those with a large number of 3D Gaussians within a certain range from the center of a particular 3D Gaussian.
[0051] Next, the rendering unit 124 extracts the 3D Gaussian within the field of view from the 3D scene model according to the camera information (S24). This is to select the 3D Gaussian to be rendered.
[0052] The rendering unit 124 then projects the 3D Gaussian selected in this manner onto a virtual camera and performs rendering (S25). Projection can be performed by using geometric transformations such as affine transformations on the 3D Gaussian. This process can utilize rasterization processing such as that of 3DGS. After that, the rendering unit 124 displays the image composed of the rendering on the display unit 125, which is configured as a monitor or an HMD or other video output device.
[0053] As described above, according to Embodiment 1, since regions are generated in response to dynamic objects, it is possible to reduce the 3D Gaussian for each object without compromising the appearance. For this reason, Embodiment 1 can achieve faster rendering and reduced memory usage.
[0054] Specifically, objects with little movement need to appear detailed because they remain in the field of view for a longer period, thus reducing the amount of 3D Gaussian reduction needed. On the other hand, objects with a lot of movement are difficult to perceive with the eye, so they do not need to be displayed in detail, allowing for a greater reduction in 3D Gaussian.
[0055] Furthermore, the technology described in Non-Patent Document 1, which targets large-scale static scenes, may not be able to efficiently maintain visual quality and reduce 3D Gaussian in dynamic scenes. This is because, in Non-Patent Document 1, regions are generated in a fixed manner, making it impossible to perform 3D Gaussian reduction that takes objects into account.
[0056] Specifically, when multiple objects are located close together, objects with different degrees of movement may be included in the same region. This can cause the 3D Gaussian of objects with less movement to be reduced, resulting in blurring. Conversely, the 3D Gaussian of objects with greater movement may not be reduced, leading to increased memory usage. Furthermore, if objects with little movement are located on the boundary of a region, the degree of 3D Gaussian reduction varies at the boundary, potentially causing partial blurring.
[0057] Embodiment 2. In Embodiment 1 described above, regions are generated from the information contained in the 3D Gaussian, but in Embodiment 2, regions can be generated more efficiently by using external data such as 3D segmentation data or 3D skeleton data.
[0058] Figure 9 is a block diagram schematically showing the configuration of the free-viewpoint video generation device 200 according to Embodiment 2. The free-viewpoint video generation device 200 includes a communication unit 210, a 3D scene model generation unit 111, a storage unit 115, a region generation unit 216, an input unit 120, a virtual camera configuration unit 121, a reduction amount determination unit 122, a reduction unit 123, a rendering unit 124, and a display unit 125.
[0059] The 3D scene model generation unit 111, storage unit 115, input unit 120, virtual camera configuration unit 121, reduction amount determination unit 122, reduction unit 123, rendering unit 124, and display unit 125 of the free-viewpoint video generation device 200 according to Embodiment 2 are the same as those of the 3D scene model generation unit 111, storage unit 115, input unit 120, virtual camera configuration unit 121, reduction amount determination unit 122, reduction unit 123, rendering unit 124, and display unit 125 of the free-viewpoint video generation device 100 according to Embodiment 1.
[0060] The communication unit 210 communicates data with other devices. In Embodiment 2, the communication unit 210 receives video data captured by multiple cameras, which are multiple imaging devices not shown, and also acquires external data from other devices not shown, such as a server.
[0061] As described above, external data refers to data that can identify objects contained in the video, such as 3D segmentation data or 3D skeletal data. 3D segmentation data is data that explicitly defines the regions of an object using labels, etc. Existing technologies such as Track Anything can be used to obtain 3D segmentation data. Furthermore, 3D skeletal data is data that represents skeletal information showing the human skeleton, such as fingers, eyes, nose, and mouth, in three-dimensional space. Existing technologies such as mediapipe can be used to obtain this 3D skeletal data.
[0062] Furthermore, the mapping between external data and video data can be achieved by methods such as alignment through calibration, or by adding this information when constructing the 3D Gaussian model. In the methods described later, 3D Gaussian construction techniques such as Feature 3DGS can be utilized.
[0063] The region generation unit 216 adaptively generates regions according to the 3D scene. In Embodiment 2, the region generation unit 216 receives external data from the communication unit 210 and generates a region for each object identified by the external data. For example, the region generation unit 216 uses external data that can identify an object to generate at least one region corresponding to that object, and adds that region to the region generated from 3D Gaussian information, similar to Embodiment 1. In the case of overlapping areas between the region generated from external data and the region generated from 3D Gaussian information, the region generation unit 216 may prioritize the region generated from external data.
[0064] Specifically, if the external data is 3D segmentation data, the region generation unit 216 only needs to generate regions based on the segments indicated by the 3D segmentation data.
[0065] Furthermore, if the external data is 3D skeletal data, 3D information can be obtained for each object, such as arms, torso, and legs. From this information, regions can be generated by grouping the 3D Gaussian values adjacent to each object into regions. For example, the region generation unit 216 can obtain information from the 3D scene model such as the center point coordinates, rotation, covariance, density, and color information of the Gaussian distribution over time, as well as the amount of movement, rotation, and density change over a predetermined period in the past, and then define a region based on the 3D Gaussian within an object identified by the 3D skeleton data and within a predetermined range from that object. The predetermined range may be defined according to the type of object.
[0066] As described above, according to Embodiment 2, since data acquired from an external source is used for region generation, the region of an object can be generated more accurately compared to Embodiment 1.
[0067] In the embodiment 2 described above, external data is acquired via the communication unit 210, but the invention is not limited to this example. For example, if external data is already stored in the storage unit 115, the area generation unit 216 may acquire the external data from the storage unit 115. Alternatively, the area generation unit 216 may receive input of external data from the input unit 120.
[0068] Embodiment 3. While embodiments 1 and 2 utilize a single-layer domain, embodiment 3 generates a domain with a hierarchical structure. In this case, the domain with a hierarchical structure refers to a domain that contains multiple layers of other domains.
[0069] Figure 10 is a block diagram schematically showing the configuration of the free-viewpoint video generation device 300 according to Embodiment 3. The free-viewpoint video generation device 300 comprises a communication unit 110, a 3D scene model generation unit 111, a storage unit 115, a region generation unit 316, an input unit 120, a virtual camera configuration unit 121, a reduction amount determination unit 322, a reduction unit 123, a rendering unit 124, a display unit 125, and an action estimation unit 326.
[0070] The communication unit 110, 3D scene model generation unit 111, storage unit 115, input unit 120, virtual camera configuration unit 121, reduction unit 123, rendering unit 124, and display unit 125 of the free-viewpoint video generation device 300 according to Embodiment 3 are the same as those of the communication unit 110, 3D scene model generation unit 111, storage unit 115, input unit 120, virtual camera configuration unit 121, reduction unit 123, rendering unit 124, and display unit 125 of the free-viewpoint video generation device 100 according to Embodiment 1.
[0071] The region generation unit 316 generates multiple regions from 3D Gaussian information, similar to the region generation unit 116 in Embodiment 1. The region generation unit 316 then identifies a target region as the final region by executing a predetermined number of times the process of generating a new target region by selecting two or more regions containing the classified 3D Gaussian as target regions and classifying them according to the characteristics of those target regions.
[0072] Here, the region generation unit 316 generates a large region, which is a new region that encompasses a predetermined time interval region. For example, the region generation unit 316 extracts predetermined information from the 3D Gaussian contained in the region. Then, the region generation unit 316 generates a higher-level region containing regions similar in the extracted information, performing this process a predetermined number of times (also referred to as a hierarchy), thereby generating the highest-level region as a large region.
[0073] Specifically, the region generation unit 316 can generate higher-level regions by selecting a representative 3D Gaussian, taking into account the average or median of the color information, center coordinates, density, or size information of the 3D Gaussians contained within the region, and the degree of influence from the 3D Gaussians contained within the region, and then performing clustering using the information itself or the representative value which is the average or median of that information. Here, the region generation unit 316 may determine that a higher number of hits with the camera rays during rendering indicates a higher degree of influence, or it may determine that a higher density or magnitude value of the 3D Gaussian indicates a higher degree of influence. The region generation unit 316 then simply selects the 3D Gaussian with the highest influence within the region as representative.
[0074] The region generation unit 316 then generates region information that indicates the label information of the large region and the 3D Gaussian index of each large region, and stores the generated region information in the storage unit 115.
[0075] The behavior estimation unit 326 estimates the behavior of objects contained within a large region, which is a region indicated by the region information. For example, the behavior estimation unit 326 estimates the behavior of an object by referring to the amount of movement or positional relationship of the areas included in the large area over a predetermined period in the past, or information such as the center coordinates, covariance, color, density, or rotation angle of the 3D Gaussian included in the areas included in the large area.
[0076] The reduction amount determination unit 322 determines the reduction amount of the 3D Gaussian based on the behavior content estimated by the behavior estimation unit 326. For example, the reduction amount determination unit 322 determines that the reduction amount is small for large domains where the behavior content is attention-grabbing, and large domains where the behavior content is similar to that of many other large domains.
[0077] Some or all of the behavior estimation unit 326 described above can also be configured, for example, as shown in Figure 6(A), with a memory 10 and a processor 11 such as a CPU that executes the program stored in the memory 10. As described above, the free-viewpoint video generation device 300 can be realized using a computer.
[0078] Furthermore, part or all of the behavior estimation unit 326 can also be composed of a processing circuit 12 such as a single circuit, a composite circuit, a program-operated processor, a program-operated parallel processor, an ASIC, or an FPGA, as shown in Figure 6(B). As described above, the behavior estimation unit 326 can be realized by a processing network.
[0079] Figure 11 is a flowchart showing the operation of the free-viewpoint video generation device 300 according to Embodiment 3 during learning. Furthermore, steps in the flowchart shown in Figure 11 that perform the same processing as the steps in the flowchart shown in Figure 7 are denoted by the same reference numerals as in Figure 7.
[0080] The processing in steps S10 to S16 in Figure 11 is the same as the processing in steps S10 to S16 in Figure 7. However, in Figure 11, after the processing in step S15, the process proceeds to step S30. In step S30, the region generation unit 316 updates the region information by generating a large region that encompasses the regions indicated by the region information generated in steps S13 to S16. This process will be explained in detail using Figure 12. After the processing in step S30, the process proceeds to step S16.
[0081] Figure 12 is a flowchart showing part of the processing in the region generation unit 316. The region generation unit 316 executes the flowchart shown in Figure 12 at predetermined intervals.
[0082] First, the region generation unit 316 determines whether the number of layers in the generated region is less than a predetermined number of layers n (S40). If the number of layers in the generated region is less than a predetermined number of layers n (Yes in S40), the process proceeds to step S41. In other words, the region generation unit 316 repeats the following process until the specified number of layers reaches n. Hereinafter, n is an integer of 2 or more. At this time, the number of layers indicates the depth of the hierarchical structure of the large region. The determination of this number of layers n may be predetermined by the user when composing a 3D scene for each scene and input via the input unit 120.
[0083] In step S41, the region generation unit 316 extracts predetermined information from the 3D Gaussian information contained in the region of the highest level. For example, the region generation unit 316 can extract information about the 3D Gaussians contained within the region of the highest level, the average amount of movement of the 3D Gaussians contained within the region of the highest level, or the number of 3D Gaussians contained within the region of the highest level. The average amount of movement can be calculated by dividing the sum of the amounts of movement of the 3D Gaussians contained within the region by the number of 3D Gaussians contained within that region.
[0084] Next, the region generation unit 316 generates higher-level regions by labeling regions with similar information as belonging to the same class from the extracted information (S42). Then, the process returns to step S40.
[0085] Figure 13 is a flowchart showing the rendering operation of the free-viewpoint video generation device 300 according to Embodiment 3. Furthermore, steps in the flowchart shown in Figure 13 that perform the same processing as the steps in the flowchart shown in Figure 8 will be denoted by the same reference numerals as the steps shown in Figure 8.
[0086] First, the reduction amount determination unit 322 identifies large regions from the 3D scene model according to the region information. Then, the reduction amount determination unit 322 and the reduction unit 123 repeat the following steps S51 to S53 for each identified large region (S50, S54).
[0087] In step S51, the behavior estimation unit 326 estimates the behavior of objects contained within the large region based on movement information or positional relationships of regions contained within the large region, or the central coordinates, covariance, color information, density, or rotation angle of the 3D Gaussian within the large region.
[0088] For example, the behavior estimation unit 326 can estimate the content of the behavior by inputting the information extracted in the manner described above into a trained machine learning model. Specifically, the behavior estimation unit 326 may obtain behavioral content with similar movements from a separately prepared machine learning model or the like, depending on how the region within the large area moved between predetermined time periods.
[0089] Next, in step S52, the reduction amount determination unit 322 determines the reduction amount of the 3D Gaussian based on the action content estimated in step S51. For example, the reduction amount determination unit 322 scores the behavior content according to predetermined criteria and determines the reduction amount of the 3D Gaussian in stages according to the scored value. This scoring method involves assigning higher scores to behaviors that are significantly different from the behavior content of other large regions. This makes it possible to retain a large number of 3D Gaussians of large regions that exhibit abnormal or attention-grabbing behavior, while significantly reducing the 3D Gaussians of large regions that exhibit the same behavior as other objects.
[0090] In step S53, the reduction unit 123 reduces the 3D Gaussian contained in the large region by the determined reduction amount. The processing here is the same as the processing in the reduction unit 123 described in Embodiment 1, except that the size of the region is different.
[0091] Then, once the processing in steps S51 to S53 is completed for all major regions, the process proceeds to step S24. The processing in steps S24 and S25 in Figure 13 is the same as the processing in steps S24 and S25 in Figure 8.
[0092] As described above, according to Embodiment 3, by generating a large region that encompasses other regions, it is possible to generate object regions in large units. This makes it possible to reduce 3D Gaussian values in a way that is closer to human perception by focusing on semantic information such as the behavior of objects. In addition, since the level of detail can be unified on an object-by-object basis, changes in appearance can be made uniform.
[0093] In Embodiment 3 described above, the region generation unit 316 first generates a large region from regions classified using 3D Gaussian information, similar to the region generation unit 116 in Embodiment 1. However, Embodiment 3 is not limited to such examples. For example, the region generation unit 316 may generate regions using external data, similar to the region generation unit 216 in Embodiment 2. In this case, the region generation unit 316 identifies a large region by selecting two or more regions containing 3D Gaussian information classified by 3D Gaussian data, and at least one region corresponding to an object as the target region, and by executing a predetermined number of times the process of generating a new target region by classifying it according to the characteristics of the target region.
[0094] Embodiment 4. In embodiments 1 to 3 described above, the amount of 3D Gaussian reduction during rendering is determined according to the amount of movement, but embodiment 4 reduces 3D Gaussian more efficiently from the information of the virtual camera during rendering.
[0095] As shown in Figure 1, the free-viewpoint video generation device 400 according to Embodiment 4 comprises a communication unit 110, a 3D scene model generation unit 111, a storage unit 115, a region generation unit 116, an input unit 120, a virtual camera configuration unit 121, a reduction amount determination unit 422, a reduction unit 123, a rendering unit 124, and a display unit 125.
[0096] The communication unit 110, 3D scene model generation unit 111, storage unit 115, region generation unit 116, input unit 120, virtual camera configuration unit 121, reduction unit 123, rendering unit 124, and display unit 125 of the free-viewpoint video generation device 400 according to Embodiment 4 are the same as those of the communication unit 110, 3D scene model generation unit 111, storage unit 115, region generation unit 116, input unit 120, virtual camera configuration unit 121, reduction unit 123, rendering unit 124, and display unit 125 of the free-viewpoint video generation device 100 according to Embodiment 1.
[0097] The reduction amount determination unit 422 determines the reduction amount such that the amount of reduction increases as the magnitude of movement in each of the multiple regions viewed from the viewpoint in the free-viewpoint video increases.
[0098] For example, the reduction amount determination unit 422 determines the reduction amount of a region based on the relative movement between the virtual camera corresponding to the user's viewpoint and the region, using the virtual camera information indicated by the camera information, the 3D Gaussian information included in the 3D scene model, and the region information.
[0099] When a user moves in a virtual 3D space, the position of the virtual camera corresponding to the user's viewpoint also moves. Therefore, the amount of movement of the visual area may vary compared to the amount of movement of the actual area, depending on the position and amount of movement of the virtual camera.
[0100] For example, an object moving near a virtual camera may appear to be moving more than it actually is, while an object moving in the same direction and at the same speed as the virtual camera may appear to be stationary.
[0101] Specifically, as shown in Figures 14(A) and (B), the reduction amount determination unit 422 can calculate the angular velocities θ1 and θ2 of the region as seen from the virtual camera positions P1 and P2, and determine the reduction amount according to those angular velocities θ1 and θ2. Here, the larger the angular velocities θ1 and θ2, the greater the reduction can be.
[0102] Furthermore, as shown in Figure 15, if the virtual camera moves from position P3 to position P4, the reduction amount determination unit 422 can calculate the relative velocity, which is the relative speed between the virtual camera and the region, and determine the reduction amount according to that relative velocity. Here, the higher the relative speed, the greater the reduction.
[0103] For example, the reduction amount determination unit 422 can pre-determine thresholds for angular velocity or relative velocity according to the reduction amount stage, and then determine the reduction amount according to the calculated angular velocity or relative velocity.
[0104] As described above, Embodiment 4 can determine the amount of 3D Gaussian reduction by taking into account the relative movement between the virtual camera and the region, by also utilizing information from the virtual camera. Therefore, rendering speed can be increased and memory usage reduced without compromising the appearance due to viewpoint changes.
[0105] Embodiment 5. When performing clustering using 3D Gaussian information, it is possible to perform clustering of regions using the 3D Gaussian information all at once, or it is possible to perform clustering multiple times using multiple pieces of information contained in the 3D Gaussian information. Embodiment 5 shows an example of performing clustering multiple times.
[0106] As shown in Figure 1, the free-viewpoint video generation device 500 according to Embodiment 5 comprises a communication unit 110, a 3D scene model generation unit 111, a storage unit 115, a region generation unit 516, an input unit 120, a virtual camera configuration unit 121, a reduction amount determination unit 122, a reduction unit 123, a rendering unit 124, and a display unit 125.
[0107] The communication unit 110, 3D scene model generation unit 111, storage unit 115, input unit 120, virtual camera configuration unit 121, reduction amount determination unit 122, reduction unit 123, rendering unit 124, and display unit 125 of the free-viewpoint video generation device 500 according to Embodiment 5 are the same as those of the communication unit 110, 3D scene model generation unit 111, storage unit 115, input unit 120, virtual camera configuration unit 121, reduction amount determination unit 122, reduction unit 123, rendering unit 124, and display unit 125 of the free-viewpoint video generation device 500 according to Embodiment 1.
[0108] The region generation unit 516 adaptively generates regions according to the 3D scene. In Embodiment 5, the region generation unit 516 generates a 3D scene model by learning 3D Gaussian information from video data in order to generate a free-viewpoint image using 3DGS. In this model, it classifies the 3D Gaussian multiple times using the learned information at predetermined time intervals, thereby identifying multiple regions that include at least one region containing a 3D Gaussian that has been classified multiple times using that information. Here, the 3D Gaussian information consists of parameters such as the center point coordinates of the Gaussian distribution, rotation, covariance, density, and color information that depends on the viewpoint direction. The predetermined time is the time at which the frame group was captured.
[0109] Furthermore, the region generation unit 516 calculates the amount of movement, rotation, and density change of the 3D Gaussian, similar to the region generation unit 116 in Embodiment 1. As described above, 3D Gaussian data contains multiple types of information.
[0110] The region generation unit 516 generates a first region containing the 3D Gaussians classified by the first type of information by classifying the 3D Gaussians using the first type of information selected from the 3D Gaussians. Furthermore, the region generation unit 516 selects from the 3D Gaussian information and uses a second type of information, which is different from the first type of information, to classify the 3D Gaussians included in the first region. This process of generating a second region containing the 3D Gaussians classified by the second type of information is repeated until the nth region containing the 3D Gaussians classified by the nth type of information, which is different from any of the types of information selected up to the (n-1)th region, is generated (where n is an integer greater than or equal to 2). The unit generates multiple regions so that the nth region is included.
[0111] Specifically, the region generation unit 516 generates a first region by classifying 3D Gaussians with similar information into groups, using a first type of information selected from multiple types of information such as 3D Gaussian parameters, displacement, rotation, and density changes.
[0112] Furthermore, the region generation unit 516 selects from multiple types of information, such as 3D Gaussian parameters, displacement, rotation, and density changes, and uses a second type of information different from the first type of information to classify 3D Gaussians with similar information into groups, thereby generating a second region.
[0113] For example, as shown in Figure 16(A), the region generation unit 516 can calculate the amount of movement from the previous time t-1 at a certain time t using each of the 3D Gaussian values 530A to 530M.
[0114] Then, the region generation unit 516 generates one region 531 using 3D Gaussians 530A to 530I with similar displacement amounts, and generates another region 532 using 3D Gaussians 530J to 530M with similar displacement amounts.
[0115] Furthermore, the region generation unit 516 performs further classification in each of the classified regions 531 and 532 using a different type of information than the amount of movement. For example, as shown in Figure 17(A), the region generation unit 516 reclassifies the 3D Gaussians 530A to 530I contained in region 531 using the center point of the Gaussian distribution.
[0116] As a result, as shown in Figure 17(B), one region 531#1 is generated by 3D Gaussians 530A to 530D that are similar at the center point of the Gaussian distribution, and another region 531#2 is generated by 3D Gaussians 530E to 530I that are similar at the center point of the Gaussian distribution.
[0117] The region generation unit 516 performs the process of selecting and classifying different information from multiple types as described above, executing a predetermined n times (where n is an integer of 2 or more), and the regions classified by this process are designated as the final classified regions. Furthermore, the classification itself can be performed using well-known clustering methods, such as k-means.
[0118] The region generation unit 516 then stores the label information of the regions (clusters) that have been ultimately generated in this manner as region information in the storage unit 115.
[0119] In this example, the first clustering stage uses migration data, and the second clustering stage uses the center point of a Gaussian distribution. However, Embodiment 5 is not limited to such examples. For example, classification may be performed using one or more types of information at each stage. Furthermore, the number of stages is not limited to two, but may be three or more.
[0120] As described above, in Embodiment 5, by using multiple types of information and performing classification in multiple stages, it is possible to suppress the excessive influence of any one type of information, and more accurate clustering of objects can be expected.
[0121] In embodiments 1 to 5 described above, the rendering unit 124 displays the generated image on the display unit 125, but embodiments 1 to 5 are not limited to such examples. For example, the rendering unit 124 may transmit the generated image to another device via the communication unit 110. This allows the other device to display, store, etc., the rendered image. In other words, the rendering unit 124 can output the generated image (in this case, a free-viewpoint image) to an output unit that serves as an output interface, such as the display unit 125 or the communication unit 110. [Explanation of Symbols]
[0122] 100,200,300,400,500 Free viewpoint video generation device, 110,210 Communication unit, 111 3D scene model generation unit, 112 Video processing unit, 113 Initial point cloud generation unit, 114 3D scene configuration unit, 115 Memory unit, 116,216,316,516 Region generation unit, 120 Input unit, 121 Virtual camera configuration unit, 122,322,422 Reduction amount determination unit, 123 Reduction unit, 124 Rendering unit, 125 Display unit, 326 Action estimation unit.
Claims
1. In order to generate free-viewpoint video using 3DGS (3D Gaussian Splatting), a 3D scene model is generated by learning 3D Gaussian information from video data, and a region generation unit generates multiple regions that include at least one region containing the 3D Gaussian classified by the information, by classifying the 3D Gaussian into multiple dynamic objects using the information learned at predetermined time intervals. A reduction amount determination unit that determines the reduction amount for reducing the 3D Gaussian in each of the aforementioned multiple regions, In each of the aforementioned multiple regions, a reduction unit reduces the 3D Gaussian according to the reduction amount, The system includes a rendering unit that generates free-viewpoint images by rendering using the remaining 3D Gaussian that has not been reduced. A free-viewpoint video generation device characterized by the following.
2. The region generation unit generates at least one region corresponding to an object using external data that can identify an object contained in the video data, and includes the at least one region in the plurality of regions. A free-viewpoint video generation device according to claim 1, characterized by the above.
3. The region generation unit identifies the target region as the multiple regions by executing a predetermined number of times the process of classifying the multiple regions according to the characteristics of the target region and generating a new target region. A free-viewpoint video generation device according to claim 1, characterized by the above.
4. The region generation unit identifies the target region as the plurality of regions by executing a predetermined number of times the process of classifying the plurality of regions according to the characteristics of the target region and generating a new target region. A free-viewpoint video generation device according to claim 2, characterized by the above.
5. The region generation unit generates a first region containing 3D Gaussians classified by the first type of information by classifying 3D Gaussians using a first type of information selected from the information, and generates a second region containing 3D Gaussians classified by the second type of information by classifying 3D Gaussians included in the first region using a second type of information selected from the information but different from the first type of information, and repeats this process until an nth region containing 3D Gaussians classified by an nth type of information different from any of the types of information selected up to the (n-1)th is generated (n is an integer of 2 or more), thereby generating a plurality of regions so as to include the nth region. A free-viewpoint video generation device according to claim 1, characterized by the above.
6. The reduction amount determination unit determines the reduction amount such that it increases as the amount of movement of each of the multiple regions increases. A free-viewpoint video generation device according to any one of claims 1 to 5, characterized by the above.
7. The reduction amount determination unit determines the reduction amount such that the amount increases as the magnitude of movement in each of the multiple regions, as viewed from the viewpoint in the free-viewpoint video, increases. A free-viewpoint video generation device according to any one of claims 1 to 5, characterized by the above.
8. Computers, In order to generate free-viewpoint video using 3DGS (3D Gaussian Splatting), a 3D scene model is generated by learning 3D Gaussian information from video data, and a region generation unit generates multiple regions that include at least one region containing the 3D Gaussian classified by the information, by classifying the 3D Gaussian into multiple dynamic objects using the information learned at predetermined time intervals. A reduction amount determination unit determines the reduction amount for reducing the 3D Gaussian in each of the aforementioned multiple regions. A reduction unit that reduces the 3D Gaussian in each of the aforementioned multiple regions according to the reduction amount, and By using the remaining 3D Gaussian that has not been reduced, the rendering unit can function to generate free-viewpoint images. A program characterized by the following.
9. In order to generate free-viewpoint video using 3DGS (3D Gaussian Splatting), a 3D scene model is generated by learning 3D Gaussian information from video data, and the 3D Gaussian is classified into multiple dynamic objects using the information learned at predetermined time intervals, thereby generating multiple regions that include at least one region containing the 3D Gaussian classified by the information. In each of the aforementioned multiple regions, the reduction amount for reducing the 3D Gaussian is determined. In each of the aforementioned multiple regions, the 3D Gaussian is reduced according to the reduction amount. By rendering using the remaining 3D Gaussian that has not been reduced, a free-viewpoint image can be generated. A free-viewpoint video generation method characterized by the following.
Citation Information
Patent Citations
Low-bit-rate free viewpoint video generation method for efficient streaming based on 3D Gaussian, computer equipment, readable storage medium and program product
CN118573978A