Image processing device, image processing method, and program
By aligning virtual cameras with real-space image capture devices, the image processing device ensures accurate reconstruction of 3D models by generating and encoding depth and texture images that maintain the original model's shape and color fidelity.
Patent Information
- Application Number
- JP2024073062
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-26
- Publication Date
- 2025-11-07
AI Technical Summary
The quality of 3D models reconstructed on receiving devices can be degraded due to improper positioning and orientation of virtual cameras, especially when the 3D model shape is complex, leading to discrepancies between the original and reconstructed models.
The image processing device sets the positions and orientations of second virtual cameras based on the optical axes of first virtual cameras corresponding to real-space image capture devices, generating depth and texture images that accurately represent the 3D model, and encodes these images for distribution to ensure accurate reconstruction.
This approach reduces the likelihood of quality degradation in reconstructed 3D models by ensuring that the depth and texture images faithfully represent the original model, maintaining shape and color accuracy.
Smart Images

Figure 2025167990000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an image processing apparatus for compressing and distributing 3D model data representing a 3D model. [Background technology]
[0002] There is a technology that generates 3D model data showing a 3D model of a subject (object) using multiple captured images taken by multiple imaging devices, and generates a virtual viewpoint image seen from a camera (virtual camera) virtually placed in a virtual space where the 3D model exists.Recently, a technology that generates 3D model data of a subject on a server, compresses and distributes the 3D model data, and allows a user to operate a virtual camera on their own local terminal (client) such as a PC or tablet to display a virtual viewpoint image has been attracting attention.
[0003] In a technology for compressing and distributing 3D model data, a distributing device arranges multiple virtual cameras (or a group of virtual cameras) to surround a 3D model and generates depth images and texture images from each of the virtual cameras. The generated depth and texture images are then compressed, and the compressed depth and texture images and information indicating the positions and orientations of the multiple virtual cameras are distributed to the user. The receiving device decodes the compressed depth and texture images, and reconstructs a 3D model based on the decoded depth and texture images and information indicating the positions and orientations of the multiple virtual cameras.
[0004] Patent Document 1 describes a method for determining the viewpoints of a group of virtual cameras so as to reduce fluctuations in the positions of main objects in depth images between frames, with the aim of improving the compression rate of depth images. By reducing the number of motion vectors included in the encoded stream when encoding depth images, it is expected that the compression rate of depth images will improve. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] International Publication No. 2018 / 079260 Summary of the Invention [Problem to be solved by the invention]
[0006] However, depending on the positions and orientations of the virtual cameras set on the transmitting device, there is a risk that the quality of the 3D model reconstructed on the receiving device may be degraded. For example, if the shape of the 3D model generated on the transmitting device is complex, depending on the positions and orientations of the virtual cameras, the complex shape of the 3D model may not be sufficiently represented in the depth image, and the shape accuracy of the 3D model reconstructed on the user side may be degraded. In other words, the shape of the 3D model reconstructed on the receiving side may differ significantly from the shape of the 3D model generated on the transmitting side. Therefore, in order to reconstruct the 3D model more accurately, it is necessary to appropriately determine the positions and orientations of the virtual cameras.
[0007] Therefore, an object of the present disclosure is to reduce the possibility of a decrease in the quality of a 3D model generated on a receiving device. [Means for solving the problem]
[0008] In order to solve the above problems, the image processing device according to the present disclosure has the following configuration: a setting unit that sets the positions and orientations of a plurality of second virtual cameras based on the optical axes of a plurality of first virtual cameras in a virtual space that correspond to the positions and orientations of a plurality of image capture devices in real space, and a generation unit that generates a plurality of depth images that indicate the distance between each of the plurality of second virtual cameras and a 3D model of a subject that is generated based on a plurality of captured images acquired by the plurality of image capture devices. [Effects of the Invention]
[0009] According to the present disclosure, it is possible to reduce the possibility of a decrease in the quality of a 3D model generated on a receiving device. [Brief explanation of the drawings]
[0010] [Figure 1] 1A is a block diagram showing an example of the configuration of an image processing system 1, and FIG. 1B is a diagram showing an example of the hardware configuration of an image processing device. [Figure 2] 2 is a diagram illustrating an example of the functional configuration of a first image processing device 20 and a second image processing device 30. FIG. [Figure 3] (a) A diagram showing an example configuration of the photography system 10, (b) A diagram showing a 3D model 304 of a subject generated by a first virtual camera group 305 that corresponds to a physical camera group 302, when the first virtual camera group 305 is set, and (c) A diagram showing a second virtual camera group 307 that is set based on the viewpoint information of the first virtual camera group 305. [Figure 4] 10A to 10C are diagrams illustrating generation of viewpoint information of a virtual camera according to the first embodiment. [Figure 5] 10 is a flowchart illustrating an example of a process for compressing and distributing 3D model data according to the first embodiment. [Figure 6] 10 is a flowchart illustrating an example of a process for receiving 3D model data and generating a virtual viewpoint image according to the first embodiment. [Figure 7] 10 is a flowchart illustrating an example of a process for generating viewpoint information of a group of virtual cameras according to the first embodiment. [Figure 8] 10 is a flowchart illustrating an example of a restoration process of a 3D model according to the first embodiment. [Figure 9] 10A and 10B are diagrams illustrating generation of viewpoint information of a group of virtual cameras according to the second embodiment. [Figure 10] 10 is a flowchart illustrating an example of a process for compressing and distributing 3D model data according to the second embodiment. [Figure 11] 10 is a flowchart illustrating an example of a restoration process of a 3D model according to the second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0011] According to a preferred embodiment of the present invention, an image processing device includes a setting unit for setting the positions and orientations of multiple second virtual cameras based on the optical axes of multiple first virtual cameras in a virtual space corresponding to the positions and orientations of multiple image capture devices in real space. The image processing device also includes a generation unit for generating multiple depth images indicating the distance between each of the multiple second virtual cameras and a 3D model of a subject generated based on multiple captured images acquired by the multiple image capture devices. Here, the multiple image capture devices are assumed to be aligned with each other. The multiple first virtual cameras are set based on the positional relationship between the multiple image capture devices in real space. In other words, the multiple first virtual cameras in the virtual space are reproductions of the multiple image capture devices in real space.
[0012] According to this aspect, the positions and orientations of the multiple second virtual cameras are set based on the positions and orientations of the multiple imaging devices. Therefore, the subject represented in the depth images of the multiple second virtual cameras is similar to the subject included in the captured images acquired from the multiple imaging devices. Specifically, the shape of the area representing the subject included in the depth images generated from the second virtual cameras is similar to the shape of the area representing the subject included in the captured images. The 3D model of the subject generated by the distribution device is generated based on the multiple captured images acquired from the multiple imaging devices. Therefore, if the depth images used to reconstruct (restore) the 3D model of the subject on the receiving device are similar to the captured images, the 3D model reconstructed on the receiving device will also be similar to the 3D model generated on the distribution device. In other words, it is possible to reduce the possibility that the shape of the 3D model generated on the receiving device will differ from that of the 3D model generated on the distribution device, i.e., the reproducibility of the 3D model will be degraded.
[0013] The image processing device also includes an encoding unit that encodes the depth images. The image processing device also includes an output unit that outputs the encoded depth images and viewpoint information indicating the positions and orientations of the second virtual cameras to another device that reconstructs the 3D model. Here, the other device is a device that reconstructs the 3D model based on the encoded depth images and the viewpoint information.
[0014] In addition, in the image processing device, the generating means generates a virtual viewpoint image including a 3D model of the subject for each of the second virtual cameras, the encoding means encodes the virtual viewpoint images, and the output means outputs the encoded virtual viewpoint images to the other device.
[0015] This configuration makes the colors of the components of the subject included in the virtual viewpoint image (texture image) generated by the second virtual camera closer to the colors of the components of the subject included in the captured image, reducing the possibility that the 3D model generated on the receiving device will have different colors from the 3D model generated on the transmitting device, i.e., reducing the possibility that the reproducibility of the 3D model will be degraded.
[0016] The generating means generates correspondence information indicating whether each pixel in the texture image corresponds to a component of the 3D model of the subject for each of the second virtual cameras, and the encoding means encodes the correspondence information, and the output means outputs the encoded correspondence information to the other device.
[0017] This allows the receiving device to identify pixels in the texture image that will be used to color the 3D model of the subject. For example, if a subject is occluded by another subject in a captured image acquired from an imaging device, the captured image is used to generate a texture image for the second virtual camera, which may result in a texture image having color information different from the color information of the subject. Therefore, by generating correspondence information that indicates areas where occlusion is not considered to occur, the possibility of different colors being generated when reconstructing a 3D model can be reduced.
[0018] Moreover, the second virtual cameras are set at positions closer to the 3D model of the subject than the first virtual cameras.
[0019] The generating means generates one depth image for one of the second virtual cameras, and when delivering a 3D model showing movement over time, generates one depth image for one of the second virtual cameras for each of a plurality of times.
[0020] According to another preferred embodiment of this embodiment, the image processing device includes an acquisition unit that acquires a plurality of encoded depth images and first viewpoint information indicating the positions and orientations of a plurality of second virtual cameras. Here, the depth images are images that indicate the distance between each of the plurality of second virtual cameras and a 3D model of the subject generated based on a plurality of captured images acquired by the imaging device. The plurality of second virtual cameras are virtual cameras set based on the optical axes of a plurality of first virtual cameras in a virtual space that correspond to the positions and orientations of the plurality of imaging devices in real space. The image processing device also includes a decoding unit that decodes the encoded depth images. The image processing device also includes a generation unit that generates a 3D model of the subject based on the decoded depth images and the first viewpoint information.
[0021] The image processing device also acquires second viewpoint information indicating the position and orientation of a third virtual camera different from the first virtual camera and the second virtual camera. The image processing device also generates a virtual viewpoint image based on the second viewpoint information and a 3D model of the subject. The third virtual camera may be set by a user operating an input device such as a joystick, or may be set based on the position of the generated 3D model.
[0022] This aspect allows a virtual viewpoint image to be generated using a 3D model generated on the receiving device.
[0023] According to another preferred embodiment of the present invention, a program causes a computer to execute the functions of the image processing device described above. By executing this program, the computer preferably functions as the image processing device described above.
[0024] <Example> Preferred embodiments of the present disclosure will be described in detail below with reference to the drawings. Note that the following embodiments do not limit the present disclosure, and not all of the combinations of features described in the embodiments are necessarily essential to the solutions of the present disclosure. Furthermore, in the accompanying drawings, the same reference numerals are used to designate the same or similar components, and redundant descriptions will be omitted.
[0025] A virtual viewpoint image is an image generated by a user freely manipulating the position and orientation of a virtual camera, and is also called a free viewpoint image, an arbitrary viewpoint image, etc. Unless otherwise specified, the term "image" will be explained as including the concepts of both moving images and still images.
[0026] The image processing system 1 is a system that generates a virtual viewpoint image representing a scene from a specified virtual viewpoint based on multiple images captured by multiple imaging devices and a specified virtual viewpoint. The virtual viewpoint image in this embodiment is also called a free viewpoint video, but is not limited to an image corresponding to a viewpoint freely (arbitrarily) specified by a user. For example, the virtual viewpoint image also includes an image corresponding to a viewpoint selected by a user from multiple candidates. In addition, although the present embodiment will mainly describe a case where the virtual viewpoint is specified by a user operation, the virtual viewpoint may also be specified automatically based on the results of image analysis, etc. In addition, in the present embodiment, the present embodiment will mainly describe a case where the virtual viewpoint image is a video, but the virtual viewpoint image may also be a still image.
[0027] The viewpoint information used to generate a virtual viewpoint image is information indicating the position and orientation (line of sight direction) of the virtual viewpoint. Specifically, the viewpoint information is a parameter set including a parameter indicating the three-dimensional position of the virtual viewpoint and a parameter indicating the orientation of the virtual viewpoint in the pan, tilt, and roll directions. Note that the content of the viewpoint information is not limited to the above. For example, the parameter set serving as viewpoint information may include a parameter indicating the size of the field of view (angle of view) of the virtual viewpoint. Furthermore, the viewpoint information may have multiple parameter sets. For example, the viewpoint information may have multiple parameter sets corresponding to multiple frames constituting a moving image of the virtual viewpoint image, and may be information indicating the position and orientation of the virtual viewpoint at each of multiple consecutive time points.
[0028] The image processing system 1 has multiple imaging devices that capture images of an imaging area from multiple directions. The imaging area may be, for example, a stadium where sports such as soccer or karate are held, or a stage where a concert or a play is held. The multiple imaging devices are installed at different positions surrounding the imaging area and capture images synchronously. Note that the multiple imaging devices do not need to be installed around the entire periphery of the imaging area; depending on installation space restrictions, they may be installed only around a portion of the periphery of the imaging area. Furthermore, the number of imaging devices is not limited to the example shown in the figure. For example, if the imaging area is a soccer stadium, approximately 30 imaging devices may be installed around the stadium. Furthermore, imaging devices with different functions, such as telephoto cameras and wide-angle cameras, may be installed.
[0029] In this embodiment, the multiple image capturing devices are cameras each having an independent housing and capable of capturing images from a single viewpoint. However, this is not limiting, and two or more image capturing devices may be configured in the same housing. For example, a single camera equipped with multiple lens groups and multiple sensors and capable of capturing images from multiple viewpoints may be installed as the multiple image capturing devices.
[0030] A virtual viewpoint image is generated, for example, by the following method. First, multiple images (multiple captured images) are obtained by capturing images from different directions using multiple imaging devices. Next, a foreground image in which a foreground region corresponding to a predetermined object, such as a person or a ball, is extracted, and a background image in which a background region other than the foreground region is extracted are obtained from the multiple captured images. Furthermore, a foreground model representing the three-dimensional shape of the predetermined object and texture data for coloring the foreground model are generated based on the foreground image, and texture data for coloring a background model representing the three-dimensional shape of a background, such as a stadium, is generated based on the background image. Then, the texture data is mapped to the foreground model and background model, and rendering is performed according to the virtual viewpoint indicated by the viewpoint information, thereby generating a virtual viewpoint image. However, the method for generating a virtual viewpoint image is not limited to this, and various methods can be used, such as a method of generating a virtual viewpoint image by projective transformation of captured images without using a three-dimensional model.
[0031] A foreground image is an image in which an object region (foreground region) is extracted from an image captured by an imaging device. An object extracted as a foreground region is a dynamic object (moving body) that moves (its absolute position and shape can change) when images are captured from the same direction in a time series. Examples of objects include players, referees, and other people on the field where a sport is being played, such as a ball in a ball game, or singers, musicians, performers, and presenters in a concert or entertainment event.
[0032] A background image is an image of at least a region (background region) different from the foreground object. Specifically, a background image is an image in which the foreground object has been removed from the captured image. Furthermore, the background refers to an imaged object that remains stationary or nearly stationary when images are captured from the same direction in chronological order. Examples of such imaged objects include a stage for a concert, a stadium where an event such as a sport is held, a structure such as a goal used in a ball game, or a field. However, the background is at least a region different from the foreground object, and the imaged object may include other objects in addition to the object and background.
[0033] A virtual camera is a virtual camera that is different from the multiple imaging devices actually installed around the imaging area, and is a concept for conveniently explaining a virtual viewpoint related to the generation of a virtual viewpoint image. That is, a virtual viewpoint image can be considered to be an image captured from a virtual viewpoint set in a virtual space associated with the imaging area. The position and orientation of the viewpoint in the virtual image capture can be expressed as the position and orientation of the virtual camera. In other words, a virtual viewpoint image can be said to be an image that simulates an image captured by a camera if it were assumed that the camera were located at the virtual viewpoint set in space.
[0034] Example 1 In this embodiment, a process for determining the viewpoints of a group of virtual cameras for generating depth images and texture images to be distributed from a server to a client based on viewpoint information of a physical camera will be described. Note that in this embodiment, the receiving device will be referred to as the client. Also, in this embodiment, the imaging device arranged in real space will be referred to as the physical camera.
[0035] <Image processing system hardware configuration> FIG. 1(a) is a diagram showing an example of the overall configuration of an image processing system 1 according to this embodiment.
[0036] The image processing system 1 includes an imaging system 10, a first image processing device 20, a second image processing device 30, an input device 40, and a display device 50. The image processing system 1 generates shape information of a 3D model of an object using images (plurality of captured images) captured by a plurality of physical cameras. The image processing system 1 then generates viewpoint information (first viewpoint information) of a second virtual camera group for generating depth images and texture images to be distributed from a server to a client. Furthermore, the image processing system 1 generates and compresses depth images and texture images based on the generated viewpoint information of the second virtual camera group, and distributes these images and information necessary to reconstruct (restore) the 3D model to a user as 3D model data.
[0037] The photography system 10 arranges multiple physical cameras at different positions and synchronously photographs a subject (object) from multiple viewpoints. Then, multiple captured images acquired by synchronous photography, viewpoint information (external / internal parameters, image size) for each physical camera of the photography system 10, and the like are transmitted to the first image processing device 20. The external parameters of a camera are parameters indicating the position and orientation of the camera (e.g., a rotation matrix and a position vector). The internal parameters of a camera are internal parameters specific to the camera, such as focal length, image center, and lens distortion parameters. The external and internal parameters of a camera are collectively referred to as camera parameters. The image size is the width and height of the image. The viewpoint information of this group of physical cameras is used when determining viewpoint information (camera parameters and image size) of a group of second virtual cameras for generating depth images and texture images.
[0038] The first image processing device 20 generates a 3D model of a foreground object based on multiple captured images input from the imaging system 10 and viewpoint information for each physical camera. The foreground object (subject) is, for example, a person or a moving object present within the imaging range of the imaging system 10. The first image processing device 20 then generates viewpoint information of a second virtual camera group for generating depth images and texture images. Furthermore, the first image processing device 20 generates and compresses depth images and texture images based on the generated viewpoint information of the second virtual camera group, and outputs the compressed images and the 3D model together with information (metadata) required to restore the 3D model to the second image processing device 30. The metadata is, for example, the generated viewpoint information of the second virtual camera group.
[0039] First, the first image processing device 20 estimates shape information (shape information) of the subject based on multiple captured images and viewpoint information for each physical camera. Estimating the shape information of the subject uses, for example, a visual hull intersection method. Multiple virtual cameras (multiple first virtual cameras) are set in a virtual space corresponding to the imaging system 10 existing in real space. These multiple first virtual cameras are referred to as a first virtual camera group. This first virtual camera group is a reproduction of a physical camera group in virtual space. In other words, the viewpoint information of the first virtual camera group corresponds to the viewpoint information of the physical camera group in real space. Then, shape information of the subject is generated by using the viewpoint information of the first virtual camera group and the multiple captured images in the visual hull intersection method. As a result of this processing, a 3D point cloud (a set of points having three-dimensional coordinates) representing the shape information of the subject is obtained. This 3D point cloud is also called a visual hull. Note that the method for deriving shape information of the subject from the captured images is not limited to this. Furthermore, the method for representing the shape information of the subject is not limited to a 3D point cloud, and meshes or voxels may also be used. Furthermore, a texture image (color information) is determined for each point in the generated 3D point cloud of the subject using multiple captured images. Therefore, 3D model data representing a 3D model of the subject includes shape information indicating the shape and color information indicating the color. In this embodiment, the 3D model is generated from shape information and color information, but this is not limited to this. For example, the 3D model includes shape information and may be managed as data separate from color information. Note that a method for setting a first virtual camera group in a virtual space corresponding to a physical camera group in a real space is a well-known technique in the field of CG, etc., and therefore will not be described here.
[0040] Next, the first image processing device 20 generates viewpoint information (positions and orientations) of multiple second virtual cameras for generating depth images and texture images to be distributed. Hereinafter, the multiple second virtual cameras are referred to as a second virtual camera group. The positions of the second virtual cameras are, for example, arranged on the optical axes of the first virtual cameras, and the orientations are set to be the same as those of the first virtual cameras. Since the viewpoint information of the first virtual cameras corresponds to the viewpoint information of the physical cameras in real space, the viewpoint information of the second virtual cameras is generated based on the viewpoint information of the physical cameras. A method for generating the viewpoint information of the second virtual cameras will be described in detail later. Hereinafter, the viewpoint information of the second virtual cameras is referred to as second viewpoint information. The first image processing device 20 generates depth images of the 3D model based on the second viewpoint information of the second virtual cameras. Depth images are generated for the number of second virtual cameras. Specifically, each point in the 3D point cloud of the subject is projected onto the same plane as the imaging plane of the second virtual cameras. For each second virtual camera group, the distance (depth) from the second virtual camera to the subject is calculated for each projected pixel, and a depth value is set for each pixel of the depth image. The first image processing device 20 also captures an image of the subject based on the second viewpoint information of the second virtual camera group and generates a texture image. The texture image is generated by blending the colors of multiple captured images while increasing the priority of pixel values of images captured by a physical camera with a line of sight close to the line of sight of the second virtual camera. This method of setting a high priority (weight) for pixel values of images captured by a physical camera with a line of sight close to the line of sight of the second virtual camera when blending multiple captured images is referred to as a virtual viewpoint-dependent texture image generation method. Details of the virtual viewpoint-dependent texture image generation method will be described later. This virtual viewpoint-dependent texture image is generated by selecting an image captured by a physical camera that determines the pixel value of the subject depending on the position and orientation of the virtual camera, so the color of the subject changes when the virtual camera moves. Therefore, a texture image generated using this method is referred to as a virtual viewpoint-dependent texture image. On the other hand, a virtual viewpoint-independent texture image is generated by a method in which the pixel values of the object do not change depending on the position and orientation of the virtual camera.In this embodiment, the generation of texture images is handled as a virtual viewpoint dependent subject, but texture images may be generated independently of a virtual viewpoint. An example of a process for generating texture images that are virtual viewpoint dependent and virtual viewpoint independent is shown below.
[0041] The process of generating a virtual viewpoint-dependent texture image includes, for example, a visibility determination process for points in a 3D point cloud that constitutes the subject and a color derivation process based on the position and orientation of a virtual camera. In the visibility determination process, a physical camera capable of capturing an image of each point in the 3D point cloud is identified based on the positional relationship between each point in the 3D point cloud and multiple physical cameras included in the physical camera group of the imaging system 10. In the color derivation process, for example, a point in the 3D point cloud is designated as a focus point, and the color of the focus point is derived. Specifically, the following process is performed for each focus point. A focus point within the imaging range of a second virtual camera is selected. Then, a first virtual camera is selected that has a line of sight similar to the line of sight of the second virtual camera whose imaging range includes the focus point and can capture the focus point. Because the first virtual camera is a reproduction of a physical camera in a virtual space, selecting the first virtual camera is the same as selecting a physical camera. The selected focus point is then projected onto the image captured by the selected physical camera. The color of the pixel at the projection destination is designated as the color of the focus point. The physical camera is selected, for example, based on whether the angle between the line of sight from the second virtual camera to the point of interest and the line of sight from the physical camera to the point of interest is equal to or less than a certain angle. If the point of interest can be captured by multiple physical cameras, multiple physical cameras with line of sight directions close to the line of sight of the second virtual camera are selected, and the point of interest is projected onto each of the images captured by those physical cameras. The pixel values of the projection destination are then acquired, and a weighted average is calculated so that pixel values from physical cameras with line of sight directions close to the line of sight of the second virtual camera are prioritized, thereby determining the color of the point of interest. This process is performed while changing the point of interest, and the color of the point of interest is projected onto the same surface as the imaging surface of the second virtual camera, thereby generating a texture image dependent on the virtual viewpoint. While the present embodiment has been described above, the present invention is not limited to the above. To generate a texture image dependent on a virtual camera other than the second virtual camera, the second virtual camera is replaced with the different virtual camera, and the above process is performed.
[0042] The process of generating a texture image independent of a virtual viewpoint includes, for example, the visibility determination process described above and a process of deriving a color independent of the position and orientation of a virtual camera. After the visibility determination process, for example, a point in a 3D point cloud is designated as a focus point, and the focus point is projected onto an image captured by a physical camera corresponding to a first virtual camera capable of capturing the image, and the color of the pixel at the projection destination is designated as the color of the focus point. Note that if the focus point can be captured by multiple first virtual cameras, the focus point is projected onto each of the images captured by the multiple physical cameras corresponding to the multiple first virtual cameras. The pixel value of the projection destination is then acquired and the average pixel value is calculated to determine the color of the focus point. This process is performed while changing the focus point, and the color of the focus point is projected onto the same surface as the imaging surface of the second virtual camera, thereby generating a texture image independent of a virtual viewpoint.
[0043] In virtual viewpoint image generation technology, a method of generating color information dependent on a virtual viewpoint is known as a method of generating images with higher image quality than a method independent of a virtual viewpoint, because pixel values of an image captured by a physical camera with a line of sight close to the line of sight of a second virtual camera are preferentially used. In addition, a virtual viewpoint-dependent virtual viewpoint image (texture image) generated by a virtual camera whose position and orientation are similar to those of a physical camera is heavily influenced by the image captured by the physical camera, making it possible to generate a texture image with image quality close to that of the captured image. On the other hand, a virtual viewpoint-dependent texture image generated from the viewpoint of a virtual camera whose position and orientation are significantly different from those of a physical camera, or a texture image generated independent of a virtual viewpoint, is generated by interpolating pixel values using multiple physical cameras. As a result, the resulting image may contain pixel values that differ from the color information of the actual subject, or may have low contrast.
[0044] Finally, the first image processing device 20 compresses (encodes) the depth image and texture image using a video compression method such as H.264 or H.265. The compression method is not limited to video compression, and any method that can encode data to a size smaller than the original data size, such as file compression, may be used. The first image processing device 20 does not need to output depth images and texture images that do not include a subject to the second image processing device 30. The second image processing device 30 back-projects the depth image into virtual space based on second viewpoint information from the second virtual camera, thereby enabling the reconstruction of shape information for a 3D model. Furthermore, by using pixel values of the texture image at the same coordinates as each depth value in the depth image as color information for the point where the depth value is back-projected, color information can be added to the shape information, thereby reconstructing a 3D model.
[0045] The first image processing device 20 may also extract a rectangular image of the subject's shooting range from the pre-compression depth image and texture image, compress this rectangular image (ROI image), and distribute it. In this case, the coordinate information of the extracted rectangular image may be included as metadata. The first image processing device 20 may also arrange rectangular images to form a single image, compress it, and distribute it. By distributing the rectangular image instead of the entire image, it is possible to reduce the amount of data.
[0046] Furthermore, the first image processing device 20 may generate depth images containing high-precision depth values, such as single-precision floating-point numbers (32 bits), which cannot be used for video compression in H.264 or H.265. In this case, the depth information is converted to a precision (8 bits or 10 bits) that allows video compression before video compression. The conversion method may involve, for example, scalar quantization, and the quantized depth images may be compressed and distributed. In this case, the minimum and maximum values of the depth range before quantization may be included as metadata. By performing scalar quantization, the client can reconstruct a 3D model of the subject or subject within the shooting range, which would otherwise be insufficiently precise for video compression.
[0047] The second image processing device 30 receives and decodes the depth image, texture image, and metadata from the first image processing device 20, and restores a 3D model. As described above, the 3D model is restored by back-projecting the depth image and texture image into the virtual space based on the second viewpoint information of the second virtual camera group included in the metadata. The second image processing device 30 also calculates viewpoint information (second viewpoint information) of a third virtual camera for generating a virtual viewpoint image viewed by the user based on input values received from an input device 40 (described later). The second image processing device 30 then generates a virtual viewpoint image based on the calculated second viewpoint information and the restored 3D model. The second image processing device 30 then outputs the generated virtual viewpoint image to the display device 50.
[0048] The input device 40 accepts input values for the user to set the third virtual camera and transmits the input values to the second image processing device 30. For example, the input device 40 has an input unit such as a joystick, a jog dial, a touch panel, a keyboard, and a mouse. The user setting the third virtual camera sets the position and orientation of the third virtual camera by operating the input unit. Note that in this embodiment, the user sets the position and orientation of the third virtual camera, but this is not limited thereto, and the position and orientation of the third virtual camera may be set using position information of a 3D model. The position information of the 3D model here may be generated by the first image processing device 20 on the distribution side or the second image processing device 30 on the receiving side.
[0049] The display device 50 displays the virtual viewpoint image generated and output by the second image processing device 30. The user looks at the virtual viewpoint image displayed on the display device 50 and sets the position and orientation of the virtual camera for the next frame via the input device 40.
[0050] 1(b) is a diagram showing an example of the hardware configuration of the first image processing device 20 of this embodiment. The hardware configuration of the second image processing device 30 is also the same as the configuration of the first image processing device 20 described below. The first image processing device 20 of this embodiment is composed of a CPU 101, a RAM 102, a ROM 103, and a communication unit 104.
[0051] The CPU 101 controls the entire first image processing device 20 using computer programs and data stored in the RAM 102 and ROM 103, thereby realizing each function of the first image processing device 20 shown in FIG. 1. Note that the first image processing device 20 may have one or more dedicated hardware components different from the CPU 101, and at least a part of the processing by the CPU 101 may be executed by the dedicated hardware components. Examples of the dedicated hardware components include an ASIC (application-specific integrated circuit), an FPGA (field-programmable gate array), and a DSP (digital signal processor). The RAM 102 temporarily stores programs and data supplied from the auxiliary storage device 214, and data supplied from the outside via the communication unit 104. The ROM 103 stores programs and the like that do not require modification.
[0052] The communication unit 104 is used for communication with an external device of the first image processing device 20. For example, when the first image processing device 20 is connected to the external device via a wired connection, a communication cable is connected to the communication unit 104. When the first image processing device 20 has a function of wirelessly communicating with the external device, the communication unit 104 includes an antenna.
[0053] <Functional configuration of image processing device> FIG. 2 is a diagram showing an example of the functional configuration of the first image processing device 20 and the second image processing device 30. As shown in FIG.
[0054] The first image processing device 20 includes a shape information generating unit 201 , a viewpoint determining unit 202 , a depth image generating unit 203 , a texture image generating unit 204 , an encoding unit 205 , and a distribution unit 206 .
[0055] The shape information generation unit 201 estimates shape information of the subject using the multiple captured images and viewpoint information of the physical cameras received from the imaging system 10, using the communication unit 104. The shape information is estimated using the volume intersection method described above. Therefore, the shape information generation unit 201 also generates viewpoint information of a first virtual camera group in virtual space that corresponds to the physical camera group in real space, from the acquired viewpoint information of the physical cameras. The shape information generation unit 201 outputs the estimated shape information to the viewpoint determination unit 202, the depth image generation unit 203, and the texture image generation unit 204. The shape information generation unit 201 also outputs the received viewpoint information of the physical cameras and the multiple captured images to the viewpoint determination unit 202 and the texture image generation unit 204.
[0056] The viewpoint determination unit 202 determines the viewpoints of the second virtual cameras and generates viewpoint information based on the input data from the shape information generation unit 201, with the aim of improving the quality of the 3D model restored by the second image processing device 30. The viewpoint determination unit 202 outputs the generated viewpoint information of the second virtual cameras to the depth image generation unit 203 and to the texture image generation unit 204 via the depth image generation unit 203.
[0057] The viewpoint information of the second virtual cameras is generated, for example, by matching the position and orientation of the second virtual camera to that of the first virtual camera corresponding to the physical camera and dolly-zooming the second virtual camera in the line of sight of the first virtual camera. Specifically, the second virtual camera, placed at the position of the first virtual camera corresponding to the physical camera, approaches the subject in the line of sight of the first virtual camera while adjusting the focal length (zoom) so that the size of the subject in the virtual viewpoint image generated from the second virtual camera is maintained. This process of deriving the viewpoint information of the second virtual camera will be described in detail below with reference to Figures 3 and 4. By performing this process while changing the physical camera, or while changing the subject if there are multiple subjects in the captured image, it is possible to generate viewpoint information of the second virtual camera for each subject, for the number of physical cameras that captured the subject.
[0058] The viewpoint determination unit 202 may also change the image size (width and height) of the second virtual camera. First, the image size of the second virtual camera is set to the same as the image size of the physical camera. After determining the position and orientation of the second virtual camera as described above, the image size of the second virtual camera is changed. For example, when the width and height of the virtual viewpoint image of the second virtual camera are changed to 1 / N of the width and height of the image captured by the physical camera, the focal length and image center of the internal parameters of the second virtual camera are also multiplied by 1 / N. In other words, when the image size of the physical camera is 4K (width 3840, height 2160) and the image size of the second virtual camera is changed to full HD (width 1920, height 1080), the width and height are each halved, the focal length and image center of the internal parameters of the second virtual camera are multiplied by 1 / 2. This allows the image size to be changed without changing the angle of view, thereby reducing the amount of data transmitted from the server to the client.
[0059] The depth image generation unit 203 performs the above-described depth image generation process based on the shape information of the subject input from the shape information generation unit 201 and the viewpoint information of the second virtual camera group input from the viewpoint determination unit 202. The depth image generation unit 203 outputs the generated depth image and the viewpoint information of the second virtual camera group to the encoding unit 205.
[0060] The texture image generation unit 204 performs the process of generating the virtual viewpoint-dependent texture image described above, based on input data from the shape information generation unit 201 and viewpoint information of the second virtual camera group input from the viewpoint determination unit 202 via the depth image generation unit 203. The texture image generation unit 204 outputs the generated texture image to the encoding unit 205. As described above, the virtual viewpoint-dependent texture image is generated by preferentially referencing pixel values of an image captured by a physical camera that is close in position and orientation to the second virtual camera. This makes it possible to generate high-quality texture images, and when the 3D model is restored on the client, a 3D model containing high-quality color information can be reproduced.
[0061] Furthermore, the texture image generation unit 204 may generate an effective pixel map based on input data from the shape information generation unit 201 and viewpoint information of the second virtual camera group input from the viewpoint determination unit 202 via the depth image generation unit 203. The texture image generation unit 204 may output the generated effective pixel map to the encoding unit 205. The effective pixel map will be described in detail in Example 2.
[0062] The encoding unit 205 acquires the depth images and viewpoint information of the second virtual camera group input from the depth image generation unit 203, and the texture images input from the texture image generation unit 204. The encoding unit 205 compresses the depth images and texture images using the compression method described above, and outputs the compressed image group and viewpoint information (metadata) of the second virtual camera group to the distribution unit 206.
[0063] The encoding unit 205 may compress not only the depth image and the texture image but also the metadata using the above-mentioned file compression method.
[0064] Furthermore, the encoding unit 205 may compress the effective pixel map input from the texture image generation unit 204 and output it to the distribution unit 206 .
[0065] The distribution unit 206 uses the communication unit 104 to transmit the compressed depth image, texture image, and viewpoint information of the second virtual camera group input from the encoding unit 205 to the reception unit 207, which will be described later.
[0066] The second image processing device 30 includes a receiving unit 207 , a decoding unit 208 , a 3D model restoring unit 209 , a virtual camera control unit 210 , and a virtual viewpoint image generating unit 211 .
[0067] The receiving unit 207 receives the compressed depth image, the compressed texture image, and the viewpoint information (metadata) of the second virtual camera group from the distribution unit 206 using the communication unit 104, and outputs them to the decoding unit 208.
[0068] The decoding unit 208 decodes the compressed depth images and compressed texture images acquired from the receiving unit 207 and outputs them to the 3D model restoration unit 209 together with viewpoint information of the second virtual camera group. The decoding unit 208 may also decode metadata in addition to the depth images and texture images. Furthermore, the decoding unit 208 may decode an effective pixel map and output it to the 3D model restoration unit 209.
[0069] The 3D model restoration unit 209 restores a 3D model using the restoration method described above, based on the decoded depth image, the decoded texture image, and the viewpoint information of the second virtual camera group acquired from the decoding unit 208. The restored 3D model is output to the virtual viewpoint image generation unit 211. Furthermore, the 3D model restoration unit 209 may generate color information of the 3D model using the effective pixel map acquired from the decoding unit 208. A method for using the effective pixel map will be described in detail in Example 2.
[0070] The virtual camera control unit 210 uses the communication unit 104 to generate viewpoint information of a third virtual camera for generating a virtual viewpoint image from input values input by the user via the input device 40, and outputs the viewpoint information of the third virtual camera to the virtual viewpoint image generation unit 211. The virtual camera control unit 210 may also output the generated viewpoint information of the third virtual camera designated by the user to the 3D model restoration unit 209.
[0071] The virtual viewpoint image generation unit 211 generates a virtual viewpoint image based on the 3D model acquired from the 3D model restoration unit 209 and viewpoint information of the third virtual camera acquired from the virtual camera control unit 210. The virtual viewpoint image is generated by arranging a 3D model of the subject, a 3D model of the background object, and the third virtual camera in a virtual space and generating an image viewed from the third virtual camera. The 3D model of the background object is, for example, a CG (Computer Graphics) model created separately to be combined with the subject, and is created in advance and stored in the second image processing device 30 (for example, stored in the ROM 103 of FIG. 1). The 3D model of the subject and the 3D model of the background object are rendered using an existing CG rendering method. The virtual viewpoint image generation unit 211 transmits the generated virtual viewpoint image to the display device 50.
[0072] <Description of an example of generating viewpoint information of a group of second virtual cameras based on viewpoint information of a group of physical cameras> Here, a method for generating viewpoint information of the second virtual camera group using viewpoint information of the physical camera group will be described with reference to Figures 3 and 4. Figure 3 is a schematic diagram for explaining an example of a method for generating viewpoint information of the second virtual camera group.
[0073] 3(a) is a diagram showing a group of physical cameras 302 arranged in real space and a subject 301. The subject 301 is photographed by the group of physical cameras 302. Note that a line 303 indicates the optical axis of each of the group of physical cameras 302.
[0074] FIG. 3(b) is a diagram showing a 3D model 304 of a subject generated by a first virtual camera group 305 that is set corresponding to the physical camera group 302. The first image processing device 20 generates the 3D model 304 of the subject using a plurality of captured images acquired from the physical camera group 302. To generate a 3D model 304 in virtual space for a subject 301 in real space, the first virtual camera group 305 is generated in virtual space that corresponds to the physical camera group 302 in real space. In other words, the first virtual camera group 305 is a reproduction in virtual space of the physical camera group 302 in real space. Therefore, the optical axis 306 of the first virtual camera group 305 corresponds to the optical axis 303 of the physical camera group 302. The first image processing device 20 then generates the 3D model 304 of the subject 301 using viewpoint information of the generated first virtual camera group 305.
[0075] FIG. 3(c) is a diagram showing a second virtual camera group 307 set based on viewpoint information of the first virtual camera group 305. The second virtual camera group 307 is set on the optical axis 306 of the first virtual camera group 305. The second virtual camera group 307 may be set at a position closer to the optical axis 306. Furthermore, the second virtual camera group 307 is set at a position closer to the 3D model 304 than the first virtual camera group 305. Generally, the second virtual camera group 307 for generating depth images and texture images for distribution places a bounding box that encompasses the 3D model 304, and places the second virtual camera group 307 on a spherical surface surrounding the bounding box so as to face the subject. This determines the position and orientation (extrinsic parameters) of the second virtual camera group 307. Then, a focal length (intrinsic parameter) that allows the second virtual camera group 307 to capture the entire 3D model 304 is determined. These are set manually, while the number and spacing of the second virtual cameras are often determined heuristically. If the 3D model 304, i.e., the subject 301, has a complex shape, occlusion occurs, whereby a portion of the 3D model 304 is hidden by another portion as viewed from a certain second virtual camera. Therefore, the number and placement of the second virtual cameras 307 must be appropriately determined. For a subject with a complex shape, if the number of second virtual cameras is small and their placement is inappropriate, resulting in frequent occlusion, the 3D model restored by the second image processing device 30 may differ significantly in shape from the 3D model 304 before distribution. However, increasing the number of second virtual cameras to accurately restore the 3D model 304 increases the number of depth images and texture images to be distributed. Therefore, if a user attempts to restore a 3D model in a low-bandwidth environment or on a local terminal with low processing performance, the frame rate may decrease. To provide a high sense of realism to the user, it is desirable for the 3D model displayed on the display device 50 to have high image quality and a smooth frame rate, such as 60 fps. Therefore, it is necessary to appropriately set the number and arrangement of the second virtual camera group 307 and deliver depth images and high-quality texture images that can accurately reconstruct the 3D model 304 with a small amount of data. Therefore, the first image processing device 20 generates viewpoint information of the second virtual camera group 307 based on viewpoint information of the physical camera group 302 described above.That is, second virtual camera group 307 is placed at the position of first virtual camera group 305 corresponding to physical camera group 302. Then, the position of second virtual camera group 307 is set by dolly-zooming second virtual camera group 307 onto optical axis 306 of first virtual camera 305 while maintaining the size of the subject on the screen so as to approach 3D model 304 of subject 301. The position to which second virtual camera group 307 is dolly-zoomed is assumed to be determined in advance. For example, dolly-zooming may be performed until the distance from 3D model 304 reaches a predetermined value. Alternatively, dolly-zooming may be performed until the size of 3D model 304 projected on the imaging surface of second virtual camera group 307 reaches a predetermined size. Note that the method for determining the viewpoint information of second virtual camera group 307 is not limited to this method. Second virtual camera group 307 may simply be placed close to optical axis 306 of first virtual camera group 305, and internal parameters such as focal length and image size may be determined manually. That is, the viewpoint information of the second virtual camera group 307 may be determined based on the viewpoint information of the physical camera group 302. In the virtual viewpoint generation technology, the physical camera group 302 is arranged in the number necessary to estimate the shape of the subject 301, such as by the volume intersection method described above, so as to surround the subject 301. Therefore, by sending depth images as observed by the physical camera group 302, the user can accurately restore the shape information of the shape-estimated 3D model 304. Furthermore, by delivering virtual viewpoint-dependent texture images with image quality similar to that of the physical camera group 302, the user can restore high-quality color information, enabling the user to play back a high-quality 3D model.
[0076] A method for determining viewpoint information for each virtual camera will be described with reference to FIG. 4. FIG. 4 is a schematic diagram illustrating a first virtual camera 401 capturing an image of a 3D model 402 of a subject, and generating viewpoint information for a second virtual camera 403 based on the viewpoint information of the first virtual camera 401. The second virtual camera 403 moves along the optical axis of the first virtual camera 401 while dolly zooming to approach the 3D model 402. The second virtual camera 403 is then positioned at a position where the 3D model 402 can be captured within an area 404 that encompasses the 3D model 402 and allows distance information (depth information) between the 3D model and the second virtual camera 403 to be expressed in 10 bits. Note that the area 404 is not limited to a cube, and may be a sphere centered on the 3D model 402. The first virtual camera 401 captures the 3D model 402 and generates a captured image 405. The first image processing device 20 estimates the 3D model 402 based on the captured images 405 captured by the physical cameras of the imaging system 10. Then, a second virtual camera 403 captures an image of the estimated 3D model 402, and generates a depth image 406 and a texture image 407. The size of the 3D model 402 captured in the depth image 406 and the texture image 407 is the same as the size of the 3D model 402 captured in the captured image 405. As described above, the determination of the viewpoint information of the second virtual camera 403 is not limited to dolly zoom, and the size of the 3D model 402 captured in the depth image 406 may be different from the size of the 3D model 402 captured in the captured image 405. Furthermore, although the second virtual camera 403 has been described as moving on the optical axis of the first virtual camera 401, the position and orientation of the second virtual camera 403 need only be close to the position and orientation of the first virtual camera 401, and it does not necessarily have to move on the optical axis of the first virtual camera 401. Furthermore, the determination of the viewpoint information of the second virtual camera 403 may be performed once when the second image processing device 30 is initialized, or may be performed for each frame in accordance with the movement of the 3D model 402.
[0077] <Compression and distribution of 3D model data and control of virtual viewpoint image generation> Fig. 5 is a flowchart showing the flow of processing for controlling compression and distribution of 3D model data in the first image processing device 20 according to this embodiment. The flow shown in Fig. 5 is realized by reading a control program stored in the ROM 103 into the RAM 102 and executing it by the CPU 101. Execution of the flow in Fig. 5 is triggered when the shape information generation unit 201 receives a plurality of captured images and viewpoint information of the physical cameras from the imaging system 10.
[0078] In S501, the shape information generation unit 201 estimates and generates shape information of the subject based on a plurality of captured images. The generated shape information, viewpoint information of the physical camera group, and a plurality of captured images are output to the viewpoint determination unit 202 and the texture image generation unit 204. The generated shape information is also output to the depth image generation unit 203. Note that viewpoint information of a first virtual camera group corresponding to the physical camera group, which is used to generate the subject, is assumed to have been generated in advance. For example, when the physical camera group is arranged as preparation before starting shooting, a first virtual camera group corresponding to the arranged physical camera group is generated. Shape information of the subject is generated using the viewpoint information of this first virtual camera group and a plurality of captured images.
[0079] In S502, the viewpoint determination unit 202 generates viewpoint information of the second virtual camera group based on the viewpoint information of the physical camera group. The generated viewpoint information of the second virtual camera group is output to the depth image generation unit 203 and to the texture image generation unit 204 via the depth image generation unit 203. The process of generating the viewpoint information of the second virtual camera group will be described with reference to FIG. 7. Note that the generated viewpoint information of the second virtual camera group may also be output to the texture image generation unit 204 without going through the depth image generation unit 203.
[0080] In S503, the depth image generation unit 203 generates a depth image of the 3D model based on the data acquired from the shape information generation unit 201 and the viewpoint determination unit 202. The generated depth image is output to the encoding unit 205.
[0081] In S504, the texture image generation unit 204 generates a texture image of the 3D model based on the data acquired from the shape information generation unit 201 and the viewpoint determination unit 202. The generated texture image is output to the encoding unit 205.
[0082] In S505, the encoding unit 205 encodes the depth image and texture image acquired from the depth image generation unit 203 and the texture image generation unit 204. The encoded data is output to the distribution unit 206.
[0083] In S506, the distribution unit 206 distributes 3D model data including the data acquired from the depth image generation unit 203 and the encoding unit 205 and the viewpoint information of the second virtual camera group, and this flow ends. In the image processing system 1, the 3D model data is distributed to the second image processing device 30, but this is not limiting. The 3D model data may also be distributed to a separate server that stores the 3D model data.
[0084] 6 is a flowchart showing a process flow for generating a virtual viewpoint image using a depth image and a texture image in the second image processing device 30 according to this embodiment. Execution of the flow in FIG. 6 is triggered by the virtual camera control unit 210 receiving an input value from the input device 40.
[0085] In S601, the virtual camera control unit 210 generates viewpoint information of a third virtual camera designated by the user based on input values from the input device 40. The generated viewpoint information of the third virtual camera is output to the virtual viewpoint image generation unit 211.
[0086] In S602, the receiving unit 207 receives the 3D model data distributed from the distribution unit 206. The received 3D model data is output to the decoding unit 208.
[0087] In S603, the decoding unit 208 decodes the depth images and texture images included in the 3D model data acquired from the receiving unit 207. If the viewpoint information of the second virtual camera group has also been encoded, the decoding unit 208 also decodes the viewpoint information of the second virtual camera group. The decoded depth images and texture images, as well as the viewpoint information of the second virtual camera group, are then output to the 3D model restoration unit 209.
[0088] In S604, the 3D model restoration unit 209 restores a 3D model of the subject based on the decoded depth image and texture image acquired from the decoding unit 208 and the viewpoint information of the second virtual camera group. The restored 3D model is output to the virtual viewpoint image generation unit 211. The 3D model restoration process will be described with reference to FIG. 8.
[0089] In S605, the virtual viewpoint image generation unit 211 generates a virtual viewpoint image based on the viewpoint information of the third virtual camera acquired from the virtual camera control unit 210 and the 3D model acquired from the 3D model restoration unit 209, and this flow ends.
[0090] After this flow is completed, the generated virtual viewpoint image is transmitted to the display device 50 and displayed on the display device 50.
[0091] The above is the content of the control of the compression and distribution of 3D model data and the generation of virtual viewpoint images according to this embodiment.
[0092] <Description of the Process for Generating Viewpoint Information of the Second Virtual Camera Group> 7 is an example of a flowchart showing the flow of processing for generating viewpoint information of the second virtual camera group according to this embodiment. Here, the method for generating viewpoint information of the second virtual camera group will be described as a method for generating viewpoint information based on the dolly zoom method of the second virtual camera described with reference to FIG. 4. That is, the second virtual camera moves on the optical axis of the first virtual camera from the position of the first virtual camera corresponding to the position of the physical camera in a direction approaching the 3D model of the subject. Furthermore, a method will be described in which the second virtual camera adjusts the focal length while maintaining the size of the subject on the imaging plane of the second virtual camera.
[0093] The flow shown in Fig. 7 is executed by the viewpoint determination unit 202. Execution of Fig. 7 is triggered by reception of a 3D model of the subject, viewpoint information of the physical cameras, and multiple captured images from the shape information generation unit 201. The flow in Fig. 7 provides a detailed explanation of the control for determining the viewpoint of the second virtual camera group based on the viewpoint information of the physical cameras in S502 of Fig. 5. When executing the flow shown in Fig. 7, it is assumed that viewpoint information of the first virtual camera group and a 3D model of the subject have been generated in the first image processing device 20.
[0094] In S701, a 3D model of the subject, viewpoint information of the first virtual camera group, and multiple captured images are acquired. Note that viewpoint information of the physical camera group may be acquired without acquiring viewpoint information of the first virtual camera group. In this case, viewpoint information of the first virtual camera group is generated using the viewpoint information of the physical camera group.
[0095] In S702, S703 and S704 are repeatedly processed for the number of physical cameras.
[0096] In S703, S704 is repeated for each subject included in the multiple captured images. Subject recognition is performed based on the results of a face detection algorithm, a person detection algorithm, etc. Alternatively, the area of each subject in the captured images can be identified by projecting a 3D model generated separately for each subject onto the same plane as the imaging surface of the physical camera using the viewpoint information of the physical camera.
[0097] In S704, viewpoint information of the second virtual camera is generated based on the 3D model and viewpoint information of the first virtual camera. The viewpoint information of the second virtual camera is generated by the method described with reference to Fig. 4. The above process is repeated as described in S702 and S703 to generate viewpoint information of the second virtual camera group.
[0098] In S705, the viewpoint information of the generated second virtual camera group is sent to the depth image generation unit 203 and the texture image generation unit 204.
[0099] <Explanation of 3D model restoration control> FIG. 8 is an example of a flowchart showing the flow of a 3D model restoration process according to this embodiment. Here, the 3D model restoration process will be described based on a method for generating a 3D model including virtual viewpoint-dependent color information. In other words, the pixel values of the texture image of the second virtual camera group, which are closest to the position and orientation of the user-specified third virtual camera, are given priority as the color information of the 3D model. The flow shown in FIG. 8 is executed by the 3D model restoration unit 209. The flow in FIG. 8 provides a detailed description of the control for restoring a 3D model of a subject based on the decoded data in FIG. 6.
[0100] In S801, the 3D model restoration unit 209 acquires a depth image, a texture image, and viewpoint information of the second virtual camera group.
[0101] In S802, the 3D model restoration unit 209 projects the depth values of each pixel of the depth image into virtual space based on the external and internal parameters of the virtual camera corresponding to the depth image, thereby generating components of the shape information of the subject. For example, if the 3D model of the subject is represented by a 3D point cloud, each point becomes a component. By performing the above process on all of the acquired multiple depth images, the shape information of the subject is restored.
[0102] In S803, if the depth image is captured from multiple texture images of the components of the shape information projected into the virtual space, the pixel values of the texture image of the virtual camera with the closest position and orientation to the user-specified virtual camera are prioritized and determined as the color information of the components. This process is repeated for all components to restore the color information corresponding to the shape information. For example, if the shape information is a point cloud and the components are each point of the point cloud, the color information corresponding to all points is restored.
[0103] In S804, the restored 3D model is output to the virtual viewpoint image generation unit 211.
[0104] As described above, in this embodiment, viewpoint information of the second virtual camera group is generated based on viewpoint information of the physical camera group, and a process of generating a depth image and a texture image using the generated viewpoint information of the second virtual camera group is performed. This process enables the receiving side to restore a high-quality 3D model without increasing the amount of data, even for 3D models with complex shapes. While this embodiment describes the delivery of compressed 3D model data, the compressed 3D model data may be stored or may be delivered in response to a client request. Furthermore, while the method of generating a 3D model including virtual viewpoint-dependent color information during 3D model restoration has been described as an example, this is not limiting. After the restoration of the 3D model's shape information, color information for the 3D model may be generated simultaneously with the generation of a virtual viewpoint image. In this case, it is not necessary to assign color information to all of the shape information of the 3D model; color information may be generated only for the shooting range (viewing angle) of the user-specified third virtual camera.
[0105] <Example 2> In the first embodiment, a process of generating viewpoint information of a second virtual camera group based on viewpoint information of a physical camera group and generating a depth image and a texture image using the generated viewpoint information of the second virtual camera group is described. Next, a mode of generating an effective pixel map in addition to a depth image and a texture image to deal with cases where a subject is occluded by another subject or object will be described as a second embodiment. Note that the description of parts common to the first embodiment, such as the hardware configuration and functional configuration of the image processing device, will be omitted or simplified.
[0106] FIG. 9 is a schematic diagram illustrating an example of a method for generating an effective pixel map according to this embodiment. A subject is photographed by a physical camera, and a photographed image 903 is generated. At this time, the subject in the photographed image 903 has a part of its body hidden by an obstruction. The obstruction is assumed to be the subject. The first image processing device 20 estimates the shapes of the subject and the obstruction based on multiple captured images taken by the physical camera of the photographing system 10, and generates shape information for each. That is, it generates a 3D model 901 of the subject and a 3D model 904 of the obstruction. Next, the first image processing device 20 generates viewpoint information for the second virtual camera group described in the first embodiment based on data such as viewpoint information and shape information of the physical camera group. The second virtual camera 905 is positioned on the optical axis of the first virtual camera 902, within an area 906 that encompasses the 3D model 901 of the subject and that can represent the depth of the subject in 10 bits. Furthermore, the focal length of the second virtual camera 905 is adjusted so that the size of the object 901 captured by the first virtual camera 902 is maintained. After generating the viewpoint information of the second virtual camera 905, the first image processing device 20 generates a depth image 907 and a texture image 908. Here, for each pixel value of the object 901 captured in the generated texture image 908, the pixel value of the image captured by the first virtual camera 902 is used preferentially in the area not obscured by the 3D model 904 of the obstructing object (the left side of the object). On the other hand, in the area obscured by the 3D model 904 of the obstructing object (the right side of the object), the pixel value of the image captured by the first virtual camera, which has a different viewpoint from the first virtual camera 902, is used. Therefore, there is a possibility that the image quality will differ significantly between the right and left sides of the object captured in the texture image 908. If the color information of the 3D model is restored using this texture image 908, the image quality of some parts of the 3D model will be reduced, and this part will stand out, which may cause the user to feel uncomfortable. Therefore, an effective pixel map 909 is generated that indicates high-quality areas of the texture image. In the effective pixel map, for example, the pixel value of a non-occluded area is set to 1, and the pixel value of an occluded area is set to 0. The image size of the effective pixel map 909 is the same as the image size of the texture image 908. Whether or not an area is an occluded area is determined, for example, based on whether or not each point constituting the shape information of the 3D model 901 of the subject is visible to the first virtual camera 902 by the visibility determination process described above.In other words, due to the visibility determination process, pixel values of the region of the occluding object 904 in the captured image 903 are not used as color information for the 3D model 901 of the subject. Therefore, a region of the texture image determined using pixel values of the captured image of the physical camera from which the viewpoint information of the second virtual camera 905 is generated is identified, and in the effective pixel map, pixel values of the identified region are set to 1, and other regions are set to 0. Note that the method of determining an occluded region is not limited to this. When restoring color information for the 3D model, the second image processing device 30 prioritizes pixel values of regions of the texture image corresponding to regions with a value of 1 in the effective pixel map and uses them as color information for the 3D model. Specifically, when generating color information for the restored shape information in a virtual viewpoint-dependent manner, the second image processing device 30 uses pixel values of the texture image corresponding to 1 in the effective pixel map and does not use pixel values of the texture image corresponding to 0 in the effective pixel map. However, if color information for shape information to which color information is to be added is stored only in the texture image corresponding to 0 in the effective pixel map, the pixel value of that texture image is used. While the pixel values of the valid pixel map have been described so far as being binary (0 or 1), they can also be multi-valued. In this case, the pixel values of the valid pixel map can be used as weighting when generating color information for the 3D model. In other words, the priority of pixel values in the texture image is determined according to the pixel values of the valid pixel map. When generating a valid pixel map with 255 levels of multi-valued values, for example, pixel values at the outline of the subject or the boundary with an obstructing object can be set to 0, and can be linearly increased to 255 as the pixel approaches a certain distance (e.g., 5px) inward toward the subject. This reduces the influence of unreliable pixel values of the texture image at the outline or boundary with an obstructing object when restoring color information for the 3D model, making it possible to generate color information for the 3D model using highly reliable pixel values.
[0107] 10 is a flowchart showing the flow of processing according to this embodiment for controlling the compression and distribution of 3D model data in the first image processing device 20. Execution of the flow in FIG. 10 is started when the shape information generation unit 201 receives a plurality of captured images and viewpoint information of the physical cameras from the imaging system 10.
[0108] S1001 to S1003 are the same as S501 to S503 in FIG.
[0109] In S1004, the texture image generation unit 204 generates a texture image and an effective pixel map of the foreground model based on the data acquired from the shape information generation unit 201 and the viewpoint determination unit 202. The generated texture image and effective pixel map are output to the encoding unit 205.
[0110] In S1005, the encoding unit 205 encodes the depth image, texture image, and valid pixel map acquired from the depth image generation unit 203 and the texture image generation unit 204. The encoded depth image, texture image, and valid pixel map are output to the distribution unit 206.
[0111] In S1006, the distribution unit 206 transmits 3D model data including the depth image, texture image, effective pixel map, and viewpoint information of the second virtual camera group acquired from the depth image generation unit 203 and the encoding unit 205 to the receiving unit 207, and this flow ends.
[0112] Fig. 11 is an example of a flowchart showing the flow of a 3D model restoration process according to this embodiment. The flow of Fig. 11 is executed by the 3D model restoration unit 209. The flow of Fig. 11 provides a detailed explanation of the control for restoring a 3D model of a subject based on the data decoded in S604 of Fig. 6.
[0113] In S1101, a depth image, a texture image, an effective pixel map, and viewpoint information of the second virtual camera group are acquired.
[0114] In S1102, the depth value of each pixel of the depth image is projected onto a virtual space based on the external parameters and internal parameters of the second virtual camera corresponding to the depth image, and shape information of the subject is restored.
[0115] In S1103, if the shape information is captured from multiple texture images, the pixel values of the texture image captured by the second virtual camera, which is closest in position and orientation to the user-specified third virtual camera, are given priority as color information. At this time, the priority of the pixel values of the texture image is determined according to the pixel values of the effective pixel map corresponding to the pixel values of the texture image. If the shape information is represented by a 3D point cloud, this process is repeated for all points in the point cloud to generate color information corresponding to the shape information.
[0116] S1104 is the same as S804 in FIG.
[0117] After this flow is completed, the virtual viewpoint image generation unit 211 generates a virtual viewpoint image based on the restored 3D model and the viewpoint information of the third virtual camera specified by the user, and the generated virtual viewpoint image is displayed on the display device 50.
[0118] The above explains the compression and distribution of 3D model data including the effective pixel map, and the control of 3D model restoration. This process makes it possible to restore a 3D model using highly reliable pixel values of texture images, even when the subject is hidden by an object.
[0119] <Other Examples> The present disclosure can also be realized by providing a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.
[0120] The disclosure of the present embodiment includes the following configurations, methods, systems, and programs.
[0121] (Configuration 1) a setting means for setting positions and orientations of a plurality of second virtual cameras based on optical axes of a plurality of first virtual cameras in a virtual space corresponding to positions and orientations of a plurality of imaging devices in a real space; a generation means for generating a plurality of depth images indicating the distance between each of the plurality of second virtual cameras and a 3D model of a subject generated based on a plurality of captured images acquired by the plurality of imaging devices; 1. An image processing device comprising:
[0122] (Configuration 2) 2. The image processing device according to configuration 1, wherein the second virtual cameras are set on the optical axes of the first virtual cameras.
[0123] (Configuration 3) 2. The image processing device according to configuration 1, further comprising encoding means for encoding the depth image.
[0124] (Configuration 4) The image processing device according to configuration 3, further comprising an output means for outputting the encoded depth images and viewpoint information indicating the positions and orientations of the second virtual cameras to another device that reconstructs the 3D model based on the encoded depth images and the viewpoint information.
[0125] (Configuration 5) the generating means generates a virtual viewpoint image including a 3D model of the subject for each of the second virtual cameras; the encoding means encodes the plurality of virtual viewpoint images; 5. The image processing device according to configuration 4, wherein the output means outputs the encoded virtual viewpoint images to the other device.
[0126] (Configuration 6) the generating means generates correspondence information indicating whether each pixel in the virtual viewpoint image corresponds to a component of the 3D model of the subject, for each of the plurality of second virtual cameras; the encoding means encodes a plurality of pieces of correspondence information; 6. The image processing device according to configuration 5, wherein the output means outputs the encoded correspondence information to the other device.
[0127] (Configuration 7) 2. The image processing device according to configuration 1, wherein the second virtual cameras are set at positions closer to the 3D model of the subject than the first virtual cameras.
[0128] (Configuration 8) 2. The image processing device according to configuration 1, wherein the generating means generates one depth image for one second virtual camera.
[0129] (Configuration 9) an acquisition means for acquiring a plurality of coded depth images indicating the distance between each of a plurality of second virtual cameras set based on the optical axes of a plurality of first virtual cameras in a virtual space corresponding to the positions and orientations of a plurality of imaging devices in real space and a 3D model of a subject generated based on a plurality of captured images acquired by the imaging devices, and first viewpoint information indicating the positions and orientations of the plurality of second virtual cameras; a decoding means for decoding the encoded depth images; a generating means for generating a 3D model of the subject based on the decoded depth images and the first viewpoint information; 1. An image processing device comprising:
[0130] (Configuration 10) acquiring second viewpoint information indicating a position and an attitude of a third virtual camera different from the first virtual camera and the second virtual camera; 10. The image processing device according to configuration 9, wherein the generating means generates a virtual viewpoint image based on the second viewpoint information and a 3D model of the subject.
[0131] (Method 1) a setting step of setting positions and orientations of a plurality of second virtual cameras based on optical axes of a plurality of first virtual cameras in a virtual space corresponding to positions and orientations of a plurality of imaging devices in a real space; a generation step of generating a plurality of depth images indicating the distance between each of the plurality of second virtual cameras and a 3D model of a subject generated based on a plurality of captured images acquired by the plurality of imaging devices; An image processing method comprising:
[0132] (Method 2) an acquisition process for acquiring a plurality of coded depth images indicating the distance between each of a plurality of second virtual cameras set based on the optical axes of a plurality of first virtual cameras in a virtual space corresponding to the positions and orientations of a plurality of imaging devices in real space and a 3D model of a subject generated based on a plurality of captured images acquired by the imaging devices, and first viewpoint information indicating the positions and orientations of the plurality of second virtual cameras; a decoding step of decoding the encoded depth images; a generation step of generating a 3D model of the subject based on the decoded depth images and the first viewpoint information; An image processing method comprising:
[0133] (program) A program for causing a computer to function as each of the means of the image processing device according to any one of configurations 1 to 10. [Explanation of symbols]
[0134] 202 Viewpoint Determination Unit 203 Depth image generation unit 204 Texture image generation unit 205 Encoding section
Claims
1. a setting means for setting positions and orientations of a plurality of second virtual cameras based on optical axes of a plurality of first virtual cameras in a virtual space corresponding to positions and orientations of a plurality of imaging devices in a real space; a generation means for generating a plurality of depth images indicating distances between each of the plurality of second virtual cameras and a 3D model of a subject generated based on a plurality of captured images acquired by the plurality of imaging devices; 1. An image processing device comprising:
2. The image processing device according to claim 1 , wherein the plurality of second virtual cameras are set on optical axes of the plurality of first virtual cameras.
3. 2. The image processing apparatus according to claim 1, further comprising encoding means for encoding the depth image.
4. The image processing device according to claim 3, further comprising an output means for outputting the encoded depth images and viewpoint information indicating the positions and orientations of the second virtual cameras to another device that reconstructs the 3D model based on the encoded depth images and the viewpoint information.
5. the generating means generates a virtual viewpoint image including a 3D model of the subject for each of the second virtual cameras; the encoding means encodes the plurality of virtual viewpoint images; 5. The image processing apparatus according to claim 4, wherein the output means outputs the encoded virtual viewpoint images to the other device.
6. the generating means generates correspondence information indicating whether each pixel in the virtual viewpoint image corresponds to a component of a 3D model of the subject, for each of the plurality of second virtual cameras; the encoding means encodes a plurality of pieces of correspondence information; 6. The image processing apparatus according to claim 5, wherein said output means outputs the encoded correspondence information to said other apparatus.
7. The image processing device according to claim 1 , wherein the second virtual cameras are set at positions closer to the 3D model of the subject than the first virtual cameras.
8. The image processing device according to claim 1 , wherein the generating means generates one depth image for one second virtual camera.
9. an acquisition means for acquiring a plurality of encoded depth images indicating the distance between each of a plurality of second virtual cameras set based on the optical axes of a plurality of first virtual cameras in a virtual space corresponding to the positions and orientations of a plurality of image capturing devices in real space and a 3D model of a subject generated based on a plurality of captured images acquired by the image capturing devices, and first viewpoint information indicating the positions and orientations of the plurality of second virtual cameras; a decoding means for decoding the encoded depth images; a generating means for generating a 3D model of the subject based on the decoded depth images and the first viewpoint information; 1. An image processing device comprising:
10. acquiring second viewpoint information indicating a position and an attitude of a third virtual camera different from the first virtual camera and the second virtual camera; The image processing apparatus according to claim 9 , wherein the generating means generates a virtual viewpoint image based on the second viewpoint information and a 3D model of the subject.
11. a setting step of setting positions and orientations of a plurality of second virtual cameras based on optical axes of a plurality of first virtual cameras in a virtual space corresponding to positions and orientations of a plurality of imaging devices in a real space; a generation step of generating a plurality of depth images indicating distances between each of the plurality of second virtual cameras and a 3D model of a subject generated based on a plurality of captured images acquired by the plurality of imaging devices; An image processing method comprising:
12. an acquisition process for acquiring a plurality of encoded depth images indicating the distance between each of a plurality of second virtual cameras set based on the optical axes of a plurality of first virtual cameras in a virtual space corresponding to the positions and orientations of a plurality of imaging devices in real space and a 3D model of a subject generated based on a plurality of captured images acquired by the imaging devices, and first viewpoint information indicating the positions and orientations of the plurality of second virtual cameras; a decoding step of decoding the encoded depth images; a generation step of generating a 3D model of the subject based on the decoded depth images and the first viewpoint information; An image processing method comprising:
13. A program for causing a computer to function as each of the means of the image processing device according to any one of claims 1 to 10.
Citation Information
Patent Citations
Image processing device and image processing method
WO2018079260A1