Image processing device, image processing method, and program
By aligning virtual cameras with real-space image capture devices, the image processing device ensures accurate 3D model reconstruction by maintaining shape and color fidelity, addressing quality degradation issues in existing technologies.
Patent Information
- Application Number
- PCT/JP2025/014112
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-26
- Filing Date
- 2025-04-09
- Publication Date
- 2025-10-30
AI Technical Summary
Existing 3D model reconstruction technologies face issues with quality degradation due to improper positioning and orientation of virtual cameras, leading to discrepancies between the generated and reconstructed 3D models.
The image processing device sets the positions and orientations of second virtual cameras based on the optical axes of first virtual cameras, aligned with real-space image capture devices, to generate depth and texture images that accurately represent the 3D model, ensuring similarity to captured images and reducing color discrepancies.
This approach enhances the reproducibility of the 3D model by maintaining shape and color accuracy during reconstruction, minimizing differences between the generated and reconstructed models.
Smart Images

Figure JP2025014112_30102025_PF_FP_ABST
Abstract
Description
Image processing device, image processing method and program
[0001] The present disclosure relates to an image processing apparatus for compressing and distributing 3D model data representing a 3D model.
[0002] There is a technology that generates 3D model data representing a 3D model of a subject (object) using multiple captured images taken by multiple imaging devices, and generates a virtual viewpoint image viewed from a camera (virtual camera) virtually placed in a virtual space in which the 3D model exists. Recently, a technology that generates 3D model data of a subject on a server, compresses and distributes the 3D model data, and allows a user to operate a virtual camera on their own local terminal (client) such as a PC or tablet to display the virtual viewpoint image has been attracting attention.
[0003] In a technology for compressing and distributing 3D model data, a distributing device arranges multiple virtual cameras (a group of virtual cameras) to surround a 3D model and generates depth images and texture images for each of the multiple virtual cameras. The generated depth images and texture images are then compressed, and the compressed depth images, texture images, and information indicating the positions and orientations of the multiple virtual cameras are distributed to a user. The receiving device decodes the compressed depth images and texture images, and reconstructs a 3D model based on the decoded depth images, texture images, and information indicating the positions and orientations of the multiple virtual cameras.
[0004] Patent Literature 1 describes a method for determining the viewpoints of a group of virtual cameras so as to reduce fluctuations in the positions of major objects in depth images between frames, with the aim of improving the compression rate of depth images. By reducing the number of motion vectors included in the encoded stream when encoding depth images, it is expected that the compression rate of depth images will improve.
[0005] International Publication No. 2018 / 079260
[0006] However, depending on the position and orientation of the virtual cameras set on the distribution device, there is a risk of a decrease in the quality of the 3D model reconstructed on the receiving device. For example, if the shape of the 3D model generated on the distribution device is complex, depending on the position and orientation of the virtual cameras, the complex shape of the 3D model may not be sufficiently represented in the depth image, and the shape accuracy of the 3D model reconstructed on the user side may decrease. In other words, the shape of the 3D model reconstructed on the receiving side may differ significantly from the shape of the 3D model generated on the distribution side. Therefore, in order to reconstruct the 3D model more accurately, it is necessary to appropriately determine the position and orientation of the virtual cameras.
[0007] Therefore, an object of the present disclosure is to reduce the possibility of a decrease in the quality of a 3D model generated on a receiving device.
[0008] In order to solve the above problems, the image processing device according to the present disclosure has the following configuration: a setting unit that sets the positions and orientations of a plurality of second virtual cameras based on the optical axes of a plurality of first virtual cameras in a virtual space that correspond to the positions and orientations of a plurality of image capture devices in real space, and a generation unit that generates a plurality of depth images that indicate the distance between each of the plurality of second virtual cameras and a 3D model of a subject that is generated based on a plurality of captured images acquired by the plurality of image capture devices.
[0009] According to the present disclosure, it is possible to reduce the possibility of a decrease in the quality of the 3D model generated on the receiving device.
[0010] 1 is a block diagram showing an example of the configuration of an image processing system 1. FIG. 1 is a diagram showing an example of the hardware configuration of an image processing device. FIG. 2 is a diagram showing an example of the functional configuration of a first image processing device 20 and a second image processing device 30. FIG. 2 is a diagram showing an example of the configuration of an imaging system 10. FIG. 3 is a diagram showing a first virtual camera group 305 corresponding to a physical camera group 302, and a 3D model 304 of a subject generated by the first virtual camera group 305. FIG. 4 is a diagram showing a second virtual camera group 307 set based on viewpoint information of the first virtual camera group 305. FIG. 5 is a diagram explaining generation of viewpoint information of virtual cameras according to Example 1. FIG. 6 is a flowchart showing an example of compression and distribution processing of 3D model data according to Example 1. FIG. 7 is a flowchart showing an example of reception processing of 3D model data and generation processing of a virtual viewpoint image according to Example 1. FIG. 8 is a flowchart showing an example of generation processing of viewpoint information of a virtual camera group according to Example 1. FIG. 9 is a flowchart showing an example of restoration processing of a 3D model according to Example 1. FIG. 10 is a diagram explaining generation of viewpoint information of a virtual camera group according to Example 2. FIG. 11 is a flowchart showing an example of compression and distribution processing of 3D model data according to Example 2. FIG. 12 is a flowchart showing an example of restoration processing of a 3D model according to Example 2.
[0011] According to a preferred embodiment of the present invention, an image processing device includes a setting unit that sets the positions and orientations of multiple second virtual cameras based on the optical axes of multiple first virtual cameras in a virtual space corresponding to the positions and orientations of multiple image capture devices in real space. The image processing device also includes a generation unit that generates multiple depth images indicating the distance between each of the multiple second virtual cameras and a 3D model of a subject generated based on multiple captured images acquired by the multiple image capture devices. Here, the multiple image capture devices are assumed to be aligned with each other. The multiple first virtual cameras are set based on the positional relationship between the multiple image capture devices in real space. In other words, the multiple first virtual cameras in the virtual space are reproductions of the multiple image capture devices in real space.
[0012] According to this aspect, the positions and orientations of the multiple second virtual cameras are set based on the positions and orientations of the multiple imaging devices. Therefore, the subject represented in the depth images of the multiple second virtual cameras is similar to the subject included in the captured images acquired from the multiple imaging devices. Specifically, the shape of the area representing the subject included in the depth images generated from the second virtual camera is similar to the shape of the area representing the subject included in the captured images. The 3D model of the subject generated by the distribution device is generated based on the multiple captured images acquired from the multiple imaging devices. Therefore, if the depth images used to reconstruct (restore) the 3D model of the subject on the receiving device are similar to the captured images, the 3D model reconstructed by the receiving device will also be similar to the 3D model generated by the distribution device. In other words, the possibility that the shape of the 3D model generated by the receiving device will differ from the 3D model generated by the distribution device, i.e., the possibility that the reproducibility of the 3D model will be deteriorated, can be reduced.
[0013] The image processing device also includes an encoding unit that encodes the depth images. The image processing device also includes an output unit that outputs the encoded depth images and viewpoint information indicating the positions and orientations of the second virtual cameras to another device that reconstructs the 3D model. Here, the other device reconstructs the 3D model based on the encoded depth images and the viewpoint information.
[0014] In addition, in the image processing device, the generating means generates a virtual viewpoint image including a 3D model of the subject for each of the plurality of second virtual cameras, the encoding means encodes the plurality of virtual viewpoint images, and the output means outputs the encoded plurality of virtual viewpoint images to the other device.
[0015] This configuration makes the colors of the components of the subject included in the virtual viewpoint image (texture image) generated from the second virtual camera similar to the colors of the components of the subject included in the captured image, thereby reducing the possibility that the 3D model generated on the receiving device will have different colors from the 3D model generated on the transmitting device, i.e., reducing the possibility that the reproducibility of the 3D model will be degraded.
[0016] The generating means generates correspondence information indicating whether each pixel in the texture image corresponds to a component of the 3D model of the subject for each of the second virtual cameras, and the encoding means encodes the plurality of pieces of correspondence information, and the output means outputs the encoded plurality of pieces of correspondence information to the other device.
[0017] This allows the receiving device to identify pixels of the texture image to be used to color the 3D model of the subject. For example, if a subject is occluded by another subject in a captured image acquired from an imaging device, the captured image is used to generate a texture image of the second virtual camera, which may result in a texture image having color information different from the color information of the subject. Therefore, by generating correspondence information indicating areas where occlusion is not considered to occur, the possibility of different colors being generated when the 3D model is reconstructed can be reduced.
[0018] The second virtual cameras are set at positions closer to the 3D model of the subject than the first virtual cameras.
[0019] The generating means generates one depth image for each of the second virtual cameras, and when delivering a 3D model showing movement over time, generates one depth image for each of the second virtual cameras for each of a plurality of time points.
[0020] According to another preferred embodiment of this embodiment, the image processing device includes an acquisition means for acquiring a plurality of encoded depth images and first viewpoint information indicating the positions and orientations of a plurality of second virtual cameras. Here, the depth images are images indicating the distance between each of the plurality of second virtual cameras and a 3D model of the subject generated based on a plurality of captured images acquired by the imaging device. The plurality of second virtual cameras are virtual cameras set based on the optical axes of a plurality of first virtual cameras in a virtual space corresponding to the positions and orientations of the plurality of imaging devices in real space. The image processing device also includes a decoding means for decoding the encoded depth images. The image processing device also includes a generation means for generating a 3D model of the subject based on the decoded depth images and the first viewpoint information.
[0021] The image processing device also acquires second viewpoint information indicating the position and orientation of a third virtual camera different from the first virtual camera and the second virtual camera. The image processing device also generates a virtual viewpoint image based on the second viewpoint information and a 3D model of the subject. The third virtual camera may be set by a user operating an input device such as a joystick, or may be set based on the position of the generated 3D model.
[0022] According to this aspect, a virtual viewpoint image can be generated using a 3D model generated on the receiving device.
[0023] According to another preferred embodiment of the present invention, a program causes a computer to execute the functions of the image processing device described above. By executing this program, the computer preferably functions as the image processing device described above.
[0024] <Examples> Preferred examples of the present disclosure will be described in detail below with reference to the drawings. Note that the following examples do not limit the present disclosure, and not all of the combinations of features described in the examples are necessarily essential to the solutions of the present disclosure. Furthermore, in the accompanying drawings, the same reference numerals are used to designate the same or similar components, and redundant explanations will be omitted.
[0025] A virtual viewpoint image is an image generated by a user freely manipulating the position and orientation of a virtual camera, and is also called a free viewpoint image, an arbitrary viewpoint image, etc. Unless otherwise specified, the term "image" will be explained as including the concepts of both moving images and still images.
[0026] The image processing system 1 is a system that generates a virtual viewpoint image representing a scene from a specified virtual viewpoint based on multiple images captured by multiple imaging devices and a specified virtual viewpoint. The virtual viewpoint image in this embodiment is also called a free viewpoint image, but is not limited to an image corresponding to a viewpoint freely (arbitrarily) specified by a user. For example, the virtual viewpoint image also includes an image corresponding to a viewpoint selected by a user from multiple options. Furthermore, in this embodiment, the description focuses on a case where the virtual viewpoint is specified by a user operation, but the virtual viewpoint may also be specified automatically based on the results of image analysis, etc. Furthermore, in this embodiment, the description focuses on a case where the virtual viewpoint image is a video, but the virtual viewpoint image may also be a still image.
[0027] The viewpoint information used to generate a virtual viewpoint image is information indicating the position and orientation (line of sight direction) of the virtual viewpoint. Specifically, the viewpoint information is a parameter set including a parameter representing the three-dimensional position of the virtual viewpoint and a parameter representing the orientation of the virtual viewpoint in the pan, tilt, and roll directions. Note that the content of the viewpoint information is not limited to the above. For example, the parameter set serving as viewpoint information may include a parameter representing the size of the field of view (angle of view) of the virtual viewpoint. Furthermore, the viewpoint information may have multiple parameter sets. For example, the viewpoint information may have multiple parameter sets corresponding to multiple frames constituting a moving image of the virtual viewpoint image, and may be information indicating the position and orientation of the virtual viewpoint at each of multiple consecutive time points.
[0028] The image processing system 1 has multiple imaging devices that capture images of an imaging area from multiple directions. The imaging area may be, for example, a stadium where sports such as soccer or karate are held, or a stage where a concert or a play is held. The multiple imaging devices are installed at different positions surrounding the imaging area and capture images synchronously. Note that the multiple imaging devices do not need to be installed around the entire periphery of the imaging area; depending on installation space restrictions, they may be installed only around a portion of the periphery of the imaging area. Furthermore, the number of imaging devices is not limited to the example shown in the figure. For example, if the imaging area is a soccer stadium, approximately 30 imaging devices may be installed around the stadium. Furthermore, imaging devices with different functions, such as telephoto cameras and wide-angle cameras, may be installed.
[0029] In this embodiment, the multiple image capturing devices are cameras each having an independent housing and capable of capturing images from a single viewpoint. However, this is not limiting, and two or more image capturing devices may be configured in the same housing. For example, a single camera equipped with multiple lens groups and multiple sensors and capable of capturing images from multiple viewpoints may be installed as the multiple image capturing devices.
[0030] A virtual viewpoint image is generated, for example, by the following method. First, multiple images (multiple captured images) are acquired by capturing images from different directions using multiple imaging devices. Next, a foreground image, in which a foreground region corresponding to a predetermined object, such as a person or a ball, is extracted, and a background image, in which a background region other than the foreground region is extracted, are acquired from the multiple captured images. Furthermore, a foreground model representing the three-dimensional shape of the predetermined object and texture data for coloring the foreground model are generated based on the foreground image, and texture data for coloring a background model representing the three-dimensional shape of a background, such as a stadium, is generated based on the background image. The texture data is then mapped to the foreground model and background model, and rendering is performed according to the virtual viewpoint indicated by the viewpoint information, thereby generating a virtual viewpoint image. However, the method for generating a virtual viewpoint image is not limited to this, and various methods can be used, such as a method of generating a virtual viewpoint image by projective transformation of captured images without using a three-dimensional model.
[0031] A foreground image is an image in which an object region (foreground region) is extracted from an image captured by an imaging device. An object extracted as a foreground region is a dynamic object (moving body) that moves (its absolute position and shape can change) when images are captured from the same direction in a time series. Examples of objects include players, referees, and other people on the field where a sport is being played, such as a ball in a ball game, or singers, musicians, performers, and presenters in a concert or entertainment event.
[0032] A background image is an image of at least an area (background area) different from the foreground object. Specifically, a background image is an image in which the foreground object has been removed from the captured image. Furthermore, the background refers to an imaged object that remains stationary or nearly stationary when images are captured from the same direction in chronological order. Examples of such imaged objects include a stage for a concert, a stadium where an event such as a sport is held, a structure such as a goal used in a ball game, or a field. However, the background is at least an area different from the foreground object, and the imaged object may include other objects in addition to the object and background.
[0033] A virtual camera is a virtual camera that is different from the multiple imaging devices actually installed around the imaging area, and is a concept for conveniently explaining a virtual viewpoint related to the generation of a virtual viewpoint image. That is, a virtual viewpoint image can be considered to be an image captured from a virtual viewpoint set in a virtual space associated with the imaging area. The position and orientation of the viewpoint in the virtual image capture can be expressed as the position and orientation of the virtual camera. In other words, a virtual viewpoint image can be said to be an image that simulates an image captured by a camera if it were assumed that the camera were located at the virtual viewpoint set in space.
[0034] <Example 1> In this example, a process for determining the viewpoints of a group of virtual cameras for generating depth images and texture images to be distributed from a server to a client based on viewpoint information of a physical camera will be described. Note that in this example, the receiving device will be referred to as the client. Also, in this example, an imaging device arranged in real space will be referred to as the physical camera.
[0035] <Hardware Configuration of Image Processing System> FIG. 1A is a diagram showing an example of the overall configuration of an image processing system 1 according to this embodiment.
[0036] The image processing system 1 includes an imaging system 10, a first image processing device 20, a second image processing device 30, an input device 40, and a display device 50. The image processing system 1 generates shape information of a 3D model of an object using images (plurality of captured images) captured by a plurality of physical cameras. The image processing system 1 then generates viewpoint information (first viewpoint information) of a second virtual camera group for generating depth images and texture images to be distributed from a server to a client. Furthermore, the image processing system 1 generates and compresses depth images and texture images based on the generated viewpoint information of the second virtual camera group, and distributes these images and information necessary to reconstruct (restore) the 3D model to a user as 3D model data.
[0037] The photography system 10 arranges multiple physical cameras at different positions and synchronously photographs a subject (object) from multiple viewpoints. The multiple captured images acquired through synchronous photography, as well as viewpoint information (external / internal parameters, image size) for each physical camera of the photography system 10, are transmitted to the first image processing device 20. The external parameters of the camera are parameters indicating the position and orientation of the camera (e.g., a rotation matrix and a position vector). The internal parameters of the camera are internal parameters specific to the camera, such as focal length, image center, and lens distortion parameters. The external and internal parameters of the camera are collectively referred to as camera parameters. The image size refers to the width and height of the image. The viewpoint information of the physical cameras is used to determine viewpoint information (camera parameters and image size) of a second virtual camera group for generating depth images and texture images.
[0038] The first image processing device 20 generates a 3D model of a foreground object based on multiple captured images input from the imaging system 10 and viewpoint information for each physical camera. The foreground object (subject) is, for example, a person or a moving object present within the imaging range of the imaging system 10. The first image processing device 20 then generates viewpoint information of a second virtual camera group for generating depth images and texture images. Furthermore, the first image processing device 20 generates and compresses depth images and texture images based on the generated viewpoint information of the second virtual camera group, and outputs the compressed images and texture images together with information (metadata) necessary to restore the 3D model to the second image processing device 30. The metadata is, for example, the generated viewpoint information of the second virtual camera group.
[0039] First, the first image processing device 20 estimates shape information (shape information) of the subject based on multiple captured images and viewpoint information for each physical camera. For example, a volume intersection method is used to estimate the shape information of the subject. Multiple virtual cameras (multiple first virtual cameras) are set in a virtual space corresponding to the imaging system 10 existing in real space. These multiple first virtual cameras are referred to as a first virtual camera group. This first virtual camera group is a reproduction of a physical camera group in virtual space. In other words, the viewpoint information of the first virtual camera group corresponds to the viewpoint information of the physical camera group in real space. Then, the viewpoint information of the first virtual camera group and the multiple captured images are used in the volume intersection method to generate shape information of the subject. As a result of this processing, a 3D point cloud (a set of points having three-dimensional coordinates) representing the shape information of the subject is obtained. This 3D point cloud is also called a visual hull. Note that the method for deriving shape information of the subject from the captured images is not limited to this. Furthermore, the method for representing the shape information of the subject is not limited to a 3D point cloud, and meshes or voxels may also be used. Furthermore, a texture image (color information) is determined for each point in the generated 3D point cloud of the subject using multiple captured images. Therefore, 3D model data representing a 3D model of the subject includes shape information indicating the shape and color information indicating the color. In this embodiment, the 3D model is generated from shape information and color information, but this is not limited to this. For example, the 3D model includes shape information and may be managed as data separate from color information. The method of setting a first virtual camera group in a virtual space corresponding to a physical camera group in a real space is a well-known technique in the field of CG, and therefore will not be described here.
[0040] Next, the first image processing device 20 generates viewpoint information (positions and orientations) of multiple second virtual cameras for generating depth images and texture images to be distributed. Hereinafter, the multiple second virtual cameras are referred to as a second virtual camera group. The positions of the second virtual cameras are, for example, arranged on the optical axes of the first virtual cameras, and the orientations are set to be the same as those of the first virtual cameras. Since the viewpoint information of the first virtual cameras corresponds to the viewpoint information of the physical cameras in real space, the viewpoint information of the second virtual cameras is generated based on the viewpoint information of the physical cameras. A method for generating the viewpoint information of the second virtual cameras will be described in detail later. Hereinafter, the viewpoint information of the second virtual cameras is referred to as second viewpoint information. The first image processing device 20 generates depth images of the 3D model based on the second viewpoint information of the second virtual cameras. Depth images are generated for the number of second virtual cameras. Specifically, each point in the 3D point cloud of the subject is projected onto the same plane as the imaging surface of the second virtual cameras. For each second virtual camera group, the distance (depth) from the second virtual camera to the subject is calculated for each projected pixel, and a depth value is set for each pixel of the depth image. Furthermore, the first image processing device 20 captures an image of the subject based on the second viewpoint information of the second virtual camera group and generates a texture image. The texture image is generated by blending the colors of multiple captured images while increasing the priority of pixel values of images captured by a physical camera with a line of sight close to the line of sight of the second virtual camera. This method of setting a high priority (weight) for pixel values of images captured by a physical camera with a line of sight close to the line of sight of the second virtual camera when blending multiple captured images is referred to as a virtual viewpoint-dependent texture image generation method. Details of the virtual viewpoint-dependent texture image generation method will be described later. This virtual viewpoint-dependent texture image is generated by selecting an image captured by a physical camera that determines the pixel value of the subject depending on the position and orientation of the virtual camera, so the color of the subject changes when the virtual camera moves. Therefore, a texture image generated using this method is referred to as a virtual viewpoint-dependent texture image. On the other hand, a virtual viewpoint-independent texture image is generated by a method in which the pixel values of the object do not change depending on the position and orientation of the virtual camera.In this embodiment, the generation of texture images is handled as a virtual viewpoint dependent subject, but texture images may be generated independently of a virtual viewpoint. An example of a process for generating texture images that are virtual viewpoint dependent and virtual viewpoint independent is shown below.
[0041] The process of generating a texture image dependent on a virtual viewpoint includes, for example, a process of determining the visibility of points in a 3D point cloud that constitutes a subject, and a process of deriving a color based on the position and orientation of a virtual camera.
[0042] In the visibility determination process, a physical camera capable of capturing an image of each point is identified based on the positional relationship between each point in the 3D point cloud and multiple physical cameras included in the group of physical cameras in the imaging system 10. In the color derivation process, for example, a point in the 3D point cloud is designated as a point of interest, and the color of the point of interest is derived. Specifically, the following process is performed for each point of interest. A point of interest included in the imaging range of the second virtual camera is selected. Then, a first virtual camera is selected that has a line of sight close to the line of sight of the second virtual camera whose imaging range includes the point of interest and that can capture the point of interest. Because the first virtual camera is a reproduction of a physical camera in a virtual space, selecting the first virtual camera is the same as selecting a physical camera. The selected point of interest is then projected onto the image captured by the selected physical camera. The color of the pixel at the projection destination is set as the color of the point of interest. The physical camera is selected, for example, based on whether the angle between the line of sight from the second virtual camera to the point of interest and the line of sight from the physical camera to the point of interest is equal to or less than a certain angle. If the point of interest can be captured by multiple physical cameras, multiple physical cameras with viewing directions close to the viewing direction of the second virtual camera are selected, and the point of interest is projected onto each of the images captured by those physical cameras.Then, pixel values of the projection destination are obtained, and a weighted average is calculated so that pixel values of physical cameras with viewing directions close to the viewing direction of the second virtual camera are used preferentially, thereby determining the color of the point of interest.
[0043] This process is performed while changing the focus point, and the color of the focus point is projected onto the same plane as the imaging plane of the second virtual camera, thereby generating a texture image that is dependent on the virtual viewpoint. Note that although the description has been given in conjunction with this embodiment, the present invention is not limited to the above, and when generating a texture image that is dependent on a virtual camera different from the second virtual camera, the second virtual camera is changed to the different virtual camera and the above process is performed.
[0044] The process of generating a texture image independent of the virtual viewpoint includes, for example, the visibility determination process described above and a process of deriving a color that is independent of the position and orientation of the virtual camera. After the visibility determination process, for example, a point in the 3D point cloud is set as a focus point, and the focus point is projected onto an image captured by a physical camera corresponding to a first virtual camera that can capture images, and the color of the pixel at the projection destination is set as the color of the focus point.
[0045] If the point of interest can be captured by multiple first virtual cameras, the point of interest is projected onto each of the images captured by multiple physical cameras corresponding to the multiple first virtual cameras. The pixel values of the projection destination are then acquired and the average of the pixel values is calculated to determine the color of the point of interest. This process is performed while changing the point of interest, and the color of the point of interest is projected onto the same surface as the imaging surface of the second virtual camera, thereby generating a texture image that is independent of the virtual viewpoint.
[0046] In virtual viewpoint image generation technology, a method of generating color information dependent on a virtual viewpoint is known as a method of generating images with higher image quality than a method independent of a virtual viewpoint, because pixel values of an image captured by a physical camera having a line of sight similar to the line of sight of a second virtual camera are preferentially used. In addition, a virtual viewpoint-dependent virtual viewpoint image (texture image) generated by a virtual camera whose position and orientation are similar to those of a physical camera is heavily influenced by the image captured by the physical camera, making it possible to generate a texture image with image quality similar to that of the captured image. On the other hand, a virtual viewpoint-dependent texture image generated from the viewpoint of a virtual camera whose position and orientation are significantly different from those of the physical camera, or a texture image generated independent of a virtual viewpoint, is generated by interpolating pixel values using multiple physical cameras. As a result, the resulting image may contain pixel values that differ from the color information of the actual subject, or may have low contrast.
[0047] Finally, the first image processing device 20 compresses (encodes) the depth image and texture image using a video compression method such as H.264 or H.265. The compression method is not limited to video compression, and any method that can encode data to a size smaller than the original data volume, such as file compression, may be used. The first image processing device 20 does not need to output depth images and texture images that do not include a subject to the second image processing device 30. The second image processing device 30 back-projects the depth image into virtual space based on second viewpoint information from the second virtual camera, thereby enabling the reconstruction of shape information for a 3D model. Furthermore, by using pixel values of the texture image at the same coordinates as each depth value of the depth image as color information for the point where the depth value is back-projected, color information can be added to the shape information, thereby reconstructing a 3D model.
[0048] The first image processing device 20 may also extract a rectangular image of the subject's shooting range from the pre-compression depth image and texture image, compress this rectangular image (ROI image), and distribute it.
[0049] In this case, the metadata may include coordinate information of the extracted rectangular image. The first image processing device 20 may also arrange the rectangular images to form a single image, compress it, and distribute it. By distributing the rectangular image instead of the entire image, it is possible to reduce the amount of data.
[0050] Furthermore, the first image processing device 20 may generate depth images containing high-precision depth values, such as single-precision floating-point numbers (32 bits), which cannot be used for video compression using H.264 or H.265. In this case, the depth information is converted to a precision (8 bits or 10 bits) that allows video compression before video compression. The conversion method may involve, for example, performing scalar quantization processing and compressing the quantized depth image for distribution. In this case, the minimum and maximum values of the depth value range before quantization may be included as metadata. By performing scalar quantization, the client can reconstruct a 3D model of the subject being photographed or the subject within the photographic range, which would otherwise be insufficiently compressible.
[0051] The second image processing device 30 receives and decodes the depth image, texture image, and metadata from the first image processing device 20 and restores a 3D model. As described above, the 3D model is restored by back-projecting the depth image and texture image into a virtual space based on second viewpoint information of a second virtual camera group included in the metadata. The second image processing device 30 also calculates viewpoint information (second viewpoint information) of a third virtual camera for generating a virtual viewpoint image viewed by the user based on input values received from an input device 40 (described later). The second image processing device 30 then generates a virtual viewpoint image based on the calculated second viewpoint information and the restored 3D model. The second image processing device 30 then outputs the generated virtual viewpoint image to a display device 50.
[0052] The input device 40 accepts input values for setting the third virtual camera from the user and transmits the input values to the second image processing device 30. For example, the input device 40 has an input unit such as a joystick, a jog dial, a touch panel, a keyboard, and a mouse. The user setting the third virtual camera sets the position and orientation of the third virtual camera by operating the input unit. Note that in this embodiment, the user sets the position and orientation of the third virtual camera, but this is not limited to this, and the position and orientation of the third virtual camera may be set using position information of a 3D model. Here, the position information of the 3D model may be generated by the first image processing device 20 on the distribution side or the second image processing device 30 on the receiving side.
[0053] The display device 50 displays the virtual viewpoint image generated and output by the second image processing device 30. The user looks at the virtual viewpoint image displayed on the display device 50 and sets the position and orientation of the virtual camera for the next frame via the input device 40.
[0054] 1B is a diagram showing an example of the hardware configuration of the first image processing device 20 of this embodiment. The hardware configuration of the second image processing device 30 is also the same as the configuration of the first image processing device 20 described below. The first image processing device 20 of this embodiment is composed of a CPU 101, a RAM 102, a ROM 103, and a communication unit 104.
[0055] The CPU 101 controls the entire first image processing device 20 using computer programs and data stored in the RAM 102 and ROM 103, thereby realizing each function of the first image processing device 20 shown in FIGS. 1A and 1B . The first image processing device 20 may have one or more dedicated hardware components different from the CPU 101, and at least a portion of the processing by the CPU 101 may be executed by the dedicated hardware components. Examples of the dedicated hardware components include an ASIC (application-specific integrated circuit), an FPGA (field-programmable gate array), and a DSP (digital signal processor). The RAM 102 temporarily stores programs and data supplied from the auxiliary storage device 214, as well as data supplied from the outside via the communication unit 104. The ROM 103 stores programs that do not require modification.
[0056] The communication unit 104 is used for communication with an external device of the first image processing device 20. For example, when the first image processing device 20 is connected to the external device via a wired connection, a communication cable is connected to the communication unit 104. When the first image processing device 20 has a function of wirelessly communicating with the external device, the communication unit 104 includes an antenna.
[0057] <Functional Configuration of Image Processing Apparatus> FIG. 2 is a diagram showing an example of the functional configuration of the first image processing apparatus 20 and the second image processing apparatus 30. As shown in FIG.
[0058] The first image processing device 20 includes a shape information generating unit 201 , a viewpoint determining unit 202 , a depth image generating unit 203 , a texture image generating unit 204 , an encoding unit 205 , and a distribution unit 206 .
[0059] The shape information generation unit 201 estimates shape information of the subject using the communication unit 104 and the viewpoint information of the physical cameras received from the imaging system 10. The shape information is estimated using the volume intersection method described above, or the like. Therefore, the shape information generation unit 201 also generates viewpoint information of a first virtual camera group in virtual space that corresponds to the physical camera group in real space, from the acquired viewpoint information of the physical cameras. The shape information generation unit 201 outputs the estimated shape information to the viewpoint determination unit 202, the depth image generation unit 203, and the texture image generation unit 204. The shape information generation unit 201 also outputs the received viewpoint information of the physical cameras and the multiple captured images to the viewpoint determination unit 202 and the texture image generation unit 204.
[0060] The viewpoint determination unit 202 determines the viewpoints of the second virtual cameras and generates viewpoint information based on the input data from the shape information generation unit 201, with the aim of improving the quality of the 3D model restored by the second image processing device 30. The viewpoint determination unit 202 outputs the generated viewpoint information of the second virtual cameras to the depth image generation unit 203 and to the texture image generation unit 204 via the depth image generation unit 203.
[0061] The viewpoint information of the second virtual camera group is generated, for example, by matching the position and orientation of the second virtual camera to the position and orientation of the first virtual camera corresponding to the physical camera and dolly-zooming the second virtual camera in the line of sight of the first virtual camera. Specifically, the second virtual camera, placed at the position of the first virtual camera corresponding to the physical camera, approaches the subject in the line of sight of the first virtual camera while adjusting the focal length (zoom) so that the size of the subject in the virtual viewpoint image generated from the second virtual camera is maintained. This process of deriving the viewpoint information of the second virtual camera will be described in detail below with reference to Figures 3A to 3C and 4. By performing this process while changing the physical camera, or while changing the subject if there are multiple subjects in the captured image, it is possible to generate viewpoint information of the second virtual camera for each subject, for the number of physical cameras capturing the subject.
[0062] The viewpoint determination unit 202 may also change the image size (width and height) of the second virtual camera. First, the image size of the second virtual camera is set to the same as the image size of the physical camera. After determining the position and orientation of the second virtual camera as described above, the image size of the second virtual camera is changed. For example, when the width and height of the virtual viewpoint image of the second virtual camera are changed to 1 / N of the width and height of the image captured by the physical camera, the focal length and image center, which are internal parameters of the second virtual camera, are also multiplied by 1 / N. In other words, when the image size of the physical camera is 4K (width 3840, height 2160) and the image size of the second virtual camera is changed to full HD (width 1920, height 1080), the width and height are each halved, the focal length and image center, which are internal parameters of the second virtual camera, are multiplied by 1 / 2. This allows the image size to be changed without changing the angle of view, thereby reducing the amount of data transmitted from the server to the client.
[0063] The depth image generation unit 203 performs the above-described depth image generation process based on the shape information of the subject input from the shape information generation unit 201 and the viewpoint information of the second virtual camera group input from the viewpoint determination unit 202. The depth image generation unit 203 outputs the generated depth image and the viewpoint information of the second virtual camera group to the encoding unit 205.
[0064] The texture image generation unit 204 performs processing to generate the virtual viewpoint-dependent texture image described above, based on input data from the shape information generation unit 201 and viewpoint information of the second virtual camera group input from the viewpoint determination unit 202 via the depth image generation unit 203. The texture image generation unit 204 outputs the generated texture image to the encoding unit 205. As described above, the virtual viewpoint-dependent texture image is generated by preferentially referencing pixel values of an image captured by a physical camera that is close in position and orientation to the second virtual camera. This makes it possible to generate high-quality texture images, and when the 3D model is restored on the client, a 3D model containing high-quality color information can be reproduced.
[0065] Furthermore, the texture image generation unit 204 may generate an effective pixel map based on input data from the shape information generation unit 201 and viewpoint information of the second virtual camera group input from the viewpoint determination unit 202 via the depth image generation unit 203. The texture image generation unit 204 may output the generated effective pixel map to the encoding unit 205. The effective pixel map will be described in detail in Example 2.
[0066] The encoding unit 205 acquires the depth images and viewpoint information of the second virtual camera group input from the depth image generation unit 203, and the texture images input from the texture image generation unit 204. The encoding unit 205 compresses the depth images and texture images using the compression method described above, and outputs the compressed image group and viewpoint information (metadata) of the second virtual camera group to the distribution unit 206.
[0067] The encoding unit 205 may compress not only the depth image and the texture image but also the metadata, using the above-mentioned file compression method.
[0068] Furthermore, the encoding unit 205 may compress the effective pixel map input from the texture image generation unit 204 and output it to the distribution unit 206 .
[0069] The distribution unit 206 uses the communication unit 104 to transmit the compressed depth image, texture image, and viewpoint information of the second virtual camera group input from the encoding unit 205 to the receiving unit 207 described later.
[0070] The second image processing device 30 is composed of a receiving unit 207 , a decoding unit 208 , a 3D model restoring unit 209 , a virtual camera control unit 210 , and a virtual viewpoint image generating unit 211 .
[0071] The receiving unit 207 uses the communication unit 104 to receive the compressed depth image, the compressed texture image, and the viewpoint information (metadata) of the second virtual camera group from the distribution unit 206 and outputs them to the decoding unit 208.
[0072] The decoding unit 208 decodes the compressed depth images and compressed texture images acquired from the receiving unit 207 and outputs them to the 3D model restoration unit 209 together with viewpoint information of the second virtual camera group. The decoding unit 208 may also decode metadata in addition to the depth images and texture images. Furthermore, the decoding unit 208 may decode an effective pixel map and output it to the 3D model restoration unit 209.
[0073] The 3D model restoration unit 209 restores a 3D model using the restoration method described above based on the decoded depth image, the decoded texture image, and the viewpoint information of the second virtual camera group acquired from the decoding unit 208. The restored 3D model is output to the virtual viewpoint image generation unit 211. Furthermore, the 3D model restoration unit 209 may generate color information of the 3D model using the effective pixel map acquired from the decoding unit 208. A method for using the effective pixel map will be described in detail in Example 2.
[0074] The virtual camera control unit 210 uses the communication unit 104 to generate viewpoint information of a third virtual camera for generating a virtual viewpoint image from input values input by the user through the input device 40, and outputs the viewpoint information of the third virtual camera to the virtual viewpoint image generation unit 211. The virtual camera control unit 210 may also output the generated viewpoint information of the third virtual camera designated by the user to the 3D model restoration unit 209.
[0075] The virtual viewpoint image generation unit 211 generates a virtual viewpoint image based on the 3D model acquired from the 3D model restoration unit 209 and viewpoint information of the third virtual camera acquired from the virtual camera control unit 210. The virtual viewpoint image is generated by arranging a 3D model of the subject, a 3D model of the background object, and the third virtual camera in a virtual space and generating an image viewed from the third virtual camera. The 3D model of the background object is, for example, a CG (Computer Graphics) model created separately to be combined with the subject, and is created in advance and stored in the second image processing device 30 (for example, stored in the ROM 103 in FIG. 1B ).
[0076] The 3D model of the subject and the 3D model of the background object are rendered by an existing CG rendering method. The virtual viewpoint image generator 211 transmits the generated virtual viewpoint image to the display device 50.
[0077] <Description of an Example of Generating Viewpoint Information of Second Virtual Camera Group Based on Viewpoint Information of Physical Camera Group> Here, a method for generating viewpoint information of the second virtual camera group using viewpoint information of the physical camera group will be described with reference to Figures 3A to 3C and 4. Figures 3A to 3C are schematic diagrams for explaining an example of a method for generating viewpoint information of the second virtual camera group.
[0078] 3A is a diagram showing a group of physical cameras 302 arranged in real space and an object 301. The object 301 is photographed by the group of physical cameras 302. Note that a line 303 indicates the optical axis of each of the group of physical cameras 302.
[0079] 3B is a diagram showing a first virtual camera group 305 corresponding to the physical camera group 302, and a 3D model 304 of the subject generated by the first virtual camera group 305. The first image processing apparatus 20 generates the 3D model 304 of the subject using a plurality of captured images acquired from the physical camera group 302. In order to generate a 3D model 304 in virtual space for the subject 301 in real space, the first virtual camera group 305 is generated in virtual space corresponding to the physical camera group 302 in real space. In other words, the first virtual camera group 305 is a reproduction in virtual space of the physical camera group 302 in real space. Therefore, the optical axis 306 of the first virtual camera group 305 corresponds to the optical axis 303 of the physical camera group 302. The first image processing apparatus 20 then generates the 3D model 304 of the subject 301 using viewpoint information of the generated first virtual camera group 305.
[0080] FIG. 3C is a diagram showing a second virtual camera group 307 set based on viewpoint information of the first virtual camera group 305. The second virtual camera group 307 is set on the optical axis 306 of the first virtual camera group 305. However, this is not limiting, and the second virtual camera 307 may be set at a position closer to the optical axis 306. Furthermore, the second virtual camera group 307 is set at a position closer to the 3D model 304 than the first virtual camera group 305. Generally, the second virtual camera group 307 for generating depth images and texture images for distribution places a bounding box that encompasses the 3D model 304, and places the second virtual camera group 307 on a spherical surface surrounding the bounding box so as to face the subject. This determines the position and orientation (extrinsic parameters) of the second virtual camera group 307. Then, a focal length (intrinsic parameters) that allows the second virtual camera group 307 to capture the entire 3D model 304 is determined. These are set manually, and the number and spacing of the second virtual cameras are often determined heuristically. If the 3D model 304, i.e., the subject 301, has a complex shape, occlusion occurs, whereby a portion of the 3D model 304 is hidden by another portion as viewed from a certain second virtual camera. Therefore, the number and placement of the second virtual cameras 307 must be appropriately determined. For a subject with a complex shape, if the number of second virtual cameras is small and their placement is inappropriate, resulting in frequent occlusion, the 3D model restored by the second image processing device 30 may differ significantly in shape from the 3D model 304 before distribution. However, increasing the number of second virtual cameras in an attempt to accurately restore the 3D model 304 increases the number of depth images and texture images to be distributed. Therefore, if a user attempts to restore the 3D model in an environment with low transmission bandwidth or on a local terminal with low processing performance, the frame rate may decrease. It is desirable for the 3D model displayed on the display device 50 to have higher image quality and a smooth frame rate, such as 60 fps, to provide the user with a high sense of realism. Therefore, it is necessary to appropriately set the number and arrangement of the second virtual camera group 307 and deliver depth images and high-quality texture images that can accurately reconstruct the 3D model 304 with a small amount of data. Therefore, the first image processing device 20 generates viewpoint information of the second virtual camera group 307 based on viewpoint information of the physical camera group 302 described above.That is, the second virtual camera group 307 is placed at the position of the first virtual camera group 305 corresponding to the physical camera group 302. Then, the position of the second virtual camera group 307 is set by dolly-zooming the second virtual camera group 307 on the optical axis 306 of the first virtual camera 305 while maintaining the size of the subject on the screen so as to approach the 3D model 304 of the subject 301. The position to which the second virtual camera group 307 is dolly-zoomed is assumed to be determined in advance. For example, the dolly-zoom may be performed until the distance from the 3D model 304 reaches a predetermined value. Alternatively, the dolly-zoom may be performed until the size of the 3D model 304 projected on the imaging surface of the second virtual camera group 307 reaches a predetermined size. Note that the method for determining the viewpoint information of the second virtual camera group 307 is not limited to this method. The second virtual camera group 307 may simply be placed near the optical axis 306 of the first virtual camera group 305, and internal parameters such as the focal length and the image size may be determined manually. That is, the viewpoint information of the second virtual camera group 307 may be determined based on the viewpoint information of the physical camera group 302. In the virtual viewpoint generation technology, the physical camera group 302 is arranged in the number necessary to estimate the shape of the subject 301 so as to surround the subject 301, such as by the visual volume intersection method described above. Therefore, by sending depth images as observed from the physical camera group 302, the user can accurately restore the shape information of the shape-estimated 3D model 304. Furthermore, by delivering virtual viewpoint-dependent texture images with image quality similar to that of the physical camera group 302, the user can restore high-quality color information, allowing the user to play back a high-quality 3D model.
[0081] A method for determining viewpoint information for each virtual camera will be described with reference to FIG. 4 . FIG. 4 is a schematic diagram illustrating a first virtual camera 401 capturing an image of a 3D model 402 of a subject and generating viewpoint information for a second virtual camera 403 based on the viewpoint information of the first virtual camera 401. The second virtual camera 403 moves along the optical axis of the first virtual camera 401 while dolly zooming to approach the 3D model 402. The second virtual camera 403 is then positioned at a position where it can capture the 3D model 402 within an area 404 that encompasses the 3D model 402 and allows distance information (depth information) between the 3D model and the second virtual camera 403 to be expressed in 10 bits. Note that the area 404 is not limited to a cube, and may be a sphere centered on the 3D model 402. The first virtual camera 401 captures the 3D model 402 and generates a captured image 405. The first image processing device 20 estimates the 3D model 402 based on the captured images 405 captured by the physical cameras of the imaging system 10. The estimated 3D model 402 is then photographed by a second virtual camera 403 to generate a depth image 406 and a texture image 407. The size of the 3D model 402 depicted in the depth image 406 and the texture image 407 is the same as the size of the 3D model 402 depicted in the photographed image 405. As described above, the determination of the viewpoint information of the second virtual camera 403 is not limited to dolly zoom, and the size of the 3D model 402 depicted in the depth image 406 may be different from the size of the 3D model 402 depicted in the photographed image 405. Furthermore, although the second virtual camera 403 has been described as moving on the optical axis of the first virtual camera 401, the position and orientation of the second virtual camera 403 need only be close to the position and orientation of the first virtual camera 401, and it does not necessarily have to move on the optical axis of the first virtual camera 401. Furthermore, the determination of the viewpoint information of the second virtual camera 403 may be performed once when the second image processing device 30 is initialized, or may be performed for each frame in accordance with the movement of the 3D model 402 .
[0082] <Control of Compression and Distribution of 3D Model Data and Generation of Virtual Viewpoint Images> Fig. 5 is a flowchart showing the flow of processing for controlling compression and distribution of 3D model data in the first image processing device 20 according to this embodiment. The flow shown in Fig. 5 is realized by reading a control program stored in the ROM 103 into the RAM 102 and executing it by the CPU 101. Execution of the flow in Fig. 5 is triggered when the shape information generation unit 201 receives a plurality of captured images and viewpoint information of the physical cameras from the imaging system 10.
[0083] In S501, the shape information generation unit 201 estimates and generates shape information of the subject based on multiple captured images. The generated shape information, viewpoint information of the physical camera group, and multiple captured images are output to the viewpoint determination unit 202 and the texture image generation unit 204. The generated shape information is also output to the depth image generation unit 203. Note that viewpoint information of a first virtual camera group corresponding to the physical camera group, which is used to generate the subject, is assumed to be generated in advance. For example, when the physical camera group is arranged as preparation before starting shooting, a first virtual camera group corresponding to the arranged physical camera group is generated. Shape information of the subject is generated using the viewpoint information of the first virtual camera group and multiple captured images.
[0084] In S502, the viewpoint determination unit 202 generates viewpoint information of a second virtual camera group based on viewpoint information of the physical camera group. The generated viewpoint information of the second virtual camera group is output to the depth image generation unit 203 and to the texture image generation unit 204 via the depth image generation unit 203.
[0085] The processing for generating the viewpoint information of the second virtual camera group will be described with reference to Fig. 7. Note that the generated viewpoint information of the second virtual camera group may be output to the texture image generation unit 204 without going through the depth image generation unit 203.
[0086] In S503, the depth image generation unit 203 generates a depth image of the 3D model based on the data acquired from the shape information generation unit 201 and the viewpoint determination unit 202. The generated depth image is output to the encoding unit 205.
[0087] In S504, the texture image generation unit 204 generates a texture image of the 3D model based on the data acquired from the shape information generation unit 201 and the viewpoint determination unit 202. The generated texture image is output to the encoding unit 205.
[0088] In S505, the encoding unit 205 encodes the depth image and texture image acquired from the depth image generation unit 203 and the texture image generation unit 204. The encoded data is output to the distribution unit 206.
[0089] In S506, the distribution unit 206 distributes 3D model data including the data acquired from the depth image generation unit 203 and the encoding unit 205 and the viewpoint information of the second virtual camera group, and this flow ends. In the image processing system 1, the 3D model data is distributed to the second image processing device 30, but this is not limiting. The 3D model data may also be distributed to a separate server that stores the 3D model data.
[0090] 6 is a flowchart showing a process flow for generating a virtual viewpoint image using a depth image and a texture image in the second image processing device 30 according to this embodiment. Execution of the flow in FIG. 6 is triggered by the virtual camera control unit 210 receiving an input value from the input device 40.
[0091] In S601, the virtual camera control unit 210 generates viewpoint information of a third virtual camera designated by the user based on input values from the input device 40. The generated viewpoint information of the third virtual camera is output to the virtual viewpoint image generation unit 211.
[0092] In S602, the receiving unit 207 receives the 3D model data distributed from the distribution unit 206. The received 3D model data is output to the decoding unit 208.
[0093] In S603, the decoding unit 208 decodes the depth images and texture images included in the 3D model data acquired from the receiving unit 207. If the viewpoint information of the second virtual camera group has also been encoded, the decoding unit 208 also decodes the viewpoint information of the second virtual camera group. The decoded depth images and texture images, as well as the viewpoint information of the second virtual camera group, are then output to the 3D model restoration unit 209.
[0094] In S604, the 3D model restoration unit 209 restores a 3D model of the subject based on the decoded depth image and texture image acquired from the decoding unit 208 and the viewpoint information of the second virtual camera group. The restored 3D model is output to the virtual viewpoint image generation unit 211. The 3D model restoration process will be described with reference to FIG. 8.
[0095] In S605, the virtual viewpoint image generation unit 211 generates a virtual viewpoint image based on the viewpoint information of the third virtual camera acquired from the virtual camera control unit 210 and the 3D model acquired from the 3D model restoration unit 209, and this flow ends.
[0096] After this flow is completed, the generated virtual viewpoint image is transmitted to the display device 50 and displayed on the display device 50 .
[0097] The above is the content of the control of the compression and distribution of 3D model data and the generation of virtual viewpoint images according to this embodiment.
[0098] <Description of Processing for Generating Viewpoint Information of Second Virtual Camera Group> Figure 7 is an example of a flowchart showing the flow of processing for generating viewpoint information of the second virtual camera group according to this embodiment. Here, the method for generating viewpoint information of the second virtual camera group will be described as a method for generating viewpoint information based on the dolly zoom method of the second virtual camera described with reference to Figure 4. That is, the second virtual camera moves on the optical axis of the first virtual camera from the position of the first virtual camera corresponding to the position of the physical camera in a direction approaching the 3D model of the subject. Furthermore, a method will be described in which the second virtual camera adjusts the focal length while maintaining the size of the subject on the imaging plane of the second virtual camera.
[0099] The flow shown in Fig. 7 is executed by the viewpoint determination unit 202. The execution of Fig. 7 is triggered by the reception of a 3D model of the subject, viewpoint information of the physical cameras, and multiple captured images from the shape information generation unit 201. The flow of Fig. 7 provides a detailed explanation of the control for determining the viewpoint of the second virtual camera group based on the viewpoint information of the physical cameras in S502 of Fig. 5. When the flow shown in Fig. 7 is executed, it is assumed that viewpoint information of the first virtual camera group and a 3D model of the subject have been generated in the first image processing device 20.
[0100] In S701, a 3D model of the subject, viewpoint information of a first virtual camera group, and multiple captured images are acquired. Note that viewpoint information of a physical camera group may be acquired without acquiring viewpoint information of the first virtual camera group. In this case, viewpoint information of the first virtual camera group is generated using the viewpoint information of the physical camera group.
[0101] In S702, S703 and S704 are repeatedly processed for the number of physical cameras.
[0102] In S703, S704 is repeated for each subject included in the multiple captured images. Subject recognition is performed based on the results of a face detection algorithm, a person detection algorithm, or the like. Alternatively, a 3D model generated separately for each subject can be projected onto the same plane as the imaging surface of the physical camera using viewpoint information from the physical camera, thereby identifying the area of each subject in the captured images.
[0103] In S704, viewpoint information of the second virtual camera is generated based on the 3D model and viewpoint information of the first virtual camera. The viewpoint information of the second virtual camera is generated by the method described with reference to Fig. 4. The above processing is repeated as described in S702 and S703 to generate viewpoint information of the second virtual camera group.
[0104] In S705, the viewpoint information of the generated second virtual camera group is transmitted to the depth image generation unit 203 and the texture image generation unit 204.
[0105] <Explanation of Control of 3D Model Restoration> FIG. 8 is an example of a flowchart showing the flow of 3D model restoration processing according to this embodiment. Here, the 3D model restoration processing method will be described based on a method for generating a 3D model including virtual viewpoint-dependent color information. In other words, the pixel values of the texture image of the second virtual camera group, which are closest to the position and orientation of the user-specified third virtual camera, are given priority as the color information of the 3D model. The flow shown in FIG. 8 is executed by the 3D model restoration unit 209. The flow in FIG. 8 provides a detailed description of the control for restoring a 3D model of a subject based on the decoded data in FIG. 6.
[0106] In S801, the 3D model restoration unit 209 acquires a depth image, a texture image, and viewpoint information of the second virtual camera group.
[0107] In S802, the 3D model restoration unit 209 projects the depth values of each pixel in the depth image into virtual space based on the external and internal parameters of the virtual camera corresponding to the depth image, thereby generating components of the shape information of the subject. For example, if the 3D model of the subject is represented by a 3D point cloud, each point becomes a component. The shape information of the subject is restored by performing the above process on all of the acquired multiple depth images.
[0108] In step S803, if the components of the shape information projected into the virtual space in the depth image are captured using multiple texture images, the pixel values of the texture image of the virtual camera with the closest position and orientation to the user-specified virtual camera are prioritized and determined as the color information of the components. This process is repeated for all components to restore the color information corresponding to the shape information. For example, if the shape information is a point cloud and the components are each point of the point cloud, the color information corresponding to all points is restored.
[0109] In S804, the restored 3D model is output to the virtual viewpoint image generation unit 211.
[0110] As described above, in this embodiment, viewpoint information of the second virtual camera group is generated based on viewpoint information of the physical camera group, and a depth image and a texture image are generated using the generated viewpoint information of the second virtual camera group. This process enables the receiving side to restore a high-quality 3D model without increasing the amount of data, even for 3D models with complex shapes. While this embodiment describes the delivery of compressed 3D model data, the compressed 3D model data may be stored or may be delivered in response to a client request. Furthermore, while the method of generating a 3D model including virtual viewpoint-dependent color information during 3D model restoration has been described as an example, this is not limiting. After the restoration of the 3D model's shape information, color information for the 3D model may be generated simultaneously with the generation of a virtual viewpoint image. In this case, it is not necessary to assign color information to all of the shape information of the 3D model; color information may be generated only for the shooting range (viewing angle) of the user-specified third virtual camera.
[0111] <Example 2> Example 1 described a process of generating viewpoint information of a second virtual camera group based on viewpoint information of a physical camera group, and generating a depth image and a texture image using the generated viewpoint information of the second virtual camera group. Next, Example 2 describes an aspect in which an effective pixel map is generated in addition to a depth image and a texture image to deal with cases in which a subject is occluded by another subject or object. Note that the description will be omitted or simplified for parts common to Example 1, such as the hardware configuration and functional configuration of the image processing device.
[0112] 9 is a schematic diagram showing an example of a method for generating an effective pixel map according to this embodiment. A subject is photographed by a physical camera, and a photographed image 903 is generated. At this time, the subject in the photographed image 903 has a part of its body hidden by an obstruction. The obstruction is assumed to be the subject.
[0113] The first image processing device 20 estimates the shapes of the subject and the obstructing object based on multiple captured images captured by the physical cameras of the imaging system 10, and generates shape information for each. That is, it generates a 3D model 901 of the subject and a 3D model 904 of the obstructing object. Next, the first image processing device 20 generates viewpoint information for the second virtual camera group described in Example 1 based on data such as viewpoint information and shape information of the physical camera group. The second virtual camera 905 is positioned on the optical axis of the first virtual camera 902, within an area 906 that encompasses the 3D model 901 of the subject and that allows the depth of the subject to be expressed in 10 bits. The focal length of the second virtual camera 905 is adjusted so that the size of the subject 901 captured by the first virtual camera 902 is maintained. After generating the viewpoint information for the second virtual camera 905, the first image processing device 20 generates a depth image 907 and a texture image 908. Here, for each pixel value of the subject 901 appearing in the generated texture image 908, the pixel values of the image captured by the first virtual camera 902 are used preferentially in the area not obscured by the 3D model 904 of the obstructing object (the left side of the subject). On the other hand, in the area obscured by the 3D model 904 of the obstructing object (the right side of the subject), the pixel values of the image captured by the first virtual camera, which has a different viewpoint from the first virtual camera 902, are used. Therefore, there is a possibility that the image quality of the right and left sides of the subject appearing in the texture image 908 will differ significantly. If the color information of the 3D model is restored using this texture image 908, the image quality of parts of the 3D model will be reduced, and this part will stand out, which may cause the user to feel uncomfortable. Therefore, an effective pixel map 909 is generated that indicates high-quality areas of the texture image. For example, the effective pixel map assigns pixel values of 1 to unobstructed areas and 0 to obstructed areas. The image size of the effective pixel map 909 is the same as the image size of the texture image 908. The determination of whether or not the area is an occluded area is made, for example, based on whether or not each point constituting the shape information of the 3D model 901 of the subject is visible to the first virtual camera 902 through the visibility determination process described above. In other words, through the visibility determination process, pixel values of the area of the occluding object 904 captured in the captured image 903 are not used as color information of the 3D model 901 of the subject.Therefore, a region of the texture image determined using pixel values of the image captured by the physical camera from which the viewpoint information of the second virtual camera 905 is generated is identified, and in the effective pixel map, pixel values of the identified region are set to 1, and other regions are set to 0. Note that the method of determining an occluded region is not limited to this. When restoring color information of a 3D model, the second image processing device 30 prioritizes pixel values of regions of the texture image corresponding to regions with a value of 1 in the effective pixel map and uses them as color information for the 3D model. Specifically, when generating color information for the restored shape information in a virtual viewpoint-dependent manner, the second image processing device 30 uses pixel values of the texture image corresponding to 1 in the effective pixel map and does not use pixel values of the texture image corresponding to 0 in the effective pixel map. However, if color information for shape information to which color information is to be added is stored only in the texture image corresponding to 0 in the effective pixel map, the pixel values of that texture image are used. Note that while the pixel values of the effective pixel map have been described so far as being binary (0 or 1), they may also be multi-valued. In this case, the pixel values of the effective pixel map may be used for weighting when generating color information for the 3D model. In other words, the priority of pixel values of the texture image is determined according to the pixel values of the effective pixel map. When generating a valid pixel map with 255 levels of multi-values, for example, the pixel values at the contour of the subject or the boundary with an obstructing object may be set to 0, and may be linearly increased to 255 as the subject approaches a certain distance (for example, 5 px) inward to the subject. This reduces the influence of pixel values of the texture image at the contour or boundary with an obstructing object that are unreliable when restoring color information of the 3D model, making it possible to generate color information of the 3D model using highly reliable pixel values.
[0114] 10 is a flowchart showing the flow of processing according to this embodiment for controlling the compression and distribution of 3D model data in the first image processing device 20. Execution of the flow in FIG. 10 is started when the shape information generation unit 201 receives a plurality of captured images and viewpoint information of the physical cameras from the imaging system 10.
[0115] S1001 to S1003 are the same as S501 to S503 in FIG.
[0116] In S1004, the texture image generation unit 204 generates a texture image and a valid pixel map of the foreground model based on the data acquired from the shape information generation unit 201 and the viewpoint determination unit 202. The generated texture image and valid pixel map are output to the encoding unit 205.
[0117] In S1005, the encoding unit 205 encodes the depth image, texture image, and valid pixel map acquired from the depth image generation unit 203 and the texture image generation unit 204. The encoded depth image, texture image, and valid pixel map are output to the distribution unit 206.
[0118] In S1006, the distribution unit 206 transmits 3D model data including the depth image, texture image, effective pixel map, and viewpoint information of the second virtual camera group obtained from the depth image generation unit 203 and the encoding unit 205 to the receiving unit 207, and this flow ends.
[0119] Fig. 11 is an example of a flowchart showing the flow of a 3D model restoration process according to this embodiment. The flow of Fig. 11 is executed by the 3D model restoration unit 209. The flow of Fig. 11 provides a detailed explanation of the control for restoring a 3D model of a subject based on the data decoded in S604 of Fig. 6.
[0120] In S1101, a depth image, a texture image, an effective pixel map, and viewpoint information of a second virtual camera group are acquired.
[0121] In S1102, the depth value of each pixel of the depth image is projected onto a virtual space based on the external parameters and internal parameters of the second virtual camera corresponding to the depth image, and shape information of the subject is restored.
[0122] In step S1103, if the shape information is captured from multiple texture images, the pixel values of the texture image captured by the second virtual camera, which is closest in position and orientation to the user-specified third virtual camera, are prioritized as color information. At this time, the priority of the pixel values of the texture image is determined according to the pixel values of the valid pixel map corresponding to the pixel values of the texture image. If the shape information is represented by a 3D point cloud, this process is repeated for all points in the point cloud to generate color information corresponding to the shape information.
[0123] S1104 is the same as S804 in FIG.
[0124] After this flow is completed, the virtual viewpoint image generation unit 211 generates a virtual viewpoint image based on the restored 3D model and the viewpoint information of the third virtual camera specified by the user, and the generated virtual viewpoint image is displayed on the display device 50.
[0125] The above explains the compression and distribution of 3D model data including the effective pixel map and the control of 3D model restoration. This process makes it possible to restore a 3D model using highly reliable pixel values of a texture image even when the subject is hidden by an obstacle.
[0126] The present disclosure can also be realized by a process in which a program that realizes one or more functions of the above-described embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in a computer of the system or device read and execute the program. The present disclosure can also be realized by a circuit (e.g., an ASIC) that realizes one or more functions.
[0127] The disclosure of the present embodiment includes the following configurations, methods, systems, and programs.
[0128] (Configuration 1) An image processing device comprising: a setting means for setting the positions and orientations of a plurality of second virtual cameras based on the optical axes of a plurality of first virtual cameras in a virtual space corresponding to the positions and orientations of a plurality of imaging devices in real space; and a generation means for generating a plurality of depth images indicating the distance between each of the plurality of second virtual cameras and a 3D model of a subject generated based on a plurality of captured images acquired by the plurality of imaging devices.
[0129] (Configuration 2) The image processing device according to Configuration 1, wherein the plurality of second virtual cameras are set on the optical axes of the plurality of first virtual cameras.
[0130] (Configuration 3) The image processing device according to configuration 1, further comprising encoding means for encoding the depth image.
[0131] (Configuration 4) The image processing device according to Configuration 3, further comprising an output means for outputting the encoded depth images and viewpoint information indicating positions and orientations of the second virtual cameras to another device that reconstructs the 3D model based on the encoded depth images and the viewpoint information.
[0132] (Configuration 5) The image processing device according to Configuration 4, wherein the generating means generates a virtual viewpoint image including a 3D model of the subject for each of the plurality of second virtual cameras, the encoding means encodes the plurality of virtual viewpoint images, and the output means outputs the encoded plurality of virtual viewpoint images to the other device.
[0133] (Configuration 6) The image processing device described in Configuration 5, wherein the generating means generates correspondence information indicating whether each pixel in the virtual viewpoint image corresponds to a component of a 3D model of the subject for each of the plurality of second virtual cameras, the encoding means encodes the plurality of pieces of correspondence information, and the output means outputs the encoded plurality of pieces of correspondence information to the other device.
[0134] (Configuration 7) The image processing device according to configuration 1, wherein the second virtual cameras are set at positions closer to the 3D model of the subject than the first virtual cameras.
[0135] (Configuration 8) The image processing device according to configuration 1, wherein the generating means generates one depth image for one second virtual camera.
[0136] (Configuration 9) An image processing device comprising: an acquisition means for acquiring encoded depth images indicating the distance between each of a plurality of second virtual cameras set based on the optical axes of a plurality of first virtual cameras in a virtual space corresponding to the positions and orientations of a plurality of imaging devices in real space, and a 3D model of a subject generated based on a plurality of captured images acquired by the imaging devices, and first viewpoint information indicating the positions and orientations of the plurality of second virtual cameras; a decoding means for decoding the encoded depth images; and a generation means for generating a 3D model of the subject based on the decoded depth images and the first viewpoint information.
[0137] (Configuration 10) The image processing device according to Configuration 9, further comprising: acquiring second viewpoint information indicating a position and an attitude of a third virtual camera different from the first virtual camera and the second virtual camera; and the generating means generating a virtual viewpoint image based on the second viewpoint information and a 3D model of the subject.
[0138] (Method 1) An image processing method comprising: a setting step of setting positions and orientations of a plurality of second virtual cameras based on the optical axes of a plurality of first virtual cameras in a virtual space corresponding to the positions and orientations of a plurality of imaging devices in real space; and a generation step of generating a plurality of depth images indicating the distance between each of the plurality of second virtual cameras and a 3D model of a subject generated based on a plurality of captured images acquired by the plurality of imaging devices.
[0139] (Method 2) An image processing method comprising: an acquisition step of acquiring encoded depth images indicating the distance between each of a plurality of second virtual cameras set based on the optical axes of a plurality of first virtual cameras in a virtual space corresponding to the positions and orientations of a plurality of imaging devices in real space, and a 3D model of a subject generated based on a plurality of captured images acquired by the imaging devices, and first viewpoint information indicating the positions and orientations of the plurality of second virtual cameras; a decoding step of decoding the encoded depth images; and a generation step of generating a 3D model of the subject based on the decoded depth images and the first viewpoint information.
[0140] (Program) A program for causing a computer to function as each of the means of the image processing device according to any one of the first to tenth aspects.
[0141] The present invention is not limited to the above-described embodiments, and various modifications and variations can be made without departing from the spirit and scope of the present invention. Therefore, the following claims are appended to apprise the public of the scope of the present invention.
[0142] This application claims priority based on Japanese Patent Application No. 2024-073062, filed April 26, 2024, the entire contents of which are incorporated herein by reference.
[0143] 202 viewpoint determination unit 203 depth image generation unit 204 texture image generation unit 205 encoding unit
Claims
1. An image processing device comprising: a setting means for setting the positions and orientations of a plurality of second virtual cameras based on the optical axes of a plurality of first virtual cameras in a virtual space corresponding to the positions and orientations of a plurality of imaging devices in real space; and a generation means for generating a plurality of depth images indicating the distance between each of the plurality of second virtual cameras and a 3D model of a subject generated based on a plurality of captured images acquired by the plurality of imaging devices.
2. The image processing device according to claim 1, wherein the plurality of second virtual cameras are set on the optical axes of the plurality of first virtual cameras.
3. An image processing apparatus according to claim 1, further comprising encoding means for encoding said depth image.
4. The image processing device described in claim 3, characterized in that it has an output means for outputting the encoded plurality of depth images and viewpoint information indicating the positions and orientations of the plurality of second virtual cameras to another device that reconstructs the 3D model based on the encoded plurality of depth images and the viewpoint information.
5. The image processing device described in claim 4, characterized in that the generation means generates a virtual viewpoint image including a 3D model of the subject for each of the plurality of second virtual cameras, the encoding means encodes the plurality of virtual viewpoint images, and the output means outputs the encoded plurality of virtual viewpoint images to the other device.
6. The image processing device described in claim 5, characterized in that the generation means generates correspondence information indicating whether each pixel in the virtual viewpoint image corresponds to a component of the 3D model of the subject for each of the multiple second virtual cameras, the encoding means encodes the multiple pieces of correspondence information, and the output means outputs the encoded multiple pieces of correspondence information to the other device.
7. The image processing device according to claim 1, wherein the second virtual cameras are set at positions closer to the 3D model of the subject than the first virtual cameras.
8. The image processing device according to claim 1, wherein said generating means generates one depth image for one second virtual camera.
9. An image processing device comprising: an acquisition means for acquiring encoded depth images indicating the distance between each of a plurality of second virtual cameras set based on the optical axes of a plurality of first virtual cameras in a virtual space corresponding to the positions and orientations of a plurality of imaging devices in real space, and a 3D model of a subject generated based on a plurality of captured images acquired by the imaging devices, and first viewpoint information indicating the positions and orientations of the plurality of second virtual cameras; a decoding means for decoding the encoded depth images; and a generation means for generating a 3D model of the subject based on the decoded depth images and the first viewpoint information.
10. The image processing device described in claim 9, characterized in that second viewpoint information indicating the position and attitude of a third virtual camera different from the first virtual camera and the second virtual camera is acquired, and the generation means generates a virtual viewpoint image based on the second viewpoint information and a 3D model of the subject.
11. An image processing method comprising: a setting step of setting the positions and orientations of a plurality of second virtual cameras based on the optical axes of a plurality of first virtual cameras in a virtual space corresponding to the positions and orientations of a plurality of imaging devices in real space; and a generation step of generating a plurality of depth images indicating the distance between each of the plurality of second virtual cameras and a 3D model of a subject generated based on a plurality of captured images acquired by the plurality of imaging devices.
12. An image processing method comprising: an acquisition step of acquiring encoded depth images indicating the distance between each of a plurality of second virtual cameras set based on the optical axes of a plurality of first virtual cameras in a virtual space corresponding to the positions and orientations of a plurality of imaging devices in real space, and a 3D model of a subject generated based on a plurality of captured images acquired by the imaging devices, and first viewpoint information indicating the positions and orientations of the plurality of second virtual cameras; a decoding step of decoding the encoded depth images; and a generation step of generating a 3D model of the subject based on the decoded depth images and the first viewpoint information.
13. A program for causing a computer to function as each means of the image processing device according to any one of claims 1 to 10.
Citation Information
Patent Citations
3D data generation device, 3D data playback device, control program, and recording medium
WO2020067441A1
Image processing device and image processing method
WO2020184174A1