Image processing device, image processing method and program
The image processing apparatus enhances virtual viewpoint image quality by setting a generation method based on object characteristics, adaptively choosing between shape-based and radiance field-based rendering methods to handle complex shapes and partial imaging effectively.
Patent Information
- Application Number
- JP2023200974
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-28
- Publication Date
- 2025-06-09
AI Technical Summary
Existing techniques for generating virtual viewpoint images struggle with accurately estimating complex object shapes, leading to deteriorated image quality, especially when the object is located in areas not fully imaged by multiple devices.
An image processing apparatus that acquires data from multiple captured images and sets a generation method for virtual viewpoint images based on object characteristics, such as shape complexity and size per pixel, to determine whether to use a rendering method based on shape or radiance field.
This approach improves the image quality of virtual viewpoint images by adaptively selecting the appropriate rendering method based on object characteristics, effectively addressing the challenges of complex shapes and partial imaging.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an image processing technique for generating a virtual viewpoint image.
Background Art
[0002] There is a technique for estimating an object shape using a plurality of captured images obtained by capturing an object from various directions, and reconstructing an image (virtual viewpoint image) corresponding to an image of the object when viewed from an arbitrary virtual viewpoint. However, due to the occurrence of a captured image that does not include an image of the object depending on the position of the object, the object shape (hereinafter referred to as "object shape") cannot be accurately estimated, and the image quality of the virtual viewpoint image may deteriorate. Patent Document 1 discloses a technique for selecting a method for generating a virtual viewpoint image output based on the position of an object. Specifically, the technique disclosed in Patent Document 1 outputs a virtual viewpoint image generated using an object shape when the object is located in an area imaged by a plurality of imaging devices. On the other hand, when the object is located in an area not imaged by some of the plurality of imaging devices, a virtual viewpoint image generated without using the object shape is output.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, even if the object is located in an area imaged by a plurality of imaging devices, if the object shape when viewed from the virtual viewpoint is complex, etc., the object shape cannot be accurately estimated, and the image quality of the virtual viewpoint image may deteriorate.
[0005] Therefore, an object of the present disclosure is to provide a technique capable of generating a high-quality virtual viewpoint image even in the case described above, for example. **Means for Solving the Problems**
[0006] An image processing apparatus according to the present disclosure includes: an image acquisition unit that acquires data of a plurality of captured images obtained by capturing an object from a plurality of directions; a viewpoint acquisition unit that acquires virtual viewpoint information regarding the position of a virtual viewpoint and the line-of-sight direction at the virtual viewpoint; a setting unit that sets a generation method of a virtual viewpoint image based on the characteristics of the object acquired based on the virtual viewpoint information; and a generation unit that generates the virtual viewpoint image based on the virtual viewpoint information and the set generation method of the virtual viewpoint image. **Advantages of the Invention**
[0007] According to the technique of the present disclosure, the image quality of the virtual viewpoint image can be improved. **Brief Description of the Drawings**
[0008]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Mode for Carrying Out the Invention
[0009] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. Note that the following embodiments do not necessarily limit the present disclosure. Also, not all combinations of the features described in the following embodiments are essential for the solution means of the present disclosure.
[0010] [Embodiment 1] In Embodiment 1, a rendering method (generation method) is set based on the characteristics of an object when viewed from a virtual viewpoint, and a virtual viewpoint image corresponding to the virtual viewpoint is generated by the set rendering method. For the generation of the virtual viewpoint image, data of a plurality of captured images (hereinafter referred to as "multi-viewpoint images") obtained by capturing from various directions by a plurality of imaging devices is used. The rendering method of the virtual viewpoint image is set based on the complexity of the object shape when the object is viewed from the virtual viewpoint and the size of the object included as an image per pixel in the virtual viewpoint image when the virtual viewpoint image corresponding to the virtual viewpoint is generated. Hereinafter, the "size of the object included as an image per pixel" will be described as the "object size per pixel".
[0011] As the rendering method of the virtual viewpoint image, either a rendering method based on the object shape or a rendering method based on the radiance field is set. When the virtual viewpoint image is generated as a moving image, the rendering method is set for each frame of the virtual viewpoint image. In addition, in Embodiment 1, it is described that the focal lengths of all the imaging devices are equal to each other.
[0012] <Configuration of Image Processing System> FIG. 1 is a diagram showing an example of the configuration of an image processing system according to Embodiment 1. The image processing system according to Embodiment 1 includes a plurality of imaging devices 101, an image processing device 102, a user interface (hereinafter referred to as "UI") panel 103, a storage device 104, and a display device 105. The plurality of imaging devices 101 perform synchronized imaging of the object 107 existing in the imaging region 106 from various directions according to the imaging conditions, and output data of the captured images obtained by the imaging (hereinafter referred to as "captured image data") to the image processing device 102. The image processing device 102 acquires a plurality of captured image data (multi-viewpoint image data) output from the plurality of imaging devices 101, and generates a virtual viewpoint image using the acquired multi-viewpoint image data. The data or signal of the virtual viewpoint image generated by the image processing device 102 is output to an external device.
[0013] The UI panel 103 includes a display device such as a liquid crystal display, and displays a user interface for presenting to the user the imaging conditions of the imaging device 101 and the processing settings of the image processing device 102, etc. The UI panel 103 may include an input device such as a touch panel or buttons. In this case, the UI panel 103 receives instructions from the user regarding changes to the above-described imaging conditions and processing settings, etc. When the input device receives an instruction from the user, the UI panel 103 transmits information indicating the instruction to the image processing device 102. Note that the input device may be provided separately from the UI panel 103, such as a mouse or a keyboard. The storage device 104 stores data of the virtual viewpoint image output by the image processing device 102. The display device 105 is configured by a liquid crystal display or the like, and receives a signal indicating the virtual viewpoint image output from the image processing device 102 and displays the virtual viewpoint image.
[0014] <Hardware Configuration of Image Processing Device> FIG. 2 is a block diagram showing an example of the hardware configuration of the image processing device 102 according to Embodiment 1. The image processing device 102 includes, as its hardware configuration, a CPU 201, a RAM 202, a ROM 203, a storage device 204, a control interface (hereinafter referred to as "I / F") 205, an input I / F 206, an output I / F 207, and a main bus 208. The CPU 201 is a processor that comprehensively controls each part of the image processing device 102. The RAM 202 functions as the main memory and work area of the CPU 201, etc. The ROM 203 stores one or more programs executed by the CPU 201. The storage device 204 is configured by a hard disk drive or the like, and stores application programs executed by the CPU 201 and data used for the processing of the CPU 201, etc.
[0015] The control I / F 205 is connected to each imaging device 101 and is a communication interface for controlling settings of imaging conditions, start of imaging, stop of imaging, etc. for each imaging device 101. The input I / F 206 is a communication interface via a serial bus such as SDI (Serial Digital Interface) or HDMI (registered trademark) (High-Definition Multimedia Interface (registered trademark)). Imaging image data is acquired from each imaging device 101 via the input I / F 206. The output I / F 207 is a communication interface via a serial bus such as USB (Universal Serial Bus) or DP (DisplayPort (registered trademark)). Data of the virtual viewpoint image or a signal indicating the virtual viewpoint image is output to the storage device 104 or the display device 105 via the output I / F 207. The main bus 208 is a transmission path that communicably connects the above-described hardware configurations of the image processing device 102 to each other.
[0016] <Functional Configuration of Image Processing Device> FIG. 3 is a block diagram showing an example of the functional configuration of the image processing device 102 according to Embodiment 1. The image processing device 102 includes, as functional components, an image acquisition unit 301, a viewpoint acquisition unit 302, a method setting unit 303, an image generation unit 304, and an image output unit 305. The image acquisition unit 301 acquires imaging image data and camera parameters (hereinafter referred to as "imaging camera parameters") corresponding to the imaging image data from each imaging device 101 that images the object 107. The viewpoint acquisition unit 302 acquires information (hereinafter referred to as "virtual camera parameters") indicating the position of the virtual viewpoint and the direction of the line of sight at the virtual viewpoint, which corresponds to the camera parameters of an imaging device (hereinafter referred to as "virtual camera") that is virtually arranged at the virtual viewpoint.
[0017] The mode setting unit 303 sets the rendering mode when generating the virtual viewpoint image based on the characteristics of the object 107 when viewed from the virtual viewpoint. Specifically, first, the mode setting unit 303 identifies the characteristics of the object 107 based on the captured image data and the imaging camera parameters corresponding to each imaging device 101 acquired by the image acquisition unit 301 and the virtual camera parameters acquired by the viewpoint acquisition unit 302. Subsequently, the mode setting unit 303 determines the rendering mode based on the identified characteristics of the object 107 and sets the determined rendering mode. The image generation unit 304 generates a virtual viewpoint image corresponding to the image when viewing the imaging region 106 from the virtual viewpoint according to the rendering mode set by the mode setting unit 303. Specifically, the image generation unit 304 uses the captured image data and the imaging camera parameters corresponding to each imaging device 101 acquired by the image acquisition unit 301 and the virtual camera parameters acquired by the viewpoint acquisition unit 302 to generate the virtual viewpoint image according to the set rendering mode. The image output unit 305 outputs the data of the virtual viewpoint image or a signal indicating the virtual viewpoint image to the storage device 104, the display device 105, or the like.
[0018] <Operation of the Image Processing Apparatus> FIG. 4 is a flowchart showing an example of the processing flow of the image processing apparatus 102 according to Embodiment 1. Note that "S" attached to the beginning of the reference numerals means steps (processes). Further, the processing of each step shown in the flowchart of FIG. 4 is realized by the CPU 201 reading a predetermined program from the ROM 203 or the storage device 204 and expanding it in the RAM 202 and then executing this by the CPU 201.
[0019] First, in S401, the image acquisition unit 301 acquires captured image data and captured camera parameters corresponding to the captured image data from each imaging device 101 via the input I / F 206. The source of the captured image data and the captured camera parameters is not limited to the imaging device 101. For example, the image acquisition unit 301 may acquire the captured image data or the captured camera parameters by reading them from a storage device 104 or a storage device 204 that stores the captured image data or the captured camera parameters in advance. Specifically, for example, when the captured camera parameters are calculated in advance by calibration or the like and stored in the storage device 204, the image acquisition unit 301 acquires the captured camera parameters by reading them from the storage device 204. The captured image data and the captured camera parameters acquired by the image acquisition unit 301 are associated with each other and held in the RAM 202.
[0020] FIG. 5 is a diagram showing an example of the arrangement of the imaging devices 101 according to the first embodiment and an example of a captured image obtained by imaging with the imaging device 101. FIG. 5(a) shows an example of the arrangement of each imaging device 101, and a plurality of imaging devices 101 are arranged so as to be able to image an object 107 existing in the imaging region 106 from various directions. Note that the imaging devices 101a, b, and c shown in FIG. 5(a) are the same as other imaging devices 101. FIG. 5(b) shows an example of a captured image 501 obtained by imaging with the imaging device 101a, FIG. 5(c) shows an example of a captured image 502 obtained by imaging with the imaging device 101b, and FIG. 5(d) shows an example of a captured image 503 obtained by imaging with the imaging device 101c. In the first embodiment, the focal lengths of all the imaging devices 101 including the imaging devices 101a, b, and c are described as being equal to each other.
[0021] After S401, in S402, the viewpoint acquisition unit 302 acquires virtual camera parameters. The virtual camera parameters may be set based on an instruction from the UI panel 103, or may be pre-set and stored in advance in the storage device 204. Next, in S403, the method setting unit 303 executes a setting process for the rendering method. By executing the setting process, the method setting unit 303 sets the rendering method for generating a virtual viewpoint image. Details of the setting process for the rendering method in S403 will be described later. After S403, in S404, the image generation unit 304 executes a generation process for the virtual viewpoint image. By executing the generation process, the image generation unit 304 generates a virtual viewpoint image corresponding to the virtual viewpoint indicated by the virtual camera parameters acquired in S402. Details of the generation process for the virtual viewpoint image in S404 will be described later. After S404, in S405, the image output unit 305 outputs data of the virtual viewpoint image or a signal indicating the virtual viewpoint image to the storage device 104 or the display device 105 via the output I / F 207. After S405, the image processing device 102 ends the processing of the flowchart shown in FIG. 4.
[0022] <Setting process for rendering method> FIG. 6 is a flowchart showing an example of the flow of the setting process for the rendering method in the method setting unit 303 according to Embodiment 1, and is a flowchart showing an example of the flow of the process in S403. By executing the processing of the flowchart, the method setting unit 303 sets the rendering method for generating a virtual viewpoint image based on the characteristics of the object 107 when the object 107 is viewed from the virtual viewpoint.
[0023] After S402 shown in FIG. 4, in S601, the method setting unit 303 acquires a schematic shape of the object 107 based on the plurality of captured image data acquired in S401 and the imaging camera parameters corresponding to each captured image data. For example, the method setting unit 303 acquires, as the schematic shape of the object 107, information represented as a set of voxels by the volume intersection method.
[0024] Specifically, in S601, first, the method setting unit 303 acquires data of a captured image (hereinafter referred to as a "background image") obtained by previously capturing only the imaging region 106 where each imaging device 101 hits the background in a state where the object 107 does not exist. Subsequently in S601, the method setting unit 303 generates a silhouette image of the object 107 corresponding to each captured image based on the difference between the background image corresponding to each imaging device 101 and the captured image corresponding to the background image acquired in S401. Subsequently in S601, the method setting unit 303 projects each voxel included in the set of voxels in the virtual space corresponding to the imaging region 106 onto each generated silhouette image based on the imaging camera parameters. Subsequently in S601, the method setting unit 303 acquires, as the approximate shape of the object 107, the set of voxels projected onto the silhouette of the object 107 for all the silhouette images.
[0025] FIG. 7 is a diagram showing an example of the approximate shape of the object 107 acquired by the visual volume intersection method according to Embodiment 1. In FIG. 7, the broken lines extending from each imaging device 101 indicate the boundaries when the outer shape of the region corresponding to the object 107 in each silhouette image is projected onto each voxel included in the set of voxels in the virtual space. Also in FIG. 7, the polygon 701 indicated by the thin solid line represents the surface of the object 107, and the polygon 702 indicated by the thick solid line represents the approximate shape of the object 107 acquired by the visual volume intersection method.
[0026] After S601, in S602, the method setting unit 303 obtains the object size per pixel in the virtual viewpoint image corresponding to the image of the object 107 when viewed from the virtual viewpoint based on the object schematic shape acquired in S601. Here, the "object size per pixel in the virtual viewpoint image corresponding to the image when viewed from the virtual viewpoint" means the "object size per pixel in the virtual viewpoint image when the virtual viewpoint image corresponding to the virtual viewpoint is generated". Hereinafter, the "object size per pixel in the virtual viewpoint image corresponding to the image of the object 107 when viewed from the virtual viewpoint" will be simply referred to as the "object size per pixel when viewed from the virtual viewpoint" for explanation. In Embodiment 1, the centroid coordinates of the schematic shape of the object 107 represented as a set of voxels are used as the representative point, and the object size per pixel when viewed from the virtual viewpoint is obtained based on the representative point and the virtual camera parameters. Specifically, for example, the object size per pixel can be calculated using Equation (1).
[0027] s = d / F ··· Equation (1) Here, s is the object size per pixel, d is the distance from the virtual camera to the representative point when the representative point of the schematic shape of the object 107 is projected in the optical axis direction of the virtual camera, that is, the direction of the line of sight at the virtual viewpoint. Also, F indicates the focal length of the virtual camera. Note that the focal length of the virtual camera is a value converted into pixel units based on the size of the pixels in the virtual viewpoint image obtained by pseudo-imaging by the virtual camera.
[0028] After S602, in S603, the method setting unit 303 obtains information indicating the complexity of the object shape when the object 107 is viewed from the virtual viewpoint based on the object approximate shape obtained in S601. The information indicating the complexity of the object shape can be obtained, for example, by the following method. In S603, first, the method setting unit 303 selects two imaging devices 101 from among the plurality of imaging devices 101 in the same direction as the line-of-sight direction at the virtual viewpoint (where "the same" includes "substantially the same"), or in the two optical axis directions closest to the line-of-sight direction at the virtual viewpoint.
[0029] Subsequently in S603, the method setting unit 303 selects one of the two captured images obtained by the two selected imaging devices 101 capturing the object 107 as the reference image and the other as the target image. Subsequently in S603, the method setting unit 303 obtains the pixel values of the corresponding pixels in the regions corresponding to the approximate shape in the reference image and the target image when the target image is projected onto the reference image based on the approximate shape of the object 107. Subsequently in S603, the method setting unit 303 obtains the difference between the obtained pixel values as a value indicating the complexity of the object shape. Specifically, for example, the value indicating the complexity of the object shape can be calculated using Equation (2).
[0030] c = ((Σ i ((y´ i -y i )) 2 )) / n) 1 / 2 ··· Equation (2) Here, c is a value indicating the complexity of the object shape. Also, y i is the luminance value of the pixel (i) included in the region corresponding to the approximate shape of the object 107 in the reference image. Also, y′ iis the luminance value of the pixel corresponding to the pixel (i) of the reference image, which is included in the region corresponding to the object 107 in the target image when the target image is projected onto the reference image based on the schematic shape. Also, n is the number of pixels included in the region corresponding to the object 107 in the reference image and the target image. In Equation (2), as an example, the root mean square is calculated using the luminance values of the pixels included in the region corresponding to the object 107 in the reference image and the target image, and this is used as the value (c) indicating the complexity of the object shape. However, the method for calculating the value (c) is not limited to this, and a simple average or the like may be used.
[0031] The more complex the actual shape of the object 107 is, the more difficult it becomes to accurately represent the actual shape of the object 107 by the schematic shape of the object 107. Therefore, the more complex the actual shape of the object 107 is, the more difficult it is to accurately project the target image onto the reference image. Accordingly, the value (c) indicating the complexity of the object shape calculated by the root mean square increases as the actual shape of the object 107 becomes more complex.
[0032] After S603, in S604, the method setting unit 303 sets the rendering method of the virtual viewpoint image based on the object size (s) per pixel and the value (c) indicating the complexity of the object shape when viewed from the virtual viewpoint. Specifically, the method setting unit 303 sets either the rendering method based on the shape or the rendering method based on the radiance field as the rendering method of the virtual viewpoint image based on the object size (s) and the value (c) indicating the complexity. Although the details of each rendering method will be described later, the quality of the virtual viewpoint image generated by the image processing apparatus 102 based on the multi-viewpoint image data is as follows according to the characteristics of the object 107 when viewed from the virtual viewpoint. When the object shape is complex and the object size per pixel is small, the quality of the virtual viewpoint image generated by the rendering method based on the shape deteriorates. In other cases, regardless of the rendering method, the quality of the virtual viewpoint image generated by the above two rendering methods is approximately the same.
[0033] It is known that the rendering method based on the shape generally requires less computation amount to generate a virtual viewpoint image compared to the rendering method based on the radiance field. Therefore, the method setting unit 303 determines and sets the rendering method as follows. First, the method setting unit 303 determines that the degree of the object size is "large" when the object size (s) per pixel when viewed from the virtual viewpoint is equal to or greater than a predetermined first threshold value (t s ). Also, the method setting unit 303 determines that the degree of the object size is "small" when the object size (s) is less than the first threshold value (t s ). Further, the method setting unit 303 determines that the degree of the complexity of the object shape is "complex" when the value (c) indicating the complexity of the object shape when viewed from the virtual viewpoint is equal to or greater than a predetermined second threshold value (t c ). Also, the method setting unit 303 determines that the degree of the complexity of the object shape is "simple" when the value (c) indicating the complexity is less than the second threshold value (t c ).
[0034] FIG. 8 is a diagram showing an example of a look-up table (hereinafter referred to as "LUT") used by the method setting unit 303 according to Embodiment 1 when setting a rendering method. In the LUT shown in FIG. 8, the rendering method, the degree of object size per pixel, and the degree of object shape complexity as viewed from the virtual viewpoint are managed in association with each other. The method setting unit 303 according to Embodiment 1 uses the LUT shown in FIG. 8(a) to determine and set the rendering method of the virtual viewpoint image according to the degree of object size and the degree of object shape complexity.
[0035] Specifically, the method setting unit 303 sets the rendering method based on the radiance field as the rendering method of the virtual viewpoint image only when the degree of object shape complexity is "complex" and the degree of object size per pixel is "small". In other cases, the method setting unit 303 sets the rendering method based on the shape as the rendering method of the virtual viewpoint image. The first threshold value (t s ) and the second threshold value (t c ) may be prescribed values, or may be values set based on an instruction from the user obtained via the UI panel 103 or the like. After the process of S604, the method setting unit 303 ends the process of the flowchart shown in FIG. 6, that is, the process of S403 shown in FIG. 4. The LUTs shown in FIGS. 8(b) and 8(c) will be described later.
[0036] <Generation process of virtual viewpoint image> The image generation unit 304 generates a virtual viewpoint image corresponding to the virtual viewpoint indicated by the virtual camera parameters by the rendering method set in S403, using the multi-viewpoint image data acquired in S401 and the virtual camera parameters acquired in S402. Hereinafter, with reference to FIG. 9, the details of the virtual viewpoint image generation process in the image generation unit 304 will be described. FIG. 9 is a flowchart showing an example of the flow of the virtual viewpoint image generation process in the image generation unit 304 according to Embodiment 1, and is a flowchart showing an example of the flow of the virtual viewpoint image generation process in S404 shown in FIG. 4.
[0037] After S403 shown in FIG. 4, in S901, the image generation unit 304 determines whether the rendering method set in S403 is a rendering method based on shape. If it is determined in S901 that the rendering method set is a rendering method based on shape, then in S902, the image generation unit 304 generates a virtual viewpoint image corresponding to the virtual viewpoint using the multi-viewpoint image data by the rendering method based on shape. In Embodiment 1, the image generation unit 304 will be described as generating a virtual viewpoint image corresponding to the virtual viewpoint based on the object approximate shape by the view volume intersection method acquired in S601.
[0038] Specifically, in S902, first, the image generation unit 304 calculates the position of the collision point between the ray corresponding to each pixel in the virtual viewpoint image to be generated and the object approximate shape based on the virtual camera parameters. Subsequently, in S902, the image generation unit 304 obtains the pixel value corresponding to the point on the captured image obtained by projecting the collision point using one or more captured images including the point in the imaging region 106 corresponding to the collision point as an image, and sets it as the pixel value of the pixel in the virtual viewpoint image. When there are a plurality of captured images including the point in the imaging region 106 corresponding to the collision point as an image, the image generation unit 304 determines the pixel value of the pixel in the virtual viewpoint image as follows. In this case, the image generation unit 304 obtains the pixel value corresponding to the point on the captured image obtained by projecting the collision point for each captured image including the point in the imaging region 106 corresponding to the collision point as an image, and sets the weighted sum of the obtained multiple pixel values as the pixel value of the pixel in the virtual viewpoint image. When calculating the weighted sum, the image generation unit 304, for example, sets a larger weight for the captured image with higher resolution near the collision point.
[0039] By setting such weights, the image generation unit 304 can generate a virtual viewpoint image having a resolution almost equivalent to that of the captured image with the highest resolution of the object 107 by a rendering method based on the shape. However, when the object approximate shape cannot represent the actual object shape, the quality of the virtual viewpoint image deteriorates. In particular, when rendering the object 107 in a digitally zoomed-up state, the deterioration of the quality of the virtual viewpoint image becomes more prominent as the object size per pixel becomes smaller.
[0040] If it is determined in S901 that the rendering method set is not a rendering method based on shape, that is, if the rendering method based on the radiance field is set in S403, the image generation unit 304 executes the process of S903. Specifically, in this case, in S903, the image generation unit 304 generates a virtual viewpoint image corresponding to the virtual viewpoint using the multi-viewpoint image data by the rendering method based on the radiance field. More specifically, in S903, the image generation unit 304 performs volume rendering based on the virtual camera parameters and the radiance field corresponding to the imaging region 106 estimated based on the multi-viewpoint image data, thereby generating a virtual viewpoint image corresponding to the virtual viewpoint.
[0041] The radiance field is a function that takes as input information indicating the position and direction within the encoded imaging region 106 and outputs information indicating color and density, and is represented using a multi-layer perceptron. By using a multi-layer perceptron, it is possible to represent color and density according to the position and direction regardless of the complexity of the object shape. In volume rendering, the pixel value corresponding to each pixel is calculated based on the color and density corresponding to the sampling points on the ray corresponding to the pixel in the generated virtual viewpoint image. The color and density corresponding to the sampling points are obtained by inputting information indicating the position of the sampling points and the direction of the ray into the radiance field. Also, the estimation of the radiance field is performed by optimizing the radiance field so that the difference between the pixel value obtained by performing volume rendering based on the imaging camera parameters and the radiance field and the pixel value of the captured image becomes small. The optimization of the radiance field is performed by a repetitive process in which a predetermined number of pixels randomly extracted from all the captured images are used as one unit without considering the resolution.
[0042] According to the rendering method using the radiance field estimated in this way, it is possible to generate a high-quality virtual viewpoint image regardless of the complexity of the object shape. However, the resolution of the virtual viewpoint image generated by the rendering method based on the radiance field is about the same as the average resolution in a plurality of captured images. Also, the amount of computation required to optimize the radiance field is large.
[0043] <Effects achieved by the image processing apparatus 102 according to Embodiment 1> As described above, the image processing apparatus 102 is configured to set either the shape-based rendering method or the radiance field-based rendering method as the rendering method for the virtual viewpoint image based on the characteristics of the object 107 when viewed from the virtual viewpoint. Specifically, the image processing apparatus 102 is configured to preferentially set the radiance field-based rendering method as the object shape when viewed from the virtual viewpoint is more complex, or the object size per pixel when viewed from the virtual viewpoint is smaller. On the other hand, the image processing apparatus 102 is configured to preferentially set the shape-based rendering method as the object shape when viewed from the virtual viewpoint is simpler, or the object size per pixel when viewed from the virtual viewpoint is larger. According to the image processing apparatus 102 configured as described above, the image quality of the virtual viewpoint image obtained by rendering can be improved, and the sense of discomfort of the viewer with respect to the image of the object 107 when viewing the virtual viewpoint image can be suppressed.
[0044] [Modification Example 1 of Embodiment 1] Although the method setting unit 303 according to Embodiment 1 has been described as obtaining the approximate object shape by the visual volume intersection method in S601, the method for obtaining the approximate object shape is not limited to the visual volume intersection method. For example, the approximate object shape may be obtained based on distance information obtained by stereo matching or the like from two captured images captured by two adjacent imaging devices 101, or distance image data obtained by measurement using a depth camera. Alternatively, the method setting unit 303 may obtain the approximate object shape by reading out an approximate object shape generated in advance and stored in a storage device 204 or the like.
[0045] Also, although the method setting unit 303 according to Embodiment 1 has been described as using the center-of-gravity coordinates of the approximate object shape as the representative point of the approximate object shape in S602, the representative point of the approximate object shape is not limited to the center-of-gravity coordinates of the approximate object shape. For example, the method setting unit 303 may select an arbitrary voxel from the set of voxels constituting the approximate object shape and use the selected voxel as the representative point. Specifically, for example, the method setting unit 303 can select any one of a plurality of voxels visible from the position of the virtual viewpoint as the representative point.
[0046] Also, the method setting unit 303 according to Embodiment 1 has been described as calculating the root mean square error of the luminance values of the pixels corresponding to the object 107 based on the reference image and the target image projected onto the reference image in S603. Also, the method setting unit 303 according to Embodiment 1 has been described as obtaining the root mean square error of the luminance values obtained by the calculation as information indicating the complexity of the object shape. However, the method for obtaining the information indicating the complexity is not limited to the above method.
[0047] For example, the method setting unit 303 can use information indicating the color of a pixel, such as an RGB value, instead of the luminance value of the pixel. Further, the method setting unit 303 may calculate, instead of the root mean square error, the mean square error or the mean absolute error, etc., and regard the value obtained by the calculation as information indicating the complexity of the object shape and acquire it. Further, the method setting unit 303 may acquire information indicating the complexity of the object shape based on some of the pixels included in the region corresponding to the object 107, instead of all the pixels included in the region corresponding to the object 107. Further, the method setting unit 303 may acquire information indicating the complexity of the object shape based on a plurality of target images.
[0048] Also, although the method setting unit 303 according to Embodiment 1 has been described as acquiring information indicating the complexity of the object shape based on the object schematic shape in S603, the method for acquiring the information indicating the complexity is not limited to the method based on the object schematic shape. For example, the method setting unit 303 may project a region with a complex shape such as a face detected from a plurality of captured images into a three-dimensional space, and regard the number of regions with a complex shape visible from a virtual viewpoint as information indicating the complexity of the object shape and acquire it. Also, although the method setting unit 303 according to Embodiment 1 has been described as setting the first threshold value (t s ) to a specified value or a value based on an instruction from the user in S604, the method for setting the first threshold value (ts) is not limited to this. For example, the method setting unit 303 may set the first threshold value (t s ) using Equation (3).
[0049] t s =α·(d´ / F´) ··· Equation (3) Here, d′ is the distance from the imaging device 101 to the center of the imaging region 106 projected in the optical axis direction of the imaging device 101, F′ is the focal length of the imaging device 101, and α is an appropriate coefficient. At this time, the method setting unit 303 sets the first threshold value t sIt is preferable to calculate. First, the method setting unit 303 selects an imaging device 101 whose position and optical axis direction are similar to the position and line-of-sight direction of the virtual camera. Subsequently, the method setting unit 303 uses the distance (d') from the selected imaging device 101 to the center of the imaging region 106 and the focal length (F´) of the imaging device 101 to calculate the first threshold value t s for calculation.
[0050] Also, the method setting unit 303 according to Embodiment 1 has been described as setting the rendering method based on the complexity of the object shape and the object size per pixel in S604. However, the method for setting the rendering method is not limited to this. FIGS. 8(b) and (c) are LUTs used by the method setting unit 303 when setting the rendering method, and show an example of a LUT different from that in FIG. 8(a). For example, the method setting unit 303 may set the rendering method according only to the degree of the object size per pixel. Specifically, as shown in FIG. 8(b), when the degree of the object size per pixel is "small", the method setting unit 303 sets a rendering method based on the radiance field, and when the degree is "large", sets a rendering method based on the shape.
[0051] Also, for example, the method setting unit 303 may set the rendering method according only to the degree of the complexity of the object shape. Specifically, as shown in FIG. 8(c), when the degree of the complexity of the object shape is "complex", the method setting unit 303 sets a rendering method based on the radiance field, and when the degree is "simple", sets a rendering method based on the shape. Also, for example, the method setting unit 303 may select any one of a plurality of LUTs as shown in FIGS. 8(a) to (c) according to a predetermined condition and set the rendering method. Also, for example, the user may select a desired setting method from among a plurality of setting methods displayed on the display device 105 by a user operation, and the method setting unit 303 may set the rendering method based on the setting method selected by the user operation.
[0052] In Embodiment 1, as an example, the method setting unit 303 was described in a form where the rendering method of the virtual viewpoint image is set based on the features of one object 107 existing in the imaging region 106 in S403. However, the number of objects to be noted when setting the rendering method of the virtual viewpoint image is not limited to 1, and the method setting unit 303 may set the rendering method of the virtual viewpoint image based on the features of each of a plurality of objects. For example, the method setting unit 303 may select a main object from a plurality of objects included in the image of the virtual viewpoint image, and set the rendering method by performing the processes of S601 to S604 based on the features of the main object when viewed from the virtual viewpoint. Further, for example, the method setting unit 303 may execute the processes of S601 to S604 for each object and set the rendering method for each object.
[0053] Referring to FIG. 10, the processing flow of the image generation unit 304 in S404 when the method setting unit 303 executes the processes of S601 to S604 for each object will be described. FIG. 10 is a flowchart showing an example of the processing flow of generating a virtual viewpoint image in the image generation unit 304 according to Modification 1 of Embodiment 1. Specifically, FIG. 10 is a flowchart showing an example of the processing flow of generating a virtual viewpoint image in S404 shown in FIG. 4 when the method setting unit 303 executes the processes of S601 to S604 for each object. The image generation unit 304 generates a virtual viewpoint image by synthesizing a plurality of virtual viewpoint images generated by different rendering methods for each object by executing the processes shown in the flowchart of FIG. 10.
[0054] First, in S1001, similar to the process of S902, the image generation unit 304 generates a virtual viewpoint image corresponding to the object for which the rendering method based on the shape is set in S604 as a temporary virtual viewpoint image by the rendering method based on the shape. Next, in S1002, similar to the process of S903, the image generation unit 304 generates a virtual viewpoint image corresponding to the object for which the rendering method based on the radiance field is set in S604 as a temporary virtual viewpoint image by the rendering method based on the radiance field. Next, in S1003, the image generation unit 304 generates a virtual viewpoint image by synthesizing a plurality of temporary virtual viewpoint images generated for each object in consideration of the front-back relationship between each other in each object. Specifically, when a plurality of objects overlap as seen from the virtual viewpoint, the plurality of temporary virtual viewpoint images are synthesized so that the object in the back is shielded by the object in the front. After the process of S1003, the image generation unit 304 ends the process of the flowchart shown in FIG. 10, that is, the process of S404.
[0055] Also, the image generation unit 304 according to Embodiment 1 has been described as generating a virtual viewpoint image corresponding to the virtual viewpoint based on the object schematic shape acquired in S601 in S902. However, the method for generating the virtual viewpoint image is not limited to this. Specifically, for example, the image generation unit 304 may newly acquire the object schematic shape again by the volume intersection method without using the object schematic shape acquired in S601. Further, the method for acquiring the object schematic shape in the image generation unit 304 is not limited to the volume intersection method.
[0056] For example, the object approximate shape may be obtained based on distance information acquired by stereo matching or the like from two captured images obtained by capturing with two imaging devices 101 adjacent to each other, or distance image data obtained by measurement with a depth camera. Also, for example, the image generation unit 304 may obtain the object approximate shape by reading out the object approximate shape generated in advance and stored in the storage device 204 or the like. The image generation unit 304 generates a virtual viewpoint image corresponding to the virtual viewpoint based on the object approximate shape newly obtained as described above without using the object approximate shape obtained in S601.
[0057] Also, in Embodiment 1, the radiance field has been described as being expressed using one multi-layer perceptron, but the method of expressing the radiance field is not limited to this. For example, the radiance field may be expressed using a plurality of multi-layer perceptrons, may be expressed using a sparse three-dimensional grid including spherical harmonic functions, or may be expressed using a tensor.
[0058] In Embodiment 1, as an example, the form in which the image processing apparatus 102 generates a virtual viewpoint image for one frame has been described. However, the image processing apparatus 102 may generate virtual viewpoint images corresponding to a plurality of consecutive frames. In this case, for example, the image processing apparatus 102 repeatedly executes the processes from S401 to S405 for each frame. Further, for example, after setting the rendering method for virtual viewpoint images corresponding to a plurality of frames, the image processing apparatus 102 may correct the rendering method and then generate virtual viewpoint images corresponding to the plurality of frames. In this case, it is preferable to correct the rendering method so that the rendering method does not change in several consecutive frames. For example, when a rendering method based on shape is set for the frame of interest and a rendering method based on radiance field is set for most of the frames around the frame of interest, the image processing apparatus 102 corrects the rendering method as follows. Specifically, in this case, the image processing apparatus 102 corrects the rendering method in the frame of interest where the rendering method based on shape is set to the rendering method based on radiance field.
[0059] Also, although the image processing apparatus 102 according to Embodiment 1 has been described as generating a virtual viewpoint image by a rendering method based on shape or a rendering method based on radiance field, the rendering method in the image processing apparatus 102 is not limited to this. For example, the image processing apparatus 102 may generate a virtual viewpoint image by a rendering method based on image deformation. The rendering method based on image deformation is a method of generating a virtual viewpoint image by synthesizing a plurality of captured images deformed based on the imaging camera parameters and the virtual camera parameters. The rendering method based on image deformation does not depend on the object schematic shape, similar to the rendering method based on radiance field. Therefore, the image processing apparatus 102 may set a rendering method based on image deformation instead of the rendering method based on radiance field and generate a virtual viewpoint image by the rendering method based on image deformation.
[0060] [Embodiment 2] In Embodiment 1, as an example, the case where the focal lengths of all the imaging devices 101 are equal (here, "equal" includes "substantially equal") was described. In Embodiment 2, the case where the image processing system includes a plurality of imaging devices with different focal lengths will be described. Note that the hardware configuration, functional configuration, and processing flow of the image processing device 102 according to Embodiment 2 are the same as those of Embodiment 1, and thus the description thereof will be omitted. However, the image processing device 102 according to Embodiment 2 (hereinafter simply referred to as "image processing device 102") is different from the image processing device 102 according to Embodiment 1 in the setting process of the rendering method in S403 shown in FIG. 4. Therefore, hereinafter, mainly the captured image and the setting process of the rendering method will be described. For the same configurations as those in Embodiment 1, the same reference numerals will be used for description.
[0061] <Captured Image> FIG. 11 is a diagram showing an example of the arrangement of the imaging devices 101 and 1101 according to Embodiment 2 and the captured image obtained by imaging with the imaging devices 101 and 1101. FIG. 11(a) shows an example of the arrangement of each of the plurality of imaging devices 101 and the plurality of imaging devices 1101. As shown in FIG. 11(a), in Embodiment 2, in addition to the imaging device 101 shown in FIG. 5(a) described in Embodiment 1, a plurality of imaging devices 1101 having a longer focal length than the imaging device 101 are arranged so as to be able to image the object 107 from various directions. Note that the imaging devices 1101a and 1101b shown in FIG. 11(a) are the same as the other imaging devices 1101.
[0062] Figs. 11(b), (c), and (d) are diagrams each showing an example of captured images 501, 502, and 503 obtained by imaging with imaging devices 101a, 101b, and 101c, respectively. Since the captured images 501, 502, and 503 are the same as the captured images 501, 502, and 503 shown in Figs. 5(b), (c), and (d), the description thereof is omitted. Figs. 11(e) and (f) are diagrams each showing an example of captured images 1102 and 1103 obtained by imaging with imaging devices 1101a and 1101b, respectively. Since the focal length of the imaging device 1101 is longer than that of the imaging device 101, the object size per pixel in the captured images 1102 to 1103 is smaller than that in the captured images 501 to 503.
[0063] <Rendering method setting process> The rendering method setting process executed by the method setting unit 303 according to Embodiment 2 differs only in the process of S604 shown in Fig. 6 as compared with the rendering method setting process executed by the method setting unit 303 according to Embodiment 1. Hereinafter, the process of S604 in the method setting unit 303 according to Embodiment 2 (hereinafter simply referred to as "method setting unit 303") will be described. In S604, the method setting unit 303 sets the rendering method of the virtual viewpoint image based on the object size per pixel and the information indicating the complexity of the object shape when viewed from the virtual viewpoint. Specifically, the method setting unit 303 sets either a shape-based rendering method or a radiance field-based rendering method as the rendering method of the virtual viewpoint image based on the object size and the information indicating the complexity.
[0064] The quality of the virtual viewpoint image generated by the image processing apparatus 102 based on the captured image is as follows according to the characteristics of the object when viewed from the virtual viewpoint. When the object shape is complex, the smaller the object size per pixel, the lower the quality of the virtual viewpoint image by the rendering method based on the shape. On the other hand, in this case, the decrease in the quality of the virtual viewpoint image by the rendering method based on the radiance field is suppressed. Also, when the object shape is simple and the object size per pixel is smaller than the average value of the captured image, the resolution of the virtual viewpoint image by the rendering method based on the radiance field decreases. On the other hand, in this case, the decrease in the quality of the virtual viewpoint image by the rendering method based on the shape is suppressed. In cases other than the above conditions, the quality of the virtual viewpoint image by the rendering method based on the shape and the quality of the virtual viewpoint image by the rendering method based on the radiance field are substantially equal to each other regardless of the rendering method. Also, compared with the rendering method based on the radiance field, the rendering method based on the shape requires less calculation amount for generating the virtual viewpoint image.
[0065] In view of these, the method setting unit 303 sets the rendering method as follows. First, when the object size per pixel (s) when viewed from the virtual viewpoint is equal to or greater than the third threshold value (t s ′), the method setting unit 303 determines that the degree of the object size is "large". Also, when the object size per pixel (s) is equal to or greater than the first threshold value t s and less than the third threshold value (t s ′), the method setting unit 303 determines that the degree of the object size is "medium". Further, when the object size per pixel (s) is less than the first threshold value (t s ), the method setting unit 303 determines that the degree of the object size is "small". Also, when the value (c) indicating the complexity of the object shape when viewed from the virtual viewpoint is equal to or greater than the second threshold value (t c ), the method setting unit 303 determines that the degree of the complexity of the object shape is "complex". Also, when the value (c) indicating the complexity of the object shape is less than the second threshold value (t c)If it is less than, it is determined that the degree of complexity of the object shape is "simple". Further, the method setting unit 303 sets the rendering method of the virtual viewpoint image according to the degree of the object size and the degree of complexity of the object shape.
[0066] FIG. 12 is a diagram showing an example of a LUT used by the method setting unit 303 according to Embodiment 2 when setting the rendering method. In the LUT shown in FIG. 12, similar to FIG. 8, the rendering method, the degree of object size per pixel, and the degree of complexity of the object shape when viewed from the virtual viewpoint are managed in association with each other. The method setting unit 303 determines and sets the rendering method of the virtual viewpoint image according to the degree of the object size and the degree of complexity of the object shape using the LUT shown in FIG. 12(a).
[0067] Specifically, the method setting unit 303 sets the rendering method based on the radiance field as the rendering method of the virtual viewpoint image only when the degree of complexity of the object shape is "complex" and the degree of object size per pixel is "medium" or less. In other cases, the method setting unit 303 sets the rendering method based on the shape as the rendering method of the virtual viewpoint image. The first threshold value (t s ), the third threshold value (t s '), and the second threshold value (t c ) may be prescribed values, or may be values set based on an instruction from the user acquired via the UI panel 103 or the like. For example, for the first threshold value (t s ) and the second threshold value (t c ), the same values as those in Embodiment 1 are set, and for the third threshold value t s ', it is preferable to set a value based on the average object size per pixel of a plurality of captured images. The LUT shown in FIG. 12(b) will be described later.
[0068] <Effects exhibited by the image processing apparatus 102 according to Embodiment 2> As described above, the image processing apparatus 102 is configured such that either a rendering method based on shape or a rendering method based on radiance field is set as the rendering method for the virtual viewpoint image based on the characteristics of the object 107 when viewed from the virtual viewpoint. Specifically, the image processing apparatus 102 is configured to preferentially set the rendering method based on the radiance field as the more complex the object shape is when viewed from the virtual viewpoint, or the smaller the object size per pixel is when viewed from the virtual viewpoint. On the other hand, the image processing apparatus 102 is configured to preferentially set the rendering method based on shape as the simpler the object shape is when viewed from the virtual viewpoint, or the larger the object size per pixel is when viewed from the virtual viewpoint. According to the image processing apparatus 102 configured as described above, even when including the imaging devices 101 and 1101 with different focal lengths, the image quality of the virtual viewpoint image obtained by rendering can be improved. Further, according to the image processing apparatus 102, the sense of discomfort of the viewer with respect to the image of the object 107 when viewing the virtual viewpoint image can be suppressed.
[0069] [Modification Example 1 of Embodiment 2] The method setting unit 303 according to Embodiment 2 set the rendering method based on the degree of complexity of the object shape and the degree of object size per pixel in S604, but the method of setting the rendering method is not limited to this. FIG. 12(b) is a LUT used by the method setting unit 303 when setting the rendering method, and shows an example of a LUT different from that in FIG. 12(a). For example, the method setting unit 303 may set the rendering method only according to the degree of complexity of the object shape. Specifically, for example, as shown in FIG. 12(b), when the degree of complexity of the object shape is "complex", the method setting unit 303 sets the rendering method based on the radiance field, and when the degree is "simple", sets the rendering method based on shape.
[0070] Also, the image processing apparatus 102 according to Embodiment 2 has been described as setting the rendering method of the virtual viewpoint image using the captured images obtained by imaging the object 107 from various directions using the imaging devices 101 and 1101 with different focal lengths. However, the method for setting the rendering method in the image processing apparatus 102 is not limited to this.
[0071] FIG. 13 is a diagram showing an example of the arrangement of the imaging devices 101 and 1101 according to Modification 1 of Embodiment 2. Specifically, in FIG. 13, the imaging devices 101 and 1101 with different focal lengths are arranged only in a specific direction. The image processing apparatus 102 is applicable, for example, even when the imaging devices 101 and 1101 with different focal lengths are arranged unevenly as shown in FIG. 13. In this case, for example, the image processing apparatus 102 switches the method for setting the rendering method according to whether the direction of the virtual viewpoint is the direction in which the imaging devices 101 and 1101 with different focal lengths are arranged. Specifically, if the direction of the virtual viewpoint is the direction in which the imaging devices 101 and 1101 with different focal lengths are arranged, the image processing apparatus 102 sets the rendering method based on the LUT shown in FIG. 12(a) or (b). On the other hand, if the direction of the virtual viewpoint is the direction in which the imaging device 101 with the same focal length is arranged, the image processing apparatus 102 sets the rendering method based on the LUT shown in any of FIGS. 8(a) to (c).
[0072] [Other Embodiments] The present disclosure can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or apparatus via a network or a storage medium and causing one or more processors in the computer of the system or apparatus to read and execute the program. It can also be realized by a circuit (for example, ASIC) that realizes one or more functions.
[0073] Note that within the scope of the present disclosure, any combination of the embodiments, modification of any component of each embodiment, or omission of any component in each embodiment is possible.
[0074] [Configuration of the Present Disclosure] The present disclosure includes the following configuration, method, and program.
[0075] <Configuration 1> Image acquisition means for acquiring data of a plurality of captured images obtained by capturing an object from a plurality of directions, Viewpoint acquisition means for acquiring the position of a virtual viewpoint and virtual viewpoint information regarding the line-of-sight direction at the virtual viewpoint, Setting means for setting a generation method of a virtual viewpoint image based on the characteristics of the object acquired based on the virtual viewpoint information, Generation means for generating the virtual viewpoint image based on the virtual viewpoint information and the set generation method of the virtual viewpoint image, An image processing apparatus characterized by comprising the above.
[0076] <Configuration 2> The setting means sets either a rendering method based on shape or a rendering method based on radiance field as the generation method of the virtual viewpoint image, The image processing apparatus according to Configuration 1, characterized by the above.
[0077] <Configuration 3> The rendering method based on shape is a method of performing rendering based on information indicating the three-dimensional shape of the object generated using the plurality of captured images by a view volume intersection method and the data of the plurality of captured images, The image processing apparatus according to Configuration 2, characterized by the above.
[0078] <Configuration 4> The radiance field is a function estimated as a result of performing iterative processing so that the difference between the pixel values in the plurality of captured images and the pixel values obtained by volume rendering based on the radiance field becomes small, The rendering method based on the radiance field is a method of performing rendering by inputting the virtual viewpoint information into the function indicating the radiance field. The image processing apparatus according to Configuration 2 or 3, characterized in that.
[0079] <Configuration 5> The feature of the object is the complexity of the shape of the object when viewed from the virtual viewpoint. The setting means preferentially sets the rendering method based on the radiance field as the generation method of the virtual viewpoint image as the shape of the object becomes more complex. The image processing apparatus according to any one of Configurations 2 to 4, characterized in that.
[0080] <Configuration 6> The feature of the object is the size of the object per pixel in the virtual viewpoint image when the object is viewed from the virtual viewpoint. The setting means preferentially sets the rendering method based on the radiance field as the generation method of the virtual viewpoint image as the size of the object per pixel becomes smaller. The image processing apparatus according to any one of Configurations 2 to 4, characterized in that.
[0081] <Configuration 7> The feature of the object is the complexity of the shape of the object and the size of the object per pixel in the virtual viewpoint image when the object is viewed from the virtual viewpoint. The setting means preferentially sets the rendering method based on the radiance field as the generation method of the virtual viewpoint image as the shape of the object becomes more complex and the size of the object per pixel becomes smaller. The image processing apparatus according to any one of Configurations 2 to 4, characterized in that.
[0082] <Configuration 8> When there are a plurality of the objects included as images in the plurality of captured images, the setting means selects one of the plurality of existing objects as a representative object, and sets a generation method of the virtual viewpoint image based on a feature of the representative object when the selected representative object is viewed from the virtual viewpoint. The image processing apparatus according to any one of Configurations 1 to 7, characterized in that.
[0083] <Configuration 9> When there are a plurality of the objects included as images in the plurality of captured images, for each object, the setting means sets a generation method of the virtual viewpoint image based on a feature of the object when the object is viewed from the virtual viewpoint. The generation means generates a plurality of provisional virtual viewpoint images by a plurality of different generation methods based on the generation methods set for each object, and generates the virtual viewpoint image by synthesizing the plurality of generated provisional virtual viewpoint images. The image processing apparatus according to any one of Configurations 1 to 7, characterized in that.
[0084] <Configuration 10> When generating the virtual viewpoint image, the generation means synthesizes the plurality of generated provisional virtual viewpoint images based on an overlap of the plurality of existing objects when viewed from the virtual viewpoint. The image processing apparatus according to Configuration 9, characterized in that.
[0085] <Method> An image acquisition step of acquiring data of a plurality of captured images obtained by capturing an object from a plurality of directions, A viewpoint acquisition step of acquiring virtual viewpoint information regarding a position of a virtual viewpoint and a line-of-sight direction at the virtual viewpoint, A setting step of setting a generation method of a virtual viewpoint image based on a feature of the object acquired based on the virtual viewpoint information. A generation step of generating the virtual viewpoint image based on the virtual viewpoint information and the set generation method of the virtual viewpoint image; An image processing method characterized by including the above.
[0086] <Program> A program for causing a computer to function as the image processing apparatus according to any one of Configurations 1 to 10.
Explanation of Signs
[0087] 102 Image processing apparatus 301 Image acquisition unit 302 Viewpoint acquisition unit 303 Method setting unit 304 Image generation unit
Claims
1. Image acquisition means for acquiring data of a plurality of captured images obtained by capturing an object from a plurality of directions; Viewpoint acquisition means for acquiring virtual viewpoint information regarding the position of a virtual viewpoint and the viewing direction at the virtual viewpoint; Setting means for setting a generation method of a virtual viewpoint image based on the characteristics of the object acquired based on the virtual viewpoint information; Generation means for generating the virtual viewpoint image based on the virtual viewpoint information and the set generation method of the virtual viewpoint image; An image processing apparatus, characterized by comprising the above.
2. The setting means sets either a rendering method based on shape or a rendering method based on radiance field as the generation method of the virtual viewpoint image; The image processing apparatus according to claim 1, characterized by the above.
3. The rendering method based on shape is a method of performing rendering based on information indicating the three-dimensional shape of the object generated using the plurality of captured images by a view volume intersection method and the data of the plurality of captured images; The image processing apparatus according to claim 2, characterized by the above.
4. The radiance field is a function estimated as a result of performing iterative processing so that the difference between the pixel values in the plurality of captured images and the pixel values obtained by volume rendering based on the radiance field becomes small; The rendering method based on the radiance field is a method of performing rendering by inputting the virtual viewpoint information into the function indicating the radiance field; The image processing apparatus according to claim 2, characterized by the above.
5. The characteristic of the object is the complexity of the shape of the object when viewed from the virtual viewpoint; The setting means preferentially sets the rendering method based on the radiance field as the generation method of the virtual viewpoint image as the shape of the object is more complex; The image processing apparatus according to claim 2, characterized by the above.
6. The characteristic of the object is the size of the object per pixel in the virtual viewpoint image when the object is viewed from the virtual viewpoint; The setting means preferentially sets the rendering method based on the radiance field as the generation method of the virtual viewpoint image as the size of the object per pixel is smaller; The image processing apparatus according to claim 2, characterized by the above.
7. The features of the object are the complexity of the shape of the object and the size of the object per pixel in the virtual viewpoint image when the object is viewed from the virtual viewpoint, wherein the setting means preferentially sets the rendering method based on the radiance field as the generation method of the virtual viewpoint image as the shape of the object is more complex and the size of the object per pixel is smaller, The image processing apparatus according to claim 2, characterized in that.
8. When there are a plurality of objects included as images in the plurality of captured images, the setting means selects one of the plurality of existing objects as a representative object, and based on the features of the selected representative object when viewed from the virtual viewpoint, sets the generation method of the virtual viewpoint image, The image processing apparatus according to claim 1, characterized in that.
9. When there are a plurality of objects included as images in the plurality of captured images, for each object, based on the features of the object when viewed from the virtual viewpoint, the generation method of the virtual viewpoint image is set, The generation means generates a plurality of virtual virtual viewpoint images by a plurality of different generation methods based on the generation method set for each object, and generates the virtual viewpoint image by synthesizing the plurality of generated virtual virtual viewpoint images, The image processing apparatus according to claim 1, characterized in that.
10. When generating the virtual viewpoint image, the generation means synthesizes the plurality of generated virtual virtual viewpoint images based on the overlap of the plurality of existing objects when viewed from the virtual viewpoint, The image processing apparatus according to claim 9, characterized in that.
11. An image acquisition step of acquiring data of a plurality of captured images obtained by imaging an object from a plurality of directions; A viewpoint acquisition step of acquiring virtual viewpoint information regarding the position of the virtual viewpoint and the viewing direction at the virtual viewpoint; A setting step of setting a generation method of a virtual viewpoint image based on the features of the object acquired based on the virtual viewpoint information; A generation step of generating the virtual viewpoint image based on the virtual viewpoint information and the set generation method of the virtual viewpoint image; An image processing method characterized by including.
12. A program for causing a computer to function as the image processing apparatus according to any one of Claims 1 to 10.
Citation Information
Patent Citations
Information processing apparatus, information processing method, and program
JP2023075859A