Image processing apparatus, image processing method, and storage medium

By adjusting learning parameters based on pixel resolution, the image processing device enhances virtual viewpoint image quality and reduces computational demands in three-dimensional information estimation.

JP2026003219APending Publication Date: 2026-01-13CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024101067
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-24
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

Conventional methods for estimating three-dimensional information result in decreased image quality of virtual viewpoint images when the number of learning parameters is insufficient relative to the captured image resolution, leading to increased computational demands.

Method used

An image processing device that sets higher pixel resolution for partial areas within the imaging area, adjusting the number of learning parameters per volume based on pixel resolution to enhance the estimation of three-dimensional information, thereby improving virtual viewpoint image quality.

Benefits of technology

The method allows for the generation of high-quality virtual viewpoint images while minimizing the number of learning parameters required, reducing computational load and maintaining accurate three-dimensional information estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026003219000001_ABST
    Figure 2026003219000001_ABST
Patent Text Reader

Abstract

To estimate three dimensional information capable of generating a high-quality virtual viewpoint image while suppressing the number of learning parameters for estimating the three dimensional information.SOLUTION: The image processing device 102 acquires a plurality of captured images obtained by imaging an imaging region from a plurality of directions, sets one or more partial regions in the imaging region, sets a learning model corresponding to the partial regions such that the number of learning parameters per volume increases as the pixel resolution corresponding to the partial regions in the captured images increases, and performs learning of the learning model using the captured images.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to techniques for estimating three-dimensional information about a space containing an object. [Background technology]

[0002] There is a technology that estimates information (hereinafter referred to as "three-dimensional information") about a space including an object using images (hereinafter referred to as "captured images") obtained by capturing images of the object from various directions. There is also a technology that uses the three-dimensional information to generate an image (hereinafter referred to as "virtual viewpoint image") that corresponds to an image of the object when viewed from an arbitrary virtual viewpoint (hereinafter referred to as "virtual viewpoint"). Patent Document 1 discloses a technology that uses captured images as training images to learn radiance fields that represent color and density according to the position and direction in a space including the object as three-dimensional information. Patent Document 1 also discloses a technology that generates a virtual viewpoint image by volume rendering using the radiance fields estimated by the learning.

[0003] Specifically, in the technology disclosed in Patent Document 1 (hereinafter referred to as the "conventional technology"), learning parameters corresponding to a radiance field are calculated by sampling points on a ray corresponding to each pixel of a teacher image and performing machine learning. More specifically, during this calculation, for the ray corresponding to each pixel of the teacher image, the sampling density within the depth of field is set higher than the sampling density outside the depth of field. In the conventional technology, by controlling the sampling density based on the depth of field, the amount of calculation required for estimating the radiance field is reduced while improving the estimation accuracy of the radiance field in space corresponding to objects within the depth of field, resulting in improved image quality of the virtual viewpoint image. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Publication No. 2023-66705 Summary of the Invention [Problem to be solved by the invention]

[0005] The conventional technology has a problem that when the number of learning parameters (hereinafter referred to as "the number of learning parameters") is small relative to the resolution of the captured image, the estimation accuracy of the radiance field of the space corresponding to the object decreases, and as a result, the image quality of the virtual viewpoint image may decrease. However, the image quality of the virtual viewpoint image is limited by the image quality of the captured image. Therefore, simply increasing the number of learning parameters may not change the image quality of the virtual viewpoint image, but may increase the amount of calculation required to estimate three-dimensional information and the amount of three-dimensional information.

[0006] Therefore, an object of the present disclosure is to provide a technique for estimating three-dimensional information that enables generation of a high-quality virtual viewpoint image while suppressing the number of learning parameters for estimating three-dimensional information. [Means for solving the problem]

[0007] The image processing device according to the present disclosure includes an image acquisition means for acquiring a plurality of captured images obtained by capturing images of an imaging area from a plurality of directions, an area setting means for setting one or more partial areas in the imaging area, a model setting means for setting a learning model corresponding to the partial area so that the higher the pixel resolution corresponding to the partial area in the imaging image, the greater the number of learning parameters per volume, and a learning means for learning the learning model using the imaging images. [Effects of the Invention]

[0008] It is possible to estimate three-dimensional information that allows generation of a high-quality virtual viewpoint image while suppressing the number of learning parameters for estimating three-dimensional information. [Brief explanation of the drawings]

[0009] [Figure 1]FIG. 1 is a diagram illustrating an example of the configuration of an image processing system according to a first embodiment. [Figure 2] 1 is a block diagram showing an example of a hardware configuration of an image processing device according to a first embodiment. [Figure 3] 1 is a block diagram showing an example of a functional configuration of an image processing device according to a first embodiment. [Figure 4] 4 is a flowchart showing an example of a processing flow in the image processing device according to the first embodiment. [Figure 5] 2A and 2B are diagrams illustrating an example of the arrangement of an imaging device and a captured image according to the first embodiment. [Figure 6] 10 is a flowchart showing an example of the flow of partial region setting processing in the region setting unit according to the first embodiment. [Figure 7] FIG. 3 is a diagram showing an example of a partial region set by a region setting unit according to the first embodiment. [Figure 8] 6 is a flowchart showing an example of the flow of a pixel resolution setting process in a resolution setting unit according to the first embodiment. [Figure 9] FIG. 4 is a diagram for explaining an example of a reference point distance according to the first embodiment. [Figure 10] 10 is a flowchart showing an example of the flow of a learning model setting process in the model setting unit according to the first embodiment. [Figure 11] FIG. 2 is a diagram for explaining an example of a learning model setting process in the model setting unit according to the first embodiment. [Figure 12] FIG. 3 is a diagram illustrating an example of a partial region according to the first embodiment. [Figure 13] FIG. 10 is a diagram showing an example of a partial region according to a modified example of the first embodiment. [Figure 14] FIG. 10 is a diagram for explaining an example of a learning model setting process in a model setting unit 304 according to a modified example of the first embodiment. [Figure 15] FIG. 10 is a diagram illustrating an example of division and integration of partial regions according to a modified example of the first embodiment. [Figure 16]10 is a flowchart showing an example of the flow of a pixel resolution setting process in a resolution setting unit according to the second embodiment. [Figure 17] 10A and 10B are diagrams for explaining an example of a pixel resolution setting process in a resolution setting unit according to the second embodiment. [Figure 18] 10 is a flowchart showing an example of the flow of a learning model setting process in a model setting unit according to the second embodiment. [Figure 19] FIG. 10 is a diagram for explaining an example of a learning model setting process in the model setting unit according to the second embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. Note that the following embodiments do not necessarily limit the means for solving the problems according to the present disclosure. Furthermore, not all of the combinations of features described in the following embodiments are necessarily essential to the means for solving the problems according to the present disclosure.

[0011] [First embodiment] In this embodiment, a mode of learning a radiance field corresponding to a space including an object is described based on data of captured images (hereinafter referred to as "captured image data") of an object captured from various directions using multiple imaging devices. Specifically, the above-described learning according to this embodiment is performed using a learning model set based on the pixel resolution of an image region in the captured images that corresponds to the space including the object.

[0012] <Image processing system configuration> 1 is a diagram showing an example of the configuration of an image processing system according to a first embodiment. The image processing system includes a plurality of image capturing devices 101, an image processing device 102, a user interface (hereinafter referred to as "UI") panel 103, a storage device 104, and a display device 105. The plurality of image capturing devices 101 are configured as digital still cameras, digital video cameras, or the like, and are arranged in different locations. Each image capturing device 101 captures an image of an object 107 present in an image capturing area 106 from various directions in accordance with predetermined image capturing conditions, and outputs captured image data obtained by the image capturing to the image processing device 102.

[0013] Note that "mutually synchronized imaging" refers to imaging performed after synchronization processing. In other words, "mutually synchronized imaging" includes not only imaging performed at exactly the same time, but also imaging performed at approximately the same time. Furthermore, the captured image data obtained by imaging using the imaging device 101 may be still image data, moving image data, or both still image and moving image data. Hereinafter, the term "image" will be explained as including the meaning of both "still image" and "moving image" unless otherwise specified.

[0014] The image processing device 102 acquires a plurality of captured image data output from a plurality of imaging devices 101, and uses the acquired plurality of captured image data to learn information (three-dimensional information) about the three-dimensional shape and color of a space including an object 107 present in an imaging area 106. The image processing device 102 also generates a virtual viewpoint image based on the three-dimensional information obtained as a result of the learning (hereinafter referred to as "learned three-dimensional information").

[0015] In this embodiment, the three-dimensional information to be learned is described as a function representing a radiance field constructed by a multilayer perceptron, as an example. However, the method of expressing the three-dimensional information to be learned varies depending on the learning content. Specifically, for example, the three-dimensional information may be expressed by InstantNGP. Furthermore, the three-dimensional information is not limited to that constructed by a multilayer perceptron, but may be expressed by Plenoxels or TensoRF (Tensorial Radiance Fields), which explicitly express three-dimensional information. Furthermore, the three-dimensional information may be expressed by NeuS, which improves the accuracy of shape estimation using SDF (Signed Distance Field). Furthermore, the three-dimensional information may be expressed by various methods, such as 3D Gaussian Splatting, which expresses three-dimensionality using a set of points with a spread.

[0016] 1, the description will be given assuming that each of the multiple imaging devices 101 and the image processing device 102 are connected to one another, but the method of connection between the imaging devices 101 and the image processing device 102 is not limited to this. Specifically, for example, the multiple imaging devices 101 may be cascade-connected by connecting adjacent imaging devices 101 to one another, and at least one of the multiple imaging devices 101 may be connected to the image processing device 102.

[0017] 1 as an example, the present embodiment will be described assuming that a plurality of image capturing devices 101 are arranged at different positions, but the number and arrangement of the image capturing devices 101 are not limited to this. For example, if the position, shape, and color of an object 107 present in the image capturing area 106, as well as the intensity or hue of ambient light, do not change over time, at least one image capturing device 101 whose position and orientation can be changed may be arranged. In this case, the image capturing device 101 may be caused to capture images at a plurality of different positions while changing the position and orientation of the image capturing device 101, and the image processing device 102 may acquire a plurality of captured image data obtained by the image capturing.

[0018] The UI panel 103 includes a display device such as a liquid crystal panel, and displays a GUI (Graphical User Interface) on the display device to present information such as the imaging conditions of the imaging device 101 and the processing settings of the image processing device 102 to the user. The UI panel 103 may also include an input device such as a touch panel or buttons, in which case the UI panel 103 accepts instructions from the user regarding changes to the imaging conditions, processing settings, etc. The input device may be provided separately from the UI panel 103, such as a mouse or a keyboard.

[0019] The storage device 104 is configured with a hard disk drive or the like, and stores data of the virtual viewpoint image output by the image processing device 102. When the image processing device 102 outputs three-dimensional information, the storage device 104 may store the three-dimensional information output from the image processing device 102. The display device 105 is configured with a liquid crystal display or the like, and acquires an image signal indicating the virtual viewpoint image output from the image processing device 102, and displays the virtual viewpoint image corresponding to the image signal. When the image processing device 102 outputs an image signal indicating three-dimensional information, the display device 105 may acquire an image signal indicating the three-dimensional information output from the image processing device 102, and display an image corresponding to the image signal. The imaging area 106 is a three-dimensional space surrounded by multiple imaging devices 101 installed in a studio or the like, and the frame indicated by a solid line in FIG. 1 indicates the outline of the imaging area 106 on the floor surface.

[0020] <Hardware configuration of image processing device> 2 is a block diagram showing an example of the hardware configuration of the image processing device 102 according to the first embodiment. The image processing device 102 has, as its hardware configuration, a CPU 201, a RAM 202, a ROM 203, a storage device 204, a control interface (hereinafter referred to as "I / F") 205, an input I / F 206, an output I / F 207, and a main bus 208. The CPU 201 is a processor that performs overall control of each unit of the image processing device 102. The RAM 202 functions as the main memory and work area of ​​the CPU 201. The ROM 203 stores a group of programs executed by the CPU 201. The storage device 204 is configured by a hard disk drive or the like, and stores application programs executed by the CPU 201, data used in processing by the CPU 201, and the like.

[0021] The control I / F 205 is connected to each image capture device 101 and is a communication interface for setting image capture conditions for each image capture device 101, starting and stopping image capture, and other controls. The input I / F 206 is a communication interface using a serial bus such as SDI (Serial Digital Interface) or HDMI (High-Definition Multimedia Interface (registered trademark)). Captured image data output from each image capture device 101 is acquired via the input I / F 206. The output I / F 207 is a communication interface using a serial bus such as USB (Universal Serial Bus) or DisplayPort (registered trademark). Data or image signals such as virtual viewpoint images are output to the storage device 104 or the display device 105 via the output I / F 207. The main bus 208 is a transmission path that connects the above-mentioned hardware components of the image processing device 102 so that they can communicate with each other.

[0022] <Functional configuration of image processing device> 3 is a block diagram showing an example of the functional configuration of the image processing device 102 according to the first embodiment. The image processing device 102 includes an image acquisition unit 301, an area setting unit 302, a resolution setting unit 303, a model setting unit 304, a learning unit 305, a viewpoint acquisition unit 306, an image generation unit 307, and an output unit 308. Each unit included in the functional configuration of the image processing device 102 is realized by the CPU 201 executing a program stored in the ROM 203 or the like, using the RAM 202 as a work memory. Note that not all of the processing of each unit included in the functional configuration of the image processing device 102, as described below, necessarily needs to be executed by the CPU 201; some or all of the processing may be executed by one or more processing circuits other than the CPU 201.

[0023] The image acquisition unit 301 acquires captured image data obtained by each imaging device 101 capturing an image of the imaging region 106, and imaging parameters (hereinafter referred to as "imaging parameters") of the imaging device 101 corresponding to the captured image. The region setting unit 302 sets a space including the object 107 in the imaging region 106 as a partial region based on the captured image data and imaging parameters acquired by the image acquisition unit 301. The resolution setting unit 303 sets a pixel resolution corresponding to the partial region based on the imaging parameters acquired by the image acquisition unit 301 and the partial region set by the region setting unit 302. The model setting unit 304 sets a learning model corresponding to the partial region based on the partial region set by the region setting unit 302 and the pixel resolution corresponding to the partial region set by the resolution setting unit 303.

[0024] The learning unit 305 learns information (three-dimensional information) about the radiance field of the space including the object 107, based on the captured image data and imaging parameters acquired by the image acquisition unit 301 and the learning model set by the model setting unit 304. Here, the three-dimensional information is, for example, network parameters in a learning model configured by a multi-layer perceptron (MLP) that expresses the radiance field of the space including the object 107.

[0025] The viewpoint acquisition unit 306 acquires information related to a virtual viewpoint (hereinafter referred to as "virtual viewpoint information"). Here, the virtual viewpoint information is information indicating the position of the virtual viewpoint and the line of sight direction at the virtual viewpoint, and is information equivalent to the imaging parameters (hereinafter referred to as "virtual camera parameters") of a virtual imaging device (hereinafter referred to as "virtual camera") arranged at the virtual viewpoint. The image generation unit 307 generates a virtual viewpoint image using trained three-dimensional information obtained as a result of training by the training unit 305, i.e., a trained model which is information related to the radiance field, and the virtual camera parameters acquired by the viewpoint acquisition unit 306. Specifically, the image generation unit 307 performs volume rendering using the trained three-dimensional information to generate a virtual viewpoint image corresponding to the view from the virtual viewpoint indicated by the virtual camera parameters.

[0026] The output unit 308 outputs data of the virtual viewpoint image generated by the image generation unit 307 to the storage device 104, and stores the data in the storage device 104. The output unit 308 may output the virtual viewpoint image as an image signal to the display device 105, and display the virtual viewpoint image on the display device 105. The output unit 308 may also output trained three-dimensional information obtained as a result of training by the training unit 305, i.e., a trained model which is information on the radiance field, to the storage device 104, etc.

[0027] <Operation of image processing device> FIG. 4 is a flowchart showing an example of a processing flow in the image processing device 102 according to the first embodiment. Hereinafter, each processing step (process) is represented by adding an "S" to the beginning of the reference numeral. When each imaging device 101 outputs moving image data as captured image data, the image processing device 102 repeatedly executes the processing of the flowchart each time data of frames obtained by synchronized imaging included in each moving image is output from the imaging device 101. First, in S401, the image acquisition unit 301 acquires multiple captured image data obtained by the imaging and imaging parameters corresponding to each captured image data. Specifically, for example, the captured image data is acquired from the imaging device 101 via the input I / F 206, and the imaging parameters are acquired by reading out data calculated in advance by performing calibration or the like and stored in the storage device 204. The captured image data and imaging parameters acquired in S401 are held in the RAM 202.

[0028] 5A and 5B are diagrams illustrating an example of the arrangement of the imaging devices 101 according to the first embodiment and captured images 501 to 503 obtained by imaging using the imaging devices 101a to 101c. FIG. 5A illustrates an example of the arrangement of the imaging devices 101, where the imaging devices 101 are arranged so as to be able to capture images of an object 107 present in an imaging area 106 from various directions. Note that in this embodiment, the imaging devices 101a and 101c shown in FIG. 5A have optical systems with short focal lengths similar to the other imaging devices 101, and the imaging device 101b has an optical system with a longer focal length than the other imaging devices 101. That is, the imaging device 101b can capture the object 107 with higher resolution than the imaging devices 101a and 101c. Figure 5(b) shows an example of an image 501 obtained by imaging device 101a, Figure 5(c) shows an example of an image 502 obtained by imaging device 101b, and Figure 5(d) shows an example of an image 503 obtained by imaging device 101c.

[0029] After S401, in S402, the region setting unit 302 sets a space including the object 107 as a partial region based on the captured image data and imaging parameters acquired in S401. Details of the partial region setting process in the region setting unit 302 will be described later. Next, in S403, the resolution setting unit 303 executes pixel resolution setting process to set a pixel resolution corresponding to the partial region based on the imaging parameters acquired in S401 and the partial region set in S402. Specifically, for example, the resolution setting unit 303 sets the highest pixel resolution among the pixel resolutions of the imaging devices 101 corresponding to the reference points in the partial region as the pixel resolution corresponding to the partial region. Details of the pixel resolution setting process in the resolution setting unit 303 will be described later. Next, in S404, the model setting unit 304 executes learning model setting process to set a learning model corresponding to the partial region based on the partial region set in S402 and the pixel resolution corresponding to the partial region set in S403. Details of the learning model setting process in the model setting unit 304 will be described later.

[0030] Next, in S405, the learning unit 305 executes a learning process for three-dimensional information. Specifically, the learning unit 305 learns a learning model representing a radiance field of a space including the object 107, based on the captured image data and imaging parameters acquired in S401 and the learning model set in S404. In this embodiment, as an example, the radiance field is described as a function that receives information indicating a position and direction within the encoded captured region 106 as input and outputs information indicating density and color, and the function is represented by a learning model configured using an MLP. In other words, the MLP according to this embodiment is configured to calculate information indicating color and density according to connection weights between nodes, based on information indicating a position and direction within the encoded captured region 106 input to the input layer, and output the calculation result from the output layer.

[0031] The learning model representing the radiance field is trained based on the difference between the pixel value (hereinafter referred to as "rendering value") acquired by volume rendering using the imaging parameters and the radiance field and the pixel value (pixel value) corresponding to the pixel in the captured image. Specifically, the learning unit 305 trains the learning model representing the radiance field by updating the connection weight values ​​of the MLP representing the radiance field so as to reduce the difference between the pixel values.

[0032] For example, the learning unit 305 first acquires ray information corresponding to each pixel of the captured image based on the imaging parameters. Each piece of ray information includes information indicating the starting point, direction, and color of the ray. Here, the color of the ray is the value (pixel value) of the pixel in the captured image corresponding to the ray. Next, for each piece of ray information, the learning unit 305 sets multiple sampling points on the ray, acquires information indicating the density and color corresponding to the position of the sampling points and the direction of the ray based on the MLP representing the radiance field, and calculates a rendering value corresponding to each ray. Specifically, for example, the learning unit 305 calculates a rendering value corresponding to each ray using Equation (1) and Equation (2).

[0033]

number

[0034]

number

[0035] where C(r) is the drawing value corresponding to the ray r, i is the index of the sampling point, σ i is the density of sampling points, c i is the color of the sampling point, and δ i is the distance to the next sampling point. iis the cumulative transmittance at each sampling point. The learning unit 305 updates the connection weights of the MLP so that the squared Euclidean distance between the rendering value C(r) corresponding to the ray and the color of the ray, i.e., the value of the pixel (pixel value) of the captured image corresponding to the ray, becomes small. The processing of S405 corresponds to the error calculation processing and error propagation processing in deep learning.

[0036] After S405, in S406, the viewpoint acquisition unit 306 acquires, as virtual viewpoint information, virtual camera parameters set based on instructions from the user using the UI panel 103. The method of acquiring the virtual camera parameters by the viewpoint acquisition unit 306 is not limited to the above-described method. For example, the viewpoint acquisition unit 306 may acquire the virtual camera parameters by reading out virtual camera parameters that have been set in advance and stored in the storage device 204 or the like.

[0037] Next, in S407, the image generation unit 307 generates a virtual viewpoint image using the virtual camera parameters acquired in S406 and the trained three-dimensional information obtained as a result of the training process in S405, i.e., the trained model representing the radiance field. Specifically, the image generation unit 307 performs volume rendering based on the virtual camera parameters on the trained three-dimensional information, which is the trained model representing the radiance field, to generate a virtual viewpoint image corresponding to the appearance from the virtual viewpoint indicated by the virtual camera parameters.

[0038] Next, in S408, the output unit 308 outputs the virtual viewpoint image generated in S407. Specifically, for example, the output unit 308 outputs data of the virtual viewpoint image or an image signal indicating the virtual viewpoint image to the storage device 104, the display device 105, or the like via the output I / F 207. After S408, the image processing device 102 ends the processing of the flowchart shown in Fig. 4. Note that, as described above, when each imaging device 101 outputs moving image data as captured image data, the image processing device 102 returns to S401 after S408 and repeatedly executes the processing of the flowchart.

[0039] <Partial area setting process> 6 is a flowchart showing an example of the flow of partial region setting processing in the region setting unit 302 according to the first embodiment, and is a flowchart showing an example of a detailed processing flow in S402 shown in Fig. 4. In S402, a space including the object 107 is set as a partial region based on the captured image data and imaging parameters acquired in S401. In this embodiment, as an example, a mode will be described in which the region setting unit 302 acquires a rough shape of the object expressed as a set of voxels by a volume intersection method, and sets a rectangular parallelepiped region that includes the rough shape of the object as a partial region.

[0040] After S401, first, in S601, the region setting unit 302 acquires a silhouette image corresponding to each captured image data acquired in S401. Here, a silhouette image is an image showing a region including an image of the object 107 in the captured image. Specifically, first, the region setting unit 302 acquires background image data (hereinafter referred to as "background image data") obtained by capturing only the background using each imaging device 101 in a state in which the object 107 is not present. The background image data can be acquired, for example, by the region setting unit 302 reading background image data that has been captured in advance by each imaging device 101 and stored in advance in the storage device 204 or the like. Next, the region setting unit 302 acquires a silhouette image of the object 107 based on the difference between the captured image data corresponding to each imaging device 101 and the background image data corresponding to the captured image data. Since the method of acquiring the silhouette image of the object 107 is well known, a detailed description thereof will be omitted.

[0041] Next, in S602, the region setting unit 302 acquires a rough shape of the object 107 based on the imaging parameters acquired in S401 and the silhouette image acquired in S601. Specifically, for example, the region setting unit 302 first projects each voxel included in the set of voxels corresponding to the imaging region 106 onto the silhouette image based on the imaging parameters acquired in S401. Next, the region setting unit 302 acquires, as a rough shape of the object, a set of voxels projected onto a silhouette region corresponding to the region of the image of the object 107 for all silhouette images. Methods for acquiring a rough shape of the object 107 using a silhouette image, such as a volume intersection method, are well known, and therefore detailed description thereof will be omitted. Furthermore, the method for acquiring a rough shape of the object 107 is not limited to the volume intersection method, and any method may be used.

[0042] Next, in S603, the region setting unit 302 sets a rectangular parallelepiped region that includes the schematic shape of the object 107 acquired in S602 as a partial region. The region setting unit 302 may set a rectangular parallelepiped region with a predetermined margin around the schematic shape of the object as a partial region, or may set a rectangular parallelepiped region that circumscribes the schematic shape of the object without a margin as a partial region. After S603, the region setting unit 302 ends the processing of the flowchart shown in FIG. 6, i.e., the processing of S402 shown in FIG. 4.

[0043] 7 is a diagram showing an example of a partial region set by the region setting unit 302 according to the first embodiment. In FIG. 7, a rectangle surrounded by a thin solid line indicates an actual contour 702 of an object 701, and a polygon surrounded by a thick solid line indicates the contour of a schematic shape 703 of the object 701. Furthermore, a polygon surrounded by a thick dashed line indicates the contour of a partial region 704 set by the region setting unit 302.

[0044] <Pixel resolution setting process> Fig. 8 is a flowchart showing an example of the flow of pixel resolution setting processing in the resolution setting unit 303 according to the first embodiment, and is a flowchart showing an example of the detailed processing flow of S403 shown in Fig. 4. In S403, the pixel resolution corresponding to the partial region is set based on the imaging parameters acquired in S401 and the partial region set in S402.

[0045] After S402, first, in S801, the resolution setting unit 303 sets a reference point for the partial region set in S402. In this embodiment, as an example, the resolution setting unit 303 sets the center position of the partial region as the reference point. However, the reference point may be any position within the partial region, such as the center of gravity of the partial region. Next, in S802, the resolution setting unit 303 calculates the pixel resolution of each image capture device 101 corresponding to the reference point set in S601 using the imaging parameters acquired in S401. First, the image capture device 101 that includes the reference point within its angle of view is identified. For example, if the reference point projected using the imaging parameters is included in the captured image, the resolution setting unit 303 determines that the reference point is included within the angle of view. Next, for each image capture device 101 that includes the reference point within its angle of view, the resolution setting unit 303 calculates the pixel resolution of the image capture device 101 that corresponds to the reference point using, for example, Equation (3).

[0046]

number

[0047] where r ij is a value indicating the pixel resolution of the image capture device 101 at the position i of the reference point (hereinafter referred to as the "resolution value"). ij is the distance in the depth direction along the optical axis of the image capturing device 101 from the position j of the image capturing device 101 to the position i of the reference point (hereinafter referred to as the "reference point distance"). FIG. 9 shows the reference point distance d ij FIG. 1 is a diagram illustrating an example of jis a value obtained by multiplying the number of light receiving elements per unit length of the image sensor of the imaging device 101 placed at position j by the focal length of the optical system of the imaging device 101, and is also called a value indicating the focal length as an internal parameter of the imaging device 101.

[0048] Resolution value r ij is a value indicating the length per light-receiving pixel of the image capturing device 101 at the reference point. ij The smaller the reference point distance d, the higher the pixel resolution at the reference point in the image capturing device 101. ij The smaller the value of , that is, the smaller the value of the distance from the position j of the image capturing device 101 to the position i of the reference point, the smaller the resolution value r ij becomes smaller and the pixel resolution becomes higher. j The larger the value of r, that is, the longer the focal length of the image capture device 101, the higher the resolution value r ij becomes smaller, and the pixel resolution becomes higher. For example, if the depth direction distances from the position of the reference point to the positions of the imaging devices 101a, 101b, and 101c are equal, the pixel resolution at the reference point of the imaging device 101b, which has a longer focal length, will be higher than the pixel resolution at the reference point of the imaging devices 101a and 101c, which has a shorter focal length. Also, if the focal lengths of the imaging devices 101a, 101b, and 101c are equal, the imaging device 101 with a smaller value of the depth direction distance from the positions of the imaging devices 101a, 101b, and 101c to the position of the reference point will have a higher pixel resolution at the reference point.

[0049] After S802, in S803, the resolution setting unit 303 sets the pixel resolution corresponding to the partial region based on the pixel resolution at the reference point in each image capture device 101. For example, the resolution setting unit 303 selects the highest pixel resolution among the pixel resolutions at the reference point in each image capture device, and sets the selected pixel resolution as the pixel resolution corresponding to the partial region. Specifically, for example, the resolution setting unit 303 uses Equation (4) to determine a value indicating the pixel resolution at the reference point in each image capture device 101 (resolution value r ij ) is selected as the resolution value corresponding to the subregion.

[0050]

number

[0051] where r i is the resolution value corresponding to the partial region. After S803, the resolution setting unit 303 ends the process of the flowchart shown in Fig. 8, that is, the process of S403 shown in Fig. 4. By this process, the resolution setting unit 303 sets one resolution value r for the partial region. i Set.

[0052] <Learning model setting process> FIG. 10 is a flowchart showing an example of the flow of a learning model setting process in the model setting unit 304 according to the first embodiment, and is a flowchart showing an example of a detailed processing flow of S404 shown in FIG. 4. In S404, a learning model corresponding to a partial region is set based on the partial region set in S402 and the pixel resolution corresponding to the partial region set in S403. In this embodiment, as an example, the model setting unit 304 controls the total number of layers in the intermediate layer of the MLP constituting the learning model according to the partial region as the setting process of the learning model corresponding to the partial region. Specifically, in this embodiment, the learning model is configured by an MLP representing density (hereinafter referred to as "density MLP") and an MLP representing color, and the model setting unit 304 controls the number of layers in the intermediate layer of the density MLP.

[0053] 11 is a diagram illustrating an example of a learning model setting process in the model setting unit 304 according to the first embodiment. Fig. 11(a) shows an example of an MLP. The MLP has an input layer having one or more nodes 1101, a hidden layer having one or more layers 1102 each having one or more nodes 1101, and an output layer having one or more nodes 1101.

[0054] After S403, first, in S1001, the model setting unit 304 sets the number of learning parameters per volume in the learning model corresponding to the partial region based on the pixel resolution corresponding to the partial region set in S403. Specifically, for example, the model setting unit 304 sets the number of learning parameters per volume in the learning model corresponding to the partial region based on the pixel resolution corresponding to the partial region set in S403. i The number of learning parameters per volume is set so that the smaller the value of , the greater the number of layers in the hidden layer per volume of the density MLP corresponding to the partial region. For example, the model setting unit 304 references a lookup table in which pixel resolution and the number of layers in the hidden layer per volume of the density MLP are previously associated. The model setting unit 304 uses the lookup table to determine the number of layers in the hidden layer per volume of the density MLP corresponding to the partial region, depending on the pixel resolution corresponding to the partial region.

[0055] For example, the resolution value r i A partial region where the pixel resolution is equal to or less than a predetermined value is defined as a partial region with high pixel resolution, and the resolution value r i A partial region where the resolution value r is larger than a predetermined value is defined in advance as a partial region where the pixel resolution is low. Specifically, for example, as shown in the look-up table in FIG. i A partial area with a pixel resolution of 4 mm (millimeters) / pix (pixel) or less is defined as a partial area with a high pixel resolution. i A partial region where the pixel resolution is greater than 4 mm / pix is ​​defined as a partial region with low pixel resolution. i is 4 mm / pix or less, that is, in a partial region where the pixel resolution is high, a large value such as "16" is set as the number of layers in the intermediate layer per volume of the density MLP. iFor partial regions where the pixel resolution is greater than 4 mm / pix, i.e., where the pixel resolution is low, a small value such as "8" is set as the number of layers in the hidden layer per volume of the density MLP. As described above, for example, the model setting unit 304 determines one of two values ​​as the number of layers in the hidden layer per volume of the density MLP corresponding to the partial region, depending on the pixel resolution corresponding to the partial region.

[0056] After S1001, in S1002, the model setting unit 304 sets a learning model, for which the number of learning parameters per volume was set in S1001, to the partial region set in S402. In this embodiment, the rectangular parallelepiped region set as the partial region will be described as the region to which the learning model is assigned (hereinafter referred to as the "learning region"). Specifically, the model setting unit 304 sets the product of the number of layers in the hidden layer per volume of the density MLP set in S1001 and the volume of the learning region, converted into an integer, as the total number of layers in the hidden layer of the density MLP constituting the learning model. After S1002, the model setting unit 304 ends the process of the flowchart shown in FIG. 10, i.e., the process of S404 shown in FIG. 4. As a result of the process of S404, a learning model constituted by a density MLP with a large number of layers in the hidden layer per volume is set to a partial region with high pixel resolution. On the other hand, a learning model constituted by a density MLP with a small number of layers in the hidden layer per volume is set to a partial region with low pixel resolution.

[0057] FIG. 12 illustrates an example of partial regions 1201 and 1202 according to the first embodiment. For example, as shown in FIG. 12(a), a partial region 1201 within the angle of view of the image capture device 101b with a long focal length has a high pixel resolution, and therefore a learning model configured with a density MLP with a large number of layers in the hidden layer per volume is set. On the other hand, as shown in FIG. 12(b), a partial region 1202 within the angle of view of only the image capture device 101 with a short focal length has a low pixel resolution, and therefore a learning model configured with a density MLP with a small number of layers in the hidden layer per volume is set. In the partial region 1201 with a high pixel resolution, the number of learning parameters per volume is large, and therefore the radiance field, i.e., three-dimensional information, can be estimated with high accuracy. In the partial region 1202 with a low pixel resolution, the number of learning parameters per volume is small, and therefore the amount of calculation required to estimate three-dimensional information and the amount of three-dimensional information can be reduced.

[0058] <Effects of image processing devices> As described above, the image processing device 102 learns three-dimensional information about a space including the object 107 based on multiple captured image data sets obtained by capturing images of the object 107 from various directions using multiple imaging devices 101. In particular, in this embodiment, the image processing device 102 is configured to acquire pixel resolution corresponding to a subspace based on imaging parameters, and to set a learning model with a large number of learning parameters per volume for a subspace with high pixel resolution. On the other hand, the image processing device 102 is configured to set a learning model with a small number of learning parameters per volume for a subspace with low pixel resolution. The image processing device 102 configured in this manner can estimate three-dimensional information with high accuracy while reducing the amount of calculation required to estimate the three-dimensional information and the amount of three-dimensional information. In particular, the image processing device 102 can estimate three-dimensional information of a space with high pixel resolution captured by the imaging devices 101 with high accuracy. As a result, the image quality of a virtual viewpoint image generated based on the estimated three-dimensional information can be improved.

[0059] [Modification of the first embodiment] In the first embodiment, the imaging device 101b is described as the only imaging device 101 with a long focal length, but the configuration of the imaging devices 101 is not limited to this. For example, all of the imaging devices 101 may have the same focal length, or multiple imaging devices 101 with longer focal lengths than the other imaging devices 101 may be included, or all of the imaging devices 101 may have different focal lengths.

[0060] Furthermore, in the processing of S402, the region setting unit 302 according to the first embodiment has been described as acquiring the outline shape of the object 107 by the volume intersection method, but the method of acquiring the outline shape of the object 107 is not limited to this. For example, the region setting unit 302 may acquire the outline shape of the object 107 based on distance information acquired by stereo matching or the like using the captured image data acquired by the image acquisition unit 301, or a distance image acquired by a depth camera. Furthermore, information representing the outline shape of the object 107 that has been generated in advance may be stored in the storage device 204 or the like, and the region setting unit 302 may acquire the outline shape of the object 107 by reading out the information.

[0061] Furthermore, in the description of the first embodiment, the region setting unit 302 sets a partial region for one object 107 in the processing of S402, but a partial region may be set for each of multiple objects. For example, if multiple objects that are not in contact with each other exist in the captured image area, in S602 the region setting unit 302 acquires multiple voxel sets corresponding to each object as multiple outline shapes using a volume intersection method. Next, in S603, for each of the multiple outline shapes acquired in S602, the region setting unit 302 sets a rectangular parallelepiped region that contains the outline shape as a partial region.

[0062] Furthermore, in the description of the first embodiment, the region setting unit 302 sets a single rectangular parallelepiped shape that encompasses the general shape of the object 107 as the partial region in the processing of S402. However, the method of setting the partial region is not limited to this. For example, the region setting unit 302 may set each of a plurality of regions obtained by dividing the imaging region as a partial region. FIG. 13 is a diagram showing an example of a partial region 1301 according to a modification of the first embodiment. Note that FIG. 13 is a diagram showing the imaging region 106 as viewed in the vertical direction. For example, as shown in FIG. 13, the region setting unit 302 may set each of the regions obtained by dividing the imaging region 106 vertically as the partial region 704 in the processing of S402.

[0063] Furthermore, although the resolution setting unit 303 according to the first embodiment has been described as setting the center of the partial region as the reference point in the processing of S801, the reference point may instead be the center of gravity of the schematic shape of the object 107. For example, the resolution setting unit 303 calculates the center of gravity of the schematic shape of the object 107 based on a set of voxels that make up the schematic shape of the object 107, and sets the calculated center of gravity as the reference point of the partial region.

[0064] Furthermore, in the processing of S802, the resolution setting unit 303 according to the first embodiment has been described as identifying an image capture device 101 that includes the reference point within its angle of view, but it may also be configured to identify an image capture device 101 that includes the reference point within its angle of view and that is not occluded. Here, the resolution setting unit 303 determines that the reference point is not occluded, for example, when there is no other partial region or outline shape of an object that does not include the reference point between the reference point and the image capture device 101.

[0065] Although the resolution setting unit 303 according to the first embodiment has been described as setting the pixel resolution corresponding to the reference point based on the imaging parameters in the process of S403, the pixel resolution may be set only in accordance with the position of the reference point. For example, the resolution setting unit 303 sets a pixel resolution for each of a plurality of regions obtained by dividing the imaging region 106 in advance, and sets the pixel resolution corresponding to the region including the reference point as the pixel resolution corresponding to the reference point.

[0066] In addition, although the model setting unit 304 according to the first embodiment has been described as controlling the total number of layers in the hidden layers of the MLP in the process of S404, it may also control the number of nodes per layer in the hidden layers of the MLP. For example, the model setting unit 304 sets the learning model so that the number of nodes per layer in the hidden layers of the MLP increases as the pixel resolution increases.

[0067] In addition, in the processing of S404, the model setting unit 304 according to the first embodiment has been described as setting a learning model based on the number of learning parameters per volume, but the method for setting a learning model is not limited to this. For example, the model setting unit 304 may set a learning model for a partial region such that the number of learning parameters per MLP is a predetermined number and the higher the pixel resolution, the smaller the learning region. Furthermore, for example, the model setting unit 304 may set a learning model for a partial region such that the learning region has a predetermined size and the higher the pixel resolution, the greater the number of learning parameters per MLP. In this case, the model setting unit 304 sets one or more learning models so that the learning region includes the entire partial region.

[0068] Furthermore, although the model setting unit 304 according to the first embodiment sets a learning model configured using MLP in the processing of S404, the configuration of the learning model is not limited to this. For example, the model setting unit 304 may set a learning model configured using a plurality of learning parameters arranged in a grid pattern at predetermined intervals in a three-dimensional space. Specifically, the model setting unit 304 may use a learning model that expresses a spatial radiance field based on a plurality of spherical harmonic functions arranged in a grid pattern. The learning model used by the model setting unit 304 is not limited to this.

[0069] For example, the model setting unit 304 may use a learning model that represents a spatial radiance field based on a matrix or vector set with a predetermined number of components, each of which has an element corresponding to each grid. In this case, the model setting unit 304 sets the learning model so that the number of grids per volume increases as the pixel resolution increases, that is, the grid spacing decreases. The model setting unit 304 may set the learning model so that the number of learning parameters per grid increases as the pixel resolution increases. Here, the number of learning parameters per grid corresponds to the number of coefficients or components of a spherical harmonic function.

[0070] 14A and 14B are diagrams for explaining an example of a learning model setting process in the model setting unit 304 according to the modified example of the first embodiment. FIG. 14A is a diagram showing an example of a grid and components according to the modified example of the first embodiment. For example, the model setting unit 304 sets a resolution value r i The learning model is set so that the smaller the resolution value r, the smaller the grid interval, and the model is associated with the subregion. Specifically, for example, as shown in the lookup table in FIG. 14(b), i The model setting unit 304 uses the lookup table to determine the resolution value r iThe model setting unit 304 then determines the number of grids based on the value obtained by dividing the length of the learning region in each axial direction by the grid interval, and sets the learning model.

[0071] In the first embodiment, the model setting unit 304 is described as setting a rectangular parallelepiped region set as a partial region as a training region in the process of S1002. However, the model setting unit 304 may set the training region based on the number of layers in the hidden layer per volume of the MLP, as described below. Specifically, the model setting unit 304 calculates the total number of layers in the hidden layer of the MLP corresponding to the partial region by multiplying the number of layers in the hidden layer per volume of the MLP by the volume of the partial region. If the calculated total number of layers in the hidden layer of the MLP corresponding to the partial region is greater than a predetermined value, the model setting unit 304 sets each of multiple regions obtained by dividing the partial region as a training region. Specifically, in this case, for example, the model setting unit 304 divides the partial region into multiple regions so that the total number of layers in the hidden layer of the MLP corresponding to each divided region is equal to or less than a predetermined value. Furthermore, for example, if the total number of layers in the hidden layer of the MLP corresponding to the partial region is smaller than a predetermined value, the model setting unit 304 sets the region obtained by integrating multiple partial regions as the learning region. Specifically, for example, in this case, the model setting unit 304 integrates multiple nearby partial regions so that the total number of layers in the hidden layer of the MLP corresponding to the region obtained by integrating multiple partial regions does not exceed a predetermined value.

[0072] 15A and 15B are diagrams illustrating an example of division and merging of partial regions according to a modification of the first embodiment. For example, as shown in FIG. 15A, the model setting unit 304 divides a partial region 1501, in which the total number of layers in the hidden layers of the MLP is greater than a predetermined value, into four regions 1502 to 1505, and sets each of the divided regions 1502 to 1505 as a partial region. Furthermore, as shown in FIG. 15B, the model setting unit 304 merges partial regions 1506 and 1507, in which the total number of layers in the hidden layers of the MLP is less than a predetermined value, into a single partial region 1508.

[0073] In addition, the model setting unit 304 according to the first embodiment has been described as selecting and determining one of two values ​​as the number of layers in the hidden layer per volume of the density MLP corresponding to the partial region in the process of S1002. However, the model setting unit 304 may select and determine one of three or more values.

[0074] Although the image processing device 102 according to the first embodiment has been described as learning a learning model representing a radiance field, the learning model to be learned is not limited to one representing a radiance field. The learning model may represent three-dimensional information that can be learned based on captured image data. For example, a radiance field represents color and density according to position and direction, but the three-dimensional information is not limited to this. Specifically, for example, the three-dimensional information may represent a color corresponding to a position in space in the three-dimensional information as a color with isotropy that is independent of direction. Furthermore, for example, the three-dimensional information may represent a density corresponding to a position in space in the three-dimensional information as a signed distance field that represents the distance to the object surface according to the position. Furthermore, the three-dimensional information may represent, for example, a density field that represents density according to position, a field represented by a bidirectional reflectance distribution function that represents the distribution characteristics of reflected light relative to incident light, or a field that represents the amount of ambient light penetration (light visibility). Furthermore, the three-dimensional information may represent a field that represents color and density corresponding to position, direction, and time. In this case, the captured image data used for learning three-dimensional information is moving image data including time-series frames.

[0075] [Second embodiment] An image processing device 102 according to a second embodiment (hereinafter simply referred to as "image processing device 102") will be described with reference to FIGS. 2 to 4 and FIGS. 16 to 19. Like the image processing device 102 according to the first embodiment, the image processing device 102 has a hardware configuration and a functional configuration as shown, for example, in the block diagram of FIG. 2 or 3. Like the image processing device 102 according to the first embodiment, the image processing device 102 also executes the processing of the flowchart shown, for example, in FIG. 4. However, in this embodiment, the processing of the resolution setting unit 303 and the model setting unit 304 differs from the processing of the resolution setting unit 303 and the model setting unit 304 according to the first embodiment. That is, the pixel resolution setting processing of S403 and the learning model setting processing of S404 according to this embodiment differs from the processing of S403 and S404 according to the first embodiment.

[0076] Specifically, in the pixel resolution setting process of S403, the resolution setting unit 303 according to the first embodiment sets the highest pixel resolution among the pixel resolutions of the image capturing devices 101 corresponding to the reference points in the partial region as the pixel resolution corresponding to the partial region. In contrast, the resolution setting unit 303 according to the present embodiment sets multiple pixel resolutions according to directions as the pixel resolution corresponding to the partial region based on the pixel resolutions of the image capturing devices 101 corresponding to the reference points in the partial region. Furthermore, the model setting unit 304 according to the present embodiment sets a learning model for the partial region based on the multiple pixel resolutions according to the directions set by the resolution setting unit 303.

[0077] The following mainly describes the pixel resolution setting process in the resolution setting unit 303 and the learning model setting process in the model setting unit 304, which are different processes in this embodiment from those in the first embodiment. Note that the same reference numerals are used to designate configurations or processing steps (steps) that perform the same processes as in the first embodiment, and descriptions thereof will be omitted.

[0078] <Pixel resolution setting process> FIG. 16 is a flowchart showing an example of the flow of pixel resolution setting processing in the resolution setting unit 303 according to the second embodiment, and is a flowchart showing an example of a detailed processing flow of S403 shown in FIG. 4. In S403 according to the second embodiment, multiple pixel resolutions corresponding to directions are set as pixel resolutions corresponding to the partial region based on the pixel resolutions of the image capture devices 101 corresponding to the reference points within the partial region. The processing of this flowchart is executed after the processing of S402 shown in FIG. 4. After S402, the resolution setting unit 303 sequentially executes the processing of S801 and S802. After S802, in S1603, the resolution setting unit 303 selects multiple pixel resolutions corresponding to directions from the pixel resolutions of the image capture devices 101 corresponding to the reference points and sets the selected multiple pixel resolutions as pixel resolutions corresponding to the partial region. After S1603, the resolution setting unit 303 ends the processing of the flowchart shown in FIG. 16, i.e., the processing of S403 according to the second embodiment.

[0079] FIG. 17 is a diagram illustrating an example of pixel resolution setting processing in the resolution setting unit 303 according to the second embodiment. In the following, the present embodiment will be described assuming that four imaging device groups 1701 to 1704 are defined, corresponding to four predetermined directions, from direction 1 to direction 4, of the imaging device 101, as shown in FIG. 17. The description will also assume that the focal lengths of all imaging devices 101 according to the present embodiment are equal to each other. The resolution setting unit 303 sets one pixel resolution for one imaging device group. Specifically, first, for each imaging device group, the resolution setting unit 303 calculates the pixel resolution for the reference point of the partial region 1705 for each imaging device 101 in the imaging device group, for example, using Equation (2). Next, the resolution setting unit 303 determines, for each imaging device group, the highest pixel resolution (resolution value r i Next, the resolution setting unit 303 associates the selected pixel resolution with the direction corresponding to the imaging device group, and sets it as the pixel resolution corresponding to the partial region.

[0080] <Learning model setting process> FIG. 18 is a flowchart showing an example of the flow of the learning model setting process in the model setting unit 304 according to the second embodiment, and is a flowchart showing an example of the detailed processing flow of S404 shown in FIG. 4. In S404 according to the second embodiment, a learning model is set based on multiple pixel resolutions according to the direction corresponding to the reference point in the partial region. The processing of this flowchart is executed after the processing of S403 according to this embodiment. FIG. 19 is a diagram for explaining an example of the learning model setting process in the model setting unit 304 according to the second embodiment.

[0081] First, in S1801, the model setting unit 304 divides the partial region set in S402. Specifically, for example, as shown in FIG. 19(a) as an example, the model setting unit 304 divides the partial region 1705 into four partial regions (hereinafter referred to as "divided partial regions") 1901 to 1904. Next, in S1802, the model setting unit 304 sets the number of learning parameters per volume for each of the divided partial regions 1901 to 1904 divided in S1801 based on a plurality of pixel resolutions according to the directions set in S1603. Specifically, first, for each divided partial region, the model setting unit 304 sets the highest pixel resolution among the pixel resolutions in the directions corresponding to the divided partial region as the pixel resolution corresponding to the divided partial region. Next, the model setting unit 304 sets the highest pixel resolution among the pixel resolutions in the directions corresponding to the divided partial region as the pixel resolution corresponding to the divided partial region. Next, the model setting unit 304 sets the highest pixel resolution among the pixel resolutions in the directions corresponding to the divided partial region as the pixel resolution corresponding to the divided partial region. i The number of learning parameters per volume is set so that the smaller is, the more layers there are in the hidden layer per volume of the density MLP.

[0082] Hereinafter, it is assumed that divided partial region 1901 corresponds to direction 1 and direction 2, divided partial region 1902 corresponds to direction 2 and direction 3, divided partial region 1903 corresponds to direction 3 and direction 4, and divided partial region 1904 corresponds to direction 4 and direction 1. For example, if partial region 1705 is located close to imaging device group 1701 and far from the other imaging device groups, a high pixel resolution is set to divided partial regions 1901 and 1904 corresponding to direction 1, as shown in FIG. 19(b). In contrast, a low pixel resolution is set to the other divided partial regions 1902 and 1903. Furthermore, if partial region 1705 is located close to imaging device groups 1701 and 1702 and far from the other imaging device groups, the following pixel resolutions are set to each divided partial region: Specifically, in this case, as shown in FIG. 19(c), high pixel resolution is set to the divided partial areas 1901, 1902, and 1904 corresponding to directions 1 and 2, and low pixel resolution is set to the other divided partial area 1903.

[0083] After S1802, in S1803, the model setting unit 304 sets a learning model for which the number of learning parameters was set in S1802 to each of the divided partial regions 1901 to 1904. The process of setting a learning model to each of the divided partial regions 1901 to 1904 in S1803 is similar to the process of setting a learning model to a partial region in S1002 shown in Fig. 10, and therefore description thereof will be omitted.

[0084] Through the above process, a learning model consisting of MLPs with a large number of layers in the hidden layer per volume is set for the divided sub-regions corresponding to directions with high pixel resolution, while a learning model consisting of MLPs with a small number of layers in the hidden layer per volume is set for the divided sub-regions corresponding to directions with low pixel resolution.

[0085] <Effects of the image processing device according to the second embodiment> As described above, the image processing device 102 according to the second embodiment sets multiple pixel resolutions for the partial region 1705 according to the direction. Furthermore, the image processing device 102 according to the second embodiment sets multiple learning models for the partial region 1705, each having a different total number of layers in the intermediate layer, based on the multiple pixel resolutions according to the direction. The image processing device 102 configured in this manner can estimate highly accurate three-dimensional information while reducing the amount of calculation required for estimating the three-dimensional information and the amount of three-dimensional information based on the pixel resolution according to the direction. In particular, the image processing device 102 can estimate highly accurate three-dimensional information for a region corresponding to a direction in which the pixel resolution of the imaging device 101 is high. As a result, the image quality of a virtual viewpoint image generated based on the estimated three-dimensional information can be improved.

[0086] [Modification of the second embodiment] In the processing of S1603, the resolution setting unit 303 according to the second embodiment has been described as dividing the multiple image capturing devices 101 into four image capturing device groups corresponding to four predetermined directions, but the method of dividing the image capturing devices 101 is not limited to this. For example, the multiple image capturing devices 101 may be divided into image capturing device groups in three or fewer directions or five or more directions. Furthermore, the resolution setting unit 303 may divide the multiple image capturing devices 101 into multiple image capturing device groups such that image capturing devices 101 with similar positions and orientations based on imaging parameters are grouped into the same image capturing device group.

[0087] Furthermore, in the processing of S1801, the model setting unit 304 according to the second embodiment has been described as dividing the partial region 1705 into four divided partial regions 1901 to 1904, but the method of dividing the partial region 1705 is not limited to this. The model setting unit 304 may divide the partial region 1705 into three or less divided partial regions, or may divide it into five or more divided partial regions.

[0088] In addition, in the description of the second embodiment, the model setting unit 304 sets a learning model for each of the divided partial regions 1901 to 1904 in the processing of S1803. However, the method of setting the learning model is not limited to this. For example, the model setting unit 304 may integrate multiple divided partial regions having the same pixel resolution and set a learning model for the integrated divided partial region. Furthermore, for the partial region 1705 before division, the model setting unit 304 may set a learning model with a different number of learning parameters for each region corresponding to the divided partial regions 1901 to 1904. Specifically, for example, the model setting unit 304 locally changes the grid spacing of a learning model configured by learning parameters arranged in a grid pattern in a three-dimensional space as shown as an example in FIG. 14(a), and sets the changed learning model for the partial region 1705. For example, when the partial region 1705 is located close to the imaging device group 1701, the model setting unit 304 reduces the grid spacing of the regions corresponding to the divided partial regions 1901 and 1904 corresponding to direction 1, as shown in Fig. 19(d) as an example. Also, a learning model in which the grid spacing of the regions corresponding to the divided partial regions 1902 and 1903 is increased is set for the partial region 1705.

[0089] [Other embodiments] The present disclosure can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.

[0090] It should be noted that within the scope of the present disclosure, the embodiments may be freely combined, any component of each embodiment may be modified, or any component of each embodiment may be omitted.

[0091] [Configuration of the present disclosure] The present disclosure includes the following configurations, methods, and programs.

[0092] <Configuration 1> an image acquisition means for acquiring a plurality of captured images obtained by capturing images of an imaging region from a plurality of directions; an area setting means for setting one or more partial areas in the imaging area; a model setting means for setting a learning model corresponding to the partial region in the captured image such that the number of learning parameters per volume increases as the pixel resolution corresponding to the partial region in the captured image increases; a learning means for learning the learning model using the captured image; 1. An image processing device comprising:

[0093] <Configuration 2> the model setting means sets the learning model corresponding to the partial region such that the total number of layers in an intermediate layer of the learning model increases as the pixel resolution corresponding to the partial region increases; 2. The image processing device according to configuration 1,

[0094] <Configuration 3> the model setting means sets the learning model corresponding to the partial region such that the higher the pixel resolution corresponding to the partial region, the greater the number of nodes included in an intermediate layer of the learning model; 3. The image processing device according to configuration 1 or 2, characterized in that:

[0095] <Configuration 4> the model setting means sets the learning model corresponding to the partial region such that the higher the pixel resolution corresponding to the partial region, the smaller the learning region in the learning model; 4. The image processing device according to any one of configurations 1 to 3, characterized in that:

[0096] <Configuration 5> a resolution setting means for setting a pixel resolution corresponding to the partial region based on an imaging parameter corresponding to the captured image; further comprising: 5. The image processing device according to any one of configurations 1 to 4, characterized in that:

[0097] <Configuration 6> the resolution setting means sets a pixel resolution corresponding to the partial region based on a pixel resolution of the captured image in which a predetermined position in the partial region is not obstructed; 6. The image processing device according to configuration 5,

[0098] <Configuration 7> the resolution setting means sets a plurality of pixel resolutions according to directions as pixel resolutions corresponding to the partial regions; 7. The image processing device according to configuration 5 or 6,

[0099] <Configuration 8> the model setting means sets the learning model having different numbers of learning parameters in a plurality of regions within the partial region based on a plurality of pixel resolutions set according to the direction; 8. The image processing device according to configuration 7,

[0100] <Configuration 9> the model setting means sets, as the learning model corresponding to the partial region, the learning model based on a plurality of pixel resolutions set according to the direction for each of a plurality of regions within the partial region; 8. The image processing device according to configuration 7,

[0101] <Configuration 10> a shape acquisition means for acquiring a rough shape of an object present in the imaging area; and the region setting means sets the partial region based on the outline shape; 10. The image processing device according to any one of configurations 1 to 9, characterized in that:

[0102] <Configuration 11> The learning model is three-dimensional information about a learning region within the imaging region; 11. The image processing device according to any one of configurations 1 to 10, characterized in that:

[0103] <Configuration 12> viewpoint acquisition means for acquiring information on a virtual viewpoint; a generating means for generating a virtual viewpoint image corresponding to the virtual viewpoint using the trained learning model; further comprising: 12. The image processing device according to any one of configurations 1 to 11,

[0104] <Method> an image acquisition step of acquiring a plurality of captured images obtained by capturing images of an imaging region from a plurality of directions; a region setting step of setting one or more partial regions in the imaging region; a model setting step of setting a learning model corresponding to the partial region in the captured image such that the number of learning parameters per volume increases as the pixel resolution corresponding to the partial region in the captured image increases; a learning step of learning the learning model using the captured image; An image processing method comprising:

[0105] <Program> 13. A program for causing a computer to function as the image processing device according to any one of configurations 1 to 12. [Explanation of symbols]

[0106] 102 Image processing device 301 Image Acquisition Unit 302 Area setting section 304 Model Setting Section 305 Learning Department

Claims

1. an image acquisition means for acquiring a plurality of captured images obtained by capturing images of an imaging region from a plurality of directions; an area setting means for setting one or more partial areas in the imaging area; a model setting means for setting a learning model corresponding to the partial region in the captured image such that the number of learning parameters per volume increases as the pixel resolution corresponding to the partial region in the captured image increases; a learning means for learning the learning model using the captured image; 1. An image processing device comprising:

2. the model setting means sets the learning model corresponding to the partial region such that the total number of layers in an intermediate layer of the learning model increases as the pixel resolution corresponding to the partial region increases; 2. The image processing device according to claim 1, wherein:

3. the model setting means sets the learning model corresponding to the partial region such that the higher the pixel resolution corresponding to the partial region, the greater the number of nodes included in an intermediate layer of the learning model; 2. The image processing device according to claim 1, wherein:

4. the model setting means sets the learning model corresponding to the partial region such that the higher the pixel resolution corresponding to the partial region, the smaller the learning region in the learning model; 2. The image processing device according to claim 1, wherein:

5. a resolution setting means for setting a pixel resolution corresponding to the partial region based on an imaging parameter corresponding to the captured image; further comprising:

2. The image processing device according to claim 1, wherein:

6. the resolution setting means sets a pixel resolution corresponding to the partial region based on a pixel resolution of the captured image in which a predetermined position in the partial region is not obstructed; 6. The image processing device according to claim 5,

7. the resolution setting means sets a plurality of pixel resolutions according to directions as pixel resolutions corresponding to the partial regions; 6. The image processing device according to claim 5,

8. the model setting means sets the learning model having different numbers of learning parameters in a plurality of regions within the partial region based on a plurality of pixel resolutions set according to the direction; 8. The image processing device according to claim 7,

9. the model setting means sets, as the learning model corresponding to the partial region, the learning model based on a plurality of pixel resolutions set according to the direction for each of a plurality of regions within the partial region; 8. The image processing device according to claim 7,

10. a shape acquisition means for acquiring a rough shape of an object present in the imaging area; and the region setting means sets the partial region based on the outline shape; 2. The image processing device according to claim 1, wherein:

11. The learning model is three-dimensional information about a learning region within the imaging region; 2. The image processing device according to claim 1, wherein:

12. viewpoint acquisition means for acquiring information on a virtual viewpoint; a generating means for generating a virtual viewpoint image corresponding to the virtual viewpoint using the trained learning model; further comprising:

2. The image processing device according to claim 1, wherein:

13. an image acquisition step of acquiring a plurality of captured images obtained by capturing images of an imaging region from a plurality of directions; a region setting step of setting one or more partial regions in the imaging region; a model setting step of setting a learning model corresponding to the partial region in the captured image such that the number of learning parameters per volume increases as the pixel resolution corresponding to the partial region in the captured image increases; a learning step of learning the learning model using the captured image; An image processing method comprising:

14. A program for causing a computer to function as the image processing device according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Image processing device, image processing method, and program

    JP2023066705A