Information processing device, information processing method, and program

The information processing device sets learning areas based on object positions and learns a radiance field for each area, addressing the challenge of distinguishing volume densities for multiple objects, enabling accurate virtual viewpoint image generation.

JP2025169574APending Publication Date: 2025-11-14CANON KK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024074382
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-01
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing technologies struggle to estimate a radiance field for multiple objects in a scene, as region segmentation does not account for positional relationships between objects, leading to difficulties in dividing images into convex polyhedral regions and failing to distinguish volume densities of individual objects.

Method used

An information processing device that sets learning areas based on object positions, associates three-dimensional space models with these areas, and learns a radiance field for each area using multi-viewpoint images and camera parameters, enabling differentiation of volume densities for each object.

Benefits of technology

Enables the estimation of a radiance field that is differentiated for each object, allowing for the generation of virtual viewpoint images that accurately represent individual objects, even when multiple objects are present in the same scene.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025169574000001_ABST
    Figure 2025169574000001_ABST
Patent Text Reader

Abstract

To estimate radiance fields distinguished for each object for a scene in which a plurality of objects exist.SOLUTION: An information processing device 100 according to the present disclosure obtains data of a plurality of photographed images obtained by imaging from a plurality of viewpoints, camera parameters when each of the plurality of photographed images was captured, and object information indicating each position of a plurality of objects included as images in the photographed images, sets a plurality of learning areas based on the object information, associates a three-dimensional space model with each of the plurality of learning areas based on the number of objects included in each of the plurality of learning areas, and performs learning on the three-dimensional space model corresponding to each of the plurality of learning areas based on the data of the plurality of photographed images, the camera parameters, and the object information.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to information processing techniques for modeling spaces of objects. [Background technology]

[0002] There is a technology that estimates radiance fields for an object existing in a target space based on multiple captured images (hereinafter referred to as "multi-viewpoint images") obtained by capturing images from multiple viewpoints. There is also a technology that uses the estimated radiance fields to generate an image (hereinafter referred to as "virtual viewpoint image") that corresponds to how the object would look when viewed from an arbitrary virtual viewpoint (hereinafter referred to as "virtual viewpoint"). In the following description, the target space for which the radiance field is to be estimated will be referred to as a "scene."

[0003] Non-Patent Document 1 discloses a technology for estimating a radiance field through deep learning using multi-viewpoint images as training data. Specifically, the technology disclosed in Non-Patent Document 1 (hereinafter referred to as "the prior art") determines pixel values ​​of a virtual viewpoint image by accumulating colors weighted using volume density along rays starting from an arbitrary viewpoint position based on the estimated radiance field. More specifically, the prior art estimates a radiance field for each of multiple convex polyhedron regions obtained by dividing an entire scene so that the regions do not overlap with each other, thereby improving the efficiency of radiance field learning and virtual viewpoint image generation even when a vast scene is used as the target space. [Prior art documents] [Non-patent literature]

[0004] [Non-Patent Document 1] Daniel Rebain, 5 others, “DeRF:Decomposed Radiance Fields”, [online], November 25, 2020, arXiv, [searched on April 8, 2020], Internet<https: / / arxiv.org / pdf / 2011.12490.pdf> Summary of the Invention [Problem to be solved by the invention]

[0005] In the prior art, region segmentation is performed based on a volume density distribution roughly estimated from multi-viewpoint images, but the region segmentation does not take into account the positional relationships between objects. Furthermore, depending on the shape or arrangement of objects, it may be difficult to divide an image into multiple convex polyhedral regions so that multiple objects are not included in the same region. Specifically, for example, when the shapes of objects in the target space are complex or the distances between objects are close, multiple objects may be included in the same region. In the prior art, even when multiple objects are included in the same region, training is performed on a radiance field that outputs one color and one volume density for any position and direction. Therefore, the volume density represented by the radiance field in the prior art is the sum of the volume densities of the multiple objects included in the region, and the prior art has the problem of being unable to obtain the volume density of each object.

[0006] Therefore, an object of the present disclosure is to provide a technique that enables estimation of a radiance field that is differentiated for each object, even when multiple objects exist in a target space. [Means for solving the problem]

[0007] The information processing device according to the present disclosure includes an image data acquisition means for acquiring data on a plurality of captured images obtained by capturing images from a plurality of viewpoints and camera parameters when capturing each of the plurality of captured images; an information acquisition means for acquiring object information indicating the position of each of a plurality of objects contained as images in the captured images; an area setting means for setting a plurality of learning areas based on the object information; an association means for associating a three-dimensional space model with each of the plurality of learning areas based on the number of objects contained in each of the plurality of learning areas; and a learning means for learning the three-dimensional space model corresponding to each of the plurality of learning areas based on the data on the plurality of captured images, the camera parameters, and the object information. [Effects of the Invention]

[0008] According to the present disclosure, for a scene in which multiple objects exist, it is possible to estimate a radiance field that is differentiated for each object. [Brief explanation of the drawings]

[0009] [Figure 1] 1 is a block diagram showing an example of a hardware configuration of an information processing device according to a first embodiment. [Figure 2] 1 is a block diagram showing an example of a logical configuration of an information processing device according to a first embodiment. [Figure 3] FIG. 2 is a diagram showing an example of the arrangement of objects and imaging devices according to the first embodiment. [Figure 4] FIG. 2 is a diagram showing an example of a multi-viewpoint image according to the first embodiment. [Figure 5] 4 is a flowchart showing an example of a processing flow of the information processing device according to the first embodiment. [Figure 6] FIG. 2 is a diagram illustrating an example of a GUI according to the first embodiment. [Figure 7] FIG. 2 is a diagram illustrating an example of a bounding box according to the first embodiment. [Figure 8] FIG. 2 is a diagram showing an example of a silhouette image according to the first embodiment. [Figure 9] FIG. 2 is a diagram for explaining an example of a learning area according to the first embodiment. [Figure 10] FIG. 3 is a diagram illustrating an example of a light beam according to the first embodiment. [Figure 11] FIG. 2 is a diagram showing an example of a virtual viewpoint image according to the first embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. Note that the following embodiments do not limit the means for solving the problems of the present disclosure, and not all of the combinations of features described in the following embodiments are necessarily essential to the means for solving the problems of the present disclosure. Note that the same components will be described with the same reference numerals.

[0011] [First embodiment] In this embodiment, a mode will be described in which a plurality of regions are set as regions (hereinafter referred to as "learning regions") for learning based on the positions of a plurality of objects. Also, in this embodiment, as an example, a mode will be described in which the radiance field in each learning region is represented by a three-dimensional space model (hereinafter simply referred to as a "model") that outputs the color and the volume density of each object included in the region.

[0012] <Hardware configuration of information processing device> FIG. 1 is a block diagram showing an example of a hardware configuration of an information processing device 100 according to the first embodiment. The information processing device 100 includes, as its hardware configuration, a CPU 101, a RAM 102, a ROM 103, a serial I / F (interface) 104, a VC (video card) 105, and a general-purpose I / F 106. The various components of the information processing device 100 are communicably connected to one another via a system bus 107. The CPU 101 uses the RAM 102 as a work memory and executes an operating system (OS) and various programs stored in the ROM 103, a storage device 111, or the like. The CPU 101 executes the various programs to control the entire information processing device 100 via the system bus 107. Note that the processing of each step shown in a flowchart described below is realized by the CPU 101 expanding program code stored in the ROM 103, the storage device 111, or the like into the RAM 102 and executing the expanded program code.

[0013] The serial I / F 104 is an interface configured by a serial ATA or the like, and connects the information processing device 100 and a storage device 111 via a serial bus 108. The storage device 111 is a large-capacity storage device such as an HDD (hard disk drive) or an SSD (solid state drive). In this embodiment, the storage device 111 is described as an external device of the information processing device 100, but the information processing device 100 may include the storage device 111 internally. The VC 105 receives a control signal from the CPU 101 and outputs a signal related to a display image to the display device 112 via the serial bus 109. The display device 112 is configured by a liquid crystal display or the like, and displays the display image based on the signal related to the display image output by the information processing device 100. The general-purpose I / F 106 is connected to an input device 113 such as a mouse or a keyboard via a serial bus 110 and receives an input signal from the input device 113.

[0014] The CPU 101 displays a GUI (Graphical User Interface) provided by a program on the display device 112 via the VC 105, and receives an input signal indicating an instruction from a user obtained via the input device 113. The information processing device 100 is realized, for example, by a desktop PC (Personal Computer). The information processing device 100 may also be realized by a notebook PC or a tablet PC integrated with the display device 112. The storage device 111 may also be realized by a medium (portable storage medium) and a drive such as a disk drive for accessing the medium, or a reader such as a memory card reader. The medium may be a flexible disk (FD), CD-ROM, DVD, USB memory, MO, flash memory, or the like.

[0015] <Logical configuration of information processing device> 2 is a block diagram showing an example of the logical configuration of the information processing device 100 according to the first embodiment. The information processing device 100 has, as its logical configuration, an imaging data acquisition unit 201, an information acquisition unit 202, a region setting unit 203, an association unit 204, a learning unit 205, a viewpoint acquisition unit 206, an image generation unit 207, and an output unit 208. Each unit of the information processing device 100 as its logical configuration is realized by the CPU 101 executing a program stored in the ROM 103 or the like, using the RAM 102 as a work memory. Note that not all of the processes described below necessarily need to be realized by the CPU 101 executing a program, and the information processing device 100 may be configured so that some or all of the processes are executed by one or more processing circuits other than the CPU 101.

[0016] The imaging data acquisition unit 201 acquires a plurality of captured image (multi-viewpoint image) data obtained by capturing images of objects present in a predetermined scene from various viewpoint positions, based on instructions from the user input via the input device 113. Hereinafter, as an example, the captured image data acquired by the imaging data acquisition unit 201 will be described as image data in RGB image format. The imaging data acquisition unit 201 may acquire captured image data output by the imaging device directly from the imaging device, or may acquire the captured image data by reading the captured image data from the storage device 111 or the like in which the captured image data is stored in advance. The acquired multi-viewpoint image data is transmitted to the information acquisition unit 202 and the learning unit 205.

[0017] FIG. 3 is a diagram showing an example of an arrangement of objects 301 to 303 present in a scene 300 according to the first embodiment, and a plurality of imaging devices including imaging devices 311 to 313 that capture images of the objects 301 to 303. FIG. 4 is a diagram showing an example of a multi-view image acquired by the imaging data acquisition unit 201 according to the first embodiment. Specifically, FIG. 4 shows examples of captured images 410, 420, and 430 acquired by imaging using the imaging devices 311 to 313, respectively. More specifically, FIG. 4(a) shows an example of a captured image 410 acquired by imaging using the imaging device 311. FIG. 4(b) shows an example of a captured image 420 acquired by imaging using the imaging device 312. FIG. 4(c) shows an example of a captured image 430 acquired by imaging using the imaging device 313. The captured images 410, 420, and 430 include images 411, 421, and 431 of the object 301, images 412, 422, and 432 of the object 302, and images 413, 423, and 433 of the object 303.

[0018] Furthermore, the imaging data acquisition unit 201 acquires camera parameters of each imaging device, including the imaging devices 311 to 313, that captured each captured image constituting the multi-view image. Hereinafter, the camera parameters acquired by the imaging data acquisition unit 201 will be described as including internal parameters, external parameters, and distortion parameters of the imaging device. The internal parameters are parameters that represent the position of the principal point of the imaging device and the focal length of the lens of the imaging device. The external parameters are parameters that represent the position of the imaging device and the optical axis direction of the imaging device, i.e., the attitude of the imaging device. The distortion parameters are parameters that represent the distortion of the lens of the imaging device.

[0019] In the following description, the imaging data acquisition unit 201 is assumed to acquire the camera parameters held by each imaging device by requesting the same from each imaging device, but the source of the camera parameters is not limited to the imaging device. For example, the imaging data acquisition unit 201 may acquire the camera parameters by reading them from the storage device 111 or the like in which the camera parameters are stored in advance. The acquired camera parameters of each imaging device are transmitted to the information acquisition unit 202 and the learning unit 205.

[0020] The information acquisition unit 202 estimates the outline shape of each of the objects 301 to 303 based on the multi-view image data and camera parameters acquired by the imaging data acquisition unit 201, thereby acquiring three-dimensional shape data of each of the objects 301 to 303. Furthermore, the information acquisition unit 202 acquires bounding box data including an identification number and information indicating the position and size of a bounding box surrounding the outline shape of each of the objects 301 to 303, based on the estimated outline shape. Details of the outline shape estimation process and the bounding box data acquisition process performed by the information acquisition unit 202 will be described later. The acquired bounding box data is transmitted to the region setting unit 203. Furthermore, the information acquisition unit 202 acquires silhouette image data indicating the silhouette of each of the objects 301 to 303 corresponding to each captured image that constitutes the multi-view image. Details of the silhouette image data acquisition process performed by the information acquisition unit 202 will be described later. The acquired silhouette image data is transmitted to the learning unit 205.

[0021] In this embodiment, the information acquisition unit 202 is described as estimating the outline shape of each of the objects 301 to 303 to acquire the three-dimensional shape data of each of the objects 301 to 303, but the method of acquiring the three-dimensional shape data is not limited to this. For example, the information acquisition unit 202 may acquire the three-dimensional shape data by receiving from an external device the three-dimensional shape data obtained by the external device estimating the outline shape of each of the objects 301 to 303 based on the multi-view image data and camera parameters. Furthermore, in this embodiment, the information acquisition unit 202 is described as acquiring the bounding box data by generating it based on the outline shape, but the method of acquiring the bounding box data is not limited to this. For example, the information acquisition unit 202 may acquire the bounding box data by receiving from the external device the bounding box data generated by the external device based on the outline shape.

[0022] The area setting unit 203 sets a learning area within a scene based on the bounding box data acquired by the information acquisition unit 202. Details of the learning area setting process by the area setting unit 203 will be described later. Information indicating the set learning area (hereinafter referred to as "learning area information") is transmitted to the association unit 204 and the learning unit 205. The association unit 204 associates, with each learning area set by the area setting unit 203, one model including at least a number of volume densities corresponding to the number of objects in the learning area as output parameters. Details of the processing by the association unit 204 will be described later. Information on the associated models is transmitted to the learning unit 205.

[0023] The learning unit 205 estimates a radiance field based on the multi-view image data and camera parameters acquired by the imaging data acquisition unit 201 and the silhouette image data acquired by the information acquisition unit 202. Specifically, the learning unit 205 estimates a radiance field for each learning region set by the region setting unit 203. In this embodiment, as an example, the learning unit 205 is described as estimating a radiance field for each learning region, the radiance field being represented by a model associated by the association unit 204. Details of the radiance field estimation process by the learning unit 205 will be described later. Information on the model indicating the radiance field estimated by the learning unit 205 is transmitted to the image generation unit 207 and the output unit 208.

[0024] The viewpoint acquisition unit 206 acquires information related to a virtual viewpoint (hereinafter referred to as "virtual viewpoint information"). The virtual viewpoint information includes at least camera parameters related to the virtual viewpoint (hereinafter referred to as "virtual camera parameters"), and the virtual camera parameters include information indicating the position of the virtual viewpoint and information indicating the line of sight at the virtual viewpoint. In the following, in order to distinguish between the camera parameters of the imaging device and the virtual camera parameters, the camera parameters of the imaging device will be simply referred to as "camera parameters." In addition to the virtual camera parameters, the virtual viewpoint information may include pixel count information indicating the number of pixels of a virtual viewpoint image generated by the image generation unit 207, and object information such as an identification number that can uniquely identify an object to be included as an image in the virtual viewpoint image. The virtual viewpoint information may also include information indicating the angle of view from the virtual viewpoint. The virtual viewpoint information is acquired based on, for example, a user instruction input via the input device 113. The virtual viewpoint information acquired by the viewpoint acquisition unit 206 is transmitted to the image generation unit 207.

[0025] The image generation unit 207 generates a virtual viewpoint image using the virtual viewpoint information acquired by the viewpoint acquisition unit 206 and the radiance field estimated by the learning unit 205. Details of the virtual viewpoint image generation process by the image generation unit 207 will be described later. Data of the virtual viewpoint image generated by the image generation unit 207 is transmitted to the output unit 208. The output unit 208 outputs the virtual viewpoint image generated by the image generation unit 207. Specifically, for example, the output unit 208 generates a display image including the virtual viewpoint image and outputs a signal related to the display image to the display device 112 to display the display image on the display device 112. The output destination of the virtual viewpoint image is not limited to the display device 112. For example, the output unit 208 may output data of the virtual viewpoint image to the storage device 111 and store the data in the storage device 111, or may output the data to an external device other than the information processing device 100. The output unit 208 also outputs information of a model indicating the radiance field estimated by the learning unit 205. Specifically, the output unit 208 may output information about the model indicating the radiance field to the storage device 111 for storage, or may output the information to an external device different from the information processing device 100.

[0026] <Operation of information processing device> 5 is a flowchart showing an example of a processing flow in the information processing device 100 according to the first embodiment. Note that the letter "S" at the beginning of a reference symbol indicates a step (process). First, in S501, the imaging data acquisition unit 201 acquires multi-view image data and camera parameters corresponding to each captured image constituting the multi-view image based on an instruction from a user.

[0027] FIG. 6 shows examples of GUIs 600 and 610 displayed on the display device 112 according to the first embodiment. An instruction from the user in S501 is accepted via the GUI 600 shown as an example in FIG. 6(a). In FIG. 6(a), data path setting fields 601 and 602 are fields for accepting input of data paths indicating the locations of files containing multi-viewpoint image data and camera parameter data, respectively. A button 603 is pressed to issue an instruction to execute processing described below. When the user presses button 603, the information processing device 100 executes processing in S501 and then executes processing in S502. FIG. 6(b) will be described later.

[0028] After S501, in S502, the information acquisition unit 202 acquires bounding box data corresponding to each object present in the scene based on the multi-view image data and camera parameters acquired in S501. Specifically, the information acquisition unit 202 first acquires, for each captured image constituting the multi-view image, a difference image indicating the difference from an image in which the object is not captured (hereinafter referred to as a "background image"). The background image data is prepared, for example, by capturing an image of a scene in which no object is present in advance. In this embodiment, the information acquisition unit 202 is described as generating the difference image, but the information acquisition unit 202 may also acquire the difference image by receiving difference image data generated by an external device.

[0029] Next, the information acquisition unit 202 estimates the three-dimensional shape of each object present in the scene based on the difference images corresponding to each captured image and the camera parameters. This three-dimensional shape estimation can be performed using known three-dimensional shape estimation techniques such as volume intersection or stereo matching. In this embodiment, the volume intersection method is used to acquire three-dimensional shape data represented by a set of voxels as data indicating the rough shape of the object. The rough shape acquired by the information acquisition unit 202 only needs to represent the position and rough shape of each object in the target space, and does not necessarily need to represent the fine irregularities and colors of each object.

[0030] Next, the information acquisition unit 202 associates an identification number corresponding to one object with each set of voxels by regarding a set of spatially continuous voxels as one object for the voxels that make up the acquired outline shape. Next, the information acquisition unit 202 calculates the position and size of a rectangular parallelepiped (bounding box) that circumscribes each set of voxels associated with an identification number. Through the above processing, the information acquisition unit 202 acquires bounding box data for each object, i.e., information indicating the position and size of the bounding box that surrounds the object and to which an identification number is assigned.

[0031] FIG. 7 is a diagram showing an example of bounding boxes acquired by the information acquisition unit 202 according to the first embodiment. Specifically, FIG. 7 shows example schematic shapes 701 to 703 of objects 301 to 303 acquired for the scene 300 shown in FIG. 3, and bounding boxes corresponding to the schematic shapes 701 to 703. In FIG. 7, schematic shapes 701, 702, and 703 associated with identification number k (1, 2, or 3) represent the schematic shapes corresponding to objects 301, 302, and 303 shown in FIG. 3, respectively. Bounding boxes BB1, BB2, and BB3 are rectangular parallelepipeds that circumscribe the schematic shapes 701, 702, and 703, respectively, and whose sides are parallel to three-dimensional coordinate axes that indicate the position of the target space. Hereinafter, the object corresponding to the schematic shape with identification number k will be referred to as OBJ. k , the total number of objects in the scene is K, and the object OBJ k Let BB be the bounding box corresponding to k This is expressed and explained as follows.

[0032] In this embodiment, the information acquisition unit 202 is described as acquiring a rough shape by estimating the three-dimensional shape of the object based on the multi-view images, but the method of acquiring the rough shape of the object is not limited to this. For example, the information acquisition unit 202 may acquire rough shape data of the object that has been estimated in advance by another external device or that has been separately prepared, by reading it from the storage device 111, based on an instruction from the user.

[0033] Furthermore, in this embodiment, the information acquisition unit 202 is described as acquiring a rough shape represented by voxels, but the information acquisition unit 202 may acquire a rough shape represented by components other than voxels. For example, the rough shape may be represented by a surface shape configured using a polygon mesh made up of a plurality of polygons. In this case, a series of polygon meshes connected by edges can be considered as the rough shape of a single object.

[0034] In addition, in this embodiment, the information acquisition unit 202 is described as calculating the position and size of a rectangular parallelepiped bounding box that circumscribes the outline shape, but the shape of the bounding box is not limited to a rectangular parallelepiped. For example, the shape of the bounding box may be a convex polyhedron or a sphere other than a rectangular parallelepiped, as long as it is a convex three-dimensional shape that contains the outline shape of each object.

[0035] After S502, in S503, the information acquisition unit 202 acquires silhouette image data indicating the silhouette of each object corresponding to each captured image that constitutes the multi-viewpoint image, based on the multi-viewpoint image data and camera parameters acquired in S501. Specifically, the information acquisition unit 202 acquires silhouette image data indicating the visibility of each object by projecting the rough shape of the object estimated in S502 onto the plane of each captured image based on the camera parameters.

[0036] More specifically, first, the information acquisition unit 202 prepares a silhouette image for each object identification number, in which all pixel values ​​are initialized to 0, for each captured image that constitutes the multi-viewpoint image. Next, the information acquisition unit 202 generates depth images for each identification number that correspond to all captured images that constitute the multi-viewpoint image by individually projecting the outline shape of each object onto an image plane using the camera parameters. Well-known computer graphics techniques can be used to generate the depth images. Next, the information acquisition unit 202 calculates the pixel value for each pixel of the generated depth image, i.e., the depth value, by calculating the pixel value for each pixel of the generated depth image. max The smallest identification number k that is less than dmin Identify the identified identification number k dmin The pixel value on the silhouette image corresponding to is set to 1. Here, the threshold d max is the maximum depth value in the scene, and is determined in advance based on the relative positional relationship between the position of the image capture device indicated by the camera parameters and the scene.

[0037] FIG. 8 is a diagram showing an example of a silhouette image acquired by the information acquisition unit 202 according to the first embodiment. In the silhouette image shown as an example in FIG. 8, pixels in an area corresponding to the silhouette of an object are represented by white, which has a pixel value of 1, and other areas are represented by black, which has a pixel value of 0. Specifically, FIGS. 8(a), 8(b), and 8(c) are silhouette images of objects corresponding to identification numbers k of 1, 2, and 3, respectively, corresponding to the captured image 410. FIGS. 8(d), 8(e), and 8(f) are silhouette images of objects corresponding to identification numbers k of 1, 2, and 3, respectively, corresponding to the captured image 420. FIGS. 8(g), 8(h), and 8(i) are silhouette images of objects corresponding to identification numbers k of 1, 2, and 3, respectively, corresponding to the captured image 430. Each silhouette image shown in FIG. 8 shows pixel areas including the image of an object in the corresponding captured image, sorted by identification number.

[0038] After S503, in S504, the area setting unit 203 sets a learning area based on the bounding box data acquired in S502. Specifically, in the target space, a certain bounding box BB k does not have an area overlapping with any other bounding box, the area setting unit 203 sets the bounding box BB k In other cases, the region setting unit 203 sets the bounding box BB k A convex three-dimensional area that encompasses the entire bounding box and one or more bounding boxes having overlapping areas and does not overlap with other learning areas is set as one of the learning areas.

[0039] 9 is a diagram illustrating an example of a learning area set by the area setting unit 203 according to the first embodiment. For three bounding boxes 901 to 903 having overlapping areas shown as an example in FIG. 9, the area setting unit 203 sets the smallest rectangular parallelepiped 904 that encompasses these three bounding boxes 901 to 903 as one of the learning areas.

[0040] Furthermore, the area setting unit 203 outputs the identification number and number of objects corresponding to each bounding box included in each set learning area, in association with information indicating the learning area. In the example shown in Fig. 7, a bounding box BB1 that contains a bounding box BB2 therein is set as a learning area ROL1 that includes bounding boxes corresponding to two objects with identification numbers k of 1 and 2. Furthermore, a bounding box BB3 that does not have an area overlapping with other bounding boxes is set as a learning area ROL2 that includes a bounding box corresponding to one object with identification number k of 3. Through this processing, two learning areas ROL1 and ROL2 are set in the scene 300.

[0041] After S504, in S505, the association unit 204 associates, with each of the learning areas set in S504, a model including at least a number of volume densities corresponding to the number of objects in the learning area as output parameters. Specifically, for example, the association unit 204 associates a model expressed by the following formula (1) with a learning area in which the number of objects in the learning area is K'. F Θ :(x,y,z,θ,φ)→(R,G,B,σ id1 ,···,σ idK´ )...Equation (1)

[0042] Here, (x, y, z) are three-dimensional coordinates that indicate the position of the object in space, (θ, φ) are the direction of the object in space, and (R, G, B) are the color determined by the position and direction of the object in space. id1 ,…,σ idK´ ) represents the volume density determined by the position of each of the K' objects whose identification numbers are id1, ..., or idK'. The function F formulated by Eq. (1) Θis a model that outputs color and volume density for each object as output parameters with respect to three-dimensional position and direction. For example, a learning area ROL1 including two objects with identification numbers k of 1 and 2 is associated with a model that outputs two volume densities σ1 and σ2 corresponding to each object. Also, a learning area ROL2 including one object with identification number k of 3 is associated with a model that outputs one volume density σ3.

[0043] After S505, in S506, the learning unit 205 estimates a radiance field represented by the model associated in S505 for each learning region set in S504. Specifically, the learning unit 205 estimates a radiance field represented by the model associated with each learning region based on the multi-viewpoint image data and camera parameters acquired in S501 and the silhouette image data acquired in S503. In this embodiment, the learning unit 205 estimates a radiance field represented by the model associated with each learning region based on the function F shown in Equation (1) as an example. Θ In this description, it is assumed that the radiance field is estimated by deep learning of a model realized by an MLP (Multi-layer Perceptron). The radiance field is expressed as MLP parameters, i.e., weight parameters related to the nodes that make up the MLP, and the weight parameters are stored in a memory area reserved in RAM 102 for each learning area.

[0044] Furthermore, in this embodiment, the learning unit 205 will be described as learning the MLP as follows. First, based on the output from the model, the learning unit 205 generates virtual viewpoint images corresponding to each captured image constituting the multi-viewpoint image acquired in S501, and silhouette images for each object based on the virtual viewpoint images (hereinafter referred to as "virtual silhouette images"). Next, the learning unit 205 optimizes the weight parameters of the MLP so that the pixel values ​​of these images approach the pixel values ​​of each captured image constituting the multi-viewpoint image acquired in S501 and the silhouette image acquired in S503. Specifically, for example, first, the learning unit 205 calculates a teacher signal CGT (r) and the predicted signal C shown in the following equation (3) based on the output value from the model associated with each learning area. pred Next, the learning unit 205 calculates the teacher signal C GT (r) and predicted signal C pred The squared Euclidean distance between (r) is calculated, and the MLP is trained using the backpropagation method with this as the loss. TIFF2025169574000002.tif44150 TIFF2025169574000003.tif37150

[0045] Here, r in equation (2) is a ray determined based on the pixel position in the captured image and the camera parameters. R (r), I G (r), and I B (r) is the pixel value of the captured image corresponding to the light ray r, and I Sk (r) is the object OBJ corresponding to the ray r k 10 is a diagram showing an example of the mutual positional relationship between a ray r, a scene 300, a position 1001 of the image capture device, an image plane 1002, and a pixel 1003 corresponding to the ray r, which are set by the learning unit 205 according to the first embodiment. Furthermore, R(r), G(r), and B(r) in equation (3) are pixel values ​​of the virtual viewpoint image corresponding to the ray r, and are calculated using, for example, the following equation (4). Furthermore, S k (r) is the object OBJ corresponding to the ray r k The pixel values ​​of the virtual silhouette image corresponding to the virtual viewpoint image are calculated using, for example, the following equation (5). TIFF2025169574000004.tif16150 TIFF2025169574000005.tif21150

[0046] Here, equation (4) corresponds to the well-known volume rendering of an RGB image. In equation (4), i is the index of the sampling point on the ray r, and N is the number of sampling points. Furthermore, R(i), G(i), and B(i) are RGB values ​​for the sampling point, and are output from a model associated with the learning region that includes the sampling point. Furthermore, T(i) is the cumulative transmittance from the position of the imaging device to the sampling point, and is calculated, for example, using the following equation (6). Furthermore, α(i) is the opacity of the sampling point summed over all objects, and is calculated, for example, using the following equation (7). Furthermore, α in equation (5) k (i) is the object OBJ k is the opacity of the sampling point with respect to the pixel value, and is calculated using, for example, the following equation (8). TIFF2025169574000006.tif18150 TIFF2025169574000007.tif15150α k (i)=1-exp(-σ k (i)δ i )...Equation (8)

[0047] Here, j in Equation (6) is the index of the sampling point before the i-th point on the ray r, and σ k (j) is the object OBJ for the sampling point k The volume density of the sampled point is the value output from the model associated with the learning region that contains the sampled point. k If the volume density of is not included, σ k The value of (j) is treated as 0. Also, δ j is the distance from the jth sampling point to the j+1th sampling point.

[0048] As described above, the learning unit 205 learns a model so that the difference between not only the RGB images of the captured image and the virtual viewpoint image but also the silhouette image and virtual silhouette image of each object is small. The model obtained by such learning makes it possible to estimate the volume density of each object. Note that, since the learning region according to this embodiment is a convex polyhedron that does not overlap with other learning regions, the information processing device 100 according to this embodiment can also generate a virtual viewpoint image using the "Painter's Algorithm" disclosed in Non-Patent Document 1. In addition, in this embodiment, the function F Θ is explained as being realized by MLP, but the function F Θ may be realized by something other than an MLP. For example, the function F Θ may be implemented using a sparse voxel grid that stores the coefficients and volume densities of spherical harmonics that represent color.

[0049] After S506, in S507, the viewpoint acquisition unit 206 acquires virtual viewpoint information based on an instruction from the user. Next, in S508, the image generation unit 207 generates a virtual viewpoint image using the virtual viewpoint information acquired in S507 and the radiance field estimated in S506. The instruction from the user in S507 is accepted via a GUI 610 displayed on the display device 112, an example of which is shown in FIG. 6(b).

[0050] In FIG. 6(b), a virtual camera parameter setting field 611 is a field for receiving input of a data path indicating the location of virtual camera parameter data used to generate a virtual viewpoint image. An image size setting field 612 is a field for receiving input of the number of pixels in the horizontal and vertical directions of the virtual viewpoint image to be generated. An object setting field 613 is a field for receiving input of an identification number or the like corresponding to an object to be included as an image in the virtual viewpoint image. A button 614 is a button that is pressed to instruct execution of the process of S508. A display area 615 is an area in which the generated virtual viewpoint image is displayed. When the user presses the button 604, the image generation unit 207 generates a virtual viewpoint image based on the values ​​input in the virtual camera parameter setting field 611, the image size setting field 612, and the object setting field 613.

[0051] The image generating unit 207 generates the virtual viewpoint image by calculating the pixel value C(r) of the virtual viewpoint image using, for example, the following equations (9) to (11).

[0052] TIFF2025169574000008.tif16150 TIFF2025169574000009.tif18150 TIFF2025169574000010.tif13150Here, K in equations (10) and (11) draw is a set of identification numbers of objects to be included as images in the virtual viewpoint image. In this embodiment, in equation (9), the volume density σ corresponding to the objects to be included as images in the virtual viewpoint image is k Color weighting is performed using (i), which generates a virtual viewpoint image in which objects other than the object to be included as an image in the virtual viewpoint image are transmitted, thereby obtaining a virtual viewpoint image that contains only the image of the desired object.

[0053] FIG. 11 is a diagram showing an example of a virtual viewpoint image 1100 generated by the image generation unit 207 according to the first embodiment. Specifically, the virtual viewpoint image 1100 is based on the multi-viewpoint image shown in FIG. 4. More specifically, the virtual viewpoint image 1100 is generated when the virtual camera parameters are the same as the camera parameters of the image capture device 313 shown in FIG. 3 and are set to include images of objects corresponding to the schematic shapes 702 and 703 of which identification numbers k are 2 and 3 in FIG. 7. The virtual viewpoint image 1100 does not include an image of an object corresponding to the schematic shape 701 of which identification number k is 1 shown in FIG. 7. Furthermore, the virtual viewpoint image 1100 only includes an image 1101 of an object corresponding to the schematic shape 702 of which identification number k is 2 and an image 1102 of an object corresponding to the schematic shape 703 of which identification number k is 3 shown in FIG.

[0054] After S508, in S509, the output unit 208 outputs the virtual viewpoint image generated in S508. For example, the output unit 208 outputs the virtual viewpoint image generated in S508 so that it is displayed in the display area 615 of the GUI 610. Next, in S510, the output unit 208 outputs information indicating the radiance field estimated in S506, i.e., information on a model representing the radiance field, to a storage device such as the storage device 111, and stores the information in the storage device. At this time, it is desirable that the output unit 208 store the model information including an identification number of an object corresponding to the volume density included in the model. In this case, by performing the processes of S507 to S509 using the stored model information, it is possible to obtain a virtual viewpoint image including only an image of the desired object without performing the processes of S501 to S506 again. After S510, the information processing device 100 ends the process of the flowchart shown in FIG. 5.

[0055] According to the information processing device 100 configured as above, when a plurality of objects are included in one learning area, it is possible to estimate a radiance field that distinguishes between the objects and enables identification of any object. Furthermore, according to the information processing device 100, it is possible to generate an image (virtual viewpoint image) that includes only one or more arbitrary objects among the plurality of objects as images, using the estimated radiance field.

[0056] In the present embodiment, the captured image has been described as an RGB image as an example, but the captured image may be an image expressed in other formats such as a grayscale image, an XYZ image, or a YUV image. Also, in the present embodiment, the color of an object has been described as being determined by its position and direction as an example, but the color of an object may be determined only by its position, regardless of its direction.

[0057] [Other embodiments] The present disclosure can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.

[0058] It should be noted that within the scope of the present disclosure, the embodiments may be freely combined, any component of each embodiment may be modified, or any component of each embodiment may be omitted.

[0059] [Configuration of the present disclosure] The present disclosure includes the following configurations, methods, and programs.

[0060] <Configuration 1> an imaging data acquisition means for acquiring data of a plurality of captured images obtained by imaging from a plurality of viewpoints and camera parameters when each of the plurality of captured images is captured; an information acquisition means for acquiring object information indicating the positions of a plurality of objects included as images in the captured image; an area setting means for setting a plurality of learning areas based on the object information; a matching means for matching a three-dimensional space model to each of the plurality of learning areas based on the number of objects included in each of the plurality of learning areas; a learning means for learning the three-dimensional space model corresponding to each of the plurality of learning regions based on the data of the plurality of captured images, the camera parameters, and the object information; An information processing device comprising:

[0061] <Configuration 2> the information acquisition means acquires, as the object information, data of a bounding box that contains each of the plurality of objects; the area setting means sets, as part of the plurality of learning areas, each of one or more of the bounding boxes that does not have an area overlapping with another of the bounding boxes among the plurality of bounding boxes that have been acquired, and sets, as part of the plurality of learning areas, an area that includes two or more of the bounding boxes whose areas at least partially overlap with each other among the plurality of bounding boxes; 2. The information processing device according to configuration 1,

[0062] <Configuration 3> the associating means associates, with each of the plurality of learning areas, the three-dimensional space model having at least parameters indicating volume density, the number of which is the same as the number of objects included in the learning area; 3. The information processing device according to configuration 1 or 2, characterized in that:

[0063] <Configuration 4> the three-dimensional spatial model represents a radiance field in a corresponding training region; 4. The information processing device according to any one of configurations 1 to 3, characterized in that:

[0064] <Configuration 5> the three-dimensional space model is a learning model configured by one or more multilayer perceptrons; 5. The information processing device according to any one of configurations 1 to 4, characterized in that:

[0065] <Configuration 6> the information acquisition means acquires three-dimensional shape data indicating three-dimensional shapes of each of the plurality of objects estimated based on the plurality of captured images and the camera parameters, and acquires the object information based on the three-dimensional shape data corresponding to each of the plurality of objects; 6. The information processing device according to any one of configurations 1 to 5,

[0066] <Configuration 7> the information acquisition means acquires the three-dimensional shape data corresponding to each of the plurality of objects by estimating a three-dimensional shape of each of the plurality of objects based on the plurality of captured images and the camera parameters; 7. The information processing device according to configuration 6,

[0067] <Configuration 8> the information acquisition means acquires the object information by regarding a set of spatially connected components that constitute the three-dimensional shape data as a three-dimensional shape corresponding to one object; 8. The information processing device according to configuration 6 or 7, characterized in that:

[0068] <Configuration 9> the information acquisition means generates a silhouette image by projecting the three-dimensional shape corresponding to each of the plurality of objects onto an image plane corresponding to each of the plurality of captured images for each of the plurality of objects using the camera parameters; the learning means learns the three-dimensional space model by calculating a loss using at least pixel values ​​in each of the plurality of captured images and pixel values ​​of the silhouette image; 9. The information processing device according to any one of configurations 6 to 8,

[0069] <Configuration 10> the information acquisition means acquires silhouette image data generated by projecting, for each of the plurality of objects, the three-dimensional shapes of the plurality of objects estimated based on the plurality of captured images and the camera parameters onto an image plane corresponding to each of the plurality of captured images using the camera parameters; the learning means learns the three-dimensional space model by calculating a loss using at least pixel values ​​in each of the plurality of captured images and pixel values ​​of the silhouette image; 6. The information processing device according to any one of configurations 1 to 5,

[0070] <Configuration 11> an image generation means for generating an image corresponding to an appearance from an arbitrary virtual viewpoint using the learning result of the three-dimensional space model; further comprising: 11. The information processing device according to any one of configurations 1 to 10, characterized in that:

[0071] <Method> an imaging data acquisition step of acquiring data of a plurality of captured images obtained by imaging from a plurality of viewpoints and camera parameters when each of the plurality of captured images is captured; an information acquisition step of acquiring object information indicating the positions of a plurality of objects included as images in the captured image; an area setting step of setting a plurality of learning areas based on the object information; a matching step of matching a three-dimensional space model to each of the plurality of learning areas based on the number of objects included in each of the plurality of learning areas; a learning step of learning the three-dimensional space model corresponding to each of the plurality of learning regions based on the data of the plurality of captured images, the camera parameters, and the object information; An information processing method comprising:

[0072] <Program> 12. A program for causing a computer to function as the information processing device according to any one of configurations 1 to 11. [Explanation of symbols]

[0073] 100 Information processing device 201 Imaging data acquisition unit 201 202 Information Acquisition Department 203 Area setting section 204 Mapping Section 205 Learning Department

Claims

1. an imaging data acquisition means for acquiring data of a plurality of captured images obtained by imaging from a plurality of viewpoints and camera parameters when each of the plurality of captured images is captured; an information acquisition means for acquiring object information indicating the positions of a plurality of objects included as images in the captured image; an area setting means for setting a plurality of learning areas based on the object information; a correspondence means for associating a three-dimensional space model with each of the plurality of learning areas based on the number of objects included in each of the plurality of learning areas; a learning means for learning the three-dimensional space model corresponding to each of the plurality of learning regions based on the data of the plurality of captured images, the camera parameters, and the object information; An information processing device comprising:

2. the information acquisition means acquires, as the object information, data of a bounding box that contains each of the plurality of objects; the area setting means sets, as part of the plurality of learning areas, one or more of the bounding boxes that do not have an area overlapping with other bounding boxes among the plurality of bounding boxes that have been acquired, and sets, as part of the plurality of learning areas, an area that includes two or more of the bounding boxes whose areas at least partially overlap with each other among the plurality of bounding boxes; 2. The information processing device according to claim 1,

3. the associating means associates, with each of the plurality of learning areas, the three-dimensional space model having at least parameters indicating volume density, the number of which is the same as the number of objects included in the learning area; 2. The information processing device according to claim 1,

4. the three-dimensional spatial model represents a radiance field in a corresponding training region; 2. The information processing device according to claim 1,

5. the three-dimensional space model is a learning model configured by one or more multilayer perceptrons; 2. The information processing device according to claim 1,

6. the information acquisition means acquires three-dimensional shape data indicating three-dimensional shapes of each of the plurality of objects estimated based on the plurality of captured images and the camera parameters, and acquires the object information based on the three-dimensional shape data corresponding to each of the plurality of objects; 2. The information processing device according to claim 1,

7. the information acquisition means acquires the three-dimensional shape data corresponding to each of the plurality of objects by estimating a three-dimensional shape of each of the plurality of objects based on the plurality of captured images and the camera parameters; 7. The information processing device according to claim 6,

8. the information acquisition means acquires the object information by regarding a set of spatially connected components that constitute the three-dimensional shape data as a three-dimensional shape corresponding to one object; 7. The information processing device according to claim 6,

9. the information acquisition means generates a silhouette image by projecting the three-dimensional shape corresponding to each of the plurality of objects onto an image plane corresponding to each of the plurality of captured images for each of the plurality of objects using the camera parameters; the learning means learns the three-dimensional space model by calculating a loss using at least pixel values ​​in each of the plurality of captured images and pixel values ​​of the silhouette image; 7. The information processing device according to claim 6,

10. the information acquisition means acquires silhouette image data generated by projecting, for each of the plurality of objects, the three-dimensional shapes of the plurality of objects estimated based on the plurality of captured images and the camera parameters onto an image plane corresponding to each of the plurality of captured images using the camera parameters; the learning means learns the three-dimensional space model by calculating a loss using at least pixel values ​​in each of the plurality of captured images and pixel values ​​of the silhouette image; 2. The information processing device according to claim 1,

11. an image generation means for generating an image corresponding to an appearance from an arbitrary virtual viewpoint using the learning result of the three-dimensional space model; further comprising:

2. The information processing device according to claim 1,

12. an imaging data acquisition step of acquiring data of a plurality of captured images obtained by imaging from a plurality of viewpoints and camera parameters when each of the plurality of captured images is captured; an information acquisition step of acquiring object information indicating the positions of a plurality of objects included as images in the captured image; an area setting step of setting a plurality of learning areas based on the object information; a matching step of matching a three-dimensional space model to each of the plurality of learning areas based on the number of objects included in each of the plurality of learning areas; a learning step of learning the three-dimensional space model corresponding to each of the plurality of learning regions based on the data of the plurality of captured images, the camera parameters, and the object information; An information processing method comprising:

13. A program for causing a computer to function as the information processing device according to any one of claims 1 to 11.