Information processing device, information processing method and program

The information processing apparatus efficiently estimates radiance fields for individual objects by using a single function to model scenes, reducing computational and memory demands and enabling focused virtual viewpoint image generation.

JP2025093503APending Publication Date: 2025-06-24CANON KK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023209192
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-12
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Existing techniques for estimating radiance fields for individual objects in a scene require significant computational resources and memory due to the need to estimate both the radiance field for each object of interest and the entire scene.

Method used

An information processing apparatus that estimates a radiance field and likelihood field using a function F or F' based on multi-viewpoint images and likelihood maps, reducing computational and memory requirements by modeling the scene with a single function that outputs color, volume density, and likelihood values for each object.

Benefits of technology

This approach allows for high-precision radiance field estimation with reduced computational and memory usage, enabling the generation of virtual viewpoint images focused on specific objects while suppressing unnecessary object representations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025093503000001_ABST
    Figure 2025093503000001_ABST
Patent Text Reader

Abstract

To suppress the computational complexity and memory consumption in estimating a radiance field of an object of interest with high accuracy.SOLUTION: An information processing device 100 is configured to: acquire data on a plurality of captured images obtained by capturing objects present in a space from a plurality of view points and camera parameters corresponding to the plurality of view points respectively; acquire, for each object of interest among the objects, information showing a likelihood that an image formed at each of pixels of the plurality of captured images is the object of interest as a likelihood value corresponding to the pixel of each of the plurality of captured images for each object of interest; and estimate, based on data on the plurality of captured images, the plurality of camera parameters, and likelihood values corresponding to the respective pixels of the plurality of captured images for respective objects of interest, information, associated with the space, including color information corresponding to the respective positions in the space and the likelihood values by the objects of interest at the respective positions in the space.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an information processing technology for modeling a target space.

Background Art

[0002] There is a technique for estimating Radiance Fields regarding an object existing in a target space based on a plurality of captured images (hereinafter referred to as "multi-viewpoint images") obtained by imaging from a plurality of different viewpoints with known camera parameters. Hereinafter, the target space for estimating the radiance field will be described as a "scene". Further, there is a technique for generating an image (hereinafter referred to as a "virtual viewpoint image") corresponding to the appearance when an object is viewed from an arbitrary virtual viewpoint (hereinafter referred to as a "virtual viewpoint"). Non-Patent Document 1 discloses a technique for estimating a radiance field representing the color and volume density of an object with respect to position and direction in a scene by deep learning using multi-viewpoint images as teachers. Further, Non-Patent Document 1 discloses a technique for determining the pixel value of a virtual viewpoint image by integrating colors weighted by volume density along a ray starting from the position of an arbitrary viewpoint based on the estimated radiance field.

[0003] In addition, there is a technique for generating a virtual viewpoint image in which editing is applied to an object image, such as paying attention to some of the plurality of objects included in a scene and including only the image corresponding to the object of interest in the virtual viewpoint image. Hereinafter, one or more objects of interest among the plurality of objects included in the scene will be described as "objects of interest". Non-Patent Document 2 discloses a technique for estimating a radiance field specialized for an object of interest based on a multi-viewpoint image and a mask image that masks regions other than the image region corresponding to the object of interest in each captured image, applying the technique disclosed in Non-Patent Document 1. Hereinafter, the image region corresponding to an object in a captured image will be referred to as an object region, and in particular, the image region corresponding to an object of interest in a captured image will be referred to as an object-of-interest region.

[0004] Specifically, in the technique disclosed in Non-Patent Document 2 (hereinafter referred to as the "conventional technique"), first, a radiance field related to each object of interest and a radiance field related to the entire scene including all objects existing in the scene are estimated. Subsequently, using the estimated radiance field related to each object of interest and the radiance field related to the entire scene, a more accurate radiance field related to each object of interest is estimated. More specifically, in the conventional technique, by deep learning using a multi-viewpoint image in which regions other than the object-of-interest region are masked using the above-described mask image as a teacher, a radiance field representing the color and volume density related only to the object of interest is estimated. Also, by deep learning using the multi-viewpoint image before masking as a teacher, a radiance field representing the color and volume density related to the entire scene is estimated. Furthermore, by using these estimated radiance fields to identify regions where the object of interest is blocked by other objects, a more accurate radiance field related to the object of interest is estimated.

[0005] By using the conventional technique, radiance fields related to a plurality of objects existing in the scene can be individually represented. Also, by changing the combination of radiance fields used for generating the virtual viewpoint image, various edits can be made to the objects existing in the scene. For example, by using only the radiance field related to a certain object of interest, a virtual viewpoint image including only the image of the object of interest can be generated.

Prior Art Documents

Non-Patent Documents

[0006]

Non-Patent Document 1

[0007] In the prior art, in order to accurately estimate the radiance field for each object of interest, it is necessary to estimate both the radiance field for each object of interest and the radiance field for the entire scene. Therefore, this technology has the problem of requiring a huge amount of computation or memory.

[0008] Therefore, an object of the present disclosure is to provide a technology capable of suppressing the amount of computation or memory usage when estimating a high-precision radiance field for an object of interest as compared with the prior art. [Means for Solving the Problems]

[0009] The information processing apparatus according to the present disclosure includes imaging data acquisition means for acquiring data of a plurality of captured images obtained by capturing one or more objects existing in a predetermined space from a plurality of viewpoints, and camera parameters corresponding to each of the plurality of viewpoints at the time of imaging, and for each of one or more target objects among the one or more objects, likelihood acquisition means for acquiring, as likelihood values corresponding to each pixel in each of the plurality of captured images, information indicating the likelihood that the image formed on each pixel in each of the plurality of captured images is the target object, and estimation means for estimating information about the space, including color information corresponding to each position in the space and likelihood values for each target object at each position in the space, based on the data of the plurality of captured images, the camera parameters corresponding to each of the plurality of viewpoints, and the likelihood values corresponding to each pixel in each of the plurality of captured images for each target object.

Effect of the Invention

[0010] According to the present disclosure, it is possible to suppress the amount of calculation or the amount of memory used when estimating a high-precision radiance field regarding a target object.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Mode for Carrying Out the Invention

[0012] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. Note that the following embodiments do not limit the means for solving the problems according to the present disclosure, and not all combinations of the features described in the following embodiments are essential for the means for solving the problems according to the present disclosure. The same components will be described with the same reference numerals.

[0013] [Embodiment 1] In Embodiment 1, based on multi-viewpoint images obtained by imaging from a plurality of different viewpoints with known camera parameters and likelihood maps corresponding to each imaging image constituting the multi-viewpoint images, a function F of the following formula (1) Θ is used to estimate information about the space modeled by

[0014] F θ :(x,y,z,θ,φ)→(R,G,B,σ,L1,L2,…,L K ) … Formula (1) Here, (x, y, z) are coordinates indicating a position in the target space (scene), and (θ, φ) are parameters indicating a direction in the scene. (R, G, B) are values indicating the color of an object determined by the position and direction in the scene (hereinafter referred to as "color values"), where R represents the value of red, G represents the value of green, and B represents the value of blue. σ represents the volume density of the object determined by the position in the scene, and L k(k = 1, 2, …, K) represents the likelihood value (hereinafter referred to as the "likelihood value") for each of the K objects of interest determined by the position in the scene.

[0015] The likelihood value L according to Embodiment 1 k is an index indicating how much an object, if it exists, resembles the k-th object of interest. The function F formulated in Equation (1) Θ is a function that outputs color values, volume density, and likelihood values for each object of interest with respect to the position and orientation in the scene. Hereinafter, the function F Θ is referred to as the radiance field for the combined information of the color value and volume density in the scene represented by it, and similarly, the information of the likelihood value in the scene is referred to as the likelihood field. By visualizing the objects existing in the scene using the estimated radiance field and likelihood field, for example, a virtual viewpoint image including only the image of a specific object can be generated.

[0016] <Hardware Configuration> FIG. 1 is a block diagram showing an example of the hardware configuration of the information processing apparatus 100 according to Embodiment 1. The information processing apparatus 100 includes, as a hardware configuration, a CPU 101, a RAM 102, a ROM 103, a serial I / F (interface) 104, a VC (video card) 105, and a general-purpose I / F 106. Each part included in the information processing apparatus 100 as a hardware configuration is communicably connected to each other via a system bus 107. The CPU 101 uses the RAM 102 as a work memory to execute an OS (operating system) and various programs stored in the ROM 103 or a storage device 111 or the like. The CPU 101 controls the entire information processing apparatus 100 via the system bus 107 by executing various programs. Note that the processing of each step shown in the flowchart described later is realized by the program code stored in the ROM 103 or the storage device 111 or the like being expanded into the RAM 102 and the CPU 101 executing this.

[0017] The serial I / F 104 is an interface configured by Serial ATA or the like, and the information processing apparatus 100 and the storage device 111 are connected via the serial bus 108. The storage device 111 is a large-capacity storage device such as an HDD (Hard Disk Drive) or an SSD (Solid State Drive). In the first embodiment, the storage device 111 is described as an external device of the information processing apparatus 100, but the information processing apparatus 100 may include the storage device 111 inside. The VC 105 receives a control signal from the CPU 101 and outputs a signal related to the display image to the display device 112 via the serial bus 109. The display device 112 is configured by a liquid crystal display or the like, and displays the display image based on the signal related to the display image output by the information processing apparatus 100. The general-purpose I / F 106 is connected to an input device 113 such as a mouse or a keyboard via the serial bus 110 and receives an input signal from the input device 113.

[0018] The CPU 101 displays a GUI (Graphical User Interface) provided by a program on the display device 112 via the VC 105 and receives an input signal indicating an instruction from the user obtained via the input device 113. The information processing apparatus 100 is realized by, for example, a desktop PC (Personal Computer). The information processing apparatus 100 may also be realized by a notebook PC integrated with the display device 112, or a tablet PC or the like. Further, the storage device 111 can also be realized by a medium (portable storage medium), a drive such as a disk drive for accessing the medium, or a reader such as a memory card reader. As the medium, an FD (Flexible Disk), a CD-ROM, a DVD, a USB memory, an MO, or a flash memory or the like can be used.

[0019] <Logical Configuration> FIG. 2 is a block diagram showing an example of the logical configuration of the information processing apparatus 100 according to Embodiment 1. The information processing apparatus 100 includes, as a logical configuration, an imaging data acquisition unit 200, a viewpoint acquisition unit 201, a likelihood acquisition unit 202, an estimation unit 203, an image generation unit 204, and an output unit 205. Each unit included in the information processing apparatus 100 as a logical configuration is realized by the CPU 101 executing a program stored in the ROM 103 or the like using the RAM 102 as a work memory. Note that not all of the processes shown below necessarily have to be executed by the CPU 101, and the information processing apparatus 100 may be configured such that part or all of the processes are executed by one or a plurality of processing circuits other than the CPU 101.

[0020] The imaging data acquisition unit 200 acquires a plurality of imaging image (multi-viewpoint image) data obtained by imaging an object existing in a predetermined scene from positions of various viewpoints based on an instruction from a user input via the input device 113. Hereinafter, the imaging image data acquired by the imaging data acquisition unit 200 will be described as being image data in the RGB image format. The imaging data acquisition unit 200 may directly acquire the imaging image data output from the imaging device from the imaging device, or may acquire the imaging image data by reading the imaging image data from a storage device 111 or the like in which the imaging image data is stored in advance. The multi-viewpoint image data acquired by the imaging data acquisition unit 200 is transmitted to the estimation unit 203.

[0021] FIG. 3 is a diagram showing an example of the arrangement of the objects 301 and 302 and the imaging devices 303 to 305 according to Embodiment 1. In FIG. 3, as an example, a spherical object 301 and a cubic object 302 are arranged as objects existing in a predetermined scene 300, and an example in which a plurality (three) of imaging devices 303 to 305 are arranged around them is shown.

[0022] FIG. 4 is a diagram showing an example of captured images 410, 420, and 430 obtained by imaging with each of imaging devices 303 to 305. Specifically, FIG. 4(a) shows an example of the captured image 410 obtained by imaging with the imaging device 303. FIG. 4(b) shows an example of the captured image 420 obtained by imaging with the imaging device 304. Further, FIG. 4(c) shows an example of the captured image 430 obtained by imaging with the imaging device 305. The captured images 410, 420, and 430 include images 411, 421, and 431 of the spherical object 301 and images 412, 422, and 432 of the cubic object 302.

[0023] In addition, the imaging data acquisition unit 200 acquires the camera parameters of the imaging device that captured each of the captured images constituting the multi-viewpoint image. Hereinafter, the camera parameters acquired by the imaging data acquisition unit 200 will be described as including the internal parameters, external parameters, and distortion parameters of the imaging device. The internal parameters are parameters representing the position of the principal point of the imaging device and the focal length of the lens of the imaging device. The external parameters are parameters representing the position of the imaging device and the optical axis direction of the imaging device, that is, the posture of the imaging device. The distortion parameters are parameters representing the distortion of the lens of the imaging device. The imaging data acquisition unit 200 may acquire the camera parameters by requesting the camera parameters held by each imaging device from the imaging device, or may acquire the camera parameters by reading the camera parameters from a storage device 111 or the like in which the camera parameters are stored in advance. The camera parameters of each imaging device acquired by the imaging data acquisition unit 200 are transmitted to the estimation unit 203.

[0024] The viewpoint acquisition unit 201 acquires information regarding a virtual viewpoint (hereinafter referred to as "virtual viewpoint information"). The virtual viewpoint information includes at least camera parameters regarding the virtual viewpoint. The camera parameters regarding the virtual viewpoint include information indicating the position of the virtual viewpoint and information indicating the direction of the line of sight at the virtual viewpoint. Hereinafter, in order to distinguish between the camera parameters of the imaging device and the camera parameters regarding the virtual viewpoint, the camera parameters of the imaging device are simply referred to as "camera parameters", and the camera parameters regarding the virtual viewpoint are referred to as "virtual camera parameters" for explanation. In addition to the virtual camera parameters, the virtual viewpoint information may include pixel number information indicating the number of pixels of the virtual viewpoint image generated by the image generation unit 204, and object information such as an identification number that can uniquely identify an object to be included as an image in the virtual viewpoint image. Further, the virtual viewpoint information may include information indicating the viewing angle from the virtual viewpoint. The virtual viewpoint information is acquired, for example, based on an instruction from the user input via the input device 113. The virtual viewpoint information acquired by the viewpoint acquisition unit 201 is transmitted to the image generation unit 204.

[0025] The likelihood acquisition unit 202 acquires data of a likelihood map for each object of interest (hereinafter referred to as "likelihood map data") based on an instruction from the user input via the input device 113. In Embodiment 1, the likelihood map is an image having, as a pixel value, the likelihood value (likelihood value) of the image formed on each pixel of the captured image corresponding to the object of interest. Hereinafter, the likelihood value is described as being able to take a real number of 0 or more and 1 or less. The likelihood map can be generated by applying a known segmentation technique to the captured image. The method for generating the likelihood map is not limited to this. For example, the likelihood map may be created by filling, with a pixel value corresponding to the likelihood value, the area corresponding to the object of interest in the captured image by manual input by the user. The likelihood map data acquired by the likelihood acquisition unit 202 is transmitted to the estimation unit 203.

[0026] FIG. 5 is a diagram showing an example of likelihood maps 510, 520, 530, 540, 550, 560 according to Embodiment 1. Specifically, FIGS. 5(a) and 5(b) show an example of likelihood maps 510, 520 corresponding to the captured image 410 shown in FIG. 4(a). FIGS. 5(c) and 5(d) show an example of likelihood maps 530, 540 corresponding to the captured image 420 shown in FIG. 4(b). FIGS. 5(e) and 5(f) show an example of likelihood maps 550, 560 corresponding to the captured image 430 shown in FIG. 4(c). In FIG. 5, as an example, pixels with a likelihood of 0 are represented in black, and pixels with a likelihood of 1 are represented in white. More specifically, likelihood maps 510, 520, 530 indicate the likelihood of the spherical object 301 at each pixel of the captured images 410, 420, 430. Also, likelihood maps 520, 540, 560 indicate the likelihood of the cubic object 302 at each pixel of the captured images 410, 420, 430.

[0027] The estimation unit 203 estimates a radiance field and a likelihood field based on the multi-viewpoint image data and camera parameters acquired by the imaging data acquisition unit 200, and the likelihood map data for each target object corresponding to each captured image acquired by the likelihood acquisition unit 202. Details of the processing of the estimation unit 203 will be described later. Information indicating the radiance field and the likelihood field estimated by the estimation unit 203 is transmitted to the image generation unit 204. The image generation unit 204 generates a virtual viewpoint image using the radiance field and the likelihood field estimated by the estimation unit 203 based on the virtual viewpoint information acquired by the viewpoint acquisition unit 201. Details of the processing of the image generation unit 204 will be described later. Data of the virtual viewpoint image generated by the image generation unit 204 is transmitted to the output unit 205.

[0028] The output unit 205 outputs the virtual viewpoint image generated by the image generation unit 204. Specifically, for example, the output unit 205 generates a display image including the virtual viewpoint image, outputs a signal related to the display image to the display device 112, and causes the display device 112 to display the display image. The output destination of the virtual viewpoint image is not limited to the display device 112. For example, the output unit 205 may output the data of the virtual viewpoint image to the storage device 111 and store the data in the storage device 111, or may output the data to another external device different from the information processing apparatus 100.

[0029] <Processing flow> FIG. 6 is a flowchart showing an example of a processing flow in the information processing apparatus 100 according to Embodiment 1. Note that “S” attached to the beginning of a reference numeral means a step (process). First, in S601, the imaging data acquisition unit 200 acquires multi-viewpoint image data and camera parameters of each captured image based on an instruction from the user. Next, in S602, the likelihood acquisition unit 202 acquires likelihood map data for each target object corresponding to each captured image data constituting the multi-viewpoint image data based on an instruction from the user.

[0030] FIG. 7 is a diagram showing an example of GUIs 700, 710, and 720 displayed on the display device 112. The instructions from the user in S601 and S602 are received via the GUI 700 shown as an example in FIG. 7(a). The GUI 700 has data path setting fields 701 to 703 and a “Run” button 704. The data path setting fields 701, 702, and 703 are fields that receive input of data paths indicating the locations of files including multi-viewpoint image data, camera parameter data, and likelihood map data as data, respectively. The “Run” button 704 is a button that receives an instruction to execute the estimation process described later. When the button 704 is pressed by the user, the information processing apparatus 100 executes the process of S603 after executing the processes of S601 and S602. FIG. 7(b) will be described later. Also, FIG. 7(c) will be described in Embodiment 2.

[0031] In S603, the estimation unit 203 executes an estimation process of estimating a radiance field and a likelihood field based on the multi-viewpoint image data and camera parameters acquired in S601 and the likelihood map acquired in S602. Specifically, for example, the function F Θ configured in advance by an MLP (Multi-layer perceptron), and the estimation unit 203 estimates the radiance field and the likelihood field by performing learning of the MLP by deep learning. Hereinafter, the function F Θ configured by the MLP is referred to as an estimation MLP. When the function F Θ is configured as the estimation MLP, each of the radiance field and the likelihood field is expressed as a parameter of the estimation MLP, that is, a weight coefficient related to each node constituting the estimation MLP.

[0032] In the learning of the estimation MLP, the predicted RGB value, which is a predicted value of the pixel value calculated based on the output of the function F Θ , and the predicted likelihood value, which is a predicted value of the likelihood value corresponding to the pixel, are approximately the same as the pixel values of the captured image and the likelihood map, and the parameters of the estimation MLP are optimized. Specifically, the learning of the estimation MLP is performed by the error backpropagation method using the squared Euclidean distance between the teacher signal C GT (r) shown in the following equation (2) and the predicted signal C Θ (r) calculated from the output value of the function F pred by the following equations (3) to (6) as a loss.

[0033] JPEG2025093503000002.jpg10373 Here, r is a ray determined based on the camera parameters of the imaging device. Also, I R (r), I G (r), and I B (r) are the pixel values of the captured image corresponding to the ray r, and are the pixel values corresponding to the components of R (red), G (green), and B (blue) in order. Also, I Lk(r)(k = 1, 2, …, K) is the pixel value of the likelihood map regarding the k-th object of interest corresponding to the ray r. FIG. 8 is a diagram showing an example of the ray r according to Embodiment 1. FIG. 8 schematically represents the positional relationship between the ray r, a predetermined scene 300 in which objects 301 and 302 are arranged, the position 801 of the imaging device, a plane 802 corresponding to the captured image, and a pixel 803 in the captured image corresponding to the ray r.

[0034] Equations (3) to (5) formulate a process corresponding to known volume rendering. i is the index of the sampling point on the ray r, and N is the number of sampling points. Also, T i is the cumulative transmittance from the position 801 of the imaging device to the sampling point, and α i is the opacity of the sampling point. Also, σ i is the volume density output by the function F Θ for the sampling point, and δ j is the distance from the j-th sampling point to the (j + 1)-th sampling point. Also, c i is a signal composed of the RGB value and the likelihood value output by the function F Θ for the sampling point, and R i , G i , and B i are, in order, the values corresponding to the R, G, and B components output from the function F Θ . Also, L k,i (k = 1, 2, …, K) is the likelihood value regarding the k-th object of interest output from the function F Θ . The predicted signal C pred (r) is a signal composed of the weighted sum of the color value and the likelihood value at the sampling point on the ray r, with the cumulative transmittance and the opacity as weighting factors.

[0035] FIG. 9 is a diagram showing an example of the radiance field and the likelihood field obtained by the learning of the estimation unit 203 according to Embodiment 1. Specifically, FIGS. 9(a) and (b) are, in order, along the rays r and r' shown in FIG. 8, the function F ΘIt is a graph plotting the output value. Here, the light ray r is a light ray passing through both the spherical object 301 and the cubic object 302, and the light ray r' is a light ray passing through only the spherical object 301. Hereinafter, the spherical object 301 will be described as having at least a red color on its surface, and the cubic object 302 will be described as having at least a green color on its surface. Further, hereinafter, the spherical object 301 will be described as the first object of interest, and the cubic object 302 will be described as the second object of interest.

[0036] In FIG. 9, the value of R becomes large in the vicinity where the light rays r and r' intersect with the surface of the spherical object 301 whose surface color is red, and becomes small in the vicinity where the light rays r and r' intersect with the surface of the cubic object 302 whose surface color is green. Also, in FIG. 9, the value of G becomes small in the vicinity where the light rays r and r' intersect with the surface of the spherical object 301, and becomes large in the vicinity where the light rays r and r' intersect with the surface of the cubic object 302. Further, in FIG. 9, the value of B becomes small overall, including the vicinity where the light rays r and r' intersect with the surfaces of the spherical object 301 and the cubic object 302.

[0037] Also, in FIG. 9, the value of σ indicating the volume density of the object becomes large in the vicinity where the light rays r and r' intersect with the surface of the spherical object 301 or the cubic object 302. Also, in FIG. 9, the value of L1 indicating the likelihood of the first object of interest becomes large in the vicinity where the light rays r and r intersect with the surface of the spherical object 301 which is the first object of interest. Also, in FIG. 9, the value of L2 indicating the likelihood of the second object of interest becomes large in the vicinity where the light ray r intersects with the surface of the cubic object 302 which is the second object of interest.

[0038] Note that it is assumed that a plurality of different objects of interest do not overlap and exist at the same position in scene 300. Also, in the learning in the estimation unit 203, a constraint condition may be provided such that the learning is performed so that the sum of the K likelihood values L k,i is 1 or less. Also, the function F Θ may be any function that outputs a color value, a volume density, and a likelihood value for each object of interest with respect to the position and direction in scene 300, and is not limited to being configured by an MLP.

[0039] After S603, in S604, the viewpoint acquisition unit 201 acquires virtual viewpoint information based on an instruction from the user. Next, in S605, the image generation unit 204 executes an image generation process of generating a virtual viewpoint image using the virtual viewpoint information acquired in S604 and the radiance field and likelihood field estimated in S603. The instruction from the user in S604 is received via the GUI 710 displayed on the display device 112, shown as an example in FIG. 7(b).

[0040] The GUI 710 has a virtual camera parameter setting field 711, an image size setting field 712, an object setting field 713, a "Render" button 714, and a display area 715. The virtual camera parameter setting field 711 is a field that accepts an input of a data path indicating the location of a file containing virtual camera parameters used for generating a virtual viewpoint image as data. The image size setting field 712 is a field that accepts an input of the number of pixels in the horizontal and vertical directions of the virtual viewpoint image to be generated. The object setting field 713 is a field that accepts an input of an identification number or the like corresponding to an object of interest to be included as an image in the virtual viewpoint image. The "Render" button 714 is a button that accepts an instruction to execute the image generation process. When the "Render" button 714 is pressed by the user, the image generation unit 204 generates a virtual viewpoint image based on the input values input in the virtual camera parameter setting field 711, the image size setting field 712, and the object setting field 713. The display area 715 is an area where the virtual viewpoint image generated by the image generation unit 204 is displayed.

[0041] The image generation unit 204 generates a virtual viewpoint image by calculating the pixel value C k (r) of the virtual viewpoint image using, for example, the following equations (7) to (9).

[0042] JPEG2025093503000003.jpg3881Here, k in equations (7) to (9) is the identification number of the object of interest to be included as an image in the virtual viewpoint image, and is the identification number input to the object setting field 713. In equations (8) and (9), the image generation unit 204 uses the likelihood value L k,i related to the object of interest to be included as an image in the virtual viewpoint image to weight the volume density σ i . As a result, the image generation unit 204 can pseudo-lower the volume density for objects of interest with a small likelihood value L k,i , that is, objects of interest not to be included as images in the virtual viewpoint image. By such processing, a virtual viewpoint image in which objects of interest other than the object of interest to be included as an image in the virtual viewpoint image are transparent is generated, and a virtual viewpoint image including only the image of the desired object of interest is obtained.

[0043] FIG. 10 is a diagram showing an example of virtual viewpoint images 1000 and 1010 generated by the image generation unit 204. Specifically, FIG. 10(a) shows an example of the virtual viewpoint image 1000 when the virtual camera parameters are the same as the camera parameters of the imaging device 305 shown in FIG. 3 and the identification number assigned to the spherical object 301 is input to the object setting field 713. Hereinafter, it will be described on the assumption that 1 is pre-assigned to the identification number of the spherical object 301 and 2 is pre-assigned to the identification number of the cubic object 302. That is, FIG. 10(a) is an example of the virtual viewpoint image 1000 generated when 1, which is the identification number assigned to the spherical object 301, is input to the object setting field 713. The virtual viewpoint image 1000 does not include the image of the cubic object 302 and includes only the image of the spherical object 301. The virtual viewpoint image 1010 shown in FIG. 10(b) will be described in Embodiment 2.

[0044] After S605, at S606, the output unit 205 outputs the virtual viewpoint image generated at S605. For example, the output unit 205 outputs so that the virtual viewpoint image generated at S605 is displayed in the display area 715 in the GUI 710. After S606, the information processing apparatus 100 ends the processing of the flowchart shown in FIG. 6. In expressions (8) and (9), the likelihood value L k,i is directly multiplied by the volume density σ i to perform weighting of the volume density σ i , but the weighting method is not limited to this. For example, based on a predetermined threshold value, the likelihood value L k,i is binarized to either 0 or 1, and the binarized likelihood value L k,i is multiplied by the volume density σ i to perform weighting.

[0045] As described above, for the radiance field of the entire object and the likelihood field for each target object, the information processing apparatus 100 is configured to model the scene using a single function F Θ and learn this. According to the information processing apparatus 100 configured in this way, the radiance field for each target object can be simultaneously learned and estimated. As a result, it is possible to suppress the amount of calculation and the amount of memory used for learning, that is, for estimating the radiance field for each target object, and it is possible to generate a virtual viewpoint image in which only the image of a specific object in the scene is extracted.

[0046] In addition, in the first embodiment, the captured image has been described as an image in the RGB image format. However, the captured image may be represented in other formats such as a grayscale image, an XYZ image, or a YUV image. Further, in the first embodiment, the color of the object has been described as being determined by the position and the direction. However, it may be determined only by the position without depending on the direction. Also, in the description of S604, an example has been described in which the virtual viewpoint image is generated such that the image of the k-th object of interest specified by the user is more transparent as the likelihood value of the k-th object of interest is lower. However, the present invention is not limited to this. For example, the virtual viewpoint image may be generated such that the image of the k-th object of interest is more transparent as the likelihood value of the k-th object of interest is higher. In this case, for example, a virtual viewpoint image in which the image of the k-th object of interest is removed is generated.

[0047] [Second Embodiment] In the first embodiment, an example has been described in which the radiance field and the likelihood field modeled by the function F of Expression (1) are estimated. In the second embodiment, a radiance field and a likelihood field that do not include volume density, which are modeled by a function F' shown in the following Expression (10), are estimated, and an aspect of calculating the volume density based on the likelihood values of the respective objects of interest that have been estimated will be described. Θ In the first embodiment, an example has been described in which the radiance field and the likelihood field modeled by the function F of Expression (1) are estimated. In the second embodiment, a radiance field and a likelihood field that do not include volume density, which are modeled by a function F' shown in the following Expression (10), are estimated, and an aspect of calculating the volume density based on the likelihood values of the respective objects of interest that have been estimated will be described. Θ An aspect of calculating the volume density based on the likelihood values of the respective objects of interest that have been estimated will be described.

[0048] F' θ :(x,y,z,θ,φ)→(R,G,B,L1,L2,…,L K ) …Expression (10) The function F' formulated by Expression (10) Θ is a function that outputs color values and likelihood values regarding each object of interest with respect to the position and the direction in the scene, and is different from the function F according to the first embodiment in that it does not output volume density. In the description of the second embodiment, the color information in the scene represented by the function F' θ is referred to as a radiance field. Θ is referred to as a radiance field.

[0049] The hardware configuration, logical configuration of the information processing apparatus 100 according to Embodiment 2 (hereinafter simply referred to as "information processing apparatus 100"), and the overall flow of processing in the information processing apparatus 100 are the same as those of the information processing apparatus 100 according to Embodiment 1. However, the processing of the information processing apparatus 100 differs from the processing of the information processing apparatus 100 according to Embodiment 1 in the estimation processing in S603 and the image generation processing in S605. Hereinafter, mainly, the processing that differs between Embodiment 2 and Embodiment 1 will be described. In the following, the same components as those in Embodiment 1 will be described with the same reference numerals.

[0050] <Estimation Processing in the Estimation Unit according to Embodiment 2> The estimation unit 203 according to Embodiment 2 (hereinafter simply referred to as "estimation unit 203") estimates a radiance field and a likelihood field. Specifically, the estimation unit 203 estimates a radiance field without volume density and a likelihood field based on multi-viewpoint image data, camera parameters corresponding to each imaging device, and likelihood map data for each target object corresponding to each captured image. When estimating the radiance field, the estimation unit 203 uses the function F' Θ It is considered that there is a high possibility that an object exists at a position where the likelihood value output by is large, and the sum of the likelihood values of each target object is used as the volume density. For example, the function F' of Equation (10) Θ is composed of an MLP. The estimation unit 203 uses the mean squared Euclidean distance between the teacher signal C GT (r) shown in Equation (2) and the predicted signal C pred '(r) calculated by the following Equations (11) to (16) as a loss, and performs learning of the MLP by the error backpropagation method.

[0051] JPEG2025093503000004.jpg10086 Here, T' i is the cumulative transmittance from the position of the imaging device to the sampling point, and α' i is the opacity of the sampling point. Also, σ' i is the volume density calculated based on the likelihood value L Θ output by the function F' for the sampling point, and L' k,i ​k,i is the likelihood value L k,i normalized by the volume density σ’ i . Also, c’ i is the RGB value output by the function F’ Θ for the sampling point, and the likelihood value L’ i after normalization by the volume density σ’ k,i . In Equation (16), as an example, the sum of the likelihood values L k,i is used as the volume density σ’ i , but the volume density σ’ i may be a large value if there is a large likelihood value among any one of the likelihood values of the K target objects. For example, the maximum value of the K likelihood values L k,i (k = 1, 2,..., K) for the sampling point may be used as the volume density σ’ i .

[0052] <Image generation process in the image generation unit according to Embodiment 2> The image generation unit 204 according to Embodiment 2 (hereinafter simply referred to as "image generation unit 204") generates a virtual viewpoint image in which only the image of a specific object in the scene is extracted by performing the same processing as the image generation unit 204 according to Embodiment 1. Specifically, the image generation unit 204 may obtain the pixel value C k,i shown in Equation (8) and the volume density σ i shown in Equation (9), and replace them with the likelihood value L’ k,i shown in Equation (15) and the volume density σ’ i shown in Equation (16), and then obtain the pixel value C k (r) shown in Equation (7).

[0053] The image generation unit 204 can also generate a virtual viewpoint image with the opacity of the image corresponding to the object of interest changed for each object of interest by using the likelihood field estimated by the estimation unit 203. In this case, for example, the information processing apparatus 100 receives an instruction from the user via the GUI 720 displayed on the display device 112, shown as an example in FIG. 7(c). The GUI 710 has an opacity setting column 721 in addition to the virtual camera parameter setting column 711, the image size setting column 712, the "Render" button 714, and the display area 715. The opacity setting column 721 is a column for receiving the input of coefficients β k (k = 1, 2,..., K) of the opacity of the image of each object of interest in the virtual viewpoint image to be generated. When the "Render" button 714 is pressed by the user, the image generation unit 204 generates a virtual viewpoint image based on the input values entered in the virtual camera parameter setting column 711, the image size setting column 712, and the opacity setting column 721. Specifically, for example, the image generation unit 204 determines the pixel value C synth (r) of the virtual viewpoint image to be generated using the following equations (17) to (19).

[0054] JPEG2025093503000005.jpg49101 FIG. 10(b) shows an example of the virtual viewpoint image 1010 generated by the image generation unit 204 when the coefficient β k relating to the opacity of the image of each object of interest in the virtual viewpoint image to be generated is set. Specifically, FIG. 10(b) is a virtual viewpoint image 1010 when the virtual camera parameters are the same as the camera parameters of the imaging device 305 shown in FIG. 3. More specifically, FIG. 10(b) is a virtual viewpoint image 1010 when the coefficient β1 for the spherical object 301 is 1.00 and the coefficient β2 for the cubic object 302 is 0.25. Since the coefficient β2 for the cubic object 302 is 0.25, the image of the cubic object 302 included in the virtual viewpoint image 1010 is represented in a semi-transparent state.

[0055] As described above, in Embodiment 2, the information processing apparatus 100 is configured to model a scene using a function F' whose output has a smaller number of dimensions than the function F according to Embodiment 1, and to learn this. According to the information processing apparatus 100 configured in this way, it is possible to simultaneously learn and estimate the radiance field for each object of interest. As a result, it is possible to suppress the amount of computation and the amount of memory used for learning, that is, for estimating the radiance field for each object of interest, and to generate a virtual viewpoint image obtained by editing a specific object in the scene. Θ than the function F Θ to model a scene, and to learn this. According to the information processing apparatus 100 configured in this way, it is possible to simultaneously learn and estimate the radiance field for each object of interest. As a result, it is possible to suppress the amount of computation and the amount of memory used for learning, that is, for estimating the radiance field for each object of interest, and to generate a virtual viewpoint image obtained by editing a specific object in the scene.

[0056] [Other Embodiments] The present disclosure can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or apparatus via a network or a storage medium, and causing one or more processors in a computer of the system or apparatus to read and execute the program. Further, it can also be realized by a circuit (for example, an ASIC) that realizes one or more functions.

[0057] Note that within the scope of the present disclosure, any combination of the embodiments, any modification of any component of each embodiment, or any omission of any component in each embodiment is possible.

[0058] [Configuration of the Present Disclosure] The present disclosure includes the following configurations, methods, and programs.

[0059] <Configuration 1> Imaging data acquisition means for acquiring data of a plurality of captured images obtained by imaging one or more objects existing in a predetermined space from a plurality of viewpoints, and camera parameters corresponding to each of the plurality of viewpoints at the time of imaging; For one or more target objects among the one or more objects, likelihood acquisition means for acquiring, for each target object, information indicating the likelihood that the image formed on each pixel in each of the plurality of captured images is the target object as a likelihood value corresponding to each pixel in each of the plurality of captured images; Estimation means for estimating information regarding the space, including color information corresponding to each position in the space and likelihood values for each target object at each position in the space, based on the data of the plurality of captured images, camera parameters corresponding to each of the plurality of viewpoints, and likelihood values corresponding to each pixel in each of the plurality of captured images for each target object; An information processing apparatus characterized by comprising:

[0060] <Configuration 2> The information regarding the space is information indicating a function that outputs color information corresponding to each position in the space and likelihood values for each target object at each position in the space; The information processing apparatus according to Configuration 1, characterized by:

[0061] <Configuration 3> The color information corresponding to each position in the space is color information for a combination of a position and a direction in the space; The information processing apparatus according to Configuration 1 or 2, characterized by:

[0062] <Configuration 4> The estimation means estimates information regarding the space further including a volume density corresponding to each position in the space; The information processing apparatus according to any one of Configurations 1 to 3, characterized by:

[0063] <Configuration 5> The information regarding the space is information indicating a function that outputs color information corresponding to each position in the space, likelihood values for each target object at each position in the space, and a volume density corresponding to each position in the space; The information processing apparatus according to Configuration 4, characterized in that...

[0064] <Configuration 6> Based on the likelihood value for each target object at each position in the space obtained by estimation, the estimation means calculates the volume density corresponding to each position in the space. The information processing apparatus according to any one of Configurations 1 to 3, characterized in that...

[0065] <Configuration 7> The estimation means... Using the data of the plurality of captured images, the camera parameters corresponding to each of the plurality of viewpoints, the color information corresponding to each position in the space obtained by estimation, and the volume density obtained by at least one of estimation and calculation, calculates an error related to color. Using the likelihood value corresponding to each pixel in each of the plurality of captured images for each target object, the likelihood value for each target object at each position in the space obtained by estimation, and the volume density obtained by at least one of estimation and calculation, calculates an error related to the likelihood value. Estimates information about the space by minimizing the error related to color and the error related to the likelihood value obtained by calculation. The information processing apparatus according to any one of Configurations 4 to 6, characterized in that...

[0066] <Configuration 8> Further comprising image generation means for generating an image for visualizing an object in the space based on the information about the space. The information processing apparatus according to any one of Configurations 1 to 7, characterized in that...

[0067] <Configuration 9> Further comprising output means for displaying and outputting the generated image to a display device. The information processing apparatus according to Configuration 8, characterized in that...

[0068] <Configuration 10> The image generation means generates the image such that the color of the image corresponding to the object of interest in the image is more transparent as the likelihood value of the object of interest in the space is smaller. The information processing apparatus according to configuration 8 or 9, characterized by the above.

[0069] <Method> An imaging data acquisition step of acquiring data of a plurality of captured images obtained by imaging one or more objects existing in a predetermined space from a plurality of viewpoints, and camera parameters corresponding to each of the plurality of viewpoints at the time of imaging; A likelihood acquisition step of acquiring, for each object of interest among the one or more objects, information indicating the likelihood that the image formed on each pixel in each of the plurality of captured images is the object of interest, as a likelihood value corresponding to each pixel in each of the plurality of captured images for each object of interest; An estimation step of estimating information about the space including color information corresponding to each position in the space and the likelihood value for each object of interest at each position in the space, based on the data of the plurality of captured images, the camera parameters corresponding to each of the plurality of viewpoints, and the likelihood values corresponding to each pixel in each of the plurality of captured images for each object of interest; An information processing method characterized by including the above.

[0070] <Program> A program for causing a computer to function as the information processing apparatus according to any one of configurations 1 to 10.

Explanation of Signs

[0071] 100 Information processing apparatus 200 Imaging data acquisition unit 202 Likelihood acquisition unit 203 Estimation unit

Claims

1. Imaging data acquisition means for acquiring data of a plurality of captured images obtained by imaging one or more objects existing in a predetermined space from a plurality of viewpoints, and camera parameters corresponding to each of the plurality of viewpoints at the time of imaging; Likelihood acquisition means for acquiring, for each of one or more target objects among the one or more objects, information indicating the likelihood that the image formed on each pixel in each of the plurality of captured images is the target object, as a likelihood value corresponding to each pixel in each of the plurality of captured images; Estimation means for estimating information about the space, including color information corresponding to each position in the space and likelihood values for each target object at each position in the space, based on the data of the plurality of captured images, the camera parameters corresponding to each of the plurality of viewpoints, and the likelihood values corresponding to each pixel in each of the plurality of captured images for each target object; An information processing apparatus characterized by comprising the above.

2. The information about the space is information indicating a function that outputs color information corresponding to each position in the space and likelihood values for each target object at each position in the space; The information processing apparatus according to claim 1, characterized by the above.

3. The color information corresponding to each position in the space is color information for a combination of a position and a direction in the space; The information processing apparatus according to claim 1, characterized by the above.

4. The estimation means estimates information about the space further including a volume density corresponding to each position in the space; The information processing apparatus according to claim 1, characterized by the above.

5. The information about the space is information indicating a function that outputs color information corresponding to each position in the space, likelihood values for each target object at each position in the space, and a volume density corresponding to each position in the space; The information processing apparatus according to claim 4, characterized by the above.

6. The estimation means calculates a volume density corresponding to each position in the space based on the likelihood values for each target object at each position in the space obtained by estimation; The information processing apparatus according to claim 1, characterized by the above.

7. The estimation means Calculate an error related to color using the data of the plurality of captured images, the camera parameters corresponding to each of the plurality of viewpoints, the color information corresponding to each position in the space obtained by estimation, and the volume density obtained by at least one of estimation and calculation. Calculate an error related to the likelihood value using the likelihood value corresponding to each pixel in each of the plurality of captured images for each of the target objects, the likelihood value for each of the target objects at each position in the space obtained by estimation, and the volume density obtained by at least one of estimation and calculation. Estimate information about the space by minimizing the error related to the color and the error related to the likelihood value obtained by calculation. The information processing apparatus according to claim 4, characterized in that.

8. Further comprising image generation means for generating an image for visualizing an object in the space based on the information about the space. The information processing apparatus according to claim 1, characterized in that.

9. Further comprising output means for displaying and outputting the generated image to a display device. The information processing apparatus according to claim 8, characterized in that.

10. The image generation means generates the image such that the color of the image corresponding to the target object in the image is transmitted more as the likelihood value of the target object in the space is smaller. The information processing apparatus according to claim 8, characterized in that.

11. An imaging data acquisition step of acquiring the data of a plurality of captured images obtained by imaging one or more objects existing in a predetermined space from a plurality of viewpoints, and the camera parameters corresponding to each of the plurality of viewpoints at the time of imaging; A likelihood acquisition step of acquiring, for each of one or more target objects among the one or more objects, information indicating the likelihood that the image formed on each pixel in each of the plurality of captured images is the target object, as a likelihood value corresponding to each pixel in each of the plurality of captured images for each of the target objects. An estimation step of estimating information about the space, including color information corresponding to each position in the space and a likelihood value for each target object at each position in the space, based on data of the plurality of captured images, camera parameters corresponding to each of the plurality of viewpoints, and likelihood values corresponding to each pixel in each of the plurality of captured images for each target object; An information processing method, characterized by including the above.

12. A program for causing a computer to function as the information processing apparatus according to any one of Claims 1 to 10.