Image processing apparatus, image processing method, and program

The image processing apparatus addresses object estimation failures and haze artifacts by setting learning rays based on object regions in captured images, ensuring accurate three-dimensional field estimation and high-precision virtual viewpoint images.

JP2025099272AActive Publication Date: 2025-07-03CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023215803
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-21
Publication Date
2025-07-03
Estimated Expiration
2043-12-21

AI Technical Summary

Technical Problem

Existing image processing techniques, such as NeRF, fail to accurately estimate objects with low occupancy rates in captured images, leading to object disappearance, and objects with high occupancy rates result in artifacts like haze in virtual viewpoint images.

Method used

An image processing apparatus that acquires captured images from multiple viewpoints, identifies object regions, and sets learning rays based on these regions to learn a three-dimensional field, controlling the number of rays used for learning to improve accuracy.

Benefits of technology

This approach enables accurate estimation of three-dimensional fields regardless of object occupancy rates, preventing object disappearance and haze artifacts in virtual viewpoint images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025099272000001_ABST
    Figure 2025099272000001_ABST
Patent Text Reader

Abstract

To estimate a three-dimensional field corresponding to an object, which is used in generating a virtual viewpoint image, with high accuracy.SOLUTION: An image processing apparatus 102 is configured to: obtain data of a plurality of captured images obtained by image capturing from a plurality of viewpoints; obtain an object area corresponding to a representation of an object in each of the plurality of captured images; set a learning ray group that is used for learning of information relating to a three-dimensional field of an image capturing space that is an image capturing target from the plurality of viewpoints, the learning ray group corresponding to pixels of each of the plurality of captured images based on the obtained object area; and perform learning of information relating to the three-dimensional field based on the set learning ray group.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an image processing technique for generating an image corresponding to an image seen from an arbitrary virtual viewpoint using a plurality of captured images obtained by capturing from a plurality of different positions.

Background Art

[0002] There is a technique for generating an image corresponding to an image seen from an arbitrary virtual viewpoint using a plurality of captured images (hereinafter referred to as "multi-viewpoint images") obtained by capturing with imaging devices having known camera parameters, each arranged at a plurality of different positions. Hereinafter, the virtual viewpoint will be referred to as a "virtual viewpoint", and the image corresponding to the image seen from the virtual viewpoint will be referred to as a "virtual viewpoint image" for explanation. Patent Document 1 discloses a technique called NeRF (Neural Radiance Fields) as a technique for generating a virtual viewpoint image corresponding to an arbitrary virtual viewpoint by inputting data of multi-viewpoint images (hereinafter referred to as "multi-viewpoint image data"). The technique called NeRF disclosed in Patent Document 1 consists of a neural network and volume rendering. Specifically, the neural network of NeRF takes as input data of a plurality of captured images (hereinafter referred to as "captured image data") constituting the multi-viewpoint image data and outputs information indicating density and color for arbitrary positions and directions. Further, the volume rendering of NeRF calculates a pixel value by accumulating colors obtained from sampling points on a ray corresponding to a pixel in the virtual viewpoint image according to the density.

[0003] The neural network of NeRF is trained by using the pixel values of multi-viewpoint images as teacher data and adjusting the network parameters to minimize the difference (loss) between the pixel values and the pixel values calculated by NeRF. Generally, in the training of the neural network of NeRF, rays corresponding to pixels in the captured images are randomly sampled, and mini-batch learning is adopted, where this sampling is used as one learning unit (hereinafter referred to as a "mini-batch") and learning is repeated. According to mini-batch learning, the usage amount of VRAM (Video Random Access Memory) can be reduced compared to batch learning that learns all rays at once. Also, since stepwise learning is possible in mini-batch learning, mini-batch learning is considered essential for implementing the learning of neural networks in technologies related to NeRF.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, in the technology disclosed in Patent Document 1 (hereinafter referred to as the "prior art"), when the occupancy rate of an object (hereinafter simply referred to as an "object") that is the target subject with respect to the angle of view of the captured image is small, the estimation of the object may fail. Here, the failure of object estimation means, for example, that part or all of the image corresponding to the object disappears in the virtual viewpoint image. Such a failure of object estimation is caused by the fact that the mini-batch does not contain rays representing the object. On the other hand, when the occupancy rate of the object with respect to the angle of view of the captured image is large, artifacts like haze called floaters may occur around the image corresponding to the object in the virtual viewpoint image.

Means for Solving the Problems

[0006] The image processing apparatus according to the present disclosure includes: an image acquisition unit that acquires data of a plurality of captured images obtained by capturing from a plurality of viewpoints; a region acquisition unit that acquires an object region corresponding to an image of an object in each of the plurality of captured images; a ray setting unit that sets a learning ray group used for learning information about a three-dimensional field of a shooting space that is a shooting target from the plurality of viewpoints, the ray setting unit setting the learning ray group corresponding to each pixel of the plurality of captured images based on the acquired object region; and a learning unit that learns information about the three-dimensional field based on the set learning ray group.

Advantages of the Invention

[0007] According to the present disclosure, regardless of the occupancy rate of an object with respect to the angle of view of a captured image, it is possible to accurately estimate a three-dimensional field corresponding to the object used in generating a virtual viewpoint image.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Modes for Carrying Out the Invention

[0009] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. Note that the following embodiments do not necessarily limit the solution means of the present disclosure. Also, not all combinations of features described in the embodiments are essential for the solution means of the present disclosure. Further, in the following embodiments, a two-dimensional region on an image is simply referred to as a "region", and a three-dimensional region in a shooting space is referred to as a "space" for explanation.

[0010] [Embodiment 1] <Configuration of Image Processing System> FIG. 1 is a diagram showing an example of the configuration of an image processing system according to Embodiment 1. The image processing system includes a plurality of imaging devices 101, an image processing device 102, a user interface (hereinafter referred to as "UI") panel 103, a storage device 104, and a display device 105. The plurality of imaging devices 101 are configured by digital still cameras, digital video cameras, etc., and are arranged at different positions. Each imaging device 101 captures an object 107 existing in the imaging space 106 in synchronization with each other according to predetermined imaging conditions, and outputs the captured image data obtained by the imaging to the image processing device 102. The captured image data obtained by the imaging by the imaging device 101 may be data of a still image, data of a moving image, or data of both a still image and a moving image. Hereinafter, the term "image" will be described as including both "still image" and "moving image" unless otherwise specified.

[0011] The image processing device 102 acquires a plurality of captured image data (multi-viewpoint image data) output from the plurality of imaging devices 101, and performs learning of a three-dimensional field that is a three-dimensional field corresponding to the object 107 in the imaging space based on the acquired multi-viewpoint image data. In addition, the image processing device 102 may control each of the plurality of imaging devices 101. Note that the three-dimensional field in the imaging space to be learned differs depending on the learning content. In Embodiment 1, as an example, the three-dimensional field to be learned will be described as a Radiance Fields.

[0012] The UI panel 103 includes a display device such as a liquid crystal display, and displays a user interface for presenting the imaging conditions in the imaging device 101 and the processing settings of the image processing device 102, etc. to the user on the display device. The UI panel 103 may include an input device such as a touch panel or buttons. In this case, the UI panel 103 receives instructions from the user regarding changes to the above-described imaging conditions and processing settings, etc. The input device may be provided separately from the UI panel 103, such as a mouse or a keyboard.

[0013] The memory device 104 is composed of a hard disk drive or the like, and stores information regarding a three-dimensional field corresponding to the object 107 output by the image processing device 102. The display device 105 is composed of a liquid crystal display or the like, acquires an image signal indicating the three-dimensional field corresponding to the object 107 output from the image processing device 102, and displays an image corresponding to the image signal. Further, the display device 105 may acquire an image signal indicating a virtual viewpoint image output from the image processing device 102 and display a virtual viewpoint image corresponding to the image signal. The imaging space 106 is a three-dimensional space surrounded by a plurality of imaging devices 101 installed in a studio or the like. The frame indicated by the solid line in FIG. 1 shows the contour of the imaging space 106 on the floor surface.

[0014] <Hardware Configuration of Image Processing Device> FIG. 2 is a block diagram showing an example of the hardware configuration of the image processing device 102 according to Embodiment 1. As a hardware configuration, the image processing device 102 includes a CPU 201, a RAM 202, a ROM 203, a storage device 204, a control interface (hereinafter referred to as "I / F") 205, an input I / F 206, an output I / F 207, and a main bus 208. The CPU 201 is a processor that comprehensively controls each part of the image processing device 102. The RAM 202 functions as the main memory and work area of the CPU 201. The ROM 203 stores one or more programs executed by the CPU 201. The storage device 204 is composed of a hard disk drive or the like, and stores application programs executed by the CPU 201 and data used for the processing of the CPU 201.

[0015] The control I / F 205 is connected to each imaging device 101 and is a communication interface for controlling settings of shooting conditions for each imaging device 101, starting and stopping shooting, etc. The input I / F 206 is a communication interface via a serial bus such as SDI (Serial Digital Interface) or HDMI (registered trademark) (High-Definition Multimedia Interface (registered trademark)). The captured image data output from each imaging device 101 is acquired via the input I / F 206. The output I / F 207 is a communication interface via a serial bus such as USB (Universal Serial Bus) or IEEE (Institute of Electrical and Electronics Engineers) 1394. Data or signals indicating the shape of the object 107 are output to the storage device 104 or the display device 105 via the output I / F 207. The main bus 208 is a transmission path that communicably connects the above-described hardware configurations of the image processing device 102 to each other.

[0016] In Embodiment 1, as an example, a form of shooting one or a plurality of objects existing in a studio from a plurality of viewpoints using eight imaging devices 101 installed in the studio will be described. Also, it is assumed that camera parameters such as internal parameters, external parameters, and distortion parameters of each imaging device 101 are stored in advance in the storage device 204. The internal parameters are information indicating coordinates corresponding to the center pixel in the captured image obtained by shooting with the imaging device 101, the focal length of the lens of the imaging device 101, etc. The external parameters are information indicating the position and orientation of the imaging device 101, etc. Note that the camera parameters of each imaging device 101 do not necessarily have to be common to each other. For example, the angle of view of the imaging device 101 may be different from that of other imaging devices 101.

[0017] <Explanation of the cause> Before giving a specific description of Embodiment 1, with reference to FIGS. 3 and 4, the causes of problems occurring in the prior art will be described. FIG. 3 is a diagram for explaining the causes of problems occurring in the prior art, and shows an example of the state of sampling of light rays when performing learning of a three-dimensional field corresponding to a thin rod-shaped object 311 using the prior art. Specifically, FIG. 3(a) shows an example of a captured image 300 obtained by capturing from a certain viewpoint. FIG. 3(b) shows an example of the relationship between the object 311 and the light rays 302, and shows an example of the relationship between the object 311 and the light rays when the imaging space 106 captured by the imaging device 101 is viewed from directly above in the vertical direction. The captured image 300 includes an image 301 corresponding to the object 311. Note that the light rays 302 shown in FIG. 3(a) are not included in the captured image 300 as an image, but explicitly show an example of the corresponding positions in the captured image 300 for each of the light rays 302 shown in FIG. 3(b).

[0018] When the object 311 is, for example, thin and rod-shaped, the occupancy rate of the object 311 in the angle of view of the imaging device 101 becomes low. Therefore, the number of samples of the light rays 302 corresponding to the pixels included in the region of the object 311 decreases. As a result, although it is the space corresponding to the object 311, it may be learned as an empty space, that is, a space where the object 311 does not exist.

[0019] On the one hand, FIG. 4 is a diagram for explaining the cause of problems occurring in the prior art, and is a diagram showing an example of the sampling state of light rays when performing learning of a three-dimensional field corresponding to a large box-shaped object 411 using the prior art. Specifically, FIG. 4(a) shows an example of a captured image 400 obtained by capturing from a certain viewpoint. FIG. 4(b) shows an example of the relationship between the object 411 and the light ray 302, and shows an example of the relationship between the object 411 and the light ray 302 when the imaging space 106 captured by the imaging device 101 is viewed from above in the vertical direction. The captured image 400 includes an image 401 corresponding to the object 411. Note that the light ray 302 shown in FIG. 4(a) is not included in the captured image 400 as an image, but explicitly shows an example of the corresponding position of each light ray 302 shown in FIG. 4(b) in the captured image 400.

[0020] When the object 411 is, for example, a large box shape, the occupancy rate of the object 411 in the angle of view of the imaging device 101 becomes high. Therefore, the sampling number of the light rays 302 corresponding to the pixels included in the region other than the region of the image 401 corresponding to the object 411 (hereinafter included in the "non-object region") decreases. As a result, the learning about the space where the object 411 does not exist becomes insufficient, and artifacts such as haze called a floater may occur around the image corresponding to the object 411 in the generated virtual viewpoint image.

[0021] <Functional Configuration of Image Processing Apparatus> With reference to FIG. 5, the functional configuration of the image processing apparatus 102 will be described. FIG. 5 is a block diagram showing an example of the functional configuration of the image processing apparatus 102 according to Embodiment 1. Specifically, FIG. 5(a) shows an example of the functional configuration in the learning phase of the image processing apparatus 102, and FIG. 5(b) shows an example of the functional configuration in the generation phase of the image processing apparatus 102. As shown in FIG. 5(a), for example, the image processing apparatus 102 has, as a functional configuration in the learning phase, a parameter acquisition unit 501, an image acquisition unit 502, a region acquisition unit 503, a light ray setting unit 504, a learning unit 505, and a model output unit 506. Further, as shown in FIG. 5(b), the image processing apparatus 102 has, for example, as a functional configuration in the generation phase, a model acquisition unit 507, a gaze acquisition unit 508, an image generation unit 509, and an image output unit 510. Each unit that the image processing apparatus 102 has as a functional configuration is realized by the CPU 201 shown in FIG. 2 executing a program stored in the ROM 203.

[0022] The parameter acquisition unit 501 acquires learning parameters (hereinafter referred to as "learning parameters"). The learning parameters are, for example, stored in advance in the storage device 204, and the parameter acquisition unit 501 acquires the learning parameters by reading them from the storage device 204. The learning parameters acquired by the parameter acquisition unit 501 are transmitted to the light ray setting unit 504. The image acquisition unit 502 acquires photographed image data. Specifically, for example, the image acquisition unit 502 acquires the photographed image data output from each imaging device 101 via the input I / F 206. The photographed image data acquired by the image acquisition unit 502 is transmitted to the region acquisition unit 503 and the learning unit 505. The region acquisition unit 503 acquires the object region in each photographed image by extracting the object region corresponding to the image of the object 107 in each of the plurality of photographed images received from the image acquisition unit 502. The region acquisition unit 503 outputs information indicating the acquired object region as an object region map. The object region map output from the region acquisition unit 503 is acquired by the light ray setting unit 504.

[0023] The ray setting unit 504 sets a group of rays (hereinafter referred to as "learning rays") used for learning the three-dimensional field of the shooting space based on the learning parameters received from the parameter acquisition unit 501 and the object region map acquired from the region acquisition unit 503. The learning unit 505 performs learning of a learning model that represents the three-dimensional field of the shooting space using the data of the learning ray group set by the ray setting unit 504 and the shooting image data transmitted from the image acquisition unit 502. The model output unit 506 outputs a learned model that represents the three-dimensional field of the shooting space obtained as a result of the learning by the learning unit 505. Specifically, for example, the model output unit 506 outputs the learned model to the storage device 204 or the storage apparatus 104 to store the learned model in the storage device 204 or the storage apparatus 104. Information regarding the learned model is, for example, the network parameters of the learned model.

[0024] The model acquisition unit 507 acquires the learned model by reading it from the storage device 204, the storage apparatus 104, or the like. The learned model acquired by the model acquisition unit 507 is transmitted to the image generation unit 509. The gaze acquisition unit 508 acquires information (hereinafter referred to as "virtual viewpoint information") indicating the position of the virtual viewpoint and the direction of the gaze at the virtual viewpoint. The virtual viewpoint information is, for example, stored in advance in the storage device 204, and the gaze acquisition unit 508 acquires the virtual viewpoint information by reading it from the storage device 204. Note that the virtual viewpoint information may be acquired by the gaze acquisition unit 508 generating it based on an input from the user using the UI panel 103. Note that the virtual viewpoint information may be so-called virtual camera path information including time-series data of the position of the virtual viewpoint or the direction of the gaze at the virtual viewpoint. The virtual viewpoint information acquired by the gaze acquisition unit 508 is transmitted to the image generation unit 509.

[0025] The image generation unit 509 receives the trained model transmitted from the model acquisition unit 507 and the virtual viewpoint information transmitted from the gaze acquisition unit 508, and generates a virtual viewpoint image using the received trained model and virtual viewpoint information. Note that the method of generating a virtual viewpoint image using the trained model representing the three-dimensional field of the shooting space and the virtual viewpoint information is the same as the conventional method of generating a virtual viewpoint image by NeRF, so the description is omitted. The image output unit 510 outputs the virtual viewpoint image generated by the image generation unit 509. Specifically, for example, the image output unit 510 outputs the data of the virtual viewpoint image to the storage device 204 or the storage unit 104, and stores the data in the storage device 204 or the storage unit 104. The output destination of the image output unit 510 is not limited to the storage device 204 or the storage unit 104, and may be, for example, the display device 105. In this case, the image output unit 510 operates as a display control unit for outputting an image signal indicating the virtual viewpoint image to the display device 105 and displaying the virtual viewpoint image on the display device 105.

[0026] <Operations of the image processing apparatus in the learning phase> With reference to FIG. 6, the operations of the image processing apparatus 102 will be described. FIG. 6 is a flowchart showing an example of the processing flow of the image processing apparatus 102 according to Embodiment 1. Specifically, FIGS. 6(a) and 6(b) show an example of the processing flow in the learning phase of the image processing apparatus 102, and FIG. 6(c) shows an example of the processing flow in the generation phase of the image processing apparatus 102. Note that the "S" attached to the beginning of the reference numerals means steps (processes). Also, the processing of each step shown in the flowchart of FIG. 6 is realized by the CPU 201 reading a predetermined program from the ROM 203 or the storage device 204 and expanding it in the RAM 202, and then the CPU 201 executing it. Further, when the captured image data input to the image processing apparatus 102 is moving image data, the processing of each step shown in the flowchart of FIG. 6 is executed for each frame constituting the moving image.

[0027] First, with reference to FIGS. 6(a) and (b), the processing flow in the learning phase of the image processing apparatus 102 will be described. In the learning phase, first, at S601, the parameter acquisition unit 501 acquires learning parameters. Specifically, the parameter acquisition unit 501 acquires, as learning parameters, the number of learning light rays (Nr) which is the learning unit (the size of the mini-batch), and the number of object light rays (Nf) which is the number of learning light rays corresponding to the object region among the number of learning light rays (Nr). Practically, the learning unit is often set to about 4096. In Embodiment 1, it will be described assuming that the number of learning light rays (Nr) is 4096 and the number of object light rays (Nf) is 2048. However, these values are hyperparameters specified by the user in advance and are not limited to the above values. If the result of learning by the learning unit 505 is not stable, the number of learning light rays (Nr) may be set to a larger value. Note that, the larger the number of learning light rays (Nr), the more the usage amount of VRAM increases. Next, at S602, the image acquisition unit 502 acquires a plurality of captured image data (multi-viewpoint image data) obtained by photographing with the plurality of imaging devices 101.

[0028] Next, at S603, the region acquisition unit 503 acquires the object region in each captured image constituting the multi-viewpoint image data acquired at S602. Specifically, the region acquisition unit 503 acquires the object region in each captured image by extracting the object region from each captured image. The region acquisition unit 503 generates an object region map indicating the acquired object region for each captured image and outputs this to the ray setting unit 504. For example, the region acquisition unit 503 specifies and extracts the object region in each captured image based on the difference between the captured image and a background image prepared in advance. The method for acquiring the object region in the region acquisition unit 503 is not limited to the above method. Note that, the region acquisition unit 503 does not necessarily need to acquire the object region in the captured image using the captured image. For example, the region acquisition unit 503 may acquire the object region in the captured image by acquiring the extraction result of the object region in each captured image by another external device.

[0029] Next, in S604, the ray setting unit 504 refers to the object area map generated in S603 and sets a group of rays (learning rays) to be used for learning the three-dimensional field of the imaging space. Specifically, the ray setting unit 504 refers to the object area map and generates a list of data of learning rays (hereinafter referred to as "object rays") corresponding to the pixels included in the object area in each captured image (hereinafter referred to as the "object ray list"). Further, the ray setting unit 504 refers to the object area map and generates a list of data of learning rays (hereinafter referred to as "non-object rays") corresponding to the pixels included in the non-object area in each captured image (hereinafter referred to as the "non-object ray list").

[0030] FIG. 7 is a diagram for explaining the processes of S602 to S604 shown in FIG. 6(a). Specifically, FIG. 7(a) is an example of a captured image 700 obtained by capturing with a certain imaging device 101, and shows an example of the captured image 700 acquired in S602. The captured image 700 includes a region of an image corresponding to the object 107, that is, an object region 701. FIG. 7(b) shows an example of an object area map 710 corresponding to the captured image 700. The object area map 710 includes an object area 711 corresponding to the object area 701 in the captured image 700 and a non-object area 712 corresponding to an area other than the object area 701 in the captured image 700. FIG. 7(c) shows an example of the distribution of the learning rays 721 corresponding to the pixels of the captured image 700. Note that in FIG. 7(c), instead of the captured image 700, the object area map 710 is used to show an example of the distribution of the learning rays 721 corresponding to the captured image 700.

[0031] After S604, in S605, the learning unit 505 executes learning processing on the learning model that represents the three-dimensional field of the shooting space. Specifically, first, the learning unit 505 samples, as learning ray data for learning the learning model that represents the three-dimensional field of the shooting space, the data of Nf object rays among the object ray data included in the object ray list. Also, the learning unit 505 samples, as learning ray data for learning the learning model that represents the three-dimensional field of the shooting space, the data of Nb non-object rays among the non-object ray data included in the non-object ray list. Here, Nb is the number of non-object rays, which is a value obtained by subtracting the number of object rays (Nf) from the number of learning rays (Nr). Therefore, the number of learning ray data sampled for learning the learning model that represents the three-dimensional field of the shooting space is the number of learning rays (Nr) obtained by adding the number of object rays (Nf) and the number of non-object rays (Nb). Subsequently, the learning unit 505 uses the sampled learning ray data to perform learning of the learning model that represents the three-dimensional field of the shooting space. Details of the learning processing in S605 will be described later with reference to FIG. 6(b).

[0032] Next, in S606, the model output unit 506 outputs the learned model obtained as a result of the learning processing in S605. Specifically, for example, the model output unit 506 outputs a file including the network parameters of the learned model as data to the storage device 204 or the storage apparatus 104, and stores the file in the storage device 204 or the storage apparatus 104. The format of the file varies depending on the learning environment of the learning model. For example, in the case of machine learning using PyTorch, a file represented by an extension such as pt or pth is often saved. After S606, the image processing apparatus 102 ends the processing of the flowchart shown in FIG. 6(a).

[0033] <Learning Processing in the Learning Unit> Referring to FIG. 6(b), the learning process in the learning unit 505 will be described. FIG. 6(b) is a flowchart showing an example of the flow of the learning process in the learning unit 505 according to Embodiment 1, and is a flowchart showing an example of the flow of the learning process in S605. The process of this flowchart is executed after S604. First, in S607, the learning unit 505 executes an initialization process. Specifically, for example, the learning unit 505 executes the following processes as the initialization process. The learning unit 505 obtains the number of object ray data included in the object ray list (hereinafter referred to as "object ray list length (Lf)") and the number of non-object ray data included in the non-object ray list (hereinafter referred to as "non-object ray list length (Lb)"). Also, the learning unit 505 sets the sampling counter (Sf) of the object ray to the object ray list length (Lf) and the sampling counter (Sb) of the non-object ray to the non-object ray list length (Lb), respectively.

[0034] Next, in S608, the learning unit 505 determines whether all the object ray data included in the object ray list have been selected (sampled). Specifically, if Sf < Lf, the learning unit 505 determines that at least some of the object ray data included in the object ray list have not been selected (No). Also, if Sf ≥ Lf, the learning unit 505 determines that all the object ray data included in the object ray list have been selected (Yes). Therefore, the first determination in S608 is always determined as Yes. If it is determined as Yes in S608, in S609, the learning unit 505 randomly rearranges the object ray data included in the object ray list and sets the sampling counter (Sf) of the object ray to 0.

[0035] After S609 or when it is determined as No in S608, in S610, the learning unit 505 determines whether data of all non-object light rays included in the non-object light ray list has been selected (sampled). Specifically, if Sb < Lb, the learning unit 505 determines that data of at least some of the non-object light rays included in the non-object light ray list has not been selected (No). On the other hand, if Sb ≥ Lb, the learning unit 505 determines that data of all non-object light rays included in the non-object light ray list has been selected (Yes). Therefore, the first determination in S610 is always determined as Yes. When it is determined as Yes in S610, in S611, the learning unit 505 randomly rearranges the data of the non-object light rays included in the non-object light ray list and sets the sampling counter (Sb) of the non-object light rays to 0.

[0036] After S611 or when it is determined as No in S610, the learning unit 505 executes the process of S612. First, in S612, the learning unit 505 selects (samples) data of Nf object light rays from the (Sf + 1)-th to the Nf-th among the data of the object light rays included in the object light ray list. Subsequently, in S612, the learning unit 505 copies the selected data of the object light rays to the learning light ray list as learning light ray data. When Sf + Nf > Lf, the learning unit 505 performs the following process. Specifically, first, the learning unit 505 selects the data of the object light rays from the (Sf + 1)-th to the Lf-th among the data of the object light rays included in the object light ray list and copies this as learning light ray data. Subsequently, the learning unit 505 selects the data of the object light rays from the 1-st to the (Sf + Nf - Lf)-th located at the upper part of the object light ray list and adds and copies this as learning light ray data. After copying to the learning light ray list, the learning unit 505 adds the number of object light rays (Nf) to the sampling counter (Sf) of the object light rays.

[0037] Subsequently, at S612, the learning unit 505 selects (samples) the data of the Nb non-object light rays from the (Sb + 1)-th to the Nb-th among the data of the non-object light rays included in the non-object light ray list. Subsequently, at S612, the learning unit 505 copies the selected data of the non-object light rays to the learning ray list as learning ray data. If Sb + Nb > Lb, the learning unit 505 performs the following processing. Specifically, first, the learning unit 505 selects the data of the non-object light rays from the (Sb + 1)-th to the Lb-th among the data of the non-object light rays included in the non-object light ray list, and copies this as learning ray data. Subsequently, the learning unit 505 selects the data of the non-object light rays from the 1st to the (Sb + Nb - Lb)-th located at the upper part of the non-object light ray list, and adds and copies this as learning ray data.

[0038] After copying to the learning ray list, subsequently, at S612, the learning unit 505 adds the number of non-object light rays (Nb) to the sampling counter (Sb) of the non-object light rays. Therefore, the learning ray list length becomes the number of learning rays (Nr) obtained by adding the number of object light rays (Nf) and the number of non-object light rays (Nb). Further, at S612, the learning unit 505 copies the pixel values of the captured images corresponding to the data of the object light rays and the data of the non-object light rays copied as learning ray data to the learning ray list to the correct answer list as correct answer data.

[0039] After S612, based on the learning ray data copied to the learning ray list and the correct answer data copied to the correct answer list, the learning process of the learning model representing the three-dimensional field of the imaging space is executed. The processes from S608 to S612 and the learning process of the learning model subsequent to S612 are repeatedly executed until it is determined as Yes at S615 described later.

[0040] FIG. 8 is a diagram showing an example of the flow of sampling of object ray data included in the object ray list according to Embodiment 1. Specifically, FIG. 8 shows, as an example, a case where the object ray list length (Lf) is 6, the learning ray list length, that is, the number of learning rays (Nr) is 4, and the number of object rays (Nf) is 2. For the sake of simplicity of explanation, FIG. 8 does not describe the sampling of non-object ray data in the learning ray list.

[0041] First, by the initialization process of S607, the object ray data is stored in the object ray list in the order of pixel positions, and the object ray sampling counter (Sf) is set to 6. Next, since it is determined in S608 that Sf≥Lf, in S609, the object ray data included in the object ray list is randomly rearranged, and the object ray sampling counter (Sf) is set to 0. Next, in S612, the data of the first two object rays among the object ray data included in the object ray list is copied to the learning ray list, and 2 is set in the object ray sampling counter (Sf). Thereafter, each time the learning process is repeatedly executed, 2 is added to the object ray sampling counter (Sf) in S612. Further, when the object ray sampling counter (Sf) reaches 6, in S609, the object ray data included in the object ray list is rearranged again, and the object ray sampling counter (Sf) is reset to 0. Thereafter, the second round of sampling for the object ray list is performed.

[0042] After S612, in S613, the learning unit 505 calculates a pixel value corresponding to each learning ray data included in the learning ray list using a learning model that represents the three-dimensional field of the imaging space. Here, when the pixel value is represented by a value indicating a color such as R (Red), G (Green), and B (Blue), the learning unit 505 calculates a color value as the pixel value. Next, in S614, the learning unit 505 updates the network parameters of the learning model that represents the three-dimensional field of the imaging space so that the difference (loss) between the calculated pixel value and the pixel value as the correct answer data included in the correct answer list becomes smaller.

[0043] Next, in S615, the learning unit 505 determines whether the learning process for the learning model that represents the three-dimensional field of the imaging space has satisfied the end condition. If it is determined in S615 that the end condition is not satisfied (No), the learning unit 505 returns to the process of S608 and repeatedly executes a series of processes from S608 to S615 until it is determined in S615 that the end condition is satisfied (Yes). If it is determined in S615 that the end condition is satisfied (Yes), the learning unit 505 ends the learning process of the flowchart shown in FIG. 6(b), that is, the learning process of S605. Note that the end condition of the learning process is, for example, to execute the learning process of the number of learning times, also called the number of iterations specified in advance, on the learning model. The end condition of the learning process is not limited to this, and for example, it may be to satisfy a predetermined convergence condition for the above-described loss. Also, for example, when a sign that the loss for a separately prepared verification image increases is confirmed, it may be determined that the end condition of the learning process is satisfied.

[0044] As described above, the image processing apparatus 102 is configured to appropriately control the number of data of learning rays corresponding to object regions in a captured image, which is used during learning of a learning model that represents a three-dimensional field of a shooting space. According to the image processing apparatus 102 configured in this way, the learning process of the learning model that represents the three-dimensional field of the shooting space can be executed at high speed without depending on the size of the object and the angle of view of the imaging apparatus 101. Further, according to the image processing apparatus 102 configured in this way, a learned model that represents the three-dimensional field of the shooting space with high accuracy can be generated without depending on the size of the object and the angle of view of the imaging apparatus 101.

[0045] <Operation in the generation phase of the image processing apparatus> With reference to FIG. 6(c), the operation in the generation phase of the image processing apparatus 102 according to Embodiment 1 will be described. First, in S616, the model acquisition unit 507 acquires a learned model that represents the three-dimensional field of the shooting space. Specifically, the model acquisition unit 507 acquires the learned model by reading out the network parameters of the learned model from the storage device 204, the storage apparatus 104, or the like. Next, in S617, the line-of-sight acquisition unit 508 acquires virtual line-of-sight information. Next, in S618, the image generation unit 509 generates a virtual viewpoint image using the learned model acquired in S616 and the virtual line-of-sight information acquired in S617. Specifically, the image generation unit 509 inputs information indicating the position of the virtual viewpoint and information indicating the direction of the line of sight at the virtual viewpoint, which are included in the virtual line-of-sight information, into the learned model. The learned model calculates the pixel values of the virtual viewpoint image based on rays corresponding to the position of the virtual viewpoint and the direction of the line of sight at the virtual viewpoint, and outputs the calculated pixel values. The image generation unit 509 generates a virtual viewpoint image by acquiring the pixel values output from the learned model and constructing a virtual viewpoint image having the pixel values.

[0046] Next, in S619, the image output unit 510 outputs the data of the virtual viewpoint image generated in S618, or the image signal for displaying the virtual viewpoint image, to the storage device 204 or the storage apparatus 104, or the display device 105. After S619, the image processing apparatus 102 ends the processing of the flowchart shown in FIG. 6(c). When the image generation unit 509 generates a moving image as the virtual viewpoint image, such as when the virtual gaze information is the virtual camera path, the image processing apparatus 102 repeatedly executes the processing of the flowchart shown in FIG. 6(b).

[0047] As described above, the image processing apparatus 102 is configured to appropriately control the number of data of the learning rays corresponding to the object regions in the captured images, which are used during the learning of the learning model representing the three-dimensional field of the shooting space. Also, the image processing apparatus 102 is configured to generate a virtual viewpoint image using the learned model that accurately represents the three-dimensional field of the shooting space, which is obtained as a result of such learning. According to the image processing apparatus 102 configured as described above, it is possible to generate a high-precision virtual viewpoint image without depending on the occupancy rate of the object regions in the captured images used during learning.

[0048] [Embodiment 2] In Embodiment 1, the image region in the captured image is divided into an object region and a non-object region, and the form of selecting (sampling) the learning ray data corresponding to each pixel based on these is described. In Embodiment 2, a form of more efficiently performing the learning model learning representing the three-dimensional field of the shooting space by dividing the image region in the captured image into a learning region and a non-learning region will be described. In Embodiment 2, the description will focus on the processes different from those described in Embodiment 1, and the description of the same processes will be omitted.

[0049] FIG. 9 is a block diagram showing an example of the functional configuration of an image processing apparatus 102 according to Embodiment 2 (hereinafter simply referred to as "image processing apparatus 102"). The image processing apparatus 102 is different from the image processing apparatus 102 according to Embodiment 1 in that it has a region setting unit 901. The region setting unit 901 acquires the object region map output from the region acquisition unit 503, and based on the acquired object region map, sets a learning region and a non-learning region in the captured image. Information indicating the learning region in the captured image set by the region setting unit 901 (hereinafter referred to as "learning region information") is output to the light ray setting unit 504. The light ray setting unit 504 acquires the learning region information output from the region setting unit 901 and the object region map output from the region acquisition unit 503. The light ray setting unit 504 sets a learning light ray group used for learning the three-dimensional field of the imaging space based on the acquired learning region information and object region map.

[0050] FIG. 10 is a flowchart showing an example of the processing flow in the learning phase of the image processing apparatus 102 according to Embodiment 2. FIG. 11 is a diagram for explaining the processing of the flowchart shown in FIG. 10. Among the processes of the steps shown in FIG. 10, steps that perform the same processes as those shown in FIG. 6(a) are denoted by the same reference numerals and the description thereof is omitted. First, the image processing apparatus 102 executes the process of S601.

[0051] After S601, in S1001, the image acquisition unit 502 acquires a plurality of captured image data (multi-viewpoint image data) obtained by photographing with a plurality of imaging devices 101. For example, the image acquisition unit 502 acquires, as the captured image data, data of a captured image with α-channel data. For example, in the α-channel, the α value is set to 1 for the object region in the captured image, and the α value is set to 0 for the non-object region. The α value is not limited to the two values of 1 or 0. For example, in the vicinity of the contour of the object region, an intermediate α value such as 0.5 may be set. FIG. 11(a) is a captured image 1100 similar to the captured image 700 shown in FIG. 7. However, the data of the captured image 1100 is attached with α-channel data.

[0052] Next, in S1002, the region acquisition unit 503 acquires the object region in each captured image acquired in S1001. Specifically, the region acquisition unit 503 extracts the α-channel data from the captured image with α-channel data, and outputs the extracted α-channel data as an object region map indicating the object region in the captured image. FIG. 11(b) is an example of the object region map 1110 output in S1002, and shows an example of the object region map 1110 corresponding to the captured image 1100 acquired in S1001. The object region map 1110 indicates the object region 1111 and the non-object region 1112 by two values.

[0053] After S1002, the region setting unit 901 executes a series of processes from S1003 to S1005 to set the learning region and non-learning region in the captured image obtained in S1001. Specifically, in S1003, the region setting unit 901 obtains the object region map corresponding to each captured image output in S1002, and obtains the three-dimensional shape of the object 107 by the volume intersection method using the obtained plurality of object region maps. Since the volume intersection method is a well-known technique, the description thereof is omitted. Next, in S1004, the region setting unit 901 obtains the learning space in the imaging space 106 based on the three-dimensional shape of the object 107 obtained in S1003. Specifically, for example, the region setting unit 901 obtains a circumscribed polyhedron such as a circumscribed rectangular parallelepiped or a circumscribed sphere that includes the three-dimensional shape of the object 107 obtained in S1003 as the learning space.

[0054] Next, in S1005, the region setting unit 901 projects the learning space obtained in S1004 onto the positions of the respective imaging devices 101, that is, onto each viewpoint, and sets the region where the projections of the learning space in each captured image intersect as the learning region. Further, the region setting unit 901 sets the region where the projections of the learning space in each captured image do not intersect as the non-learning region. FIG. 11(c) shows an example of the learning region 1121 and the non-learning region 1122 in the captured image 1100. Note that in FIG. 11(c), an example of the learning region 1121 and the non-learning region 1122 in the captured image 1100 is shown using the object region map 1110 instead of the captured image 1100.

[0055] Next, in S1006, with reference to the object region map output in S1002 and the learning region set in S1005, the ray setting unit 504 sets a group of rays (learning rays) to be used for learning the three-dimensional field of the imaging space. Specifically, the ray setting unit 504 refers to the object region map and generates a list (object ray list) of data of learning rays (object rays) corresponding to each pixel group included in the object region in each captured image. Also, the ray setting unit 504 refers to the object region map and the learning region and generates a list (non-object ray list) of data of learning rays (non-object rays) corresponding to the pixel groups included in the non-object regions of the learning regions of each captured image.

[0056] That is, the ray setting unit 504 generates a non-object ray list so as not to include the data of the learning rays corresponding to the pixel groups included in the non-learning regions in the non-object ray list. FIG. 11(d) shows an example of the distribution of the learning rays 1131 corresponding to the pixels of the captured image 1100. Note that in FIG. 11(d), instead of the captured image 1100, an object region map 1110 is used to show an example of the distribution of the learning rays 1131 corresponding to the captured image 1100.

[0057] Next, in S605, the learning unit 505 executes a learning process for the learning model representing the three-dimensional field of the imaging space. Since the process of S605 according to the second embodiment is the same as the process of S605 according to the first embodiment, detailed description thereof is omitted. Next, in S606, the model output unit 506 outputs the learned model. Since the process of S606 according to the second embodiment is the same as the process of S606 according to the first embodiment, detailed description thereof is omitted. After S606, the image processing apparatus 102 ends the processing of the flowchart shown in FIG. 10.

[0058] According to the image processing apparatus 102 configured as described above, by setting a non-learning region in the captured image, it is possible to efficiently remove the redundant light rays that are unnecessary for the learning of the learning model representing the three-dimensional field of the captured space. As a result, according to the image processing apparatus 102, the learning of the learning model can be performed at a higher speed.

[0059] In the above description, the non-object light ray list has been described as being generated so as not to include the data of the learning light rays corresponding to the pixels included in the non-learning region, but it is not limited thereto. For example, the light ray setting unit 504 may generate a non-object light ray list so as to include the data of the learning light rays corresponding to the pixels included in the non-learning region, similar to the light ray setting unit 504 according to the first embodiment. In this case, when executing the learning process, the learning unit 505 may not sample the data of the non-object light rays corresponding to the pixels included in the non-learning region among the data of the non-object light rays included in the non-object light ray list.

[0060] [Embodiment 3] In the first and second embodiments, a form has been described in which the predetermined number of learning light rays (Nr) and the number of object light rays (Nf) are acquired as learning parameters, and sampling of the object light rays and non-object light rays is performed. In the third embodiment, a form in which the number of object light rays is determined according to each captured image will be described. With reference to FIG. 12, the operation of the image processing apparatus 102 according to the third embodiment will be described. Since the functional configuration of the image processing apparatus 102 according to the third embodiment is the same as the functional configuration of the image processing apparatus 102 according to the second embodiment shown as an example in FIG. 9, detailed description thereof will be omitted.

[0061] FIG. 12 is a flowchart showing an example of the processing flow of the image processing apparatus 102 (hereinafter simply referred to as "image processing apparatus 102") according to Embodiment 3. Among the processes of the steps shown in FIG. 12, steps that perform the same processes as the steps shown in FIG. 6(a) or FIG. 10 are denoted by the same reference numerals and the description thereof is omitted. First, in S1210, the parameter acquisition unit 501 acquires learning parameters.

[0062] Specifically, the parameter acquisition unit 501 acquires an object ray ratio lower limit value (Rf min ) indicating the lower limit value of the ratio of the number of object rays (Nf) to the number of learning rays (Nr), and an object ray ratio upper limit value (Rf max ) indicating the upper limit value of the ratio. The parameter acquisition unit 501 may acquire the number of learning rays (Nr), and the lower limit number of object rays (Nf min ) and the upper limit number of object rays (Nf max ) indicating the lower limit value of the number of object rays. In this case, the parameter acquisition unit 501 can acquire the object ray ratio lower limit value (Rf min ) by dividing the lower limit number of object rays (Nf min ) by the number of learning rays (Nr). Also, the parameter acquisition unit 501 can acquire the object ray ratio upper limit value (Rf max ) by dividing the upper limit number of object rays (Nf max ) by the number of learning rays (Nr).

[0063] After S1210, the image processing apparatus 102 executes a series of processes from S1001 to S1005. After S1005, at S1220, the light ray setting unit 504 refers to the object area map output at S1002 and the learning area set at S1005, and sets a group of light rays (learning light rays) to be used for learning the three-dimensional field of the imaging space. Details of the learning light ray group setting process in S1220 will be described later with reference to FIG. 12(b). After S1220, at S1230, the learning unit 505 executes a learning process for the learning model representing the three-dimensional field of the imaging space. Details of the learning process in S1230 will be described later. After S1230, the image processing apparatus 102 executes the process of S606. After S606, the image processing apparatus 102 ends the process of the flowchart shown in FIG. 12(a).

[0064] FIG. 12(b) is a flowchart showing an example of the flow of the learning light ray group setting process in the light ray setting unit 504 according to Embodiment 3, and is a flowchart showing an example of the flow of the learning light ray group setting process in S1220. The process of this flowchart is executed after S1005. First, at S1201, the light ray setting unit 504 selects an arbitrary captured image from among the plurality of captured images acquired at S1001. Next, at S1202, the light ray setting unit 504 refers to the learning area set at S1005, and calculates learning light rays corresponding to the pixels included in the learning area in the captured image selected at S1201.

[0065] Next, at S1203, the light ray setting unit 504 generates an object light ray list and a non-object light ray list corresponding to the captured image selected at S1201 based on the object area acquired at S1002 and the learning light rays calculated at S1202. Also, at S1203, the light ray setting unit 504 obtains the list length (object light ray list length (Lf i )) of the generated object light ray list and the list length (non-object light ray list length (Lb i )) of the generated non-object light ray list. Note that Lf i and Lb iThe "i" included in [it] indicates the index of the captured image selected in S1201.

[0066] Next, in S1204, the light ray setting unit 504 determines whether all the captured images have been selected in S1201. If it is determined in S1204 that at least some of the captured images have not been selected, the light ray setting unit 504 returns to the process of S1201 and repeatedly executes the processes from S1201 to S1204 until it is determined in S1204 that all the captured images have been selected. When repeating the processes from S1201 to S1204, in S1201, the light ray setting unit 504 selects an arbitrary captured image from one or more unselected captured images among the plurality of captured images.

[0067] If it is determined in S1204 that all the captured images have been selected, the light ray setting unit 504 executes the process of S1205. Specifically, in S1205, the light ray setting unit 504 calculates the object light ray number (Nf i ) and the non-object light ray number (Nb i ) of each captured image from the object light ray list length (Lf i ) and the non-object light ray list length (Lb i ) corresponding to each captured image. For example, the object light ray number (Nf i ) and the non-object light ray number (Nb i ) of each captured image can be calculated using the following expressions (1) to (5).

[0068] L i =ΣLf i +ΣLb i ··· Expression (1) Rf i =(ΣLf i ) / L i ··· Expression (2) Rf i ´=min(max(Rf i ,Rf min ),Rf max ) ··· Expression (3) Nf i =L i ·Rf i··· Formula (4) Nb i =L i ·(1 - Rf i ′) ··· Formula (5) Here, L i is the number of candidate learning rays, and Rf i is the ratio of the number of object rays to the number of candidate learning rays (L i ). max() is the max function, and min() is the minimum function. Rf i ′ is the value obtained by applying the upper limit value (Rf i ) and the lower limit value (Rf i ) of the object ray ratio to the ratio of the number of object rays (Rf max ) to the number of candidate learning rays (L min ). Nf i is the number of object rays sampled for each captured image, and Nb i is the number of non-object rays sampled for each captured image.

[0069] After S1205, the ray setting unit 504 ends the process of the flowchart shown in FIG. 12(b), that is, the process of S1220 shown in FIG. 12(a). After S1220, at S1230, the learning unit 505 uses the number of object rays (Nf i ) and the number of non-object rays (Nb i ) of each captured image calculated at S1205 to perform a learning process on the learning model representing the three-dimensional field of the imaging space. Specifically, the learning unit 505 replaces the number of object rays (Nf) in the process of the flowchart shown in FIG. 6(b) with the number of object rays (Nf i ) corresponding to the captured image for each captured image, and executes the process of the flowchart. Similarly, the learning unit 505 replaces the number of non-object rays (Nb) in the process of the flowchart shown in FIG. 6(b) with the number of non-object rays (Nb i ) corresponding to the captured image for each captured image, and executes the process of the flowchart.

[0070] According to the image processing apparatus 102 configured as described above, while maintaining learning according to the ratio occupying the object region for each captured image, it is possible to suppress the bias in learning associated with an extreme ratio bias between captured images. Further, according to the image processing apparatus 102, for a specific captured image in which the ratio of the object region is small, learning of the object region can be performed without omission. As a result, according to the image processing apparatus 102 configured as described above, a more stable learning result can be obtained as compared with the case of randomly sampling learning rays from the entire captured image. That is, according to the image processing apparatus 102 configured as described above, a learned model that expresses the three-dimensional field of the imaging space with high accuracy can be generated without depending on the size of the object and the angle of view of the imaging device 101. Further, according to the image processing apparatus 102 configured as described above, a virtual viewpoint image with high accuracy can be generated without depending on the occupancy rate of the object region in the captured image used in learning.

[0071] [Other Modifications] <Object Region Extraction> As a method for extracting the object region in the captured image, a method using the grabCut algorithm or other methods such as a method using a learned model for object extraction obtained as a result of machine learning may be used.

[0072] The present disclosure can also be realized by supplying a program that realizes one or more functions of the above-described embodiment to a system or apparatus via a network or a storage medium, and causing one or more processors in a computer of the system or apparatus to read and execute the program. Further, it can also be realized by a circuit (for example, ASIC) that realizes one or more functions.

[0073] In addition, the present disclosure is not limited to the above-described embodiments, and the configuration may be modified without departing from the gist thereof at the implementation stage. Further, various embodiments can be formed by appropriately combining a plurality of configurations disclosed in the above-described embodiments. For example, some configurations may be deleted from all the configurations shown in the above-described embodiments. Also, the configurations described in a plurality of embodiments may be appropriately combined.

[0074] Note that, in the above-described embodiments, the expression "at least one of Configuration A and Configuration B" means that it may be only Configuration A, only Configuration B, or both Configuration A and Configuration B.

[0075] [Configuration of the Present Disclosure] The present disclosure includes the following configurations, methods, and programs.

[0076] [Configuration 1] Image acquisition means for acquiring data of a plurality of captured images obtained by capturing from a plurality of viewpoints, Region acquisition means for acquiring an object region corresponding to an image of an object in each of the plurality of captured images, Ray setting means for setting a group of learning rays used for learning information about a three-dimensional field of a captured space that is a capture target from the plurality of viewpoints, the ray setting means setting the group of learning rays corresponding to each pixel of the plurality of captured images based on the acquired object region, Learning means for performing learning of information about the three-dimensional field based on the set group of learning rays, An image processing apparatus characterized by comprising the above.

[0077] [Configuration 2] The learning means controls the number of selected learning rays corresponding to the object region when selecting learning rays used for learning information about the three-dimensional field from among the group of learning rays. The image processing apparatus according to Configuration 1, characterized by the above.

[0078] [Configuration 3] Parameter acquisition means for acquiring, as a learning parameter, information regarding the number of selected learning rays when selecting learning rays to be used for learning information regarding the three-dimensional field from among the learning ray group, The image processing apparatus according to Configuration 1 or 2, characterized by this.

[0079] [Configuration 4] The parameter acquisition means acquires, as the learning parameter, the number of selected learning rays corresponding to the object region when selecting learning rays to be used for learning information regarding the three-dimensional field from among the learning ray group, When selecting the learning rays to be used for learning information regarding the three-dimensional field from among the set learning ray group, the learning means selects the learning rays based on the number of selected learning rays corresponding to the object region acquired as the learning parameter, The image processing apparatus according to Configuration 3, characterized by this.

[0080] [Configuration 5] The parameter acquisition means acquires, as the learning parameter, the ratio of the learning rays corresponding to the object region, Based on the acquired ratio, the learning means determines the number of selected learning rays corresponding to the object region when selecting the learning rays to be used for learning information regarding the three-dimensional field, and selects the learning rays to be used for learning information regarding the three-dimensional field based on the determined number of selected learning rays, The image processing apparatus according to Configuration 3, characterized by this.

[0081] [Configuration 6] The learning means determines the number of selected learning rays corresponding to the object region based on the ratio of the number of pixels included in the object region and the number of pixels included in a non-object region, which is a region other than the object region, in each of the plurality of captured images, The image processing apparatus according to Configuration 2, characterized by this.

[0082] [Configuration 7] The parameter acquisition means acquires, as the learning parameter, at least one of a lower limit value and an upper limit value of the number of selections of the learning light beam corresponding to the object region when selecting the learning light beam used for learning information about the three-dimensional field, The learning means determines the number of selections of the learning light beam corresponding to the object region when selecting the learning light beam used for learning information about the three-dimensional field based on at least one of the acquired lower limit value and the upper limit value. The image processing apparatus according to Configuration 5, characterized in that.

[0083] [Configuration 8] The learning means determines the number of selections of the learning light beam corresponding to the object region for each of the captured images in the plurality of captured images when selecting the learning light beam used for learning information about the three-dimensional field from among the set learning light beam groups. The image processing apparatus according to any one of Configurations 1 to 7, characterized in that.

[0084] [Configuration 9] Region setting means for setting a learning region in each of the plurality of captured images, further comprising The learning means selects a learning light beam used for learning information about the three-dimensional field from among the learning light beam groups corresponding to the learning region, and performs learning of information about the three-dimensional field using the selected learning light beam. The image processing apparatus according to any one of Configurations 1 to 8, characterized in that.

[0085] [Configuration 10] Region setting means for setting a learning region in each of the plurality of captured images, further comprising The light beam setting means sets the learning light beam group corresponding to the pixels included in the learning region, The learning means performs learning of information about the three-dimensional field based on the set learning light beam group. The image processing apparatus according to any one of Configurations 1 to 9, characterized in that...

[0086] [Configuration 11] The region setting means acquires learning region information indicating the learning region in each of the plurality of captured images, and sets the learning region in each of the plurality of captured images based on the learning region information. The image processing apparatus according to Configuration 9 or 10, characterized in that...

[0087] [Configuration 12] The region setting means sets the learning region in each of the plurality of captured images by calculating the learning region in each of the plurality of captured images based on the object region in each of the plurality of captured images. The image processing apparatus according to any one of Configurations 9 to 11, characterized in that...

[0088] [Configuration 13] A viewpoint acquisition means for acquiring virtual viewpoint information including at least viewpoint position information indicating the position of a virtual viewpoint and line-of-sight direction information indicating the direction of the line of sight at the virtual viewpoint; An image generation unit that generates a virtual viewpoint image corresponding to the virtual viewpoint based on the acquired virtual viewpoint information and information on the three-dimensional field obtained as a result of learning by the learning means. The image processing apparatus according to any one of Configurations 1 to 12, characterized in that...

[0089] [Method] An image acquisition step of acquiring data of a plurality of captured images obtained by capturing from a plurality of viewpoints; A region acquisition step of acquiring an object region corresponding to an image of an object in each of the plurality of captured images; A ray setting step of setting a group of learning rays used for learning information about a three-dimensional field of a shooting space that is a shooting target from the plurality of viewpoints, the ray setting step of setting the group of learning rays corresponding to each pixel of the plurality of shooting images based on the acquired object region; A learning step of learning information about the three-dimensional field based on the set group of learning rays; An image processing method characterized by including the above.

[0090] [Program] A program for causing a computer to function as the image processing apparatus according to any one of Configurations 1 to 13.

Explanation of Signs

[0091] 102 Image processing apparatus 502 Image acquisition unit 503 Region acquisition unit 504 Ray setting unit 505 Learning unit

Claims

1. Image acquisition means for acquiring data of a plurality of captured images obtained by capturing from a plurality of viewpoints; Region acquisition means for acquiring an object region corresponding to an image of an object in each of the plurality of captured images; Ray setting means for setting a learning ray group used for learning information about a three-dimensional field of a shooting space that is a shooting target from the plurality of viewpoints, the ray setting means setting the learning ray group corresponding to each pixel of the plurality of captured images based on the acquired object region; Learning means for learning information about the three-dimensional field based on the set learning ray group; An image processing apparatus comprising the same.

2. The learning means controls the number of selected learning rays corresponding to the object region when selecting learning rays used for learning information about the three-dimensional field from among the learning ray group. The image processing apparatus according to claim 1, characterized in that.

3. Parameter acquisition means for acquiring, as a learning parameter, information regarding the number of selected learning rays when selecting learning rays used for learning information about the three-dimensional field from among the learning ray group, further comprising the same. The image processing apparatus according to claim 1, characterized in that.

4. The parameter acquisition means acquires, as the learning parameter, the number of selected learning rays corresponding to the object region when selecting learning rays used for learning information about the three-dimensional field from among the learning ray group. When the learning means selects the learning rays used for learning information about the three-dimensional field from among the set learning ray group, the learning means selects the learning rays based on the number of selected learning rays corresponding to the object region acquired as the learning parameter. The image processing apparatus according to claim 3, characterized in that.

5. The parameter acquisition means acquires, as the learning parameter, the ratio of the learning rays corresponding to the object region. Based on the acquired ratio, the learning means determines the number of selected learning rays corresponding to the object region when selecting the learning rays used for learning information about the three-dimensional field, and selects the learning rays used for learning information about the three-dimensional field based on the determined number of selected learning rays. The image processing apparatus according to claim 3, characterized in that.

6. The learning means determines the number of selected learning light rays corresponding to the object region based on the ratio between the number of pixels included in the object region and the number of pixels included in a non-object region, which is a region other than the object region, in each of the plurality of captured images. The image processing apparatus according to claim 2, characterized in that.

7. The parameter acquisition means acquires, as the learning parameter, at least one of a lower limit value and an upper limit value of the number of selected learning light rays corresponding to the object region when selecting the learning light rays used for learning information about the three-dimensional field. The learning means determines the number of selected learning light rays corresponding to the object region when selecting the learning light rays used for learning information about the three-dimensional field based on at least one of the acquired lower limit value and the upper limit value. The image processing apparatus according to claim 5, characterized in that.

8. The learning means determines the number of selected learning light rays corresponding to the object region for each captured image in the plurality of captured images when selecting the learning light rays used for learning information about the three-dimensional field from among the set learning light ray groups. The image processing apparatus according to claim 1, characterized in that.

9. Region setting means for setting a learning region in each of the plurality of captured images. Further comprising. The learning means selects learning light rays used for learning information about the three-dimensional field from among the learning light ray groups corresponding to the learning region, and performs learning of information about the three-dimensional field using the selected learning light rays. The image processing apparatus according to claim 1, characterized in that.

10. Region setting means for setting a learning region in each of the plurality of captured images. Further comprising. The light ray setting means sets the learning light ray group corresponding to the pixels included in the learning region. The learning means performs learning of information about the three-dimensional field based on the set learning light ray group. The image processing apparatus according to claim 1, characterized in that.

11. The region setting means acquires learning region information indicating the learning region in each of the plurality of captured images, and sets the learning region in each of the plurality of captured images based on the learning region information. The image processing apparatus according to claim 9, characterized in that.

12. The area setting means sets the learning area for each of the plurality of captured images by calculating the learning area for each of the plurality of captured images based on the object area in each of the plurality of captured images. The image processing apparatus according to claim 9, characterized in that.

13. A viewpoint acquisition means for acquiring virtual viewpoint information including at least viewpoint position information indicating the position of a virtual viewpoint and line-of-sight direction information indicating the direction of the line of sight at the virtual viewpoint; An image generation unit that generates a virtual viewpoint image corresponding to the virtual viewpoint based on the acquired virtual viewpoint information and the information regarding the three-dimensional field obtained as a result of learning by the learning means. The image processing apparatus according to claim 1, characterized in that.

14. An image acquisition step of acquiring data of a plurality of captured images obtained by capturing from a plurality of viewpoints; An area acquisition step of acquiring an object area corresponding to an image of an object in each of the plurality of captured images; A ray setting step of setting a learning ray group used for learning information regarding a three-dimensional field of a shooting space that is a shooting target from a plurality of viewpoints, the ray setting step of setting the learning ray group corresponding to each pixel of the plurality of captured images based on the acquired object area; A learning step of learning information regarding the three-dimensional field based on the set learning ray group; An image processing method characterized by including.

15. A program for causing a computer to function as the image processing apparatus according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • View synthesis robust to unconstrained image data

    US11308659B2