Image processing apparatus, image processing method, and program
By focusing on high-frequency pattern regions and adjusting ray density, the image processing device enhances the estimation of three-dimensional information, improving virtual viewpoint image quality for complex objects.
Patent Information
- Application Number
- JP2024032877
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-05
- Publication Date
- 2025-09-18
AI Technical Summary
Existing technologies struggle to accurately estimate three-dimensional information for complex objects, leading to degraded image quality in virtual viewpoint images due to uniform low sampling density outside the depth of field.
An image processing device that extracts high-frequency pattern regions from captured images and generates groups of rays with higher density in these regions, using a learning mechanism to estimate three-dimensional information.
This approach allows for the generation of high-quality virtual viewpoint images while reducing the computational burden, particularly for objects with complex shapes and patterns.
Smart Images

Figure 2025135192000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to techniques for estimating three-dimensional information about a space containing an object. [Background technology]
[0002] There is a technology that estimates information about a space containing an object (hereinafter referred to as "three-dimensional information") using images (hereinafter referred to as "captured images") obtained by capturing an object from various directions. There is also a technology that uses the three-dimensional information to generate an image (hereinafter referred to as "virtual viewpoint image") that corresponds to the image of the object when viewed from an arbitrary virtual viewpoint (hereinafter referred to as "virtual viewpoint"). However, if the three-dimensional shape, color, or reflection characteristics of the object are complex, it may be difficult to accurately estimate the three-dimensional information about the space containing the object, and as a result, the image quality of the virtual viewpoint image may be degraded.
[0003] Patent Document 1 discloses a technology for learning radiance fields (radiance fields) representing color and density according to the position and direction in a space including an object as three-dimensional information using a captured image as a training image. Patent Document 1 also discloses a technology for generating a virtual viewpoint image by volume rendering using the radiance field estimated by the learning. Specifically, the technology disclosed in Patent Document 1 calculates learning parameters corresponding to the radiance field by sampling points on a light ray corresponding to each pixel of the training image and performing machine learning. More specifically, the technology disclosed in Patent Document 1 calculates learning parameters corresponding to the radiance field by setting the sampling density within the depth of field for the light ray corresponding to each pixel of the training image higher than the sampling density outside the depth of field. The technology disclosed in Patent Document 1 controls the sampling density based on the depth of field to reduce the amount of calculation required for estimating the radiance field, while improving the estimation accuracy of the radiance field in the space corresponding to an object within the depth of field, thereby improving the image quality of the virtual viewpoint image. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Publication No. 2023-66705 Summary of the Invention [Problem to be solved by the invention]
[0005] However, the technology disclosed in Patent Document 1 has the problem that the sampling density outside the depth of field of the light rays is uniformly low, making it impossible to estimate the radiance field in space corresponding to objects outside the depth of field with high accuracy.
[0006] Therefore, an object of the present disclosure is to provide a technique for estimating three-dimensional information that enables generation of a high-quality virtual viewpoint image while suppressing the amount of calculation required for estimating the three-dimensional information. [Means for solving the problem]
[0007] The image processing device according to the present disclosure includes an acquisition means for acquiring data of a captured image, an extraction means for extracting a high-frequency pattern region from the captured image, a generation means for generating a plurality of groups of rays based on rays corresponding to each of a plurality of pixels included in the captured image, the generation means generating the plurality of groups of rays including one or more groups of rays for a high-frequency pattern region, which are groups of rays in which the density of rays corresponding to pixels included in the high-frequency pattern region is higher than the density of rays corresponding to pixels included in regions other than the high-frequency pattern region, and a learning means for learning three-dimensional information of a space using the plurality of groups of rays. [Effects of the Invention]
[0008] According to the technology disclosed herein, it is possible to estimate three-dimensional information that allows for generating a high-quality virtual viewpoint image while suppressing the amount of calculation required for estimating the three-dimensional information. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 is a diagram illustrating an example of the configuration of an image processing system according to a first embodiment. [Figure 2] 1 is a block diagram showing an example of a hardware configuration of an image processing device according to a first embodiment. [Figure 3] 1 is a block diagram showing an example of a functional configuration of an image processing device according to a first embodiment. [Figure 4] 4 is a flowchart showing an example of a processing flow in the image processing device according to the first embodiment. [Figure 5] 2A and 2B are diagrams illustrating an example of the arrangement of an imaging device and a captured image according to the first embodiment. [Figure 6] 5A to 5C are diagrams for explaining extraction processing of a high-frequency pattern region in an extraction unit according to the first embodiment. [Figure 7] 10 is a flowchart showing an example of the flow of ray group generation processing in a ray group generation unit 303 according to the first embodiment. [Figure 8] 4A and 4B are diagrams for explaining an example of a group of rays generated by a group of rays generating unit according to the first embodiment. [Figure 9] 4A to 4C are diagrams for explaining an example of light rays corresponding to each piece of light ray information included in a group of light rays according to the first embodiment. [Figure 10] 10 is a flowchart showing an example of the flow of a learning process for three-dimensional information in a learning unit according to the first embodiment. [Figure 11] 10 is a flowchart showing an example of the flow of extraction processing of a high-frequency pattern region in an extraction unit according to the second embodiment. [Figure 12] 10A and 10B are diagrams for explaining extraction processing of a high-frequency pattern region in an extraction unit according to the second embodiment. [Figure 13] FIG. 10 is a diagram showing an example of a rough shape of an object obtained by a volume intersection method. [Figure 14] 10 is a flowchart showing an example of the flow of a ray group generation process in a ray group generation unit according to the second embodiment. [Figure 15]10A and 10B are diagrams for explaining an example of a group of rays generated by a group of rays generating unit according to the second embodiment. [Figure 16] 10A and 10B are diagrams for explaining an example of light rays corresponding to each piece of light ray information included in a group of light rays according to the second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. Note that the following embodiments do not necessarily limit the means for solving the problems according to the present disclosure. Furthermore, not all of the combinations of features described in the following embodiments are necessarily essential to the means for solving the problems according to the present disclosure.
[0011] [First embodiment] In this embodiment, a form will be described in which learning of a radiance field corresponding to a space including an object is performed based on data of captured images (hereinafter referred to as "captured image data") obtained by capturing images of an object from various directions using multiple imaging devices. Specifically, in the above-described learning according to this embodiment, a group of rays is used that is selected so that the light rays are highly dense in the space including the high-frequency pattern based on an image region (hereinafter referred to as a "high-frequency pattern region") that includes an image of a high-frequency pattern extracted from each captured image.
[0012] <Image processing system configuration> FIG. 1 is a diagram showing an example of the configuration of an image processing system according to a first embodiment. The image processing system includes multiple image capture devices 101, an image processing device 102, a user interface (hereinafter referred to as "UI") panel 103, a storage device 104, and a display device 105. The multiple image capture devices 101 are configured as digital still cameras, digital video cameras, or the like, and are arranged in different locations. Each image capture device 101 captures an object 107 present in an image capture area 106 from various directions in accordance with predetermined image capture conditions, and outputs captured image data obtained by the image capture to the image processing device 102. The captured image data obtained by image capture by the image capture device 101 may be still image data, moving image data, or both still image and moving image data. Hereinafter, the term "image" will be explained as including both "still image" and "moving image" unless otherwise specified.
[0013] The image processing device 102 acquires a plurality of captured image data output from a plurality of imaging devices 101, and performs learning of information (three-dimensional information) relating to the three-dimensional shape and color of a space including an object 107 present in an imaging area 106, which is generated using the acquired plurality of captured images. The image processing device 102 also generates a virtual viewpoint image based on the three-dimensional information obtained as a result of the learning. Note that the three-dimensional information within the imaging area 106 to be learned differs depending on the learning content. In the first embodiment, as an example, the three-dimensional information to be learned will be described as a radiance field.
[0014] 1, the description will be given assuming that each of the multiple imaging devices 101 and the image processing device 102 are connected to one another, but the method of connection between the imaging devices 101 and the image processing device 102 is not limited to this. Specifically, for example, the multiple imaging devices 101 may be cascade-connected by connecting adjacent imaging devices 101 to one another, and at least one of the multiple imaging devices 101 may be connected to the image processing device 102.
[0015] Furthermore, in the present embodiment, the description is given assuming that the multiple image capturing devices 101 are arranged at different positions, but the number and arrangement of the image capturing devices 101 are not limited to this. For example, if the position, shape, and color of the object 107 present in the image capturing area 106, as well as the intensity or hue of the ambient light, do not change over time, at least one image capturing device 101 whose position and orientation can be changed may be arranged. In this case, the image capturing device 101 may be caused to capture images at each of multiple different positions while changing the position and orientation of the image capturing device 101, and the image processing device 102 may acquire multiple captured image data obtained by the image capturing.
[0016] The UI panel 103 includes a display device such as a liquid crystal panel, and displays a user interface on the display device to present information such as the image capture conditions of the image capture device 101 and the processing settings of the image processing device 102 to the user. The UI panel 103 may also include an input device such as a touch panel or buttons, in which case the UI panel 103 accepts instructions from the user regarding changes to the image capture conditions, processing settings, etc. The input device may be provided separately from the UI panel 103, such as a mouse or a keyboard.
[0017] The storage device 104 is configured with a hard disk drive or the like, and stores data of the virtual viewpoint image output by the image processing device 102. When the image processing device 102 outputs three-dimensional information, the storage device 104 may store the three-dimensional information output from the image processing device 102. The display device 105 is configured with a liquid crystal display or the like, and acquires an image signal indicating the virtual viewpoint image output from the image processing device 102 and displays the virtual viewpoint image corresponding to the image signal. When the image processing device 102 outputs an image signal indicating three-dimensional information, the display device 105 may acquire an image signal indicating the three-dimensional information output from the image processing device 102 and display an image corresponding to the image signal. The imaging area 106 is a three-dimensional space surrounded by multiple imaging devices 101 installed in a studio or the like, and the frame indicated by a solid line in FIG. 1 indicates the outline of the imaging area 106 on the floor surface.
[0018] <Hardware configuration of image processing device> 2 is a block diagram showing an example of the hardware configuration of the image processing device 102 according to the first embodiment. The image processing device 102 has, as its hardware configuration, a CPU 201, a RAM 202, a ROM 203, a storage device 204, a control interface (hereinafter referred to as "I / F") 205, an input I / F 206, an output I / F 207, and a main bus 208. The CPU 201 is a processor that performs overall control of each unit of the image processing device 102. The RAM 202 functions as the main memory and work area of the CPU 201. The ROM 203 stores a group of programs executed by the CPU 201. The storage device 204 is configured by a hard disk drive or the like, and stores application programs executed by the CPU 201, data used in processing by the CPU 201, and the like.
[0019] The control I / F 205 is connected to each image capture device 101 and is a communication interface for setting image capture conditions for each image capture device 101, starting and stopping image capture, and other controls. The input I / F 206 is a communication interface using a serial bus such as SDI (Serial Digital Interface) or HDMI (High-Definition Multimedia Interface (registered trademark)). Captured image data output from each image capture device 101 is acquired via the input I / F 206. The output I / F 207 is a communication interface using a serial bus such as USB (Universal Serial Bus) or DisplayPort (registered trademark). Data or image signals such as virtual viewpoint images are output to the storage device 104 or the display device 105 via the output I / F 207. The main bus 208 is a transmission path that connects the above-mentioned hardware components of the image processing device 102 so that they can communicate with each other.
[0020] <Functional configuration of image processing device> 3 is a block diagram showing an example of the functional configuration of the image processing device 102 according to the first embodiment. The image processing device 102 includes an image acquisition unit 301, an extraction unit 302, a ray group generation unit 303, a learning unit 304, a viewpoint acquisition unit 305, an image generation unit 306, and an output unit 307. Each unit included in the functional configuration of the image processing device 102 is realized by the CPU 201 executing a program stored in the ROM 203, using the RAM 202 as a work memory. Note that not all of the processing of each unit included in the functional configuration of the image processing device 102, as described below, necessarily needs to be executed by the CPU 201; some or all of the processing may be executed by one or more processing circuits other than the CPU 201.
[0021] The image acquisition unit 301 acquires captured image data obtained by each imaging device 101 capturing an image of the imaging region 106, and imaging parameters (hereinafter referred to as "imaging parameters") of the imaging device 101 corresponding to the captured image. The extraction unit 302 extracts a high-frequency pattern region from the captured image acquired by the image acquisition unit 301. Details of the extraction process of the high-frequency pattern region by the extraction unit 302 will be described later. The ray group generation unit 303 generates a plurality of ray groups including ray groups that will result in a high density of light rays in a space including a high-frequency pattern, based on the captured image and imaging parameters acquired by the image acquisition unit 301 and the high-frequency pattern region extracted by the extraction unit 302. Details of the ray group generation process by the ray group generation unit 303 will be described later.
[0022] The learning unit 304 learns information (three-dimensional information) relating to the radiance field of the space including the object 107 based on the ray group generated by the ray group generation unit 303. Here, the three-dimensional information is, for example, network parameters in a learning model configured by a multi-layer perceptron (MLP) that expresses the radiance field of the space including the object 107. Details of the learning process for the three-dimensional information in the learning unit 304 will be described later.
[0023] The viewpoint acquisition unit 305 acquires information related to a virtual viewpoint (hereinafter referred to as "virtual viewpoint information"). Here, the virtual viewpoint information is information indicating the position of the virtual viewpoint and the line of sight direction at the virtual viewpoint, and is information equivalent to the imaging parameters (hereinafter referred to as "virtual camera parameters") of a virtual imaging device (hereinafter referred to as "virtual camera") arranged at the virtual viewpoint. The image generation unit 306 generates a virtual viewpoint image using three-dimensional information obtained as a result of learning by the learning unit 304, i.e., a learned model which is information related to the radiance field, and the virtual camera parameters acquired by the viewpoint acquisition unit 305. Specifically, the image generation unit 306 generates a virtual viewpoint image corresponding to the virtual viewpoint indicated by the virtual camera parameters by performing volume rendering using the three-dimensional information.
[0024] The output unit 307 outputs data of the virtual viewpoint image generated by the image generation unit 306 to the storage device 104, and stores the data in the storage device 104. The output unit 307 may output the virtual viewpoint image as an image signal to the display device 105, and display the virtual viewpoint image on the display device 105. The output unit 307 may also output three-dimensional information obtained as a result of learning by the learning unit 304, i.e., a learned model which is information about the radiance field, to the storage device 104, etc.
[0025] <Operation of image processing device> FIG. 4 is a flowchart showing an example of a processing flow in the image processing device 102 according to the first embodiment. Hereinafter, each processing step (process) is represented by adding an "S" to the beginning of the reference numeral. When each imaging device 101 outputs moving image data as captured image data, the image processing device 102 repeatedly executes the processing of the flowchart each time data of frames obtained by synchronized imaging included in each moving image is output from the imaging device 101. First, in S401, the image acquisition unit 301 acquires multiple captured image data obtained by the imaging and imaging parameters of the imaging device 101 corresponding to each captured image. Specifically, for example, the captured image data is acquired from the imaging device 101 via the input I / F 206, and the imaging parameters are acquired by reading out data calculated in advance by performing calibration or the like and stored in the storage device 204. The captured image data and imaging parameters acquired in S401 are held in the RAM 202.
[0026] 5A and 5B are diagrams showing an arrangement of the imaging devices 101 according to the first embodiment and examples of captured images 501 to 503 obtained by imaging using the imaging devices 101a to 101c. FIG. 5A shows an example of the arrangement of the imaging devices 101, where the imaging devices 101 are arranged so that they can capture images of an object 107 present in an imaging area 106 from various directions. Note that the imaging devices 101a, 101b, and 101c shown in FIG. 5A are the same as the other imaging devices 101. FIG. 5B shows an example of a captured image 501 obtained by imaging using the imaging device 101a, FIG. 5C shows an example of a captured image 502 obtained by imaging using the imaging device 101b, and FIG. 5D shows an example of a captured image 503 obtained by imaging using the imaging device 101c.
[0027] After S401, in S402, the extraction unit 302 extracts a high-frequency pattern region from each captured image acquired in S401 by performing extraction processing of a high-frequency pattern region on each captured image acquired in S401. The extraction unit 302 uses, for example, a differential filter to identify an image region from each captured image where the density of edges, such as stripes, is higher than, for example, a predetermined density, and extracts the identified image region as a high-frequency pattern region. Specifically, for example, the extraction unit 302 applies a Laplacian filter to the captured image to first extract, as an edge region, a plurality of pixels whose absolute value of the filter output is greater than a predetermined threshold. Next, the extraction unit 302 sequentially applies a filling process by a closing process and a small component removal process by an opening process to the extracted edge region, and extracts the result obtained as a high-frequency pattern region.
[0028] FIG. 6 is a diagram illustrating the extraction process of a high-frequency pattern region in the extraction unit 302 according to the first embodiment. FIG. 6(a) shows an example of a captured image 600 obtained by capturing an image using a certain imaging device 101. The captured image 600 shown in FIG. 6(a) includes, as an example, images 601 to 603 of three objects present in the captured image region 106. FIG. 6(b) shows an example of an image 610 (hereinafter referred to as an "edge image") showing an edge region 611 extracted from the captured image 600 by the extraction unit 302. In the edge image 610 shown as an example in FIG. 6(b), pixels corresponding to the edge region 611 are shown in white, and pixels corresponding to regions other than the edge region 611 are shown in black. FIG. 6(c) shows an example of an image 620 (hereinafter referred to as a "high-frequency pattern image") showing a high-frequency pattern region 621 extracted from the edge image 610 by the extraction unit 302. In the high-frequency pattern image 620 shown as an example in FIG. 6(c), pixels corresponding to the high-frequency pattern region 621 are shown in white, and pixels corresponding to regions other than the high-frequency pattern region 621 are shown in black.
[0029] After S402, in S403, the ray group generation unit 303 executes a ray group generation process to generate multiple ray groups based on the captured image and imaging parameters acquired in S401 and the high-frequency pattern area extracted in S402. Specifically, the ray group generation unit 303 generates multiple ray groups including a ray group that will result in a high density of light rays in a space including a high-frequency pattern. Details of the ray group generation process in the ray group generation unit 303 will be described later. Next, in S404, the learning unit 304 executes a three-dimensional information learning process to learn three-dimensional information of a space including an object, i.e., a learning model that represents the radiance field of the space, based on the multiple ray groups generated in S403. Details of the three-dimensional information learning process in the learning unit 304 will be described later.
[0030] Next, in S405, the learning unit 304 determines whether or not to end the learning process for the three-dimensional information. Specifically, the learning unit 304 compares the captured image acquired in S401 with an image generated by volume rendering based on the imaging parameters corresponding to the captured image and the three-dimensional information (learning model representing the radiance field) being learned in S404. If the result of the comparison shows that the difference between the captured image and the image generated by volume rendering is smaller than a predetermined threshold, the learning unit 304 determines to end the learning process for the three-dimensional information. If the difference is greater than the predetermined threshold, the learning unit 304 determines not to end the learning process for the three-dimensional information. If it is determined not to end the learning process for the three-dimensional information in S405, the image processing device 102 returns to the process of S403 and repeatedly executes the processes from S403 to S405 until it is determined to end the learning process for the three-dimensional information in S405.
[0031] If it is determined in S405 that the learning process for three-dimensional information is to be terminated, in S406, the viewpoint acquisition unit 305 acquires the virtual camera parameters set based on instructions from the user using the UI panel 103. The method of acquiring the virtual camera parameters by the viewpoint acquisition unit 305 is not limited to the above-described method. For example, the viewpoint acquisition unit 305 may acquire the virtual camera parameters by reading out virtual camera parameters that have been set in advance and stored in the storage device 204 or the like.
[0032] After S406, in S407, the image generation unit 306 generates a virtual viewpoint image using the virtual camera parameters acquired in S406 and the three-dimensional information (trained model representing the radiance field) obtained as a result of the learning process in S404. Specifically, the image generation unit 306 generates a virtual viewpoint image corresponding to the virtual viewpoint indicated by the virtual camera parameters by performing volume rendering of the three-dimensional information (trained model representing the radiance field) based on the virtual camera parameters.
[0033] Next, in S408, the output unit 307 outputs the virtual viewpoint image generated in S407. Specifically, for example, the output unit 307 outputs data of the virtual viewpoint image or an image signal indicating the virtual viewpoint image to the storage device 104, the display device 105, or the like via the output I / F 207. After S408, the image processing device 102 ends the processing of the flowchart shown in Fig. 4. Note that, as described above, when each imaging device 101 outputs moving image data as captured image data, the image processing device 102 returns to S401 after S408 and repeatedly executes the processing of the flowchart.
[0034] <Generation of ray groups> 7 is a flowchart showing an example of the flow of ray group generation processing in the ray group generation unit 303 according to the first embodiment, and is a flowchart showing an example of the detailed processing flow of S403 shown in Fig. 4. In S403, multiple ray groups including ray groups that will result in a high density of rays in a space including a high-frequency pattern are generated based on the captured image and imaging parameters acquired in S401 and the high-frequency pattern area extracted in S402. The processing of this flowchart is executed after the processing of S402.
[0035] After S402, first, in S701, the ray group generation unit 303 acquires information about light rays corresponding to each pixel of the captured image (hereinafter referred to as "ray information") based on the captured image and imaging parameters acquired in S401. In this embodiment, the light ray information includes values related to the color, starting point, and direction of the light ray, and the ray group generation unit 303 acquires the light ray information by calculating these values. The ray group generation unit 303 acquires a value related to the color of the light ray, assuming that the color corresponds to the value of the corresponding pixel. Furthermore, the ray group generation unit 303 acquires a value related to the starting point of the light ray, assuming that the starting point is the position of the corresponding imaging device 101. Furthermore, the ray group generation unit 303 acquires a value related to the direction of the light ray by performing a calculation using, for example, the following equation (1) based on the imaging parameters corresponding to the captured image and the coordinates of the pixel in the captured image (hereinafter referred to as "image coordinates").
[0036] d=((uc x ) / f x ,(vc y ) / f y ,1) Equation (1) Here, d is a vector representing the direction of the ray corresponding to the image coordinates, where the two axes that are orthogonal to each other in the image are the X-axis and the Y-axis, and the axis that is orthogonal to the X-axis and the Y-axis is the Z-axis. (u, v) are the image coordinates, and (c x ,c y ) is the image coordinate corresponding to the center position of the image. xis a value obtained by multiplying the number of light receiving elements per unit length in the X-axis direction of the image sensor of the image pickup device 101 by a value indicating the focal length of the optical system of the image pickup device 101. y is a value obtained by multiplying the number of light receiving elements per unit length in the Y-axis direction of the image sensor of the image pickup device 101 by a value indicating the focal length of the optical system of the image pickup device 101. (f x ,f y ) is also called a focal length as an internal parameter of the image capture device 101. Note that the internal parameters will be described as being included in the image capture parameters.
[0037] After S701, in S702, the ray group generation unit 303 generates multiple ray groups based on ray information corresponding to the image coordinates of each pixel included in the high-frequency pattern area (hereinafter simply referred to as "ray information corresponding to the high-frequency pattern area"). Note that one ray group includes multiple pieces of ray information. In this embodiment, multiple ray groups are generated so that identical ray information corresponding to the high-frequency pattern area does not overlap with each other among the ray groups. Specifically, the ray group generation unit 303 randomly selects a predetermined number of pieces of ray information from the ray information corresponding to the high-frequency pattern area to form one ray group. More specifically, the ray group generation unit 303 generates multiple ray groups so that identical ray information is not selected multiple times among the multiple pieces of ray information corresponding to the high-frequency pattern area. The ray group generation unit 303 repeatedly executes the above process until the number of unselected ray information in the ray information corresponding to the high-frequency pattern area becomes less than the predetermined number. As a result, multiple ray groups are generated so that identical ray information corresponding to the high-frequency pattern area does not overlap with each other among the ray groups. Note that unselected light ray information in the light ray information corresponding to the high frequency pattern region may be used as other light ray information in the process of S703 described later.
[0038] FIG. 8 is a diagram illustrating an example of a ray group generated by the ray group generation unit 303. Specifically, FIG. 8(a) is a diagram illustrating an example of a ray group according to the first embodiment, and FIG. 8(b) is a diagram illustrating an example of a ray group according to Modification 1 of the first embodiment, which will be described later. In FIG. 8, a gray rectangle 841 represents ray information corresponding to a high-frequency pattern region. A white rectangle 842 represents ray information (hereinafter simply referred to as "other ray information") corresponding to the image coordinates of each pixel included in a region other than the high-frequency pattern region (hereinafter referred to as "other region"). A set of ray information 851 is a set of all ray information corresponding to the high-frequency pattern region, and represents a set of ray information that the ray group generation unit 303 can select when generating multiple ray groups in the process of S702. A set of ray information 852 represents a set of all other ray information. A set of light ray information 850 represents a set of light ray information corresponding to the image coordinates of all pixels in the captured image, including all light ray information corresponding to high frequency pattern regions and all other light ray information.
[0039] By executing the process of S702, the ray group generation unit 303 generates a plurality of ray groups 801 to 803, each of which includes only ray information corresponding to a high-frequency pattern region, as shown as an example in Fig. 8(a). The ray corresponding to each piece of ray information included in the ray groups 801 to 803 thus generated intersects with an object including a high-frequency pattern.
[0040] FIG. 9 is a diagram illustrating an example of a light ray 904 corresponding to each piece of light ray information included in the light ray groups 801 to 803 shown in FIG. 8(a). Specifically, FIG. 9(a) is a diagram illustrating the imaging region 106 viewed vertically from above, and illustrates an example of a light ray 904 corresponding to each piece of light ray information included in the light ray groups 801 to 803 shown in FIG. 8(a). In FIG. 9(a), object 901 is an object that includes a high-frequency pattern, and objects 902 and 903 are objects that do not include a high-frequency pattern. FIG. 9(b) illustrates an example of a captured image 910 obtained by imaging using the imaging device 101c. The captured image 910 shown in FIG. 9(b) includes images 911 to 913 of three objects 901 to 903 that exist in the imaging region 106.
[0041] As described above, the ray 904 corresponding to each piece of ray information included in the ray groups 801 to 803 intersects with the object 903 including a high-frequency pattern. Therefore, as shown in FIG. 9( a), the density of the rays passing through the space including the high-frequency pattern is higher than the density of the rays passing through other spaces. Therefore, it is possible to perform learning using a high-density ray group for the space corresponding to the object 901 including a high-frequency pattern. As a result, it is possible to estimate the three-dimensional information (radiance field) of the space corresponding to the object 901 including a high-frequency pattern with high accuracy.
[0042] After S702, in S703, the ray group generation unit 303 generates multiple ray groups based on the other ray information. In this embodiment, multiple ray groups are generated so that identical ray information in the other ray information does not overlap with each other among the ray groups. Specifically, a predetermined number of ray information pieces randomly selected from the other ray information are used as one ray group. More specifically, the ray group generation unit 303 generates multiple ray groups so that identical ray information is not selected multiple times among the other ray information. The ray group generation unit 303 repeatedly performs the above process until the number of unselected ray information pieces in the other ray information falls below a predetermined number. This generates multiple ray groups in which identical ray information in the other ray information does not overlap with each other among the ray groups. Note that ray groups 804 to 812 shown in FIG. 8(a) are examples of ray groups including only other ray information, generated by the ray group generation unit 303 performing the process of S703.
[0043] After S703, the ray group generation unit 303 ends the processing of the flowchart shown in Fig. 7, i.e., the processing of S403 shown in Fig. 4. That is, by executing the processing of S403, the ray group generation unit 303 generates a plurality of ray groups 801 to 812 from a set 850 of ray information corresponding to the image coordinates of all pixels included in the captured image, as shown as an example in Fig. 8(a).
[0044] <Learning processing of three-dimensional information> FIG. 10 is a flowchart showing an example of the flow of learning processing for three-dimensional information in the learning unit 304 according to the first embodiment, and is a flowchart showing an example of the detailed processing flow of S404 shown in FIG. 4. In S404, a rendering value of each ray group is calculated by volume rendering based on the multiple ray groups generated in S403, and the three-dimensional information is updated so that the difference between the rendering value of the ray and the value indicating the color of the ray is reduced. In this embodiment, the three-dimensional information is described as a function (learning model) representing a radiance field configured by a multilayer perceptron, which receives information indicating a position and a vector encoding a direction as input, and outputs values indicating the corresponding density and color. The processing of this flowchart is executed after the processing of S403.
[0045] After S403, first, in S1001, the learning unit 304 selects an arbitrary ray group from the multiple ray groups generated in S403. Specifically, the learning unit 304 randomly selects an arbitrary ray group from the multiple ray groups that have not yet been selected. Hereinafter, the ray group selected in S1001 will be referred to as the "ray group of interest." Next, in S1002, the learning unit 304 calculates rendering values of rays corresponding to each piece of ray information by volume rendering using multiple pieces of ray information included in the ray group of interest. Specifically, in S1002, first, for each piece of ray information included in the ray group of interest, the learning unit 304 sets multiple sampling points on the ray based on a value related to the start point of the ray and a value related to the direction of the ray, which are included in the ray information. Next, in S1002, the learning unit 304 performs calculations using, for example, the following equations (2) and (3), to obtain the density and color corresponding to the position of the sampling point and the direction of the ray based on the learning model representing the radiance field, and calculates the drawing value of the ray corresponding to the ray information.
[0046] TIFF2025135192000002.tif20150 where C(r) is the drawing value of the ray r, i and j are the indices of the sampling points, and N is the total number of sampling points. i is a value indicating the density of sampling points, and ci is a value indicating the color of the sampling point, and δ i is a value indicating the distance to the next sampling point.
[0047] After S1002, in S1003, the learning unit 304 updates the three-dimensional information for each piece of light ray information included in the target light ray group so as to reduce the difference between the rendering value of the light ray calculated in S1002 and the value related to the color of the light ray included in the light ray information. Specifically, the learning unit 304 updates the three-dimensional information by updating the parameters of a function (learning model) representing the radiance field so as to reduce the difference. The process of calculating the difference between the rendering value of these light rays and the value related to the color of the light ray and the process of updating the learning model correspond to the error calculation process and error propagation process in deep learning. Note that the difference between the rendering value of a light ray and the value related to the color of the light ray is defined by the squared Euclidean distance of color values defined by, for example, R (Red), G (Green), and B (Blue), etc.
[0048] Next, in S1004, the learning unit 304 determines whether or not all ray groups have been selected in S1001. If it is determined in S1004 that not all ray groups have been selected, that is, that there are ray groups that have not yet been selected, the learning unit 304 returns to S1001 and repeatedly executes the processes from S1001 to S1004. If it is determined in S1004 that all ray groups have been selected, the learning unit 304 ends the process of the flowchart shown in Fig. 10, that is, the process of S404 shown in Fig. 4.
[0049] <Advantages of the image processing device according to the first embodiment> As described above, the image processing device 102 learns a radiance field of a space including an object based on multiple captured image data obtained by capturing images of the object from various directions using multiple imaging devices. In particular, in this embodiment, the image processing device 102 is configured to learn a radiance field, which is three-dimensional information of the captured region, using a group of rays selected based on a high-frequency pattern region extracted from the captured image so that the light rays are densely packed in the space including the high-frequency pattern. The image processing device 102 configured in this manner can estimate three-dimensional information that enables generation of a high-quality virtual viewpoint image while suppressing the amount of calculation required for estimating the three-dimensional information. In particular, the image processing device 102 can estimate three-dimensional information of a space corresponding to an object including a high-frequency pattern with high accuracy. As a result, the image quality of a virtual viewpoint image generated based on the estimated three-dimensional information can be improved.
[0050] [Modification 1 of the First Embodiment] In the process of S402, the extraction unit 302 according to the first embodiment extracts a high-frequency pattern region from a captured image by performing filtering using a Laplacian filter, closing processing, and opening processing. However, the method for extracting a high-frequency pattern region by the extraction unit 302 is not limited to the above-described method. For example, the extraction unit 302 may extract edges from a captured image using a differential filter such as a Sobel filter instead of a Laplacian filter. Furthermore, instead of performing closing processing and opening processing, the extraction unit 302 may identify a local region containing a high proportion of edge regions and extract the identified local region as a high-frequency pattern region. Furthermore, for example, an image including an image corresponding to a specific object defined as a high-frequency pattern may be prepared as a template image in advance, and the extraction unit 302 may extract a high-frequency pattern region by template matching using the template image. Furthermore, the extraction unit 302 may use a detector that detects an image corresponding to a specific object defined as a high-frequency pattern from an input image and extract the detection result of the detector as a high-frequency pattern region.
[0051] Furthermore, in the processing of S702, the ray group generation unit 303 according to the first embodiment generates a ray group consisting of only ray information corresponding to high-frequency pattern regions, but the method of generating a ray group in the ray group generation unit 303 is not limited to the above-described method. For example, the ray group generation unit 303 may generate a ray group including ray information corresponding to high-frequency pattern regions and other ray information so that the number of pieces of ray information corresponding to high-frequency pattern regions is equal to or greater than a predetermined number.
[0052] Specifically, for example, the number of pieces of light ray information corresponding to high-frequency pattern regions and the number of pieces of other light ray information to be included in one light ray group are set in advance. The light ray group generation unit 303 generates a light ray group by randomly selecting light ray information from all of the light ray information corresponding to high-frequency pattern regions and all of the other light ray information based on these preset numbers. For example, the number of pieces of light ray information corresponding to high-frequency pattern regions to be included in one light ray group is set in advance so that the ratio of the number of pieces of light ray information corresponding to high-frequency pattern regions to the number of all light ray information to be included in one light ray group is equal to or greater than a predetermined ratio. Here, the predetermined ratio is, for example, the ratio of the number of pixels included in the high-frequency pattern region to the total number of pixels in the captured image.
[0053] Fig. 8(b) shows examples of ray groups 813 to 816 that include ray information corresponding to a high-frequency pattern region and other ray information that are generated by the above-described processing by the ray group generation unit 303. Note that ray groups 817 to 824 shown in Fig. 8(b) are examples of ray groups that include only other ray information, and are generated by the ray group generation unit 303 executing the processing of S703 after the ray group generation unit 303 generates the ray groups 813 to 816.
[0054] Furthermore, in the processes of S702 and S703, the ray group generation unit 303 according to the first embodiment generates multiple ray groups so that the same ray information does not overlap among the ray groups. However, the ray group generation unit 303 may generate multiple ray groups while allowing the same ray information to overlap among the ray groups. In this case, the total number of ray groups may be the same as the number of cases in which the same ray information does not overlap among the ray groups.
[0055] Furthermore, the ray group generation unit 303 according to the first embodiment calculates ray information corresponding to each pixel of the captured image in S701. However, the ray group generation unit 303 may store the ray information calculated in the first execution of S701 in the RAM 202 or the like, and acquire the ray information stored in the RAM 202 or the like when executing S701 for the second or subsequent times.
[0056] Furthermore, although the three-dimensional information according to the first embodiment is a radiance field, and the radiance field has been described as a function (learning model) represented by a multilayer perceptron, the method of representing the radiance field is not limited to this. For example, the radiance field may be represented using multiple multilayer perceptrons, a sparse three-dimensional grid including spherical harmonic functions, or a tensor.
[0057] Although the learning unit 304 according to the first embodiment has been described as training a learning model representing a radiance field, the learning unit 304 may also train a learning model representing three-dimensional information that can be learned based on a group of light rays. For example, a radiance field represents color and density depending on position and direction, but the three-dimensional information is not limited to this. Specifically, for example, a color corresponding to a position in space in the three-dimensional information may be an isotropic color that is independent of direction. Furthermore, the three-dimensional information may be, for example, a density field representing density depending on position, a field represented by a bidirectional reflectance distribution function representing the distribution characteristics of reflected light relative to incident light, or a field representing the amount of ambient light penetration (light visibility). Furthermore, the three-dimensional information may be a field representing color and density depending on position, direction, and time. In this case, the group of light rays used to train the three-dimensional information is generated based on captured images as a moving image including time-series frames.
[0058] [Second embodiment] An image processing device 102 according to the second embodiment will be described with reference to FIGS. 2 to 4 and 11 to 16. Like the image processing device 102 according to the first embodiment, the image processing device 102 according to this embodiment has a hardware configuration and a functional configuration as shown, for example, in the block diagram of FIG. 2 or 3. Like the image processing device 102 according to the first embodiment, the image processing device 102 according to this embodiment executes the processing of the flowchart shown, for example, in FIG. 4. However, in this embodiment, the processing of the extraction unit 302 and the ray group generation unit 303 differs from that of the extraction unit 302 and the ray group generation unit 303 according to the first embodiment. That is, the processing of extracting a high-frequency pattern region in S402 and the processing of generating a ray group in S403 according to this embodiment differs from that of S402 and S403 according to the first embodiment.
[0059] Specifically, the extraction unit 302 according to the first embodiment extracts high-frequency pattern regions from a captured image using a differential filter such as a Laplacian filter in the high-frequency pattern region extraction process of S402. In contrast, the extraction unit 302 according to the present embodiment extracts, from the captured image, an image region corresponding to the outline shape of each object containing a high-frequency pattern, as a high-frequency pattern region, based on the outline shape of the object. Furthermore, the ray group generation unit 303 according to the present embodiment generates a ray group for each high-frequency pattern region extracted by the extraction unit 302. The following mainly describes the high-frequency pattern region extraction process in the extraction unit 302 and the ray group generation process in the ray group generation unit 303, which are different from those in the first embodiment. Note that the same reference numerals are used to designate configurations or processing steps (steps) that perform the same processes as those in the first embodiment, and their description will be omitted.
[0060] <Extraction of high-frequency pattern areas> 11 is a flowchart showing an example of the flow of processing for extracting a high-frequency pattern area in the extraction unit 302 according to the second embodiment, and is a flowchart showing an example of a detailed processing flow in S402 shown in Fig. 4. In S402 according to the second embodiment, an image area corresponding to a rough shape including a high-frequency pattern is extracted as a high-frequency pattern area from the captured image acquired in S401 based on the rough shape of the object. The processing of this flowchart is executed after the processing of S401 shown in Fig. 4.
[0061] After S401, first, in S1101, the extraction unit 302 extracts an area with a high edge density from each captured image acquired in S401, as in the first embodiment. FIG. 12 is a diagram for explaining the extraction process of a high-frequency pattern area by the extraction unit 302 according to the second embodiment. FIG. 12(a) shows an example of a captured image 1200 obtained by capturing an image using a certain imaging device 101. As an example, the captured image 1200 shown in FIG. 12(a) includes images 1201 to 1203 of three objects present in the imaging area 106. Furthermore, among the images 1201 to 1203 included in the captured image 1200, the object corresponding to image 1201 (hereinafter referred to as "object A") and the object corresponding to image 1203 (hereinafter referred to as "object B") include high-frequency patterns.
[0062] Fig. 12(b) shows an example of an image 1210 showing the region with high edge density extracted in S1101. As an example, the image 1210 shown in Fig. 12(b) is shown as a binary image in which the pixel value of the region with high edge density is set to 1 and the pixel value of the region with low edge density is set to 0. In the above example, in S1101, a region 1211 corresponding to the image 1201 of object A and a region 1213 corresponding to the image 1203 of object B are extracted as regions with high edge density. Figs. 12(c) and 12(d) will be described later.
[0063] After S1101, in S1102, the extraction unit 302 acquires a rough shape of an object based on the captured image acquired in S401 and the imaging parameters. In this embodiment, the extraction unit 302 acquires a rough shape of an object expressed as a set of voxels using a volume intersection method. In this case, for example, for each captured image acquired in S401, the extraction unit 302 first acquires a silhouette image of the object based on the difference between the captured image and a background image obtained by capturing only the background in a state in which no object is present. Here, the background image may be captured in advance, and the extraction unit 302 may acquire the background image data stored in advance in the storage device 204 or the like. Since the method of acquiring a silhouette image of an object is well known, a detailed description thereof will be omitted.
[0064] Next, the extraction unit 302 projects each voxel included in the set of voxels corresponding to the imaging region onto a silhouette image based on the imaging parameters acquired in S401. Next, the extraction unit 302 acquires a set of voxels projected onto the silhouette region of the object for all silhouette images as a rough shape of the object. FIG. 13 is a diagram showing an example of a rough shape of an object 1301 acquired by the visual volume intersection method. Note that the rectangle surrounded by a thin solid line indicates the actual contour 1302 of the object 1301, and the polygon surrounded by a thick solid line indicates the contour 1303 of the rough shape of the object 1301. Methods for acquiring a rough shape of an object based on a silhouette image of the object, such as the visual volume intersection method, are well known, and therefore will not be described in detail.
[0065] After S1102, in S1103, the extraction unit 302 extracts a schematic shape of an object including a high-frequency pattern from the schematic shapes of the multiple objects acquired in S1102. Specifically, for example, first, the extraction unit 302 projects each voxel corresponding to the surface of the schematic shape onto each captured image using the corresponding imaging parameters. Next, the extraction unit 302 extracts, as the schematic shape of the object including a high-frequency pattern, a schematic shape in which the number of voxels projected onto an area of high edge density in the captured image without the schematic shape of the object being occluded by the schematic shapes of other objects is equal to or greater than a predetermined value.
[0066] Next, in S1104, the extraction unit 302 extracts, as a high-frequency pattern area, an image area in the captured image that corresponds to the outline shape of the object including the high-frequency pattern extracted in S1103. Specifically, the extraction unit 302 extracts, as a high-frequency pattern area, a set of pixels where a ray corresponding to each pixel intersects with the outline shape of the object including the high-frequency pattern without being blocked by the outline shape of another object. After S1104, the extraction unit 302 ends the processing of the flowchart shown in FIG. 11, i.e., the processing of S402 according to this embodiment.
[0067] Fig. 12(c) shows an example of the relationship between schematic shapes 1224 and 1225 of objects containing high-frequency patterns extracted in S1103 and high-frequency pattern areas 1226 and 1227 extracted in S1104. In Fig. 12(c), objects 1221 to 1223 are objects present in the imaging area 106, and schematic shapes 1224 and 1225 are, respectively, the schematic shapes of objects 1221 and 1223 containing high-frequency patterns. Also in Fig. 12(c), high-frequency pattern areas 1226 and 1227 are, respectively, high-frequency pattern areas corresponding to the schematic shapes 1224 and 1225 of objects containing high-frequency patterns. Note that objects 1221 and 1223 correspond, respectively, to objects A and B described above.
[0068] Fig. 12(d) shows an example of an image 1230 representing the high-frequency pattern regions 1226 and 1227 extracted in S1104. In the image 1230 shown in Fig. 12(d), the high-frequency pattern region 1226 corresponds to the outline shape 1224 of the object 1221 including the high-frequency pattern. Furthermore, the high-frequency pattern region 1227 corresponds to the outline shape 1225 of the object 1223 including the high-frequency pattern. In the following description, the high-frequency pattern region 1226 will be referred to as "high-frequency pattern region A" and the high-frequency pattern region 1227 will be referred to as "high-frequency pattern region B."
[0069] <Generation of ray groups> Fig. 14 is a flowchart showing an example of the flow of ray group generation processing in the ray group generation unit 303 according to the second embodiment, and is a flowchart showing an example of the detailed processing flow of S403 shown in Fig. 4. In S403 according to this embodiment, a ray group is generated for each high-frequency pattern region extracted in S402 according to this embodiment. The processing of this flowchart is executed after the processing of S402 according to this embodiment.
[0070] After S402 according to this embodiment, first, the ray group generation unit 303 executes processing similar to S701 shown in FIG. 7 to acquire light ray information corresponding to each pixel of the captured image based on the captured image and imaging parameters acquired in S401. Next, in S1401, the ray group generation unit 303 selects an arbitrary high-frequency pattern area from one or more high-frequency pattern areas extracted in S402. Specifically, the ray group generation unit 303 randomly selects an arbitrary high-frequency pattern area from high-frequency pattern areas that have not yet been selected among the one or more high-frequency pattern areas. Hereinafter, the high-frequency pattern area selected in S1401 will be referred to as a "region of interest."
[0071] Next, in S1402, the ray group generation unit 303 generates multiple ray groups based on ray information corresponding to the image coordinates of each pixel included in the region of interest (hereinafter referred to as "ray information corresponding to the region of interest"). Note that, as in the first embodiment, one ray group includes multiple pieces of ray information. Furthermore, in this embodiment, as in the first embodiment, multiple ray groups are generated so that identical ray information corresponding to the region of interest does not overlap with each other among the ray groups. Specifically, the ray group generation unit 303 randomly selects a predetermined number of pieces of ray information from the ray information corresponding to the region of interest as one ray group. More specifically, the ray group generation unit 303 generates multiple ray groups so that identical ray information is not selected multiple times in the ray information corresponding to the region of interest. The ray group generation unit 303 repeatedly executes the above process until the number of unselected ray information in the ray information corresponding to the region of interest falls below a predetermined number. As a result, multiple ray groups are generated so that identical ray information corresponding to the region of interest does not overlap with each other among the ray groups. Note that unselected ray information in the ray information corresponding to the region of interest may be used as another ray group in the process of S1405 described later.
[0072] Fig. 15 is a diagram illustrating an example of a ray group generated by the ray group generating unit 303. Specifically, Fig. 15(a) is a diagram illustrating an example of a ray group according to the second embodiment, and Fig. 15(b) is a diagram illustrating an example of a ray group according to Modification 1 of the second embodiment, which will be described later. In Fig. 15, a rectangle 1541 filled in dark gray represents ray information corresponding to the high-frequency pattern region A, and a rectangle 1542 filled in light gray represents ray information corresponding to the high-frequency pattern region B. Furthermore, a rectangle 1543 filled in white represents ray information (other ray information) corresponding to the image coordinates of each pixel included in regions other than the high-frequency pattern regions A and B (other regions).
[0073] Furthermore, the set of light ray information 1551 is a set of all light ray information corresponding to the high-frequency pattern region A. Specifically, for example, the set of light ray information 1551 represents a set of light ray information that the light ray group generation unit 303 can select when generating multiple light ray groups in the processing of S1402 when the target region is the high-frequency pattern region A. Furthermore, the set of light ray information 1552 is a set of all light ray information corresponding to the high-frequency pattern region B. Specifically, for example, the set of light ray information 1552 represents a set of light ray information that the light ray group generation unit 303 can select when generating multiple light ray groups in the processing of S1402 when the target region is the high-frequency pattern region B. Furthermore, the set of light ray information 1553 represents a set of all other light ray information. Furthermore, the set of light ray information 1550 represents a set of light ray information corresponding to the image coordinates of all pixels included in the captured image, including all light ray information corresponding to the high-frequency pattern regions A and B and all other light ray information.
[0074] When the region of interest is high-frequency pattern region A, the ray group generation unit 303 executes the process of S1402 to generate multiple ray groups 1501 to 1503, each of which includes only ray information corresponding to high-frequency pattern region A, as shown as an example in Fig. 15(a). When the region of interest is high-frequency pattern region B, the ray group generation unit 303 executes the process of S1402 to generate multiple ray groups 1504 to 1506, each of which includes only ray information corresponding to high-frequency pattern region B, as shown in Fig. 15(a). Light rays corresponding to each piece of ray information included in the ray groups 1501 to 1506 generated in this manner intersect with object A or object B, which includes a high-frequency pattern.
[0075] FIG. 16 is a diagram illustrating an example of light rays 1605 corresponding to each piece of light ray information included in the light ray groups 1501 to 1503 shown in FIG. 15(a). Specifically, FIG. 16(a) is a diagram illustrating the imaging region 106 viewed vertically from above, and illustrates an example of light rays 1605 corresponding to each piece of light ray information included in the light ray groups 1501 to 1503 shown in FIG. 15(a). In FIG. 15(a), objects 1601 and 1603 are objects that include high-frequency patterns, and object 1602 is an object that does not include a high-frequency pattern. FIG. 15(b) illustrates an example of a captured image 1610 obtained by imaging using the imaging device 101c. The captured image 1610 shown in FIG. 15(b) includes images 1611 to 1613 of three objects 1601 to 1603 that exist in the imaging region 106.
[0076] As described above, the ray 1605 corresponding to each piece of ray information included in the ray groups 1501 to 1503 intersects with the object 1601 containing a high-frequency pattern. Furthermore, although not shown in FIG. 16( a), the ray corresponding to each piece of ray information included in the ray groups 1504 to 1506 intersects with the object 1603 containing a high-frequency pattern. Therefore, even if multiple objects containing high-frequency patterns exist, the density of the ray passing through the space containing the high-frequency pattern is higher than the density of the ray passing through other spaces. Therefore, learning using a high-density ray group is possible for the spaces corresponding to the objects 1601 and 1603 containing high-frequency patterns. As a result, it is possible to estimate the three-dimensional information (radiance field) of the space corresponding to the object containing a high-frequency pattern with high accuracy.
[0077] After S1402, in S1403, the ray group generation unit 303 determines whether all high-frequency pattern areas extracted in S402 according to this embodiment have been selected in S1402. If it is determined in S1403 that not all high-frequency pattern areas have been selected, that is, that there are high-frequency pattern areas that have not yet been selected, the ray group generation unit 303 returns to S1401 and repeatedly executes the processes from S1401 to S1403. If it is determined in S1403 that all high-frequency pattern areas have been selected, the ray group generation unit 303 executes the process of S1404. In S1404, the ray group generation unit 303 executes the same process as S703 shown in FIG. 7 to generate multiple ray groups based on other ray information. Note that ray groups 1507 to 1512 shown in FIG. 15(a) are examples of ray groups that include only other ray information, generated by the ray group generation unit 303 executing the process of S1404.
[0078] After S1404, the ray group generation unit 303 ends the processing of the flowchart shown in Fig. 14, i.e., the processing of S403 according to this embodiment. That is, by executing the processing of S403, the ray group generation unit 303 according to this embodiment generates multiple ray groups 1501 to 1512 from a set 1550 of ray information corresponding to the image coordinates of all pixels included in the captured image, as shown in Fig. 15(a).
[0079] [Modification 1 of the second embodiment] In the process of S1402, the ray group generation unit 303 according to the second embodiment generates a ray group consisting of only ray information corresponding to the high-frequency pattern region set as the region of interest. However, the method of generating a ray group in the ray group generation unit 303 is not limited to the method described above. For example, the ray group generation unit 303 may generate a ray group consisting of ray information corresponding to the region of interest and other ray information so that the number of pieces of ray information corresponding to the region of interest is equal to or greater than a predetermined number.
[0080] Specifically, for example, the number of pieces of light ray information corresponding to the region of interest and the number of other pieces of light ray information to be included in one light ray group are set in advance. The light ray group generation unit 303 generates a light ray group by randomly selecting light ray information from all of the light ray information corresponding to the region of interest and all of the other light ray information based on these preset numbers. For example, the number of pieces of light ray information corresponding to the region of interest to be included in one light ray group is set in advance so that the ratio of the number of light ray information corresponding to the region of interest to the number of all light ray information to be included in one light ray group is equal to or greater than a predetermined ratio. Here, the predetermined ratio is, for example, the ratio of the number of pixels included in the region of interest to the total number of pixels in the captured image.
[0081] Fig. 15(b) shows an example of light ray groups 1513 to 1516 that include light ray information corresponding to high-frequency pattern region A and other light ray information, which are generated by the above-described processing by the light ray group generation unit 303. Similarly, Fig. 15(b) shows an example of light ray groups 1517 to 1520 that include light ray information corresponding to high-frequency pattern region B and other light ray information, which are generated by the above-described processing by the light ray group generation unit 303. Note that light ray groups 1521 to 1524 shown in Fig. 15(b) are examples of light ray groups that include only other light ray information, which are generated by the light ray group generation unit 303 executing the processing of S1404 after the light ray group generation unit 303 generates light ray groups 1513 to 1520.
[0082] Furthermore, in the processes of S1402 and S1404, the ray group generation unit 303 according to the second embodiment generates multiple ray groups so that the same ray information does not overlap between the ray groups. However, the ray group generation unit 303 may generate multiple ray groups while allowing the same ray information to overlap between the ray groups. In this case, the total number of ray groups may be the same as the number of cases in which the same ray information does not overlap between the ray groups.
[0083] Furthermore, although the extraction unit 302 according to the second embodiment acquires the outline shape of the object by the volume intersection method in the processing of S1102, the method of acquiring the outline shape of the object is not limited to this. For example, the extraction unit 302 may acquire the outline shape of the object based on distance information acquired by stereo matching using the captured images acquired by the image acquisition unit 301, or a distance image acquired by a depth camera. Furthermore, information representing the outline shape of the object that has been generated in advance may be stored in the storage device 204, etc., and the extraction unit 302 may acquire the outline shape of the object by reading out the information.
[0084] Alternatively, the extraction unit 302 may identify a local space containing the object from among a plurality of local spaces obtained by dividing the imaging area 106, instead of the outline shape of the object, and acquire a collection of the identified local spaces containing the object as the outline shape of the object. In this case, for example, the extraction unit 302 may acquire a local space containing the outline shape of the object as the local space containing the object. Alternatively, for example, the extraction unit 302 may acquire a local space containing the object by the following method. Specifically, the extraction unit 302 first extracts feature points from each captured image acquired by the image acquisition unit 301, and projects the extracted feature points into space while associating them with each other between the captured images. Next, the extraction unit 302 acquires a local space with a high density of projected feature points as the local space containing the object.
[0085] <Effects of the image processing device according to the second embodiment> As described above, the image processing device 102 in the second embodiment extracts, from the captured image, an image region corresponding to the outline shape of each object including a high-frequency pattern, as a high-frequency pattern region, based on the outline shape of the object. The image processing device 102 also generates a group of rays for each high-frequency pattern region. The image processing device 102 configured in this manner can estimate three-dimensional information that enables generation of a high-quality virtual viewpoint image while suppressing the amount of calculation required for estimating the three-dimensional information. In particular, the image processing device 102 can estimate three-dimensional information of a space corresponding to multiple objects including high-frequency patterns with high accuracy. As a result, the image quality of a virtual viewpoint image generated based on the estimated three-dimensional information can be improved.
[0086] [Other embodiments] The present disclosure can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.
[0087] It should be noted that within the scope of the present disclosure, the embodiments may be freely combined, any component of each embodiment may be modified, or any component of each embodiment may be omitted.
[0088] [Configuration of the present disclosure] The present disclosure includes the following configurations, methods, and programs.
[0089] <Configuration 1> an acquisition means for acquiring captured image data; an extraction means for extracting a high frequency pattern region from the captured image; a generation means for generating a plurality of light ray groups based on light rays corresponding to each of a plurality of pixels included in the captured image, the generation means generating the plurality of light ray groups including one or more light ray groups for a high frequency pattern area, which are light ray groups in which the density of light rays corresponding to pixels included in the high frequency pattern area is higher than the density of light rays corresponding to pixels included in areas other than the high frequency pattern area; a learning means for learning three-dimensional information of a space using the plurality of groups of rays; 1. An image processing device comprising:
[0090] <Configuration 2> the generating means generates a group of light rays for the high frequency pattern area by randomly selecting a plurality of light rays from light rays corresponding to each of a plurality of pixels included in the high frequency pattern area; 2. The image processing device according to configuration 1,
[0091] <Configuration 3> the generating means generates the group of light rays for the high frequency pattern area so that a ratio of a total number of light rays corresponding to pixels included in the high frequency pattern area to a total number of light rays included in the group of light rays for the high frequency pattern area is equal to or greater than a predetermined ratio; 3. The image processing device according to configuration 1 or 2, characterized in that:
[0092] <Configuration 4> the predetermined ratio is equal to or greater than the ratio of the total number of pixels included in the high-frequency pattern region to the total number of pixels included in the captured image; 4. The image processing device according to configuration 3,
[0093] <Configuration 5> the generating means generates the group of light rays for the high frequency pattern area so that a predetermined number or more of light rays corresponding to pixels included in the high frequency pattern area are included; 5. The image processing device according to any one of configurations 1 to 4, characterized in that:
[0094] <Configuration 6> the predetermined number is a number whose ratio to the total number of light rays included in the light ray group for the high frequency pattern area is equal to or greater than the ratio of the total number of pixels included in the high frequency pattern area to the total number of pixels included in the captured image; 6. The image processing device according to configuration 5,
[0095] <Configuration 7> the generating means generates a group of light rays for the high frequency pattern area including only light rays corresponding to pixels included in the high frequency pattern area; 7. The image processing device according to any one of configurations 1 to 6,
[0096] <Configuration 8> the extraction means extracts the high-frequency pattern region based on a local space corresponding to an object including the high-frequency pattern; 8. The image processing device according to any one of configurations 1 to 7, characterized in that:
[0097] <Configuration 9> the extraction means extracts, for each of a plurality of objects including the high-frequency pattern, the high-frequency pattern region based on the local space corresponding to the object; 9. The image processing device according to configuration 8,
[0098] <Configuration 10> the local space is a rough shape of an object included as an image in the captured image; 10. The image processing device according to configuration 8 or 9,
[0099] <Configuration 11> the high-frequency pattern region is an image region in the captured image where the density of edges is higher than a predetermined density; 11. The image processing device according to any one of configurations 1 to 10, characterized in that:
[0100] <Configuration 12> the generating means generates a group of light rays for the high frequency pattern area by selecting a plurality of light rays that do not overlap with each other from among light rays corresponding to pixels included in the high frequency pattern area; 12. The image processing device according to claim 1, wherein:
[0101] <Method> an acquisition step of acquiring captured image data; an extraction step of extracting a high-frequency pattern region from the captured image; a generation step of generating a plurality of light ray groups based on light rays corresponding to each of a plurality of pixels included in the captured image, the generation step including one or more light ray groups for a high frequency pattern area, which are light ray groups in which the density of light rays corresponding to pixels included in the high frequency pattern area is higher than the density of light rays corresponding to pixels included in areas other than the high frequency pattern area; a learning step of learning three-dimensional information of a space using the plurality of groups of rays; An image processing method comprising:
[0102] <Program> 11. A program for causing a computer to function as the image processing device according to any one of configurations 1 to 10. [Explanation of symbols]
[0103] 102 Image processing device 301 Image Acquisition Unit 302 Extraction part 303 Ray group generator 304 Learning Department
Claims
1. an acquisition means for acquiring captured image data; an extraction means for extracting a high frequency pattern region from the captured image; a generation means for generating a plurality of light ray groups based on light rays corresponding to each of a plurality of pixels included in the captured image, the generation means generating the plurality of light ray groups including one or more light ray groups for a high frequency pattern area, which are light ray groups in which the density of light rays corresponding to pixels included in the high frequency pattern area is higher than the density of light rays corresponding to pixels included in areas other than the high frequency pattern area; a learning means for learning three-dimensional information of a space using the plurality of groups of rays; 1. An image processing device comprising:
2. the generating means generates a group of light rays for the high frequency pattern area by randomly selecting a plurality of light rays from light rays corresponding to each of a plurality of pixels included in the high frequency pattern area; 2. The image processing device according to claim 1, wherein:
3. the generating means generates the group of light rays for the high frequency pattern area so that a ratio of a total number of light rays corresponding to pixels included in the high frequency pattern area to a total number of light rays included in the group of light rays for the high frequency pattern area is equal to or greater than a predetermined ratio; 2. The image processing device according to claim 1, wherein:
4. the predetermined ratio is equal to or greater than the ratio of the total number of pixels included in the high-frequency pattern region to the total number of pixels included in the captured image; 4. The image processing device according to claim 3, wherein:
5. the generating means generates the group of light rays for the high frequency pattern area so that a predetermined number or more of light rays corresponding to pixels included in the high frequency pattern area are included; 2. The image processing device according to claim 1, wherein:
6. the predetermined number is a number whose ratio to the total number of light rays included in the light ray group for the high frequency pattern area is equal to or greater than the ratio of the total number of pixels included in the high frequency pattern area to the total number of pixels included in the captured image; 6. The image processing device according to claim 5,
7. the generating means generates a group of light rays for the high frequency pattern area including only light rays corresponding to pixels included in the high frequency pattern area; 2. The image processing device according to claim 1, wherein:
8. the extraction means extracts the high-frequency pattern region based on a local space corresponding to an object including the high-frequency pattern; 2. The image processing device according to claim 1, wherein:
9. the extraction means extracts, for each of a plurality of objects including the high-frequency pattern, the high-frequency pattern region based on the local space corresponding to the object; The image processing device according to claim 8 ,
10. the local space is a rough shape of an object included as an image in the captured image; The image processing device according to claim 8 ,
11. the high-frequency pattern region is an image region in the captured image where the density of edges is higher than a predetermined density; 2. The image processing device according to claim 1, wherein:
12. the generating means generates a group of light rays for the high frequency pattern area by selecting a plurality of light rays that do not overlap with each other from among light rays corresponding to pixels included in the high frequency pattern area; 2. The image processing device according to claim 1, wherein:
13. an acquisition step of acquiring captured image data; an extraction step of extracting a high-frequency pattern region from the captured image; a generation step of generating a plurality of light ray groups based on light rays corresponding to each of a plurality of pixels included in the captured image, the generation step including one or more light ray groups for a high frequency pattern area, which are light ray groups in which the density of light rays corresponding to pixels included in the high frequency pattern area is higher than the density of light rays corresponding to pixels included in areas other than the high frequency pattern area; a learning step of learning three-dimensional information of a space using the plurality of groups of rays; An image processing method comprising:
14. A program for causing a computer to function as the image processing device according to any one of claims 1 to 12.
Citation Information
Patent Citations
Image processing device, image processing method, and program
JP2023066705A