Image processing device, imaging device, image processing method and program
The imaging device automatically determines the direction of a virtual light source using shape and area detection to generate appropriate shadows on subjects, enhancing shadow generation accuracy.
Patent Information
- Application Number
- JP2021094028
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-06-04
- Publication Date
- 2025-08-28
- Estimated Expiration
- 2041-06-04
AI Technical Summary
Existing image processing techniques for relighting fail to generate appropriate shadows on subjects due to the inability to automatically determine the direction of a virtual light source.
An imaging device that includes a shape acquisition means for acquiring normal information, a first and second area detection means to identify regions for shadow generation and projection, and a direction setting means to set the virtual light source direction based on these regions, excluding areas with low reliability from shadow casting.
Enables the imaging device to set a virtual light source direction that allows a main subject to cast appropriate shadows on other subjects, improving shadow generation accuracy.
Smart Images

Figure 0007730668000001 
Figure 0007730668000002 
Figure 0007730668000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image processing device, an imaging device, an image processing method, and a program. [Background technology]
[0002] There is an image processing technique called relighting (shadow casting) that generates a shaded image by applying shading through image processing. In relighting, for example, a virtual light source is set, and a shadow area generated by the virtual light source is calculated using the direction of the set virtual light source and shape information of the subject, and a shaded image is generated by applying shading to the calculated area. Patent Document 1 discloses a technique for setting the position of the virtual light source based on the positional relationship between a person and an obstructing object so that, when there is an obstructing object between the set virtual light source and the person's face, unnecessary shadows are not cast on the person's face by the relighting process. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2020-10168 Summary of the Invention [Problem to be solved by the invention]
[0004] Although Patent Document 1 can prevent a person's face from being cast in the shadow of an occluding object, it may not be possible to generate the appropriate shadow desired by the user. In order to generate an appropriate shadow in the relighting process, it is necessary to automatically determine the direction of an appropriate virtual light source.
[0005] An object of the present invention is to provide an imaging device capable of setting the direction of a virtual light source that allows a main subject to cast an appropriate shadow on other subjects. [Means for solving the problem]
[0006] In order to solve the above problem, the image processing device of the present invention is At least normal information and reliability of the normal information a shape acquisition means for acquiring a shape of the object, a first area detection means for detecting a first area where a shadow of the object is generated, and a second area detection means for detecting a second area where the shadow is projected; Normal information, the reliability of the normal information; a direction setting means for setting a direction of a virtual light source that causes the first region to cast the shadow onto the second region based on the first region and the second region; the direction setting means sets the direction of the virtual light source so that an area in the second area where the reliability of the normal information is lower than a threshold is excluded from the area where the shadow is cast. . In addition, a second form of the image processing device of the present invention includes a shape acquisition means for acquiring at least normal information of the subject as shape information of the subject, a first area detection means for detecting a first area that generates a shadow of the subject, a second area detection means for detecting a second area onto which the shadow is projected, and a direction setting means for setting the direction of a virtual light source that projects the shadow from the first area onto the second area based on the shape information, the first area, and the second area, wherein the second area detection means calculates a histogram of the surface normals of the subject from the normal information of the subject, and detects the area with the surface normal having the highest frequency in the histogram as the second area. [Effects of the Invention]
[0007] According to the present invention, it is possible to set the direction of a virtual light source that allows a main subject to cast an appropriate shadow on other subjects. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a diagram illustrating the overall configuration of an imaging device. [Figure 2] FIG. 2 is a diagram illustrating an imaging unit and an imaging optical system. [Figure 3] 10 is a flowchart showing a relighting process. [Figure 4] 10 is a flowchart showing a shape information acquisition method. [Figure 5] 10 is a flowchart showing a method for calculating distance information. [Figure 6] FIG. 10 is a diagram illustrating a minute block. [Figure 7] 10A and 10B are diagrams illustrating the relationship between the amount of image shift and a correlation value. [Figure 8] 10 is a flowchart showing a method for determining the direction of a virtual light source. [Figure 9] FIG. 10 is a diagram illustrating a method for determining the direction of a virtual light source. [Figure 10] 10A and 10B are diagrams illustrating a method for determining the direction of a virtual light source based on the approximate direction of the virtual light source specified by a user. [Figure 11] FIG. 10 is a diagram illustrating a method for generating an image using a single light source. [Figure 12] FIG. 1 is a diagram illustrating a method for generating an image using multiple light sources. DETAILED DESCRIPTION OF THE INVENTION
[0009] (First embodiment) 1 is a diagram illustrating the overall configuration of an imaging device. The imaging device 110 includes an image processing device 100 and an imaging unit 112. A lens device 111 is detachably connected to the imaging device 110. The image processing device 100 of this embodiment performs relighting processing on an acquired image to generate a shaded image by adding shading. The image processing device 100 includes a control unit 101, a memory 102, a parameter setting unit 103, an image acquisition unit 104, a shape information acquisition unit 105, a first region detection unit 106, a second region detection unit 107, a virtual light source direction setting unit 108, and an image generation unit 109.
[0010] The lens device 111 has an imaging optical system. The imaging optical system includes a plurality of lenses 119, such as zoom lenses and focus lenses, an aperture, and a shutter, and forms an optical image of a subject on an imaging element. The imaging unit 112 captures an image of the subject. The subject is a target for image processing by the image processing device 100. The imaging unit 112 is an imaging element having a photoelectric conversion element such as a CMOS or CCD, and outputs an output signal (analog signal) corresponding to the optical image. Note that, although an example in which the lens device 111 is detachably connected to the imaging device 110 will be described in this embodiment, an imaging device in which the imaging device 110 and the lens device 111 are integrated may also be used.
[0011] The control unit 101 controls the overall operation of the imaging device 110 including the image processing device 100. The control unit 101 includes, for example, a CPU (Central Processing Unit). The CPU executes a program stored in a nonvolatile memory such as a ROM (Read Only Memory), thereby realizing the functions of the imaging device 110 and various processes described below.
[0012] The image (hereinafter also referred to as image information) output by the imaging unit 112 is supplied to the image acquisition unit 104. The image acquisition unit 104 acquires an image captured by the imaging unit 112 or an image captured by a device other than the imaging device 110, and stores it in the memory 102. The memory 102 stores the image. The memory 102 also reads out information necessary for processing in each module and stores the processing results of each module. The parameter setting unit 103 accepts settings of various parameters related to imaging, image processing, etc. from the user, and stores the parameter information input by the user in the memory 102. The parameter setting unit 103 also accepts a designation from the user of an approximate direction of a virtual light source designated by the user, which will be described later.
[0013] The shape information acquisition unit 105 is a shape acquisition means that calculates and acquires shape information from an image captured by the imaging unit 112. The shape information acquisition unit 105 also acquires the calculated shape information or shape information input by a user to the image processing device 100. The shape information acquisition unit 105 stores the acquired shape information in the memory 102.
[0014] The first area detection unit 106 detects a first area from at least one of image information, shape information, and second area information. In this embodiment, the first area is an object that casts a shadow during relighting, i.e., an area where a shadow is generated by relighting. The first area detection unit 106 saves information about the detected first area in memory 102. The second area detection unit 107 detects a second area from at least one of image information, shape information, and first area information. In this embodiment, the second area is an area where a shadow of the first area is cast (applied) during relighting. The second area detection unit 107 saves information about the detected second area in memory 102.
[0015] Virtual light source direction setting unit 108 sets the direction of the virtual light source when performing relighting processing, based on the first region, the second region, and shape information. Virtual light source direction setting unit 108 saves information about the set virtual light source direction in memory 102. Image generation unit 109 generates a shaded image by adding a shadow calculated based on the virtual light source direction and shape information to the image. Image generation unit 109 saves the generated shaded image in memory 102. As described above, in this embodiment, the region of the subject where the virtual light source will generate a shadow (first region) is calculated, and the shade is added to the calculated region (second region) such as the floor, thereby generating a shaded image.
[0016] FIG. 2 is a diagram illustrating the imaging unit 112 and the imaging optical system. FIG. 2(A) is a diagram illustrating the configuration of an imaging element, which is the imaging unit 112. Pixels 115 are regularly arranged two-dimensionally in the imaging element. Each pixel 115 has one microlens 114 and a pair of photoelectric conversion units (photoelectric conversion unit 116A and photoelectric conversion unit 116B). The pair of photoelectric conversion units receives light beams passing through different pupil regions of the imaging optical system via the one microlens 114. Multiple viewpoint images (a pair of viewpoint images) are generated from the light beams received by the photoelectric conversion units. Hereinafter, the viewpoint image captured by the photoelectric conversion unit 116A will be referred to as image A, and the viewpoint image captured by the photoelectric conversion unit 116B will be referred to as image B. An image obtained by combining the images A and B will be referred to as image A+B. Since each pixel of the imaging unit 112 has a pair of photoelectric conversion units, it is possible to acquire a pair of image data (image A, image B) based on light beams passing through different pupil regions of the imaging optical system. Based on the parallax between image A and image B, it is possible to calculate distance information between the subject and image capture device 110 by using parameters such as a conversion coefficient determined by the magnitude of the opening angle of the center of gravity of the light beams passing through a pair of distance measurement pupils.
[0017] 2(B) is a diagram showing the configuration of an imaging optical system included in the lens device 111. The imaging unit 112 forms an image of light emitted from an object 118 on an imaging plane 120 using a lens 119, and receives the light on a sensor plane 121 of an imaging element. The imaging unit 112 performs photoelectric conversion and outputs an image (analog signal) corresponding to the optical image. In this embodiment, the imaging device 110 configured as described above performs relighting processing by setting the direction of a virtual light source that can add an appropriate shadow to the subject.
[0018] 3 is a flowchart showing the relighting process. In step S301, the parameter setting unit 103 acquires parameters related to the relighting process. The parameter setting unit 103 acquires at least the following information (a) to (d) as parameters: (a) Distance from the imaging device to the focal point of the subject (b) Conversion coefficient determined by the magnitude of the opening angle of the center of gravity of the light beam passing through a pair of measuring pupils (c) Distance from the image-side principal point of the lens of the imaging device to the sensor surface (d) Focal length of the imaging device The parameter setting unit 103 may acquire these parameters from the memory 102 or from each module in the image capturing device 110 .
[0019] In step S302, the image acquisition unit 104 acquires an image. The image acquisition unit 104 may acquire an image captured by the imaging device 110 from the imaging unit 112, or may acquire a previously captured and saved image from the memory 102. In step S303, the shape information acquisition unit 105 acquires shape information. Here, the shape information is information about the shape of the subject. Specifically, the shape information in this embodiment is information including shape position information indicating the position of the shape of the subject using a point cloud, normal information indicating the inclination of the surface of a local region in the point cloud, and reliability information of the position information and normal information. The shape position information can be acquired from distance information from the imaging device 110 to the subject. The inclination of the surface of a local region is information such as the amount of displacement in the shape of the local region and the surface normal.
[0020] Details of the shape information acquisition process will be described with reference to Fig. 4. Fig. 4 is a flowchart showing the shape information acquisition process shown in step S303. In the shape information acquisition process, distance information, normal information, and reliability information are acquired.
[0021] (Distance information acquisition process) First, in step S401, the shape information acquisition unit 105 calculates distance information from the image capture device 110 to the subject. For example, the shape information acquisition unit 105 calculates the amount of image shift from the parallax between multiple viewpoint signals by a correlation calculation or the like, and converts it into a defocus amount. Then, the shape information acquisition unit 105 can calculate the distance information based on the defocus amount. Details of the distance information calculation process will be described with reference to FIG. 5.
[0022] FIG. 5 is a flowchart showing the distance information acquisition process. In step S501, the shape information acquisition unit 105 generates multiple images (viewpoint images, pupil-divided images) from the images acquired in step S302. In this embodiment, a pair of images is generated as the multiple images. The pair of images is, for example, an image A output from the photoelectric conversion unit 116A and an image B output from the photoelectric conversion unit 116B. It is preferable that the pair of images have as few shadow areas as possible. This is because shadow areas have low contrast, which reduces the accuracy of parallax calculation using stereo matching. Images with few shadow areas can be acquired by capturing images under lighting conditions such as full illumination or no illumination. Furthermore, if the imaging device 110 has a light-emitting unit, a pattern such as a grid can be placed in the projection system of the light-emitting unit and an image projected with the pattern on the subject can be captured, thereby adding a pattern to a subject without texture and improving the accuracy of parallax calculation.
[0023] In step S502, the shape information acquisition unit 105 sets minute blocks for each pair of image data. In this embodiment, minute blocks of the same size are set for the images A and B generated in step S501. A minute block is generally synonymous with a window set when performing template matching. The setting of minute blocks will be described with reference to FIG. 6.
[0024] FIG. 6 is a diagram illustrating minute blocks. FIG. 6(A) is a diagram illustrating minute blocks set in step S502. In order to calculate the defocus amount of a pixel of interest 604, a minute block 603 centered on the pixel of interest 604 is set in the image A 601, and a minute block 605 of the same size as the minute block 603 is set in the image B 602. In this way, in this embodiment, a pixel of interest that is the center of a minute block is set for each pixel, and a minute block centered on the pixel of interest is set. FIG. 6(B) is a diagram illustrating the shape of a minute block. While the minute block 603 has a size of 9 pixels centered on the pixel of interest 604, the minute block 606 has a size of 25 pixels centered on the pixel of interest 604. In this way, the size of the minute blocks to be set can be changed.
[0025] In step S503, the shape information acquisition unit 105 calculates the amount of image shift. The shape information acquisition unit 105 performs correlation calculation processing in the minute block set in step S502 to calculate the amount of image shift at each point. In the correlation calculation, a pair of pixel data in the minute block is generalized and represented as E and F, respectively. For example, while relatively shifting the data series F corresponding to image B with respect to the data series E corresponding to image A, the correlation amount C(k) at a shift amount k between these two data series is calculated using the following mathematical formula (1). C(k)=Σ|E(n)−F(n+k)|···(1) In formula (1), C(k) is calculated for the number n of the data series. The shift amount k is an integer and is the relative shift amount in units of the data interval of the image data. Note that the shift amount k is synonymous with the amount of parallax in the stereo matching method.
[0026] An example of the calculation result of Equation (1) will be described with reference to FIG. 7. FIG. 7 is a diagram illustrating the relationship between the image shift amount and the correlation value. In FIG. 7, the horizontal axis represents the image shift amount, and the vertical axis represents the correlation amount. Furthermore, the correlation amount C(k) represents the discrete correlation amount at the shift amount k, and the correlation amount C(x) represents the continuous correlation amount at the image shift amount x. In FIG. 7, the correlation amount C(k) is minimum at an image shift amount where the correlation between a pair of data series is high.
[0027] The image shift amount x at which the continuous correlation amount C(x) is minimum is calculated using, for example, a three-point interpolation method using the following equations (2) to (5). x=kj+D / SLOP (2) C(x)=C(kj)-|D| (3) D={C(kj-1)-C(kj+1)} / 2···(4) SLOP=MAX{C(kj+1)-C(kj),C(kj-1)-C(kj)} ···(5) Here, kj is the k at which the discrete correlation amount C(k) is minimum. The shift amount x calculated by Equation (2) is the image shift amount between a pair of image data. The unit of the image shift amount x is pixels.
[0028] In step S504, the shape information acquisition unit 105 converts the image shift amount into a defocus amount. The magnitude of the defocus amount represents the distance from the imaging plane 120 of the subject image to the sensor plane 121. Specifically, the shape information acquisition unit 105 can calculate the defocus amount DEF using the following formula (6) based on the image shift amount x calculated using formula (2). DEF=KX·x···(6) In equation (6), KX is a conversion coefficient determined by the magnitude of the opening angle of the center of gravity of the light beams passing through a pair of distance measurement pupils. In this way, by repeating the processes of steps S502 to S504 while shifting the pixel of interest position by one pixel at a time, the defocus amount for each pixel position can be calculated.
[0029] In step S505, the shape information acquisition unit 105 calculates the distance z from the sensor surface 121 of the imaging device 110 to the subject based on the defocus amount calculated in step S504. The distance z can be calculated using the following formulas (7) and (8). dist=1 / (1 / (dist_d+DEF)-1 / f)...(7) z=length-dist (8) dist is the distance from the focal position to the subject, dist_d is the distance from the image-side principal point of the lens of the imaging unit 112 of the imaging device 110 to the sensor plane 121, f is the focal length, and length is the distance from the sensor plane 121 of the imaging device 110 to the focal position. The focal position corresponds to the in-focus position. Note that, in this embodiment, an example has been described in which the shape information acquisition unit 105 calculates the distance z at each point from the image shift amount x using equations (6) to (8), but this is not limiting, and the distance z at each point may be calculated using other calculation methods.
[0030] The distance (length) from the sensor surface 121 of the imaging device 110 to the focal position can be measured, for example, by a laser distance measuring means (not shown). Also, by having a data table that indicates the relationship between the lens position and the focal position during shooting, it is possible to estimate the distance to the focal position corresponding to the lens position during shooting. The data table indicating the relationship between the lens position and the focal position can reduce the effort required to measure the distance from the sensor surface 121 of the imaging device 110 to the focal position.
[0031] As described above, the distance z from the sensor surface 121 of the imaging device 110 to the subject can be calculated from the defocus amount or image shift amount obtained from multiple images. Because information can be calculated by generating a pair of image data from a single image, the imaging device does not need to be a compound eye camera and can be a monocular camera, simplifying the configuration of the imaging device. Furthermore, calibration processing when installing multiple cameras can be simplified or eliminated. While the example in step S501 has been described in which the image processing device 100 calculates the distance z from image information captured by the imaging device 110, this is not limiting. For example, the distance z can also be calculated using a stereo camera. Alternatively, the distance z can be calculated by a device external to the imaging device 110, and the distance z calculated by the external device can be acquired by the image processing device 100. Information about the distance to the subject may be acquired using a device such as a LiDAR, rather than calculated from an image. Note that the distance information between the image capture device 110 and the subject calculated in S401 has been described using the distance from the sensor surface 121 as an example, but is not limited to this and may be the distance from any position, such as the tip of the lens of the image capture device 110.
[0032] (Normal information acquisition process) Returning to the description of FIG. 4, in step S402, the shape information acquisition unit 105 acquires normal information. The normal information is the inclination of the surface of the local region. The inclination of the surface of the local region can be acquired by a method of calculating the amount of displacement by differentiating the acquired distance information in the in-plane direction, or by a method of calculating the surface normal by photometric stereo. In this embodiment, as an example, the surface normal is used as the inclination of the surface of the local region, i.e., as the surface normal information.
[0033] (Reliability information acquisition process) In step S403, the shape information acquisition unit 105 acquires reliability information. The reliability can be calculated by determining an evaluation value of the reliability. The reliability can be acquired for each of the distance information and the normal information.
[0034] The reliability of distance information is explained below. The reliability of distance information is information used to determine whether the acquired distance information can be used. Low reliability also reduces the accuracy of the calculated distance information. The reliability is calculated as an evaluation value, quantified, and compared with a threshold to evaluate the reliability. The reliability evaluation value can be calculated from, for example, image brightness, image contrast, and the amount of defocus when acquiring distance information. The reliability obtained from the evaluation value calculated from image brightness, image contrast, and the amount of defocus when acquiring distance information is applicable to a method of acquiring distance information by generating a pair of image data from a single image and calculating the amount of defocus. On the other hand, when acquiring distance information using a device such as LiDAR, if the amplitude of the reflected laser signal is lower than a specified value, a possible method is to lower the evaluation value of the reliability of the distance information. In this way, it is possible to change the evaluation items used to calculate the evaluation value of the reliability of distance information depending on the method of acquiring distance information.
[0035] First, the evaluation value of image brightness will be described. In this embodiment, when acquiring distance information, the amount of image shift is calculated from a pair of images, but the accuracy of calculating the amount of image shift decreases in areas where the image signal is too bright or too dark. If the median brightness value when expressing image brightness is Lm and the brightness of the pixel to be evaluated is Lp, the evaluation value L of brightness can be calculated using the following formula (9): L = -|Lm-Lp| (9) By setting the evaluation value L as in Equation (9), it is possible to decrease the evaluation value L of brightness as the brightness of a pixel deviates from the median brightness value when expressing brightness. The evaluation value L of brightness makes it possible to take into account the influence of image brightness on the reliability of distance information.
[0036] Next, the evaluation value of image contrast will be described. In this embodiment, when distance information is acquired, the amount of image shift is calculated from a pair of images, but the accuracy of calculating the minimum value from the correlation amount C(k) decreases in areas where the contrast of the image signal is low. The evaluation value of image contrast can be calculated by calculating the variance of brightness from the brightness of each pixel and its surroundings, and setting the variance as the contrast evaluation value B. The contrast evaluation value B makes it possible to take into account the effect of image contrast on the reliability of distance information.
[0037] Finally, we will explain the evaluation value of the defocus amount. Distance information can be calculated from the defocus amount, but as the absolute value of the defocus amount increases, the image becomes more blurred, which reduces distance measurement accuracy. In other words, as the absolute value of the defocus amount increases, the evaluation value of reliability decreases. For the defocus amount DEF calculated using equation (6), if the evaluation value of the defocus amount is D, then the evaluation value D of the defocus amount can be calculated using equation (10). D = |DEF| (10) The evaluation value D of the defocus amount makes it possible to take into consideration the influence of the defocus amount when acquiring the distance information on the reliability of the distance information.
[0038] If the evaluation value of the reliability of the distance information is M, the evaluation value of the reliability M can be calculated using the following formula (11) based on the evaluation value of the image brightness, the evaluation value of the image contrast, and the evaluation value of the defocus amount when the distance information is obtained. M = L + B + D (11) The shape information acquisition unit 105 calculates an evaluation value M of the reliability of the distance information for each pixel. Then, the shape information acquisition unit 105 compares the evaluation value M of each pixel with an arbitrary threshold to determine the reliability of the distance information for each pixel. If the evaluation value M of the reliability of the distance information is less than the threshold, the shape information acquisition unit 105 determines that the distance information is low-reliable, and if the evaluation value M of the reliability of the distance information is equal to or greater than the threshold, the shape information acquisition unit 105 determines that the distance information is high-reliable.
[0039] Next, the reliability of normal information will be described. The reliability of normal information is information for determining whether acquired normal information can be used. One method for calculating the reliability of normal information is, for example, to use images captured when calculating normal information using photometric stereo. Photometric stereo involves capturing multiple images of a scene in which light is irradiated onto an object from different light source directions, and calculating the surface normal of the object's surface from the brightness ratio of the object captured in each light source direction. When capturing multiple images with different light source directions, the positional relationship between the object and the camera must be fixed, and the light source direction at the time of light emission must be known. Therefore, if the object or camera moves during multiple image capture, or if strong external light enters, this is detected and processed so that the image is not used in normal calculation. At this time, the ratio between the number of captured images and the number of images actually used in the calculation is calculated, and if the number of images used is small, the reliability is reduced. If the number of captured images is Gn, the number of images actually used to calculate normal information is Rn, and the evaluation value for the number of images used is Cn, the reliability evaluation value Cn can be calculated using equation (12). Cn = Rn / Gn (12)
[0040] Furthermore, if the number of images in which the luminance of the pixel of interest is saturated is equal to or greater than a threshold, the reliability of the pixel of interest is reduced. If the number of images in which the luminance of the pixel of interest is saturated is Ln and the evaluation value for the number of saturated images is H, the evaluation value H of the reliability can be calculated using Equation (13). H = Ln / Gn (13) If the evaluation value of the reliability of the normal vector information is N, then the evaluation value N of the reliability can be calculated using equation (14). N = Cn × H (14)
[0041] The shape information acquisition unit 105 calculates an evaluation value N of the reliability of the normal vector information for each pixel. The shape information acquisition unit 105 then calculates the reliability of the normal vector information for each pixel by setting an arbitrary threshold and comparing it with the evaluation value N of the reliability of the normal vector information. If the evaluation value N of the reliability of the normal vector information is less than the threshold, the shape information acquisition unit 105 determines that the normal vector information is low-reliable, and if the evaluation value N of the reliability of the normal vector information is equal to or greater than the threshold, the shape information acquisition unit 105 determines that the normal vector information is high-reliable.
[0042] Moreover, an evaluation value R that evaluates both the reliability of the distance information and the reliability of the normal vector information can be calculated using equation (15). R = M × N (15) The shape information acquisition unit 105 can determine the reliability of each pixel by setting an arbitrary threshold for the evaluation value R and comparing it with the evaluation value R. The shape information acquisition unit 105 determines that the evaluation value R is low reliable if it is less than the threshold, and determines that the evaluation value R is high reliable if it is equal to or greater than the threshold.
[0043] By limiting the range of image processing in this embodiment based on the reliability calculated as described above, it is possible to obtain a shadow-added image with more accurate shape information. Furthermore, by utilizing the reliability of each of the distance information and normal information used in area detection, which will be described in step S304 below, it is possible to improve the accuracy of area detection. Furthermore, it is also possible to interpolate distance information and normal information by performing a fill-in interpolation process on low-reliability distance information and normal information using surrounding high-reliability distance information and normal information.
[0044] (area detection) Returning to the description of Figure 3, in step S304, the first area detection unit 106 and the second area detection unit 107 perform area detection. The first area detection unit 106 detects a first area that generates a shadow. The second area detection unit 107 detects two areas that are second areas onto which the shadow is cast.
[0045] First, detection of a first region that generates a shadow will be described. Here, two detection methods for detecting the first region performed by the first region detection unit 106 will be described. The first first region detection method is a method in which an area near the center of image height within the angle of view of an acquired image is set as the first region. By setting an area near the center of image height as the first region, it is possible to set the first region so that it becomes an area in which a subject placed at the center of the image by the user generates a shadow. This makes it possible to set the first region without performing particularly complicated processing.
[0046] The second method for detecting the first area is to designate an area near the focal point (focus position) as the first area. By designating an area near the focal point as the first area, the user can set the first area so that it is the area where the subject on which the user focuses casts a shadow. Three methods are described as examples of methods for determining whether an area is near the focal point. The first method refers to the defocus amount DEF calculated using equation (6). A threshold is set for the defocus amount DEF, and areas below the threshold are designated as areas near the focal point. The second method is based on contrast. Areas near the focal point generally have less image blur, resulting in stronger contrast than other areas. Therefore, local contrast within the image is evaluated and compared with a threshold, and areas above the threshold are designated as areas near the focal point. The third method uses position information from the AF frame used when capturing an image. The area of the AF frame used when capturing an image is designated as the area near the focal point. The second and third methods allow for detection of areas near the focal point without calculating the defocus amount DEF.
[0047] The first region detection method may be performed using the above two methods independently, or may be performed by combining the methods. For example, by combining the first first region detection method and the second first region detection method, a region near the center of the image height and having high contrast can be detected as the first region, thereby improving the detection accuracy of the first region. Furthermore, the subject region may be detected as the first region using a method other than the three detection methods described above.
[0048] Next, detection of the second region onto which a shadow is cast will be described. Here, two detection methods for detecting the second region performed by the second region detection unit 107 will be described. The first second region detection method is a method of detecting the second region from acquired normal information. Generally, the region onto which a shadow is to be cast is the floor or ground. Therefore, the floor is detected as the second region using a surface normal calculated from the normal information. Specifically, the second region detection unit 107 takes a histogram of the directions of surface normals across the entire angle of view of the image based on the normal information acquired in step S303, and determines the region of the surface normal with the highest frequency in the histogram as the second region. This makes it possible to detect the floor using the surface normal, particularly in scenes where the floor occupies a large portion of the angle of view, and therefore detect the floor as the second region.
[0049] The second method of detecting the second region is a method of detecting the second region from acquired distance information. Generally, the region onto which a shadow is to be cast is the floor or the ground. On the floor, shape information calculated from distance information changes continuously in a fixed direction. Therefore, the second region detection unit 107 detects, as the second region, a region where shape information calculated from the distance information acquired in step S303 changes continuously in a fixed direction within a certain range. Two detection methods for detecting the second region have been described above. The above two detection methods for the second region may be performed independently, or the two methods may be combined to perform detection. Furthermore, the floor may be detected as the second region by a method other than the two detection methods described above.
[0050] Furthermore, by detecting the first region and the second region as mutually exclusive regions, it is possible to improve the detection accuracy of the first region and the second region. For example, it is possible to set a candidate for the first region as a region that is not the second region. This means that if the second region is detected, the other region becomes a candidate for the first region, and since the candidates for the first region are limited, it is possible to reduce the calculation cost for calculating the first region.
[0051] Furthermore, it is possible to detect both the first and second regions by performing image recognition on them. If, for example, a floor surface is detected as a result of image recognition, that region is designated as the second region. Then, the region of the object recognized in a region other than the floor surface is designated as the first region. Since image recognition allows for region detection based on the recognized object, it is possible to detect the region with higher accuracy than detection near the center of the image height or near the focus. Furthermore, region detection by image recognition can be performed in combination with region detection by the first region detection method and second region detection method described above. For example, by calculating a histogram of surface normals of the region recognized as the floor and re-detecting the region of surface normals with a high frequency in the histogram as the floor, it is possible to detect a floor that was not detected due to a failed image recognition.
[0052] When using the distance information or normal information acquired in step S303 in detecting the first and second regions in step S304, the information to be used may be limited based on the reliability information of the information. For example, the reliability of the distance information is compared with a threshold, and distance information with a reliability lower than the threshold is not used for region detection. By using the respective reliability of the distance information and normal information used in region detection, it is possible to improve the accuracy of region detection.
[0053] (Virtual light source direction determination process) In step S305, virtual light source direction setting unit 108 determines the direction of the virtual light source. The virtual light source is a surface light source, and parallel light is irradiated onto the subject from the direction of the determined virtual light source. The position of the virtual light source in this embodiment is determined as the direction of the deflection angle when the radius is infinite in a polar coordinate system. Therefore, in this embodiment, the direction of the virtual light source is determined rather than the position of the virtual light source. By automatically determining the direction of the virtual light source, it is possible for the main subject within the angle of view to generate appropriate shadows on other subjects without the user having to adjust the distance between the virtual light source and the subject or the light intensity and light direction of the virtual light source.
[0054] The method for determining the direction of the virtual light source will be described with reference to Figs. 8 to 10. Fig. 8 is a flowchart showing an example of the virtual light source direction determination process in step S305. Fig. 9 is a diagram explaining the method for determining the direction of the virtual light source. In Fig. 9, the horizontal direction of the image captured by imaging device 110 is the x direction, the vertical direction is the y direction, and the depth direction is the z direction.
[0055] 9(A) is a diagram showing the positional relationship between the imaging device 110 and each subject when viewed from the x direction. The x direction value increases horizontally from the left edge of the image to the right edge. The y direction value increases vertically from the top edge of the image to the bottom edge. The z direction value increases from the depth value on the near side to the depth value on the far side. Region 901 is the region detected as the first region, and region 902 is the region detected as the second region.
[0056] In step S801, the virtual light source direction setting unit 108 calculates a representative value of the first region detected in step S304. For example, the virtual light source direction setting unit 108 calculates the representative value of the first region by calculating the average values of the x, y, and z coordinate values of the first region. Furthermore, if the focus position is included in the first region, the virtual light source direction setting unit 108 may set the x, y, and z coordinate values of the focus position as the representative value of the first region. Furthermore, the virtual light source direction setting unit 108 may set the center of gravity of the first region as the representative value of the first region.
[0057] In step S802, virtual light source direction setting unit 108 calculates a straight line connecting the point in the second region where the z-direction value is smallest and the representative value of the first region calculated in step S801. For example, in Fig. 9(A), it is possible to obtain straight line 905 by connecting representative value 903 of the first region and point 904 in the second region that is closest to the camera and is the point in the second region where the z-direction value is smallest.
[0058] In step S803, virtual light source direction setting unit 108 obtains a straight line connecting the point in the second region with the largest z-direction value and the representative value of the first region. In Fig. 9(A), line 907 can be obtained by connecting representative value 903 of the first region with point 906 in the second region that is the largest z-direction value in the second region and is the farthest from the camera.
[0059] In step S804, virtual light source direction setting unit 108 determines the direction defined by line 905 calculated in step S802 and line 907 calculated in step S803 as the direction of the virtual light source. In FIG. 9A, direction 908 is the direction of the virtual light source defined by line 905 and line 907.
[0060] Steps S801 to S804 limit the direction of the virtual light source in the two-dimensional yz space defined by the y and z directions when viewed from the x direction. Similarly, it is possible to limit the direction of the virtual light source in the xz space when viewed from the y direction. By setting the virtual light source within the range of the limited virtual light source direction, virtual light source direction setting unit 108 sets the virtual light source direction that ensures that the first region casts a shadow on the second region. In this way, by automatically setting the virtual light source direction within the limited range of the virtual light source direction, it is possible to easily obtain an image in which the first region casts a shadow on the second region.
[0061] Next, an example of determining the direction of a virtual light source using a defocus amount (effective ranging range) will be described. In the process of determining the direction of a virtual light source, it is also possible to determine the direction of the virtual light source using an effective ranging range according to the defocus amount calculated in step S504. FIG. 9(B) is a diagram illustrating a method of determining the direction of a virtual light source using an effective ranging range. Like FIG. 9(A), FIG. 9(B) shows the positional relationship between image capture device 110 and each subject when viewed from the x direction. Line 909 indicates the distance from image capture device 110 to the focus position. A focus plane in the x and y directions is found from the focus position, and the focus plane when viewed from the x direction in the x and y directions is represented by line 910.
[0062] As the absolute value of the defocus amount calculated in step S504 increases, the image becomes more blurred and the distance measurement accuracy decreases. Therefore, by setting a threshold value for the defocus amount, the direction of the virtual light source is limited so that a shadow is not cast on the second region having a defocus amount equal to or greater than the threshold. Lines 911 and 912 are lines when a plane having a defocus amount that is the threshold away from line 910 is viewed from the x direction. The virtual light source direction setting unit 108 obtains lines connecting the points where each line intersects with the second region and the representative value 903 of the shape of the first region.
[0063] Specifically, point 913 where line 911 closest to image capture device 110 intersects with the second region is point 913 closest to image capture device 110 within the effective ranging range of the second region. Virtual light source direction setting unit 108 obtains line 914 by connecting representative value 903 of the shape of the first region with point 913 in the second region. Point 915 where line 912 farthest from image capture device 110 intersects with the second region is point 915 farthest from image capture device 110 within the effective ranging range of the second region. Virtual light source direction setting unit 108 obtains line 916 by connecting representative value 903 of the shape of the first region with point 915 in the second region. By performing the same process as in step S804 based on line 914 and line 916, virtual light source direction setting unit 108 can set direction 917 of a virtual light source limited by line 914 and line 916. The virtual light source direction setting unit 108 sets a virtual light source within the range of the limited virtual light source direction, thereby setting a virtual light source direction that ensures that the first region casts a shadow on the second region. In this way, by limiting the range of the virtual light source direction based on the defocus amount (effective ranging range), it is possible to improve the accuracy of the relighting process.
[0064] Furthermore, in the process of determining the direction of the virtual light source, virtual light source direction setting unit 108 may limit the range of the direction of the virtual light source based on the reliability information acquired in step S403 so that a shadow is not cast on a second region with low reliability. By excluding a second region with reliability lower than a threshold from the region to be shaded, it is possible to obtain a shadow-added image with higher accuracy. Note that the reliability information used in the process of determining the direction of the virtual light source may be an evaluation value M of the reliability of the distance information or an evaluation value N of the reliability of the normal information, or an evaluation value R that evaluates both the reliability of the distance information and the reliability of the normal information.
[0065] Next, a process for determining the direction of a virtual light source when the user specifies the approximate direction of the virtual light source will be described. In this embodiment, the direction of the virtual light source can be determined based on the approximate direction of the virtual light source specified by the user. The approximate direction of the virtual light source that the user can specify is, for example, whether the light from the virtual light source is "front-lit / backlit / top-lit" with respect to the subject, or whether the direction of the virtual light source is "right / center / left" of the subject as seen from image capture device 110.
[0066] A method for determining the direction of a virtual light source based on the approximate direction of the virtual light source specified by the user will be described with reference to Fig. 10. Fig. 10 is a diagram for explaining a method for determining the direction of a virtual light source based on the approximate direction of the virtual light source specified by the user. In Fig. 10, as in Fig. 9, the horizontal direction of the image captured by imaging device 110 is defined as the x direction, the vertical direction as the y direction, and the depth direction as the z direction.
[0067] FIG. 10A is a diagram illustrating the direction of a virtual light source when a user specifies "backlit, right side" as the general direction of the virtual light source in the positional relationship of each subject when viewed from the z direction. In order for the light from the virtual light source to "backlit" the subject, the direction of the virtual light source needs to be set further back than region 901, which is the subject, with respect to image capture device 110. In addition, in order for the direction of the virtual light source to be "right side," the direction of the virtual light source needs to be set to the right of first region 901, which is the subject. In this embodiment, with line 921 connecting representative value 903 of the shape of first region 901 and image capture device 110 as the boundary, the "right side" of the direction of the virtual light source specified by the user is defined as the region to the right of line 921, and the "left side" of the direction of the virtual light source specified by the user is defined as the region to the left of line 921. Therefore, when a user specifies "backlit, right side" as the general direction of the virtual light source, the virtual light source is set to the right and further back than region 901.
[0068] Furthermore, if a virtual light source is located in the approximate direction of "backlit, right side" specified by the user, the shadow of the subject cast by the virtual light source is cast on the left front side of the subject. Therefore, virtual light source direction setting unit 108 selects point 918 on the left front side in second region 902 within the angle of view for "backlit, right side" specified by the user as the approximate direction of the virtual light source. Virtual light source direction setting unit 108 obtains line 919 connecting selected point 918 and representative value 903 of the first region. Then, virtual light source direction setting unit 108 determines direction 920 of the virtual light source defined by line 919 and line 921. In this way, by defining the direction of the virtual light source by line 919, it is possible to determine the direction of the virtual light source that allows the first region to cast a shadow on the second region in the approximate direction of the virtual light source specified by the user.
[0069] The general direction of the virtual light source specified by the user can be set to "front-lit, center" or the like in addition to "backlit, right side," and the direction of the virtual light source can be determined in the same way as in the case of "backlit, right side." In this way, by determining the direction of the virtual light source based on the general direction of the virtual light source specified by the user, it is possible to generate an image with shadows from the direction desired by the user.
[0070] A case where the approximate direction of the virtual light source specified by the user is a top light will be described with reference to FIG. 10(B). FIG. 10(B) is a diagram illustrating the positional relationship between the image capture device 110 and each subject when viewed from the x direction. When the approximate direction of the virtual light source is a top light, the virtual light source direction setting unit 108 determines, as the direction of the virtual light source, a line 922 drawn perpendicularly from the representative value 903 of the region 901 detected as the first region to the second region 902. The direction of the top light in the first region 901, which is the subject, is a direction perpendicular to the surface (second region 902) on which the subject is placed, such as a floor, and the direction of the second region 902 can be calculated from the normal information acquired in step S303. In this way, even when the approximate direction of the virtual light source specified by the user is a top light, it is possible to determine the direction of the virtual light source. This makes it possible to generate a shaded image from the top light direction desired by the user. In this embodiment, an example has been described in which the user indicates the general direction of the virtual light source, but this is not limited to this. For example, the user may specify the direction of a shadow to be cast by the relighting process.
[0071] Returning to the explanation of Fig. 3, in step S306, image generation unit 109 performs image processing (relighting processing) to add shading of the subject to the image based on the direction of the virtual light source determined in step S305 and the shape information acquired in step S303, thereby generating a shaded image. The method of generating a shaded image will be described with reference to Fig. 11.
[0072] FIG. 11 is a diagram illustrating a method for adding a shadow of an object to an image based on the direction of the determined virtual light source. Virtual light source 1101 indicates the direction of the virtual light source uniquely set from the direction of the virtual light source limited in step S305. Image generation unit 109 generates a shadow 1102 based on shape information of object 1103, which is a first region, and the direction of virtual light source 1101, and adds the shadow to a second region on the image. Methods for generating a shadow based on shape information and the direction of the virtual light source include, for example, ray tracing and shadow mapping. Using these methods, an image with a shadow can be generated. Adding a shadow in this manner makes it possible to generate a shaded image in which the first region projects a shadow onto the second region. Furthermore, it is also possible to add a shadow to the first region by using normal information of the first region and the direction of virtual light source 1101.
[0073] When generating a shaded image, it is also possible to generate a shaded image using multiple virtual light sources by setting the directions of multiple virtual light sources. A method for generating a shaded image using multiple virtual light sources will be described with reference to FIG. 12. FIG. 12 is a diagram illustrating a method for adding shadows of a subject to an image according to the directions of multiple virtual light sources. Any number of virtual light sources can be set as long as they are within the range of the virtual light source directions limited in step S305. FIG. 12 describes an example in which three virtual light sources (virtual light source 1201 to virtual light source 1203) are set.
[0074] Shadow 1204 of subject 1103 generated by each virtual light source can be generated by using the ray tracing method or the shadow map method, just as in the case of one virtual light source. Also, just as in the case of one virtual light source, it is possible to add a shadow from each virtual light source to subject 1103 itself, which is the first region, by using normal information of the first region and the direction of each virtual light source. By adding shadows and shades in this way, it is possible to generate a shaded image even in the case of multiple virtual light sources.
[0075] As described above, according to this embodiment, by limiting the direction of a virtual light source that can cast the shadow of a first area on a second area, it is possible to set the direction of a virtual light source that allows the main subject to generate appropriate shadows on other subjects.
[0076] (Other embodiments) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program.The present invention can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.
[0077] Although the preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments and various modifications and changes are possible within the scope of the gist of the present invention. [Explanation of symbols]
[0078] 100 Image processing device 101 Control section 102 memory 103 Parameter setting section 104 Image acquisition unit 105 Shape information acquisition diagram 106 First area detection unit 107 Second area detection unit 108 Virtual light source direction setting unit 109 Image Generation Unit 110 Imaging device
Claims
1. a shape acquisition means for acquiring at least normal information and a reliability of the normal information as shape information of the object; a first area detection means for detecting a first area that generates a shadow of a subject; a second area detection means for detecting a second area onto which the shadow is cast; a direction setting means for setting a direction of a virtual light source in which the first region projects the shadow onto the second region, based on the normal information, the reliability of the normal information, the first region, and the second region; The image processing device is characterized in that the direction setting means sets the direction of the virtual light source so that areas in the second area where the reliability of the normal information is lower than a threshold value are excluded from the area onto which the shadow is cast.
2. The image processing device according to claim 1 , wherein the shape information includes distance information of the subject.
3. The image processing device according to claim 2 , wherein the shape information includes a reliability of the distance information.
4. 4. The image processing device according to claim 3, wherein the direction setting means sets the direction of the virtual light source so that an area in the second area where the reliability of the distance information is low is excluded from the area onto which the shadow is cast.
5. 5. The image processing apparatus according to claim 1, wherein the first area detection means detects the first area from an area near the center of the image height of the image.
6. 5. The image processing apparatus according to claim 1, wherein the first area detection means detects the first area from an area in the vicinity of a focus position.
7. A shape acquisition means for acquiring at least normal information of an object as shape information of the object; a first area detection means for detecting a first area that generates a shadow of a subject; a second area detection means for detecting a second area onto which the shadow is cast; a direction setting means for setting a direction of a virtual light source that causes the first region to project the shadow onto the second region based on the shape information, the first region, and the second region; The second region detection means calculates a histogram of the surface normals of the subject from the normal information of the subject, and detects the region having the surface normal with the highest frequency in the histogram as the second region.
8. 8. The image processing apparatus according to claim 1, wherein the first area detection means and the second area detection means detect the first area and the second area by image recognition.
9. 9. The image processing device according to claim 1, wherein the first area and the second area are mutually exclusive areas.
10. The image processing device according to any one of claims 1 to 9, characterized in that the direction setting means sets the direction of a virtual light source based on a representative value of the first region and shape information of the second region so that the first region can cast a shadow onto the second region.
11. 11. The image processing device according to claim 1, wherein the direction setting means sets the direction of the virtual light source based on an effective distance measurement range calculated from a defocus amount.
12. 12. The image processing device according to claim 1, wherein the direction setting means sets the direction of the virtual light source in accordance with an approximate direction of the virtual light source designated by a user.
13. An image processing device described in any one of claims 1 to 12, further comprising an image generation means for generating an image to which the shadow has been added based on the shape information and the set direction of the virtual light source.
14. 14. The image processing apparatus according to claim 13, wherein the image generating means generates the image to which the shadow is added based on the shape information and the directions of a plurality of virtual light sources.
15. 15. An imaging device comprising: the image processing device according to claim 1; and imaging means for capturing a plurality of images by receiving light beams passing through different pupil regions of an imaging optical system, The imaging apparatus is characterized in that the shape acquisition means calculates distance information of the subject from the amount of defocus or the amount of image shift obtained from the plurality of images.
16. An image processing method for adding a shadow to an image, comprising: a shape acquisition step of acquiring at least normal information and reliability of the normal information as shape information of the object; a first region detection step of detecting a first region that generates a shadow of the subject; a second area detection step of detecting a second area onto which the shadow is cast; a direction setting step of setting a direction of a virtual light source in which the first region projects the shadow onto the second region, based on the normal information, the reliability of the normal information, the first region, and the second region; An image processing method characterized in that, in the direction setting step, the direction of the virtual light source is set so that areas in the second region where the reliability of the normal information is lower than a threshold value are excluded from the area where the shadow is cast.
17. An image processing method for adding a shadow to an image, comprising: a shape acquisition step of acquiring at least normal information of the object as shape information of the object; a first region detection step of detecting a first region that generates a shadow of the subject; a second area detection step of detecting a second area onto which the shadow is cast; a direction setting step of setting a direction of a virtual light source in which the first region projects the shadow onto the second region based on the shape information, the first region, and the second region; an image generating step of generating an image to which the shadow is added based on the shape information and the set direction of the virtual light source, An image processing method characterized in that in the second region detection step, a histogram of the surface normals of the subject is calculated from the normal information of the subject, and the region having the surface normal with the highest frequency in the histogram is detected as the second region.
18. A program for causing a computer to function as each of the means of the image processing apparatus according to any one of claims 1 to 14.
Citation Information
Patent Citations
Image processing apparatus, image processing program, and camera
JP2009053748A
Image processing apparatus and image processing method, imaging apparatus, program
JP2018117211A
Image processor, image processing method and program
JP2019082958A
Image processing apparatus
JP2020010168A