Image processing device, imaging apparatus, image processing method, program, and storage medium

The image processing device addresses the issue of photographer's shadow in generating three-dimensional models by masking and adjusting shooting positions, ensuring accurate 3D shape and texture estimation.

JP2025152640APending Publication Date: 2025-10-10CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024054630
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-28
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Conventional techniques for generating three-dimensional models from multiple viewpoints can be degraded by the photographer's shadow, especially under strong light sources, leading to erroneous estimation and texture differences in the generated 3D model.

Method used

An image processing device that acquires and processes multiple images from different viewpoints, extracts and masks the photographer's shadow regions, calculates missing areas, and adjusts shooting positions to generate a three-dimensional model without using shadow area information, optionally using additional lighting to capture missing areas.

Benefits of technology

Reduces the influence of the photographer's shadow, resulting in a high-quality three-dimensional model by accurately estimating the 3D shape and texture without shadow-related errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025152640000001_ABST
    Figure 2025152640000001_ABST
Patent Text Reader

Abstract

To reduce the influence of a photographer's shadow appearing on a subject when generating a three-dimensional model by capturing the subject from multiple directions.SOLUTION: Included are acquisition unit that acquires multiple images of a subject captured from multiple viewpoints, a generation unit that generates a three-dimensional model of the subject using multiple images, and an extraction unit that extracts areas on the subject where the position changes at each of the multiple viewpoints using the multiple images. The generation unit generates the three-dimensional model from the multiple images without using the image information of the areas extracted by the extraction unit.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image processing device that performs processing to generate a three-dimensional model using a plurality of images obtained by capturing images from a plurality of viewpoints. [Background technology]

[0002] There is known a technique for generating a three-dimensional model using multiple captured images obtained by placing multiple imaging devices at different positions to capture images, or by moving a single imaging device to different positions to capture images multiple times.

[0003] Specifically, the system estimates the camera's shooting position and angle from still images, extracts common feature points from multiple images as point cloud data, and estimates the 3D shape. It also generates a 3D model by applying a texture corresponding to the 3D shape from the input image.

[0004] When generating such a three-dimensional model, the quality of the generated three-dimensional model may be reduced depending on the conditions of the captured image. For example, Patent Document 1 discloses a technology that uses a depth image of the captured scene to remove areas affected by blur or diffraction from the captured image, and estimates the three-dimensional shape based on the captured image after removal. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Japanese Patent Application Publication No. 2020-166498 Summary of the Invention [Problem to be solved by the invention]

[0006] However, when capturing images from multiple viewpoints by moving a single imaging device, the photographer's shadow may be cast on the subject under a strong light source, such as outdoors on a sunny day. In this case, the photographer's shadow moves depending on the shooting position. This results in erroneous estimation and differences in texture between images, degrading the quality of the generated 3D model when estimating 3D shapes using conventional techniques.

[0007] The present invention has been made in consideration of the above-mentioned problems, and its purpose is to reduce the influence of the photographer's shadow reflected on a subject when photographing the subject from multiple directions to generate a three-dimensional model. [Means for solving the problem]

[0008] The image processing device of the present invention comprises an acquisition means for acquiring a plurality of images of a subject taken from a plurality of viewpoints, a generation means for generating a three-dimensional model of the subject using the plurality of images, and an extraction means for extracting an area on the subject whose position changes for each of the plurality of viewpoints using the plurality of images, wherein the generation means generates the three-dimensional model from the plurality of images without using image information of the area extracted by the extraction means. [Effects of the Invention]

[0009] According to the present invention, when a subject is photographed from a plurality of directions to generate a three-dimensional model, it is possible to reduce the influence of the photographer's shadow reflected on the subject. [Brief explanation of the drawings]

[0010] [Figure 1] 1 is a diagram showing an example of the arrangement of an imaging apparatus according to a first embodiment of the present invention. [Figure 2] FIG. 2 is a diagram showing an example of the arrangement of an image processing unit according to the first embodiment. [Figure 3A] 4 is a flowchart showing the operation of the imaging apparatus according to the first embodiment. [Figure 3B]4 is a flowchart showing the operation of the imaging apparatus according to the first embodiment. [Figure 3C] 4 is a flowchart showing the operation of the imaging apparatus according to the first embodiment. [Figure 4] 3A and 3B are diagrams showing examples of subjects and shooting positions in the first embodiment. [Figure 5] FIG. 4 is a diagram showing an example of a shooting position displayed to a user in the first embodiment. [Figure 6] FIG. 2 is a diagram showing an example of a captured image in the first embodiment. [Figure 7] FIG. 3 is a diagram showing an example of a mask image according to the first embodiment. [Figure 8] FIG. 4 is a diagram showing an example of a missing area in the first embodiment. [Figure 9] FIG. 4 is a diagram showing an example of an additional shooting position in the first embodiment. [Figure 10] FIG. 10 is a diagram showing an example of the arrangement of an imaging apparatus according to a second embodiment. [Figure 11] 10 is a flowchart showing the operation of an imaging apparatus according to a second embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0011] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the invention claimed. Although multiple features are described in the embodiments, not all of these multiple features are necessarily essential to the invention, and multiple features may be combined arbitrarily. Furthermore, in the accompanying drawings, the same reference numerals are used to designate the same or similar components, and redundant explanations will be omitted.

[0012] Furthermore, unless otherwise specified, the contents, relative positions, etc. of the described components are not intended to limit the scope of the present invention to those. In the following, an imaging device will be described as an example of an image processing device, but the image processing device of the present invention is not limited to an imaging device and may be, for example, a personal computer (PC), etc.

[0013] (First embodiment) In this embodiment, a process for generating a three-dimensional model of a subject based on a plurality of images obtained by a user photographing the subject using a single imaging device while changing the photographing position, i.e., the viewpoint, will be described.

[0014] FIG. 1 is a block diagram showing the configuration of an image capturing apparatus 100 which is a first embodiment of an image processing apparatus according to the present invention.

[0015] The optical system 101 includes a lens group consisting of a zoom lens and a focus lens, an aperture adjustment device, and a shutter device. The optical system 101 adjusts the magnification, focus position, or light amount of the subject image that reaches the imaging unit 102.

[0016] The image capturing unit 102 includes an image capturing element 102a, such as a CCD or CMOS sensor, that photoelectrically converts the light beam of the subject that has passed through the optical system 101 into an electrical signal.

[0017] The A / D converter 103 converts the analog image signal output from the imaging unit 102 into a digital image signal.

[0018] The image processing unit 104 performs image processing such as shadow area detection processing and mask processing in this embodiment in addition to normal signal processing. The image processing unit 104 can perform similar image processing not only on images output from the A / D conversion unit 103 but also on images read out from the recording unit 107.

[0019] The control unit 105 is, for example, a CPU, and controls the operation of each block included in the imaging device 100 by executing a control program stored in the memory 109. For example, the control unit 105 calculates the amount of exposure during shooting to obtain an input image with appropriate brightness, and to achieve this, controls the aperture, shutter speed, analog gain of the image sensor 102a, and the like in the optical system 101 and the imaging unit 102.

[0020] The display unit 106 functions as an electronic viewfinder (EVF) by sequentially displaying images output from the image processing unit 104 on a display member such as an LCD.

[0021] The recording unit 107 has a function of recording images, and may include, for example, an information recording medium such as a memory card equipped with a semiconductor memory or a package containing a rotary recording medium such as a magneto-optical disk.

[0022] The operation unit 108 includes input members such as buttons and receives operations by the user. If the display unit 106 is equipped with a touch panel, touch operations are also treated as user operations on the operation unit 108.

[0023] FIG. 2 is a diagram showing functional blocks of the image processing unit 104.

[0024] The signal processing unit 201 performs normal signal processing such as noise reduction and development, as well as processing such as tone compression using gamma conversion to compress tone into a predetermined output range.

[0025] A camera position and orientation calculation unit 202 calculates the position and orientation of the image capturing device 100 at the time of capturing an image. A region extraction unit 203 extracts a shadow region of the photographer from the input image based on the position and orientation of the image capturing device 100 calculated by the camera position and orientation calculation unit 202. A mask processing unit 204 masks the region extracted by the region extraction unit 203 from the input image.

[0026] A three-dimensional model generation unit 205 generates a three-dimensional model of the photographed scene by estimating a three-dimensional shape and generating a texture from the input image masked by the mask processing unit 204. A photographing position calculation unit 206 calculates the photographing position required for generating the three-dimensional model.

[0027] The missing area calculation unit 207 calculates an area that is missing for generating a three-dimensional model from the input image masked by the mask processing unit 204 and the shooting position calculated by the shooting position calculation unit 206. The shooting position calculation unit 206 calculates a new shooting position that is required based on the missing area calculated by the missing area calculation unit 207.

[0028] The display image generating unit 208 generates an image for notifying the user of the image capturing position calculated by the image capturing position calculating unit 206 .

[0029] 3A to 3C are flowcharts showing a process for generating a three-dimensional model of a subject using multiple images obtained by a user photographing the subject from various viewpoints while moving. Here, it is assumed that the internal parameters of the camera of the image capture device 100 have already been acquired by photographing a known pattern and by a known calibration method. The operations of the flowcharts shown in FIGS. 3A to 3C are realized by the control unit 105 executing a control program stored in the memory 109.

[0030] 4 is a diagram showing a shooting scene, in which a subject 401 is shot under the environment of a light source 411.

[0031] First, the overall operation of the imaging device 100 will be described with reference to FIG. 3A.

[0032] In S301, the control unit 105 of the imaging device 100 uses the imaging position calculation unit 206 to calculate imaging positions for capturing images of the subject from multiple directions. Here, an initial imaging position is calculated. For example, like imaging positions 421 and 422 shown in FIG. 4, imaging positions are arranged at equal intervals with the subject 401 at the center. Note that the dashed line in FIG. 4 indicates the orientation of the imaging device 100.

[0033] In S302, the control unit 105 performs a photographing process. This photographing process will be described with reference to FIG.

[0034] In S311, the control unit 105 acquires an image for live view to be displayed on the display unit .

[0035] In S312, the control unit 105 calculates the position and orientation of the imaging device 100 that acquired the image in S311. For example, the calculation is performed using a known technique that uses image features, such as SfM (Structure from Motion), based on the image acquired in S311. Furthermore, the calculation is not limited to this, and an acceleration sensor or an angular velocity sensor may also be used. Furthermore, the calculation may also be performed using a combination of these.

[0036] In S313, the control unit 105 instructs the user on the image capturing position where the image capturing device 100 should capture an image, based on the position and orientation of the image capturing device 100 calculated in S301. An example of instructions to the user is shown in Fig. 5. Fig. 5 shows an image displayed on the display unit 106. Here, the display image generation unit 208 generates an image in which a rectangular virtual object 501 is superimposed on the input image. In Fig. 5, the virtual object 501 represents the position of the image capturing device 100.

[0037] The user can acquire an image based on this shooting position by moving the position of the imaging device 100 so that it overlaps with the virtual object 501. Here, only the virtual object 501 is displayed, but in consideration of the case where the virtual object 501 is not displayed on the screen, an auxiliary arrow pointing to the position of the virtual object 501 may be displayed on the screen. Furthermore, in addition to a graphical UI, distance information from the current position to the virtual object 501 may be displayed as a numerical value. Furthermore, audio information or tactile information may be used in addition to image information.

[0038] In S314, the control unit 105 determines whether or not the user has performed an image capturing operation. The processes of S311 to S313 are repeated until the user has performed an image capturing operation.

[0039] In S315, the control unit 105 acquires images for generating a three-dimensional model in response to the user's imaging operation. Here, it is described that the user actively performs the imaging operation, but the imaging device may automatically capture an image when the position and orientation of the imaging device 100 matches the shooting position.

[0040] This concludes the description of the photographing process S302.

[0041] Returning to the explanation of FIG. 3A, in S303, the control unit 105 extracts the photographer's shadow region from the image for generating a three-dimensional model. An example of the photographer's shadow region is shown in FIG. 6. Images 610 and 620 are images captured at shooting positions 421 and 422, respectively. Due to light source 411, the photographer's shadow is present in the image as shadow region 611 and shadow region 621, respectively. This photographer's shadow region leads to a decrease in the accuracy of the three-dimensional model, so the photographer's shadow region is extracted.

[0042] For example, the position of the photographer's shadow region in real space (on the subject) moves depending on the shooting position. Here, a projection matrix from image 620 to image 610 is calculated based on the position and orientation of image capturing device 100 when image 610 was captured and the position and orientation of image capturing device 100 when image 620 was captured. Image 630 is obtained by projecting image 620 using this projection matrix. Images 630 and 610 have approximately the same angle of view. By calculating the difference between image 630 and image 610, only the photographer's shadow region can be extracted as a difference region. In this way, multiple images projected based on the position and orientation of image capturing device 100 are compared, and regions with significantly different brightness are extracted as the photographer's shadow region.

[0043] Furthermore, for example, a motion vector representing the movement of the image capture device 100 and a motion vector representing the movement of a shadow indicate different movements. Here, a motion vector for each pixel or region of the image is calculated from the displacement for each frame of the display image acquired in S311. Similarly, the displacement for each frame of the position and orientation of the image capture device 100 acquired in S312 is calculated. From these, a region where the displacement of the position and orientation of the image capture device 100 and the motion vector move differently may be extracted as the shadow region of the photographer.

[0044] In S304, the control unit 105 generates a mask image based on the shadow area of ​​the photographer. An example of the mask image is shown in Fig. 7. Mask areas 711 and 721 in mask images 710 and 720 are generated based on shadow areas 611 and 621, respectively.

[0045] In S305, it is determined whether or not photography has been completed at each of the multiple photography positions calculated in S301. S302 to S304 are repeated until photography has been completed.

[0046] In S306, the control unit 105 additionally captures the missing images. This will be described with reference to the flowchart in FIG.

[0047] In S321, the control unit 105 calculates the areas that are insufficient for generating a 3D model. The concept of the insufficient areas is shown in Fig. 8. A group of rectangular areas 801 is formed by projecting captured mask images based on the position and orientation of the image capture device 100 and arranging them on the same plane. Here, it is determined whether there are enough images necessary for generating a 3D model.

[0048] For ease of understanding, it is assumed that at least one unmasked area is required for each region to generate a 3D model. For example, although a masked area exists in region 811, images 801a and 801b exist that have unmasked areas that cover region 811, making it a region for which a 3D model can be generated. However, the cross-hatched region in region 812 cannot be covered by the unmasked regions of other images, such as image 801a, image 801c, image 801d, image 801e, and image 801f. Therefore, it is a region for which a 3D model cannot be generated. In this embodiment, the cross-hatched region in region 812 is referred to as a missing region. In this manner, the missing region for 3D model generation is calculated.

[0049] In S322, the control unit 105 determines whether or not there is a missing area. If there is no missing area, the additional image capturing process ends, and if there is a missing area, the process proceeds to S323 to execute the additional image capturing process.

[0050] In S323, the control unit 105 calculates an additional shooting position (additional imaging position) based on the shooting position where the missing area occurs. The additional shooting position is shown in FIG. 9. For example, it is considered effective to acquire images surrounding the image where the missing area occurs. For this reason, shooting positions 901 and 902 are added near the shooting position 421 where the missing area occurs. Also, for example, it is considered effective to capture an image of the missing area from another direction. For this reason, shooting positions 911 and 912 are added from shooting positions surrounding the shooting position 421 where the missing area occurred, with the line of sight of the imaging device 100 facing the missing area. Also, for example, it is considered effective to prevent a shadow from entering the angle of view by having the photographer take the image from a distance. Therefore, shooting position 921 is added, which is farther from the subject in the same direction as the shooting position 421 where the missing area occurs. The control unit 105 notifies the user by displaying the calculated additional shooting position on the display unit 106.

[0051] The photographing process in S324, the extraction of the photographer's shadow area in S325, and the generation of the mask image in S326 are similar to the photographing process in S302, the extraction of the photographer's shadow area in S303, and the generation of the mask image in S304, respectively, and therefore will not be described here.

[0052] Returning to the explanation of FIG. 3A, in S307, the control unit 105 estimates the three-dimensional shape of the photographed subject. For example, a known technique such as MVS (Multi-View Stereo) is used to estimate the three-dimensional shape. In estimating the three-dimensional shape, an image for generating a three-dimensional model and a mask image are input, and the masked area is not used in estimating the three-dimensional shape. This reduces erroneous estimation of the three-dimensional shape due to the shadow area of ​​the photographer.

[0053] In S308, the control unit 105 generates a texture and applies it to the estimated three-dimensional shape. Here, as in S307, the image for generating the three-dimensional model and the mask image are input to generate the texture, and the masked area is not used to generate the texture. This makes it possible to reduce degradation of the texture quality due to the shadow area of ​​the photographer.

[0054] This completes the description of the process for generating a three-dimensional shape model of a subject using a plurality of images obtained by photographing the subject from various viewpoints while the user is moving.

[0055] With the above configuration, it is possible to generate a high-quality three-dimensional shape model by appropriately masking the shadow area of ​​the photographer in the captured image and capturing sufficient additional images.

[0056] In this embodiment, we have described rule-based 3D model generation using MVS, but other methods are also effective. For example, neural rendering such as NeRF (Neural Radiance Fields) is also effective. For example, by inputting a training image in which the photographer's shadow area has been removed using a mask image and learning the relationship with the 3D model, higher-quality neural rendering becomes possible.

[0057] Furthermore, although the present embodiment has been described as a process for reducing the influence of the photographer's shadow, the present invention is also effective in other areas besides processing the photographer's shadow. For example, highlight components (highlight areas) such as specular reflections on the subject also change depending on the shooting position, so the influence of these on the generation of a 3D model can be reduced using a similar technique.

[0058] (Second embodiment) Next, a second embodiment of the present invention will be described. In this second embodiment, a case where shooting cannot be performed at an additional shooting position even when a missing area as described in the first embodiment occurs due to limitations in the shooting environment will be described. Note that parts similar to those in the first embodiment will be assigned the same reference numerals and descriptions thereof will be omitted. Also, descriptions of processes similar to those in the first embodiment and processes using known techniques will be omitted. Next, a description will be given of the process that is the main feature of this embodiment.

[0059] FIG. 10 is a block diagram showing the configuration of an image capturing apparatus 1000 which is a second embodiment of the image processing apparatus of the present invention.

[0060] 10, the configuration is the same as that of the first embodiment in Fig. 1 except for the light emitting unit 1009. The light emitting unit 1009 is, for example, a strobe or LED light, and emits light to illuminate the subject when capturing an image.

[0061] Fig. 11 is a flowchart showing the process of generating a three-dimensional shape of a subject using multiple images obtained by a user photographing the subject from various viewpoints while moving around in this embodiment. The process in Fig. 11 corresponds to the process of S306 in Fig. 3A of the first embodiment. Here, the overall flow and photographing process are the same as those in Figs. 3A and 3B, and therefore their explanations are omitted. Furthermore, the processes in S321 to S323 and S324 to S326 are the same as those in Fig. 3C of the first embodiment, and therefore their explanations are omitted.

[0062] In S1107, the control unit 105 displays a message prompting the user to determine whether the additional image capturing position calculated in S323 is a position where image capturing is possible. The user performs an operation based on this message. For example, if the user cannot move to the specified position due to limitations in the image capturing environment, the user inputs to the image capturing device 1000 using the operation unit 1008 that image capturing is not possible.

[0063] In S1108, the control unit 105 determines whether the light emitting unit 1009 can emit light. For example, it determines whether the light intensity of the light emitting unit 1009 is sufficient for the shooting environment. Also, for example, it determines whether there is enough battery power left for the light emitting unit 1009 to emit light. Also, for example, the user inputs whether or not the current shooting environment prohibits light emission.

[0064] In S1109, the control unit 105 instructs the user to take an image at the same photographing position where the missing area occurs, and acquires an image by emitting light from the light emitting unit 1009. This makes it possible to capture an image without the photographer's shadow.

[0065] In S1110, the control unit 105 performs color correction on the image captured by emitting light from the light emitting unit 1009. Since the image captured by emitting light has different colors compared to images captured under other shooting conditions, color correction must be performed to match the colors. At this time, color correction may be performed using an unmasked area of ​​an image captured at a nearby shooting position.

[0066] In S1111, the control unit 105 performs processing when light emission is not possible. Here, only the area determined to be a missing area, i.e., the area where the photographer is in shadow, is extracted. The reason for extracting only the shadow area is to reduce the influence of the shadow edge. Because the color of this area differs from other areas due to the shadow, color correction must be performed to match the colors. At this time, color correction may be performed using an unmasked area of ​​an image captured at a nearby shooting position, as in S1110.

[0067] In estimating the three-dimensional shape, as in the first embodiment, an image for generating a three-dimensional model and a mask image are input, and the masked area is not used to estimate the three-dimensional shape. Furthermore, in this embodiment, if there are insufficient feature points for estimating the three-dimensional shape, the three-dimensional shape is estimated by adding the color-corrected luminescent image in S1110 and the color-corrected missing area in S1111. This reduces erroneous estimation of the three-dimensional shape due to the shadow area of ​​the photographer.

[0068] In texture generation, as in the first embodiment, an image for generating a 3D model and a mask image are input, and the masked area is not used to generate the texture. Also, in this embodiment, if there is insufficient information for texture generation, the texture is generated by adding the luminescent image color-corrected in S1110 and the missing area color-corrected in S1111. This makes it possible to reduce degradation of texture quality due to the photographer's shadow area.

[0069] This concludes the explanation of the process when a missing area occurs and additional imaging cannot be performed at the additional imaging position. The above process makes it possible to reduce the deterioration of the accuracy of the 3D shape model by correcting images and areas with different conditions.

[0070] It should be noted that some or all of the image processing described in the above embodiments may be executed by a device (such as a computer) external to the device (such as a camera) used to capture the image.

[0071] The disclosure of this specification includes the following image processing device, imaging device, image processing method, program, and storage medium.

[0072] (Item 1) an acquisition means for acquiring a plurality of images of a subject taken from a plurality of viewpoints; a generating means for generating a three-dimensional model of the subject using the plurality of images; an extraction means for extracting an area where a position on the subject changes from each of the plurality of viewpoints using the plurality of images, The image processing device is characterized in that the generating means generates the three-dimensional model from the plurality of images without using image information of the region extracted by the extracting means.

[0073] (Item 2) 2. The image processing device according to item 1, further comprising a processing unit that generates a mask image by masking the region extracted by the extraction unit, wherein the generation unit generates the three-dimensional model without using image information of the masked region in the mask image.

[0074] (Item 3) 3. The image processing device according to item 2, further comprising a first calculation means for calculating, based on the mask image, a missing region where image information for generating the three-dimensional model is missing.

[0075] (Item 4) 4. The image processing device according to item 3, wherein the first calculation means further calculates an additional imaging position, which is a position of an imaging device that captures an image to supplement the image information of the missing area.

[0076] (Item 5) 5. The image processing device according to item 4, further comprising a notification unit that notifies a user of the additional imaging position.

[0077] (Item 6) 6. The image processing device according to item 4 or 5, wherein the generating means generates the three-dimensional model using image information captured at the additional imaging position and the plurality of images.

[0078] (Item 7) 7. The image processing device according to any one of items 1 to 6, wherein the area on the subject whose position has changed is an area of ​​the user's shadow.

[0079] (Item 8) 7. The image processing device according to any one of items 1 to 6, wherein the region on the subject whose position has changed is a highlight region caused by specular reflection on the subject.

[0080] (Item 9) 9. The image processing device according to any one of items 1 to 8, further comprising a second calculation means for calculating a position and orientation of an image capturing device that captured the plurality of images, wherein the extraction means extracts an area whose position on the subject is changing based on the position and orientation of the image capturing device.

[0081] (Item 10) 9. The image processing device according to any one of items 1 to 8, further comprising a third calculation means for calculating a motion vector for each pixel or region of the plurality of images, wherein the extraction means uses the motion vector to extract a region whose position on the subject is changing.

[0082] (Item 11) 4. The image processing device according to item 3, wherein the generating means generates the three-dimensional model by further using an image captured by illuminating the missing region with a light emitting means.

[0083] (Item 12) Item 12. The image processing device according to item 11, further comprising a correction unit that corrects the color of an image captured by illuminating the missing area with a light emitting unit based on the color of an area that is not the missing area.

[0084] (Item 13) 3. The image processing device according to item 2, wherein the generating means generates the three-dimensional model by learning the relationship between the mask image and the three-dimensional model.

[0085] (Item 14) an imaging means for imaging a subject; An image processing device according to any one of items 1 to 13, An imaging device comprising:

[0086] (Item 15) an acquisition step of acquiring a plurality of images of a subject captured from a plurality of viewpoints; a generation step of generating a three-dimensional model of the subject using the plurality of images; an extraction step of extracting an area whose position on the subject changes at each of the plurality of viewpoints using the plurality of images, The image processing method is characterized in that in the generating step, the three-dimensional model is generated from the plurality of images without using image information of the region extracted in the extracting step.

[0087] (Item 16) A program for causing a computer to function as each of the means of the image processing device according to any one of items 1 to 13.

[0088] (Item 17) A computer-readable storage medium storing a program for causing a computer to function as each of the means of the image processing device according to any one of items 1 to 13.

[0089] (Other embodiments) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program.The present invention can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.

[0090] The invention is not limited to the above-described embodiments, and various changes and modifications can be made without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]

[0091] 100: imaging device, 101: optical system, 102: imaging section, 102a: imaging element, 103: A / D conversion section, 104: image processing section, 105: control section, 106: display section, 107: recording section, 108: operation section

Claims

1. an acquisition means for acquiring a plurality of images of a subject taken from a plurality of viewpoints; a generating means for generating a three-dimensional model of the subject using the plurality of images; an extraction means for extracting an area where a position on the subject changes from each of the plurality of viewpoints using the plurality of images, The image processing device is characterized in that the generating means generates the three-dimensional model from the plurality of images without using image information of the region extracted by the extracting means.

2. 2. The image processing device according to claim 1, further comprising a processing unit that generates a mask image by masking the region extracted by the extraction unit, wherein the generation unit generates the three-dimensional model without using image information of the masked region in the mask image.

3. 3. The image processing apparatus according to claim 2, further comprising a first calculation means for calculating a missing region where image information for generating the three-dimensional model is insufficient based on the mask image.

4. 4. The image processing apparatus according to claim 3, wherein the first calculation means further calculates an additional image capture position, which is a position of an image capture device that captures an image to supplement the image information of the missing area.

5. 5. The image processing apparatus according to claim 4, further comprising a notification unit that notifies a user of the additional image capturing position.

6. 5. The image processing apparatus according to claim 4, wherein the generating means generates the three-dimensional model using image information captured at the additional imaging position and the plurality of images.

7. 2. The image processing device according to claim 1, wherein the area on the subject whose position has changed is an area of ​​the user's shadow.

8. 2. The image processing apparatus according to claim 1, wherein the region on the subject whose position has changed is a highlight region caused by specular reflection on the subject.

9. 2. The image processing device according to claim 1, further comprising: a second calculation unit that calculates a position and orientation of an image capturing device that captured the plurality of images; and wherein the extraction unit extracts an area where the position on the subject is changing based on the position and orientation of the image capturing device.

10. 2. The image processing device according to claim 1, further comprising a third calculation means for calculating a motion vector for each pixel or region of the plurality of images, wherein the extraction means uses the motion vector to extract a region whose position on the subject is changing.

11. 4. The image processing apparatus according to claim 3, wherein the generating means generates the three-dimensional model by further using an image captured by illuminating the missing region with a light emitting means.

12. 12. The image processing device according to claim 11, further comprising a correction unit that corrects the color of an image captured by illuminating the missing area with a light emitting unit, based on the color of an area other than the missing area.

13. 3. The image processing apparatus according to claim 2, wherein the generating means generates the three-dimensional model by learning the relationship between the mask image and the three-dimensional model.

14. an imaging means for imaging a subject; An image processing device according to any one of claims 1 to 13; An imaging device comprising:

15. an acquisition step of acquiring a plurality of images of a subject captured from a plurality of viewpoints; a generation step of generating a three-dimensional model of the subject using the plurality of images; an extraction step of extracting an area whose position on the subject changes at each of the plurality of viewpoints using the plurality of images, The image processing method is characterized in that in the generating step, the three-dimensional model is generated from the plurality of images without using image information of the region extracted in the extracting step.

16. A program for causing a computer to function as each of the means of the image processing apparatus according to any one of claims 1 to 13.

17. 13. A computer-readable storage medium storing a program for causing a computer to function as each of the means of the image processing apparatus according to claim 1.

Citation Information

Patent Citations

  • Information processing apparatus, three-dimensional model generating method, and program

    JP2020166498A