Image processing apparatus, control method of image processing apparatus
The imaging device addresses the challenge of focusing on virtual objects in mixed reality by using an image processing apparatus to synthesize virtual objects and adjust the lens position, resulting in improved focus accuracy in mixed reality environments.
Patent Information
- Application Number
- JP2021058480
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-03-30
- Publication Date
- 2025-06-23
- Estimated Expiration
- 2041-03-30
AI Technical Summary
Existing imaging devices struggle to appropriately focus on virtual objects in mixed reality spaces, as they primarily rely on real-space information for focus detection and depth of field adjustment.
An imaging device equipped with an image processing apparatus that acquires image signals from an image sensor with pixels receiving light from different pupil regions, synthesizes virtual objects with these image signals, and adjusts the lens position based on image displacement between composite reality images.
Enables appropriate focusing on desired subjects in mixed reality spaces by considering virtual objects, improving the accuracy of focus detection and depth of field adjustment.
Smart Images

Figure 0007696742000001 
Figure 0007696742000002 
Figure 0007696742000003
Abstract
Description
Technical Field
[0001] The present invention relates to an imaging device, and particularly to an imaging device capable of photographing a mixed reality space.
Background Art
[0002] In recent years, technologies for synthesizing virtual objects in the real space for photographing and displaying have been known. These technologies are called augmented reality (AR) and mixed reality (MR), and are used in various applications such as industry and entertainment.
[0003] Patent Document 1 discloses an imaging device that generates real coordinates representing an imaging space imaged through a lens and a space outside thereof, and processes a synthesis target according to a synthesis position in the real coordinates, thereby appropriately synthesizing a virtual object arranged outside the angle of view of the imaging device.
[0004] In such a mixed reality space configured by fusing the real space and the virtual space, real objects and virtual objects are mixed.
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0006] In photographing a mixed reality space, a photographer takes a photograph while observing the mixed reality space. Therefore, the object to be focused can be determined to include virtual objects. However, since the imaging device disclosed in Patent Document 1 is configured to determine focus detection and depth of field based on information on the real space obtained through an optical system, there is a problem that an appropriate focus position and depth of field considering virtual objects cannot be adjusted.
[0007] In view of the above problems, an object of the present invention is to provide an imaging device capable of appropriately focusing in the shooting of a composite reality space.
Means for Solving the Problems
[0008] To solve the above problems, an image processing apparatus of the present invention includes: acquisition means for acquiring an image signal from an image sensor in which a plurality of pixels for receiving light passing through different pupil regions of an imaging optical system are arranged; and synthesis means for synthesizing a virtual object with each of the image signals corresponding to the different pupil regions using the image signal acquired by the acquisition means to generate a pair of composite reality images, and focus adjustment means for adjusting a lens position of the imaging optical system based on an amount of image displacement between the pair of composite reality images.
Effects of the Invention
[0009] It becomes possible to appropriately focus on a desired subject in the shooting of a composite reality space.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Mode for Carrying Out the Invention
[0011] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings.
[0012] (First Embodiment) In the first embodiment, in an imaging device having a so-called imaging surface phase difference type distance measuring means in which pixels having a phase difference detection function are arranged in an imaging element, an embodiment of focusing when imaging a composite reality space will be described.
[0013] FIG. 1 is a block diagram showing the configuration of an imaging device 100 according to the present invention.
[0014] The imaging optical system 101 guides light from the subject to the imaging element 106. The imaging optical system 101 includes a focusing lens 102, a diaphragm 103, and a lens group (not shown). The focusing lens 102 is driven in the optical axis direction by a drive control command from the lens drive means 104. The diaphragm 103 is driven by the diaphragm drive means 105 to have a predetermined aperture diameter to adjust the amount of light. The imaging element 106 is a pixel array in which unit pixels described later are two-dimensionally arranged, photoelectrically converts the received light beam, and outputs a pair of signals having a parallax. In the present embodiment, the imaging element 106 converts the analog signal output from the photoelectric conversion unit 203 into a digital signal via an A / D converter and outputs it. However, the present invention is not limited to this, and the imaging element 106 may output the analog signal in the state of the analog signal and may have an A / D converter separately from the imaging element 106. That is, the A-image signal and the B-image signal output from the imaging element 106 may be analog signals or digital signals.
[0015] Here, referring to FIG. 2, the details of the imaging device 106 will be described. FIG. 2(a) is a view of a partial region of the imaging device 106 seen from above. As shown in FIG. 2(a), the imaging device 106 is configured by arranging a plurality of unit pixels 200 two-dimensionally.
[0016] As shown in FIG. 2(b), the pixel 200 has two photoelectric conversion units 202 and 203 with respect to the microlens 201. The photoelectric conversion units 202 and 203 are designed to receive light beams that have passed through different pupil regions of the imaging optical system 101, respectively, and a pair of signals obtained by the photoelectric conversion units 202 and 203 have parallax. By detecting the phase difference between the pair of signals, focus detection and distance measurement can be performed. Hereinafter, for the sake of explanation, the signal that can be obtained by the photoelectric conversion unit 202 is called an A-image signal, and the signal that can be obtained by the photoelectric conversion unit 203 is called a B-image signal.
[0017] In this embodiment, the configuration is such that an A-image signal and a B-image signal obtained from a pair of photoelectric conversion units are output. However, not limited to this, the signals (charges) obtained by the photoelectric conversion unit 202 and the photoelectric conversion unit 203 may be mixed by floating diffusion and output as an A + B image. A pair of signal groups (A-image signal, B-image signal) output from the imaging device 106 are stored in a storage area provided in the CPU 114 that controls the entire imaging device, and transferred to various processing units via the CPU 114.
[0018] The real image generation unit 107 performs development processing including defective pixel correction and color conversion processing on the A-image signal and the B-image signal (a pair of image signals) output from the imaging device 106, and generates a pair of real space images.
[0019] The virtual object generation unit 108 generates a pair of virtual object models to be synthesized with each of the pair of real space images. The details of the virtual object model generation process will be described later.
[0020] The composite reality image generation unit 109 overlays and synthesizes a real space image and a virtual object model to generate a composite reality image. The composite reality image generation unit 109 generates an A-image composite reality image obtained by synthesizing a real image corresponding to the A-image signal and a virtual object model, and a B-image composite reality image obtained by synthesizing a real image corresponding to the B-image signal and a virtual object model. Alternatively, an A + B-image composite reality image obtained by adding and synthesizing the A-image composite reality image and the B-image composite reality image may be generated.
[0021] The display unit 110 is a display device composed of a liquid crystal display or the like, and performs live view display, a shooting setting screen, and display of a reproduced image. In the imaging device 100, during live view display, any one of the composite reality images generated by the composite reality image generation unit 111 is displayed.
[0022] The recording unit 111 is a storage medium such as an SD card, and records the generated A + B-image composite reality image.
[0023] The focus adjustment unit 112 calculates a defocus amount for each region within the angle of view based on the phase shift between the A-image composite reality image and the B-image composite reality image.
[0024] The calculated defocus amount is treated as depth information for autofocus or composite reality image synthesis.
[0025] The instruction unit 113 is a physical switch configured on the body of the imaging device 100, and is used for switching the shooting mode, specifying the focus detection position during autofocus, instructing the autofocus operation, instructing the start of exposure for actual shooting, and the like. Note that the instruction unit 113 may be a touch panel built into the display unit 112.
[0026] The CPU 114 is a CPU that controls the overall operation of the imaging device 100.
[0027] It performs overall flow control of shooting described later, and issues commands to the lens driving means 104 and the aperture driving means 105.
[0028] The posture detection unit 115 is composed of a gyro sensor and an acceleration sensor, and outputs the position and posture information of the imaging device.
[0029] Next, the shooting of the composite reality space by the imaging device 100 will be described.
[0030] FIG. 3 is a flowchart of the shooting operation of the composite reality space by the imaging device 100 controlled by the CPU 114. The operations of each step are performed by each processing unit according to the instruction of the CPU 114 or the CPU 114.
[0031] When shooting is started, it is determined whether the power is turned off in step S301.
[0032] If the power is not off, the process proceeds to step S302. If the power is off, the shooting process ends.
[0033] In step S302, the CPU 114 performs the live view display process of the composite reality image. Details of the display process of the composite reality image will be described later.
[0034] In step S303, it is determined whether SW1 configured in the instruction unit 113 is pressed. If SW1 is pressed, the process proceeds to step S304 to perform the focus adjustment process and then proceeds to step S305. If SW1 is not pressed, the process proceeds to step S305 without performing the focus adjustment process.
[0035] Details of the focus adjustment process will be described later. Also, an embodiment in which the focus adjustment process of step S304 is always performed even when SW1 is not pressed may be used.
[0036] In step S305, it is determined whether SW2 configured in the instruction unit 113 is pressed. If SW2 is pressed, the process proceeds to step S306 to perform the recording process of the composite reality image. If SW2 is not pressed, the process returns to step S301 and the series of operations is repeated.
[0037] Details of the recording process of the composite reality image will be described later. After executing the recording process of the composite reality image in step S306, the process returns to step S301 and a series of operations are repeated.
[0038] Next, referring to FIG. 4, details of the display process of the composite reality image will be described. FIG. 4 is a flowchart for explaining the display process of the composite reality image. The operations of each step are performed by each processing unit under the instruction of the CPU 114 or the CPU 114.
[0039] In step S401, exposure is performed at a period corresponding to the display rate, and the A-image signal and the B-image signal generated by imaging by the imaging device 106 are sequentially acquired. Next, in step S402, the position and orientation information is detected by the attitude detection unit 115. The detected position and orientation information is stored in association with the A-image signal and the B-image signal acquired in step S401 by the CPU 114. Next, in step S403, the depth information of the space within the viewing angle is detected. The depth information can be obtained by dividing the area within the viewing angle and detecting the phase difference between the A-image signal and the B-image signal for each divided area. The detected depth information is stored in association with the A-image signal and the B-image signal acquired in step S401 by the CPU 114. Subsequently, in step S404, a composite reality image (A-image composite reality image) corresponding to the A-image signal is generated, and in step S405, a composite reality image (B-image composite reality image) corresponding to the B-image signal is generated.
[0040] Here, referring to FIG. 5, details of the generation process of the composite reality image will be described. FIG. 5 is a flowchart showing the generation process of the composite reality image. The operations of each step are performed by each processing unit under the instruction of the CPU 114 or the CPU 114.
[0041] In step S501, development by various image processes is performed on the input predetermined image signal. Next, in S502, virtual object model information arranged in the real image is acquired. For example, in the method of arranging a marker for arranging a 3D model in the real space, the virtual object model information is acquired based on the marker detection information within the viewing angle. Subsequently, in step S503, the shape of the virtual object model as seen from the imaging device, that is, the posture of the virtual object model, is determined using the posture information of the imaging device that has been previously detected and stored in association with the image signal. Subsequently, in step S504, the size of the virtual object model as seen from the imaging device is determined based on the depth information of the position where the virtual object model is to be placed. Finally, in step S505, the virtual object model as seen from the imaging device is projected onto the real image and composite processing is performed to generate a composite reality image.
[0042] When the A-image composite reality image and the B-image composite reality image are generated in steps S404 and S405, next, in step S406, the pixel values of the same coordinates of the A-image composite reality image and the B-image composite reality image are added. Further, a composite image (A + B-image composite reality image) of the A-image and the B-image is generated, and the process proceeds to step S407. In step S407, the CPU 114 reshapes the A + B-image composite reality image to a size suitable for display and performs a live view display on the display unit 110. In this way, by synthesizing a virtual object at a position and / or size corresponding to the position and orientation of the imaging device for the A-image signal and the B-image signal sequentially acquired from the image sensor 106, generating a composite reality image, and sequentially displaying it on the display unit 110, a live view display of the composite reality image becomes possible.
[0043] Next, the details of the focus adjustment process will be described with reference to FIG. 6.
[0044] FIG. 6 is a flowchart of the focus adjustment process. The operations of each step are performed by each processing unit under the instruction of the CPU 114 or the CPU 114.
[0045] When the focus detection process starts, focus detection position information indicating the subject area to be focused on is acquired at step S601. The focus detection position information may be determined based on a user operation on the instruction unit 113, or may be the position information of the area corresponding to the subject determined as the main subject by image analysis. Subsequently, at step S602, the defocus amount at the focus detection position of the A-image composite reality image and the B-image composite reality image is calculated. Here, in the present embodiment, the correlation operation and the defocus amount are calculated not from the images indicated by the A-image signal and the B-image signal output from the imaging device 106, but from the composite reality image with a virtual object synthesized. Thereby, it is possible to acquire distance information (defocus amount) adjusted to the composite reality space after the virtual object is arranged.
[0046] As a method for detecting the defocus amount of the image signal, for example, there is a method using a correlation operation method called the SAD method (Sum of Absolute Difference). While shifting the relative positional relationship between the A-image and the B-image, the sum of the absolute differences between the image signals at each shift position is obtained, and the defocus amount can be detected by detecting the shift position where the sum of the absolute differences is the smallest.
[0047] At step 603, the CPU 114 calculates the lens driving amount using a predetermined coefficient for converting the calculated defocus amount into the lens driving amount, and issues a driving instruction to the lens driving means 104. The lens driving means 104 moves the focusing lens 102 according to the lens driving amount, and ends the focus adjustment process.
[0048] Next, with reference to FIG. 7, the recording process of the composite reality image will be described. FIG. 7 is a flowchart for explaining the recording process of the composite reality image. The operations of each step are performed by each processing unit under the instruction of the CPU 114 or the CPU 114.
[0049] In step S701, exposure is performed using the aperture value and shutter speed set in advance as shooting parameters, and the A-image signal and B-image signal generated by the imaging device 106 are acquired. The processing from step S702 to step S706 performs the same processing as steps S402 to S406 in FIG. 4, except that the input data is different. In FIG. 4, the input data is an image captured under the conditions of display exposure in step S401, and in FIG. 7, the input data is an image captured under the conditions of main shooting exposure (for recording) in step S701. Therefore, for example, the number of pixels and the number of bits of the input data may be different.
[0050] In step S707, the A + B image composite reality image generated in step S706 is written to the recording unit 111 to end the recording process of the composite reality image.
[0051] As described above, in the imaging device 100 according to the first embodiment, a composite reality image corresponding to each of a pair of image signals that have passed through different pupil regions of the imaging optical system 101 is generated, and focus adjustment is performed using the composite reality image. By doing so, an appropriate defocus amount can be obtained even when a virtual object exists.
[0052] (Second Embodiment) In the first embodiment, a configuration for focusing on a single subject has been described.
[0053] In the second embodiment, a configuration for focusing on a plurality of subjects in different depth directions will be described using an imaging device capable of designating a plurality of focus detection positions.
[0054] FIG. 8 shows the imaging flow of the imaging device 100 according to the second embodiment. The imaging device 100 focuses on a plurality of subjects in different depth directions by the focus adjustment process shown in step S804 and the aperture value determination process shown in step S805.
[0055] Here, the focus adjustment process of the second embodiment shown in step S804 will be described. The operations of each step are performed by each processing unit under the instruction of CPU 114 or CPU 114.
[0056] FIG. 9 is a flowchart of the focus adjustment process. When the focus adjustment process starts, focus detection position information is acquired in step S901. In this embodiment, it is assumed to include area information of two or more points.
[0057] Next, in step S902, a defocus amount is calculated for each of the plurality of focus detection positions. In step S903, a position for internally dividing the plurality of focus detection positions is determined. Specifically, the minimum value and the maximum value of the plurality of defocus amounts are detected, and a lens position for internally dividing the minimum value and the maximum value is obtained. At this time, the internally divided defocus amount is stored in the storage unit of CPU 114 and referred to in the aperture value determination process described later.
[0058] Next, in step S904, the lens is driven to the position for internally dividing the plurality of focus detection positions, and the focus adjustment process ends.
[0059] After the focus adjustment process of step S804, the aperture value determination process according to step S805 is executed. Specifically, the aperture value at the time of shooting is calculated by dividing the internally divided defocus amount obtained in step S903 by the allowable confusion circle diameter.
[0060] When SW2 is pressed in step S305, a composite reality image is generated based on the signal obtained by performing exposure with the aperture value obtained in step S805, and the image is written to the recording unit 111 in step S307.
[0061] As described above, a defocus amount corresponding to a plurality of focus detection positions is calculated using the composite reality image, and the lens is driven to the position for internally dividing the plurality of defocus amounts. Furthermore, by performing shooting with an aperture value such that the images of a plurality of subjects are within the allowable confusion circle at the position after lens driving, it becomes possible to focus on all the subjects corresponding to the plurality of focus detection positions.
[0062] (Third Embodiment) In the first and second embodiments, a method of calculating the defocus amount by performing phase difference method focus detection on a pair of composite reality images was shown. In the third embodiment, a method of updating the depth information in which a virtual object is arranged at the time of generating a composite reality image and focusing on a plurality of subjects based on the updated depth information will be described.
[0063] Hereinafter, for the content already described, the same reference numerals will be used and the description will be omitted, and only the differences will be described.
[0064] FIG. 10 is a flowchart of the composite reality image generation process in the third embodiment.
[0065] When a composite reality image is generated from a real image by a series of processes from step S501 to step S505, the depth information update process is performed in step S1006.
[0066] More specifically, the depth information of the region where the virtual object is arranged is updated to the depth information after the virtual object is arranged. Note that it is preferable to update the depth information in consideration of the depth of the virtual object itself with respect to the arrangement position of the virtual object.
[0067] FIG. 11 is a diagram for explaining the update of depth information.
[0068] 1101 is an object in the real space, and 1102 is an object in the virtual space arranged in front of (on the imaging device side) the real object 1101.
[0069] First, the depth information 1103 corresponding to the object 1101 in the real space is detected in step S403. Subsequently, the arrangement position 1104 of the virtual object 1102 is determined in step S503. 1104 is, for example, the barycentric coordinate position of the marker. In step S505, the virtual object 1102 is arranged so that its barycentric position coincides with 1104. Subsequently, in step S906, the depth information 1105 is determined in view of the depth information of the virtual object 1102 itself, and the depth information 1103 is updated to the new depth information 1105.
[0070] Next, the focus adjustment process according to the third embodiment will be described.
[0071] FIG. 12 is a flowchart of the focus adjustment process in the third embodiment.
[0072] After acquiring a plurality of focus detection position information in step S601, depth information of the focus detection position is acquired in step S1202. The depth information acquired here is the depth information updated in step S906 of FIG. 9. Therefore, it is depth information in which a virtual object is reflected.
[0073] Next, in step S1203, the closest depth information and the farthest depth information are detected from the depth information corresponding to each of the plurality of focus detection positions, and the driving position of the lens is determined so as to internally divide them.
[0074] In this way, by using the depth information after arranging the virtual object, the driving position of the lens can be determined without performing a correlation operation.
[0075] Here, an example of acquiring depth information of the real space using an imaging device equipped with an image sensor configured with divided pixels has been shown for the sake of explanation. However, the method of acquiring the depth information of the real space is not limited to this. For example, it may be acquired by a distance measurement sensor such as a TOF (Time of Flight) method. In that case, it is not necessary for the image sensor to constitute divided pixels.
[0076] After the focus adjustment process, the process transitions to the aperture value determination process for shooting in the same manner as step S705 in FIG. 7, and the aperture value at the time of shooting is determined.
[0077] As described above, in the third embodiment, the driving position of the lens and the aperture value are determined based on the depth information in which the virtual object is reflected.
Explanation of Reference Numerals
[0078] Virtual object generation unit 103 Image sensor 106 Real image generation unit 107 Composite reality image generation unit 109 Focus adjustment unit 112 CPU 114 Posture detection unit 115
Claims
1. Acquisition means for acquiring an image signal from an image sensor in which a plurality of pixels that receive light passing through different pupil regions of an imaging optical system are arranged; Synthesis means for synthesizing a virtual object with the image signals corresponding to the different pupil regions respectively using the image signal acquired by the acquisition means, and generating a pair of composite reality images; Focus adjustment means for adjusting the lens position of the imaging optical system based on the amount of image shift between the pair of composite reality images; An image processing apparatus, characterized by comprising the above.
2. Acquisition means for acquiring an image of the real space through an imaging optical system; Synthesis means for synthesizing a virtual object with the image of the real space and generating a composite reality image; Focus adjustment means for adjusting the lens position of the imaging optical system based on a plurality of depth information in the composite reality image; Aperture adjustment means for adjusting the aperture value at the time of shooting based on a plurality of depth information in the composite reality image; An image processing apparatus, characterized by comprising the above.
3. The imaging apparatus according to claim 2, characterized in that the image of the real space is captured by an image sensor in which a plurality of pixels that receive light passing through different pupil regions of an imaging optical system are arranged.
4. The imaging apparatus according to claim 3, characterized in that the depth information is the amount of image shift between a pair of composite reality images obtained by synthesizing a virtual object with the image signals corresponding to different pupil regions respectively.
5. The imaging apparatus according to claims 2 and 3, characterized in that the depth information is relative distance information between the imaging apparatus and the subject.
6. An acquisition step of acquiring an image signal from an image sensor in which a plurality of pixels that receive light passing through different pupil regions of an imaging optical system are arranged; A synthesizing step of synthesizing a virtual object with each of the image signals corresponding to different pupil regions by using the image signal obtained in the obtaining step to generate a pair of composite reality images; A focus adjustment step of adjusting the lens position of the imaging optical system based on the amount of image displacement between the pair of composite reality images; A control method for an image processing apparatus, characterized by comprising:
7. An obtaining step of obtaining an image of the real space through an imaging optical system; A synthesizing step of synthesizing a virtual object with the image of the real space to generate a composite reality image; A focus adjustment step of adjusting the lens position of the imaging optical system based on a plurality of depth information in the composite reality image; An aperture adjustment step of adjusting the aperture value at the time of shooting based on a plurality of depth information in the composite reality image; A control method for an image processing apparatus, characterized by comprising:
Citation Information
Patent Citations
Image processing method and device, and computer readable storage medium
CN111462337A
Information processor, and information processing method and program
JP2016152023A
Imaging device
JP2019054463A
Information processor, information processing method, program, and system
JP2019125345A
Imaging device and control method thereof
JP6685814B2