Image processing apparatus, and image processing method

JP2024084218A5Pending Publication Date: 2025-12-10CANON KK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2022198374
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-12-13
Publication Date
2025-12-10

AI Technical Summary

Technical Problem

Existing image processing technologies in MR systems fail to account for image quality deterioration due to optical factors like blur and light falloff, leading to unnatural composite images when combining virtual and captured images.

Method used

An image processing device and method that adjusts virtual object images based on the optical deterioration of captured images, using motion detection and image processing to align and correct optical factors such as peripheral dimming and blur, ensuring a natural composite image.

Benefits of technology

Reduces user discomfort by creating a composite image that seamlessly integrates virtual and captured images, minimizing the unnatural appearance caused by optical distortions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To reduce discomfort for a user in viewing a composite image obtained by combining a virtual object with a picked-up image.SOLUTION: An image processing apparatus has: display control means that controls a display to display a composite image obtained by combining an image of a virtual object with a picked-up image picked up by an imaging apparatus; acquisition means that acquires a composition position in the picked-up image where the image of the virtual object is combined; image processing means that executes, on the image of the virtual object, image processing based on information on a deterioration of the picked-up image at the composition position; and control means that controls the image processing according to the amount of movement of the display.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an image processing device and an image processing method. [Background technology]

[0002] In recent years, MR (Mixed Reality) technology has become known as a technology that seamlessly fuses the real world and the virtual world in real time. For example, a video see-through type HMD (Head Mounted Display) is used to capture an image of a subject that approximately matches the subject observed from the pupil position of the HMD user using a video camera or the like. Then, an image (composite) that combines the captured image with CG (Computer Graphics) is provided to the HMD user.

[0003] Here, when compositing a CG image (virtual object image) with a captured image, if the CG image and the captured image are not consistent, the composite image after composition will be unnatural. For example, when a user shakes his or her head, if the captured image captured by the camera is blurred but the CG is not blurred, a composite image that looks unnatural will be generated. Therefore, Patent Document 1 describes a technology that adds blur to a CG image according to the movement of the HMD and the distance between the HMD and the CG. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] JP 2018-28572 A Summary of the Invention [Problem to be solved by the invention]

[0005] However, in the technology described in Patent Document 1, when degradation of image quality caused by the imaging optical system occurs in the captured image (more specifically, when blurring or loss of light occurs), the blurring or loss of light is not taken into account in the CG. Therefore, if the images are synthesized as is, the image will look unnatural.

[0006] Therefore, an object of the present invention is to reduce the sense of discomfort felt by the user when a virtual object is synthesized into a captured image. [Means for solving the problem]

[0007] One aspect of the present invention is a method for producing a composition comprising the steps of: a display control means for controlling the display device so as to display a composite image obtained by combining an image of a virtual object with an image captured by the imaging device; an acquisition means for acquiring a synthesis position for synthesizing an image of the virtual object in the captured image; an image processing means for performing image processing based on information on image deterioration of the captured image at the synthesis position on the image of the virtual object; a control means for controlling the image processing in accordance with an amount of motion of the display device; The image processing device is characterized by having the following features.

[0008] One aspect of the present invention is a method for producing a composition comprising the steps of: a display control step of controlling the display device so as to display a composite image obtained by combining an image of a virtual object with an image captured by the imaging device; acquiring a synthesis position for synthesizing an image of the virtual object in the captured image; an image processing step of performing image processing based on information of image deterioration of the captured image at the synthesis position on the image of the virtual object; a control step of controlling the image processing in accordance with an amount of movement of the display device; The image processing method according to the present invention is characterized in that: Effect of the Invention

[0009] According to the present invention, when a virtual object is synthesized into a captured image, it is possible to reduce the sense of discomfort felt by the user when visually recognizing the image. [Brief description of the drawings]

[0010] [Figure 1] FIG. 1 is a diagram illustrating an example of the configuration of a system according to a first embodiment. [Diagram 2] 11A and 11B are diagrams illustrating the sense of incongruity that a synthetic image gives to a user. [Diagram 3] 3 is a diagram illustrating the internal configuration of a CG rendering unit according to the first embodiment. FIG. [Figure 4] 4 is a flowchart of a synthetic image generation process according to the first embodiment. [Diagram 5] 5 is a flowchart of a process for determining an amount of optical degradation according to the first embodiment. [Figure 6] 4A to 4C are diagrams for explaining the necessity of changing the amount of vignetting according to the first embodiment. [Figure 7] 4A to 4C are diagrams for explaining calculation of the amount of peripheral light falloff according to the first embodiment. [Figure 8] 4A to 4C are diagrams for explaining a simple calculation of the amount of vignetting in the first embodiment. [Figure 9] 4A to 4C are diagrams for explaining calculation of the amount of peripheral light falloff according to the first embodiment. [Figure 10] FIG. 2 is a diagram illustrating a composite image according to the first embodiment. [Figure 11] 10 is a flowchart of a process for determining an amount of optical degradation according to the second embodiment. [Figure 12] FIG. 11 is a diagram illustrating a filter according to a second embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0011] First, a composite image that may give a user a sense of incongruity will be described. Fig. 2A to Fig. 2C show a composite image obtained by combining a CG 200 (virtual object) with captured images 201, 202, and 203, which are the background, and which the user visually recognizes on the HMD.

[0012] 2A shows a composite image in which there is almost no aberration in the captured image 201. In this case, even if the CG 200 to be synthesized is used as is for synthesis, a composite image that does not give a sense of discomfort to the user can be generated. However, since the imaging device mounted on the HMD is required to be small and lightweight, the captured image often suffers from image quality degradation due to various optical causes.

[0013] For example, Fig. 2B shows a composite image when vignetting occurs in the captured image 202. Vignetting is a phenomenon in which the brightness of the outer edge of an image is lower than that of the center of the image due to the optical system of the imaging device. In this case, if the CG 200 is used for composition as is, a composite image that gives a strong sense of discomfort to the user will be generated.

[0014] 2C shows an image obtained by combining a CG 200 with a captured image 203 when the captured image 203 is blurred due to aberration of the imaging optical system. In this case, the resolution of the CG 200 is higher than that of the captured image 203 in the background, resulting in a composite image that gives a strong sense of discomfort to the user.

[0015] Although it is possible to digitally correct these optically-induced image degradations, it is difficult to completely remove them. Therefore, by using CG200, which also takes into account the optically-induced image degradation that occurs in the background image, the resulting composite image often looks natural.

[0016] <Embodiment 1> 1A, a configuration example of a system according to the embodiment 1 will be described. The system according to the embodiment 1 includes an HMD 110, a controller 120, a display device 130, and an image processing device 140.

[0017] The HMD 110 is an example of a display device that can be worn on the head of a user. The controller 120 relays (controls) data communication between the image processing device 140 and the HMD 110. The display device 130 is a display device for confirming an image processed by the image processing device 140. The image processing device 140 generates an image of a virtual object (CG) and a composite image (an image obtained by combining an image of a virtual object with a captured image).

[0018] Data communication between the HMD 110 and the controller 120 is performed using a wireless LAN (WLAN). The data communication between the HMD 110 and the controller 120 is performed via a small-scale network such as a Wireless Local Area Network (WPAN) or a Wireless Personal Area Network (WPAN). The data communication between the HMD 110 and the controller 120 is not limited to a specific communication format, and can be realized in any format, regardless of whether it is wireless communication or wired communication. In addition, the controller 120 is a separate device from the image processing device 140 in FIG. 1A, but it may be incorporated into the image processing device 140 to be integrated.

[0019] An example of the functional configuration of each of the image processing device 140 and the HMD 110 will be described with reference to the block diagram of Fig. 1B. In reality, as described above, the image processing device 140 and the HMD 110 perform data communication via the controller 120, but the controller 120 is not shown in Fig. 1B.

[0020] First, a description will be given of the HMD 110. The HMD 110 includes an imaging unit 122, a motion detection unit 123, a display unit 124, a control unit 125, an image correction unit 126, and an interface (I / F) 127.

[0021] The imaging unit 122 captures moving images captured in real space (a range that approximately coincides with the field of view of a user wearing the HMD 110 on his / her head). Since the position and orientation of the imaging unit 122 change in accordance with the movement of the HMD 110, the imaging unit 122 captures an image in a range that corresponds to the movement of the HMD 110. Images of each frame that constitutes the moving image (captured images) are sequentially transmitted to the image correction unit 126. The imaging unit 122 has an imaging optical system (one or more lenses). For this reason, the moving images (images of each frame that constitute the moving image) captured by the imaging unit 122 are subject to image degradation due to the characteristics of the imaging optical system.

[0022] The motion detection unit 123 is a sensor for detecting the direction and amount of motion of the HMD 110. The motion detection unit 123 includes, for example, at least one of a gyro sensor and an acceleration sensor. The motion detection unit 123 outputs data indicating the detected direction and amount of motion as motion vector data to the image processing device 140 via the I / F 127.

[0023] The display unit 124 is attached to the HMD 110 so as to be located in front of the user's eyes. The display unit 124 displays images or characters received from the image processing device 140 via the I / F 127. As a result, the images or characters received from the image processing device 140 are displayed in front of the user's eyes.

[0024] The control unit 125 controls the overall operation of the HMD 110 .

[0025] The image correction unit 126 performs various correction processes on the captured image transmitted from the image capture unit 122. Here, the image correction unit 126 may correct the influence of image deterioration (such as vignetting or blurring) caused by the image capture unit 122 as much as possible. However, when high-sensitivity image capture is performed, In such a case, if the correction is performed, it will have a large effect on the captured image in terms of noise degradation, and therefore in such a case, it is necessary to intentionally weaken the correction.

[0026] Next, a description will be given of the image processing device 140. The image processing device 140 includes an interface (I / F) 141, a CG drawing unit 142, a content database (DB) 143, and a calculation unit 144.

[0027] The CG rendering unit 142 constructs a virtual object by using data registered in the content DB 143. Then, the CG rendering unit 142 generates an image of the virtual object seen by the imaging unit 122 in the position and orientation calculated by the calculation unit 144 as a "virtual object image."

[0028] Here, the CG rendering unit 142 blurs the virtual object image based on various information received from the HMD 110 via the I / F 141, for example. Then, the CG rendering unit 142 synthesizes (superimposes) the virtual space image on the captured image received from the HMD 110 to generate a composite image (superimposed image). Then, the CG rendering unit 142 outputs the generated composite image. For example, the CG rendering unit 142 may output the composite image to the display device 130, or may output the composite image to the HMD 110 (display unit 124) via the I / F 141. When the composite image is output to the HMD 110 (display unit 124), the composite image is presented in front of the user's eyes.

[0029] Data necessary for drawing a virtual object is registered in the content DB 143. For example, data defining the shape of the virtual object, data defining the color and texture of the virtual object, data related to the position in real space where the virtual object is to be placed, texture map data, and the like are registered in the content DB 143.

[0030] The calculation unit 144 calculates the position and orientation of the imaging unit 122 (HMD 110). Various methods have been proposed for calculating the position and orientation of the imaging unit 122, and any of these methods may be adopted. For example, the calculation unit 144 may calculate the position and orientation of the imaging unit 122 based on the motion vector data received from the HMD 110 and the captured image. In addition, when the HMD 110 is equipped with a position and orientation sensor, the calculation unit 144 may calculate the position and orientation of the imaging unit 122 based on the measurement result by the position and orientation sensor. In addition, when a marker whose position in the real space is known in advance is placed, the calculation unit 144 may calculate the position and orientation of the imaging unit 122 based on the coordinates of the marker in the captured image and the position of the marker in the real space. Note that a natural feature in the real space (such as a corner of a desk) may be used instead of the marker.

[0031] (Internal structure of the CG drawing section) An example of the internal configuration of the CG rendering unit 142 will be described with reference to Fig. 3. The CG rendering unit 142 includes an image processing unit 302, an addition amount determining unit 303, a synthesis unit 304, and a rendering unit 305.

[0032] The image processing unit 302 adds an amount of optical degradation such as peripheral illumination or blur (blur due to aberration) to a virtual object image (virtual object) according to a synthesis position (a position where a virtual object image is arranged in a captured image). Specifically, when the image processing unit 302 acquires an instruction to add the amount of optical degradation to a virtual object image from the addition amount determination unit 303, the image processing unit 302 adds the amount of optical degradation to the virtual object image as correction information for correcting the virtual object image. Note that "adding the amount of optical degradation to a virtual object image" means, for example, "performing a correction such as adding the amount of optical degradation (pixel value indicating the amount of optical degradation) to the pixel value of the virtual object image."

[0033] On the other hand, when the image processing unit 302 acquires an instruction not to add the amount of optical degradation from the addition amount determination unit 303, the image processing unit 302 outputs the virtual object image generated by the drawing unit 305 as it is to the synthesis unit 304. At this time, image processing other than the addition of the amount of optical degradation may be performed.

[0034] The addition amount determination unit 303 determines whether or not to add the amount of optical degradation to the virtual object image. When the addition amount determination unit 303 adds the amount of optical degradation to the virtual object image, the addition amount determination unit 303 determines the amount of optical degradation to be added to the virtual object image.

[0035] The synthesis unit 304 generates a synthetic image by synthesizing the captured image received from the HMD 110 with the virtual object image acquired from the image processing unit 302. Then, the synthesis unit 304 outputs the generated synthetic image to the HMD 110 (display unit 124) via the I / F 141. In this way, the synthesis unit 304 controls the HMD 110 to display the synthetic image (performs display control).

[0036] The rendering unit 305 constructs a virtual object based on data registered in the content DB 143, and places the virtual object in a virtual space. The rendering unit 305 then generates an image of the virtual object viewed from the imaging unit 122 in the position and orientation calculated by the calculation unit 144 as a "virtual object image." Since a technology for generating an image of a virtual object viewed from an object in a specified position and orientation is well known, a description of this technology will be omitted. The rendering unit 305 also calculates the distance between the imaging unit 122 and the virtual object (the distance in the three-dimensional space represented by the composite image) based on the position of the imaging unit 122 calculated by the calculation unit 144 and the data of the position where the virtual object is to be placed.

[0037] (About the generation process) The process of generating a composite image (generation process) will be described with reference to the flowchart of Fig. 4. The process of the flowchart of Fig. 4 is realized by controlling each component of the image processing device 140 by a control unit (such as a CPU) not shown. It is assumed that the generation of a virtual object image by the drawing unit 305 and the calculation of a synthesis position of the virtual object image (a position at which the virtual object image is synthesized in a captured image) by the calculation unit 144 or the like are completed before the process of the flowchart of Fig. 4 starts. The synthesis position of the virtual object image is a position in the composite image that corresponds to the range captured by the imaging unit 122. This is because the synthesis position is fixed to a specific position in real space.

[0038] In step S401, the addition amount determination unit 303 determines the amount of optical degradation (optical degradation information) to be added to each pixel of the virtual object image. Details of this process will be described later with reference to the flowchart in FIG.

[0039] In step S402, the image processing unit 302 adds the amount of optical degradation (optical degradation information) calculated in step S401 to the virtual object image. That is, the image processing unit 302 corrects (adjusts) the optical degradation of the virtual object image so that it matches the optical degradation of the captured image. Here, since gamma may be applied to each virtual object image, it is necessary to correct the amount of optical degradation calculated in step S401 based on that gamma.

[0040] In step S403, the synthesis unit 304 synthesizes the virtual object image with the amount of optical degradation added to the captured image. This makes it possible to use a virtual object image for synthesis that gives the user less discomfort even when the user moves his / her head abruptly.

[0041] The process of determining the amount of optical degradation in step S401 will be described below. Note that the following will be described in the case where the amount of vignetting is used as the amount of optical degradation. In this case, the amount of vignetting corresponding to the state of light loss caused by the characteristics of the imaging optical system at the synthesis position (the position where the virtual object image is synthesized in the captured image) is added.

[0042] In general, the relationship of the amount of light with respect to peripheral light falloff in a captured image is defined according to the cosine fourth power law (COS4 law) as shown in the following formula (1).

number

[0043] In formula (1), I indicates the amount of light (illuminance) after it enters the imaging optical system, and I0 indicates the amount of light before it enters. Therefore, the amount of vignetting is the result of subtracting the amount of light after it enters, I, from the amount of light before it enters, I0. Furthermore, as shown in Figure 7, θ indicates the angle between the incident light and the optical axis, and can generally be defined as the following formula (2).

number

[0044] In formula (2), h indicates the distance from the optical axis on the imaging surface (image height), and f indicates the focal length of the imaging optical system, as shown in Fig. 7. In addition, θ changes according to the aperture of the imaging optical system, so the amount of light I changes according to the focal length, aperture, and image height.

[0045] FIG. 9 is a diagram for explaining an example of calculation of the amount of vignetting. Point P in FIG. 9 indicates the optical axis center position of the imaging optical system. For example, to synthesize a virtual object image at the position of point Q, the distance between P and Q (that is, (x0-x) 2 +(y0-y) 2 The amount of vignetting at point Q is calculated by taking into account the square root of the vignetting amount at point Q. This calculation is usually performed for all pixels of the CG image of the virtual object. However, since the amount of vignetting changes depending on the focal length, aperture, and image height, the calculation load is large when this calculation is performed for all pixels.

[0046] In addition, the captured image may have undergone partial processing to correct the effects of vignetting. In this case, the amount of vignetting to be added to each pixel of the virtual object must be determined taking into account the correction information.

[0047] Here, in the correction applied to the captured image, since the correction causes noise deterioration, it is often impossible to completely correct the influence of vignetting in the imaging optical system. Therefore, if a calculation is performed to correct the virtual object image, taking into consideration the fact that the influence of vignetting in the captured image cannot be completely corrected, the calculation takes more time, and the calculation of the vignetting amount (correction value) for correcting the virtual object image takes even more time. If it takes a long time to calculate the vignetting amount, a virtual object with an inappropriate vignetting amount is displayed until the calculation of the vignetting amount is completed. Furthermore, at the time when the calculation of the vignetting amount is completed, the virtual object is adjusted to a desired brightness, but at that time, the brightness of the virtual object changes while the brightness of the captured image does not change, so the user recognizes flickering.

[0048] (Optical degradation amount determination process; S401) Therefore, in step S401, the addition amount determination unit 303 quickly changes the amount of vignetting to be added when the amount of vignetting to be added changes significantly, and then improves the accuracy of the amount of vignetting to be added, thereby reducing the sense of incongruity felt by the user. Details of such a method of calculating the amount of vignetting will be described with reference to the flowchart in Fig. 5. This process is particularly intended to reduce the sense of incongruity felt by the user when the composition position changes suddenly due to the user shaking his head.

[0049] In step S501, the addition amount determination unit 303 judges whether or not it is necessary to add vignetting to the virtual object image. For example, if the vignetting amount of the pixel of the captured image corresponding to the position of the center of gravity of the virtual object image (if the captured image has been corrected for the effect of vignetting, the vignetting amount after correction) is smaller than a predetermined threshold value set in advance, it is not necessary to add vignetting to the virtual object image. If it is judged that vignetting amount needs to be added to the virtual object image, If so, the process proceeds to step S502. If it is determined that there is no need to add vignetting to the virtual object image, the process of this flowchart ends.

[0050] In step S502, the additional amount determination unit 303 determines whether the amount of movement of the HMD 110 (the amount of rotation or movement of the HMD 110) in one frame between the current frame and the previous frame is smaller than a threshold. If it is determined that the amount of movement of the HMD 110 is smaller than the threshold, the process proceeds to step S505. If it is determined that the amount of movement of the HMD 110 is equal to or greater than the threshold, the process proceeds to step S503.

[0051] In addition, since the movement amount of the HMD 110 and the change in the synthesis position are mutually related, the additional amount determination unit 303 may determine whether the difference between the synthesis position of the current frame and the synthesis position of the previous frame is smaller than a threshold value. In this case, if it is determined that the difference between the two synthesis positions is smaller than the threshold value, the process proceeds to step S505. If it is determined that the difference between the two synthesis positions is equal to or greater than the threshold value, the process proceeds to step S503.

[0052] Here, the threshold value in step S502 may be changed according to the characteristics of the imaging optical system. For example, when the difference in the amount of vignetting between the center of the captured image and the outer edge (periphery) of the captured image is smaller than a predetermined value, the threshold value may be large because the user is unlikely to feel uncomfortable even if the amount of movement of the HMD 110 is large, even if the amount of vignetting of the current frame is used in the next frame.

[0053] In step S503, the addition amount determination unit 303 calculates (obtains) the amount of vignetting to be added to each pixel of the virtual object image of the next frame by a simple method (processing with a smaller amount of calculation than the processing in step S506). Specifically, the addition amount determination unit 303 selects one representative pixel (a pixel at the center position or the center of gravity position, etc.) of the virtual object image, and calculates the amount of vignetting occurring in a pixel (position) of the captured image corresponding to the position of that pixel. Then, the addition amount determination unit 303 determines the calculated amount of vignetting as the amount of vignetting to be added to all pixels of the virtual object image. Note that in step S503, any method may be used as long as it can calculate (obtain) the amount of vignetting to be added to each pixel of the virtual object image of the next frame by processing with a smaller amount of calculation than the processing for calculating the amount of vignetting of all pixels in the range corresponding to the virtual object image in the captured image. For example, the addition amount determination unit 303 may select four representative pixels (e.g., the upper right, upper left, lower right, and lower left pixels) of the virtual object image, and calculate the amount of vignetting occurring at four pixels (positions) of the captured image corresponding to the positions of these pixels.The addition amount determination unit 303 may then determine the average value of the calculated amounts of vignetting for the four pixels as the amount of vignetting to be added to all pixels of the virtual object image.

[0054] In step S504, the additional amount determination unit 303 assigns a flag indicating that the target has been achieved (target achievement flag) to the next frame.

[0055] In step S505, the additional amount determination unit 303 judges whether or not the current frame is a computation frame (a frame to be processed in step S506). One of several consecutive frames is set in advance as a computation frame. A frame with a target achievement flag attached also corresponds to a computation frame. If it is judged that the current frame is a computation frame, the process proceeds to step S506. If it is judged that the current frame is not a computation frame, the process proceeds to step S507. Note that the process of step S505 may be skipped and the process of step S506 may be performed. In other words, the process of step S506 may be performed regardless of the type of frame the current frame is.

[0056] In step S506, the addition amount determination unit 303 calculates the addition amount of a plurality of pixels (corresponding to the positions of all pixels of the virtual object image) corresponding to the virtual object image in the captured image according to the formula (1). The addition amount determination unit 303 calculates the amount of vignetting occurring in each of the pixels (multiple pixels of the captured image to be added). The addition amount determination unit 303 sets the calculated amount of vignetting for each pixel as “target correction information” which is the ideal amount of vignetting to be added to each pixel of the virtual object image.

[0057] The process of step S506 does not need to be performed in real time. It may take a long time to calculate accurate peripheral illumination amounts for all of the pixels corresponding to the virtual object image in the captured image. Then, the process of step S507 may start at the same time as the process of step S506 starts.

[0058] As described above, the process performed in step S506 imposes a high computational load on the image processing device 140. For this reason, when there is little need for recalculation (that is, when the HMD 110 has not moved significantly or the composite position has not changed significantly), the process in step S506 is executed once every few frames.

[0059] In step S507, the addition amount determination unit 303 calculates the amount of vignetting to be added to each pixel of the virtual object image in the next frame, based on the amount of vignetting added to each pixel of the virtual object image in the current frame (present) (hereinafter referred to as “current correction information”) and the target correction information.

[0060] Specifically, when the target correction information and the current correction information are different, the addition amount determination unit 303 calculates the amount of vignetting to be added to each pixel of the virtual object image of the next frame by correcting the current correction information so as to approach the target correction information. For example, when the target correction information and the current correction information are different, the addition amount determination unit 303 calculates correction information included between the target correction information and the current correction information (such as an average of the target correction information and the current correction information) as the amount of vignetting to be added to the virtual object image. For example, when the process of step S503 is performed (when the amount of movement of the HMD 110 changes from a state in which it is larger than a threshold to a state in which it is smaller), an amount of vignetting different from the target correction information is calculated, so that an amount of vignetting approaching the target correction information is calculated. Note that the amount of vignetting may reach the target correction information in the next frame, but if the amount of vignetting suddenly changes significantly, a flicker will be visually recognized by the user. For this reason, the amount of vignetting to be changed at one time may be limited to gradually approach the target correction information.

[0061] On the other hand, when the target correction information and the current correction information match (are identical), the addition amount determination unit 303 determines to use the target correction information for the amount of peripheral light falloff to be added to the virtual object image (each pixel of the virtual object image) of the next frame.

[0062] (Necessity of changing the amount of optical degradation) When the synthesis position (position on the captured image) where the virtual object image is synthesized changes suddenly, for example, when the user moves his / her head, the amount of optical degradation to be added must also be changed according to the movement, otherwise the synthesized image will become unnatural. Therefore, the necessity of changing the amount of optical degradation will be described with reference to Figs. 6A to 6C.

[0063] 6A is a diagram showing an environment 601 in which a user is located. Areas 603 and 604 are areas that the user views through the HMD 110. Specifically, a case will be described in which, after area 603 is displayed on the HMD 110, area 604 is displayed on the HMD 110 as a result of the user shaking his head.

[0064] 6B shows the state in which the area 603 is displayed on the HMD 110. At this point, the same amount of vignetting (amount of optical degradation) as that occurring in the captured image is added to the virtual object image, the flower 602. Therefore, the virtual object image can be viewed as a natural image that blends in with the background.

[0065] On the other hand, Fig. 6C shows an image (area 604) that is displayed when the user turns his / her head to the right in the state shown in Fig. 6B. At this time, the calculation process of the amount of vignetting that has occurred in the captured image has not been completed, and the amount of vignetting that has been applied in the state of Fig. 6B remains added to the flower 602, even though the flower 602 is located near the center of the captured image. This causes a mismatch in brightness between the background and the virtual object.

[0066] Furthermore, after that, the brightness of the virtual object image is adjusted to an appropriate brightness by completing the calculation process of the amount of vignetting occurring in the captured image. However, during this adjustment, the brightness of only the flower 602 changes. This causes the user to feel flickering (unnaturalness). In this way, when the composite position of the virtual object image changes significantly, it is necessary to change the amount of vignetting added to the virtual object image as quickly as possible so that it blends in with the surroundings.

[0067] (Necessity of processing in step S503) 8A to 8C, the calculation of the peripheral shading amount as the amount of optical degradation in step S503 using a simple method (a method with a small amount of calculation) will be described. As described using FIG. 6A to FIG. 6C, when the user moves his / her head, the amount of optical degradation to be added to the virtual object image must be changed quickly, otherwise the user will feel uncomfortable. On the other hand, when the resolution of the virtual object image is high (the number of pixels is large), the calculation load is large if the amount of optical degradation with high accuracy is calculated for all pixels. For this reason, it is difficult to quickly calculate the amount of optical degradation with high accuracy (high-precision degradation amount). Therefore, even if a low-precision amount of optical degradation (low-precision degradation amount) is used, it is necessary to minimize flickering by quickly updating the amount of optical degradation to be added.

[0068] For example, when it takes time to calculate the amount of optical degradation with high accuracy, it is better to uniformly add the amount of optical degradation to be added to each pixel at a representative position of the virtual object image of the current frame, rather than displaying a virtual object image to which the amount of optical degradation of the previous frame has been added. The representative position is, for example, the center of gravity or the upper left position of the virtual object image.

[0069] For example, Fig. 8A shows a composite image (an image corresponding to Fig. 6C) in the case where the amount of optical degradation added in the previous frame is also added to each pixel of the virtual object in the current frame (when the addition of the high-precision degradation amount is not complete). On the other hand, Fig. 8B shows a composite image in the case where the amount of optical degradation to be added to the pixel at the center of gravity position 803 of the virtual object image is uniformly added to all pixels of the virtual object image.

[0070] Since the virtual object image 802 in FIG. 8A is affected by the characteristics of the vignetting of the previous frame, the virtual object image 802 appears dark, which gives a user who sees the composite image a great sense of incongruity. On the other hand, the virtual object image 802 in FIG. 8B does not include a highly accurate amount of optical degradation, but is corrected to a brightness that takes into account the brightness of the captured image near the center of gravity position 803 of the virtual object image 802. Therefore, the composite image shown in FIG. 8B gives a user less sense of incongruity than the composite image shown in FIG. 8A. And, when correcting all pixels with a uniform amount of optical degradation, it is necessary to calculate the amount of optical degradation for only one pixel, so that the computation load (amount of computation) can be significantly reduced. Therefore, the amount of optical degradation can be applied in real time without delay.

[0071] Thus, according to the first embodiment, when a virtual object (virtual object image) is synthesized with a captured image, the amount of optical degradation (amount of vignetting) added to the virtual object can be quickly changed (adjusted), thereby reducing the sense of discomfort felt by the user when viewing the image.

[0072] In the first embodiment, the HMD 110 has been taken as an example for explanation, but instead of the HMD 110, a device equipped with an imaging device (such as a smartphone) may be used.

[0073] <Embodiment 2> Next, a description will be given of embodiment 2. In the following, in embodiment 2, the description of the same contents as in embodiment 1 will be omitted. In embodiment 2, an example will be described in which a degradation component related to blur caused by an imaging optical system (spherical aberration, etc.) is added to a virtual object image.

[0074] Due to the effects of aberrations and other factors that occur in the imaging optical system, light generated from one point on the subject does not converge to one point, but spreads out slightly. This distribution with a small spread is expressed by the point spread function (PSF). Since the captured image is formed by applying the PSF to the subject image, the resolution of the captured image is degraded due to the effects of the imaging optical system (the captured image becomes blurred).

[0075] That is, blurring occurs in the captured image according to the PSF, but this blurring does not occur in the virtual object image. Therefore, in order to make the composite image look natural, it is advisable to add blurring (blurring as correction information) according to the composite position to the virtual object image. This makes it possible to provide the user with a composite image (video) that feels less unnatural.

[0076] 10A to 10C show a composite image visually recognized by a user when a virtual object image is composited with a captured image. A captured image 1001 in Fig. 10A is an image in a state in which the imaging optical system has no aberration that affects blur. In this case, a virtual object image 1002 is used for composition as it is without being subjected to any particular processing.

[0077] 10B, when the imaging optical system has an aberration that affects blur (the captured image 1003 has blur), if the virtual object image 1004 is used for synthesis as is, the synthesized image will look unnatural. Here, in the captured image 1003, the blur increases as one approaches the outer edge. However, since the distribution tendency of the blur depends on the imaging optical system, there are also cases where the position closer to the center of the image becomes blurred.

[0078] Fig. 10C shows a composite image obtained by adding blurring that matches the composition position to virtual object image 1004. As a result, the composite image shown in Fig. 10C can be viewed as a composite image including a virtual object that blends in with the background more than the composite image shown in Fig. 10B.

[0079] In the second embodiment, the process of adding blur, which is the amount of optical degradation, to a virtual object image (blurring process) can be realized in the same manner as the process of the flowchart shown in Fig. 4. Specifically, the image processing device 140 calculates blur (amount of blur to be added) as the amount of optical degradation in step S401, adds the blur to the virtual object image in step S402, and then performs synthesis in step S403. In this way, the image processing device 140 realizes the generation of a synthetic image shown in Fig. 10C.

[0080] At this time, as a method for calculating the blur to be added (step S401), there is a method in which a plurality of filters defining the blur are prepared in advance, and a filter to be applied to the virtual object image is selected from among the filters.

[0081] 12A shows the regions of a captured image. Here, the region of the captured image is divided into three regions, regions 1201 to 1203, according to the characteristics (characteristics related to blur) of the imaging optical system. Since the characteristics related to blur also change according to the image height (image height direction) around the optical axis, the region is divided according to the image height.

[0082] Then, a filter that realizes a blur similar to the blur that occurs in each region is designed in advance. For example, as shown in FIG. 12B, since the region 1201 is near the center of the optical axis, the blur is Since there is no need to add a reference filter to region 1201, the reference filter is set to "none." Region 1202 is a region where some blurring begins to occur due to aberration, so filter A that adds blur is set to region 1202. For example, filter A is a 3×3 Gaussian filter as shown in FIG. 12C. Region 1203 is a region where the image height is high and the amount of aberration is large, so filter B that adds stronger blurring (blurring) is set to region 1203. For example, filter B is a 5×5 Gaussian filter as shown in FIG. 12D.

[0083] The above-described filter and area division method are merely examples, and other methods may be used for the area division method and the blurring method. For example, blurring may be added by a recovery process using the inverse function of the PSF of the actual optical system.

[0084] Here, the process of step S401 according to the second embodiment (processing of determining the blur to be added to the virtual object image based on the information on the blur of the captured image at the synthesis position) will be described in detail with reference to FIG.

[0085] In step S1101, the addition amount determination unit 303 determines whether or not it is necessary to add blur to the virtual object image. For example, when the positions of all pixels of the current frame of the virtual object image are included in the area 1201 in FIG. 12A, the addition amount determination unit 303 determines that it is not necessary to add blur to the virtual object image. In other cases, the addition amount determination unit 303 determines that it is necessary to add blur to the virtual object image. When it is determined that it is necessary to add blur to the virtual object image, the process proceeds to step S1102. When it is determined that it is not necessary to add blur to the virtual object image, the process of this flowchart ends.

[0086] In step S1102, similar to step S502, the additional amount determination unit 303 determines whether the amount of motion of the HMD 110 in one frame between the current frame and the previous frame is smaller than a threshold. If it is determined that the amount of motion of the HMD 110 is smaller than the threshold, the process proceeds to step S1105. If it is determined that the amount of motion of the HMD 110 is equal to or greater than the threshold, the process proceeds to step S1103.

[0087] In step S1103, the addition amount determination unit 303 determines a filter to be applied to each pixel of the virtual object image of the next frame by a simple method (processing with a smaller amount of calculation than the processing in step S1106). If the filter is changed for each pixel of the virtual object image, the processing becomes complicated and it becomes difficult to immediately reflect the filter. Therefore, the addition amount determination unit 303 determines a filter (blur) to be applied to the image height corresponding to the center of gravity position of the virtual object image as a filter (blur) to be applied to all pixels of the virtual object image.

[0088] In step S1104, similarly to step S504, the additional amount determining unit 303 assigns a flag indicating that the target has been achieved (target achievement flag) to the next frame.

[0089] In step S1105, similarly to step S505, the additional amount determination unit 303 judges whether or not the current frame is a computed frame (a frame to be processed in step S1106). If it is judged that the current frame is a computed frame, the process proceeds to step S1106. If it is judged that the current frame is not a computed frame, the process proceeds to step S1107.

[0090] In step S1106, the addition amount determination unit 303 sets target correction information. In this embodiment, the addition amount determination unit 303 determines which of the areas 1201 to 1203 shown in FIG. 12A each pixel of the virtual object image is included in, and determines the filter (blur) to be applied to each pixel. Then, the addition amount determination unit 303 determines the filter (blur) to be applied to each pixel. The filter coefficient is set as the target correction information.

[0091] In step S1107, the addition amount determination unit 303 determines a filter to be applied to each pixel of the virtual object image of the next frame. First, the addition amount determination unit 303 judges whether or not the filter (current correction information) applied to each pixel of the virtual object image of the current frame matches the target correction information. If the addition amount determination unit 303 judges that the two correction information (filters) do not match, it determines to use the target correction information as a filter to be applied to each pixel of the virtual object image of the next frame. Here, if there is a possibility that flickering will occur by applying the target correction information to the virtual object image, the addition amount determination unit 303 may prepare intermediate correction information (filter) between the target correction information and the current correction information and determine to apply it to the virtual object image.

[0092] According to the second embodiment, even if the HMD is moved vigorously, it is possible to generate a composite image including a virtual object that has little flicker and has a high degree of fusion with the captured image.

[0093] Although the present invention has been described in detail based on the preferred embodiments, the present invention is not limited to these specific embodiments, and various forms within the scope of the gist of the present invention are also included in the present invention. Parts of the above-described embodiments may be combined as appropriate.

[0094] Also, in the above, "If A is equal to or greater than B, proceed to step S1, and if A is smaller (lower) than B, proceed to step S2" may be read as "If A is greater (higher) than B, proceed to step S1, and if A is equal to or less than B, proceed to step S2." Conversely, "If A is greater (higher) than B, proceed to step S1, and if A is equal to or less than B, proceed to step S2" may be read as "If A is greater (higher) than B, proceed to step S1, and if A is smaller (lower) than B, proceed to step S2." Therefore, unless a contradiction occurs, "equal to or greater than A" may be read as "equal to or greater than A (high; long; many)," and "equal to or less than A" may be read as "equal to or less than A (low; short; few)." And, "equal to or greater than A" may be read as "equal to or greater than A," and "equal to or less than A" may be read as "equal to or less than A."

[0095] Each functional unit in each of the above embodiments (variations) may or may not be individual hardware. The functions of two or more functional units may be realized by common hardware. Each of a plurality of functions of one functional unit may be realized by individual hardware. Two or more functions of one functional unit may be realized by common hardware. Furthermore, each functional unit may or may not be realized by hardware such as an ASIC, FPGA, or DSP. For example, the device may have a processor and a memory (storage medium) in which a control program is stored. Then, the functions of at least some of the functional units of the device may be realized by the processor reading and executing the control program from the memory.

[0096] (Other embodiments) The present invention can also be realized by a process in which a program for implementing one or more of the functions of the above-described embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in a computer of the system or device read and execute the program. The present invention can also be realized by a circuit (e.g., ASIC) for implementing one or more of the functions.

[0097] The disclosure of the above embodiments includes the following configurations, methods, and programs. (Configuration 1) a display control means for controlling the display device so as to display a composite image obtained by combining an image of a virtual object with an image captured by the imaging device; an acquisition means for acquiring a synthesis position for synthesizing an image of the virtual object in the captured image; an image processing means for performing image processing based on information on image deterioration of the captured image at the synthesis position on the image of the virtual object; a control means for controlling the image processing in accordance with an amount of motion of the display device; 13. An image processing device comprising: (Configuration 2) the imaging device captures an image of a range corresponding to the movement of the display device; The synthesis position is a position corresponding to the range captured by the imaging device. 2. The image processing device according to claim 1, (Configuration 3) The information on image degradation is information caused by characteristics of the optical system of the imaging device. 3. The image processing device according to configuration 1 or 2. (Configuration 4) The information on image degradation is information on peripheral light falloff. 4. The image processing device according to configuration 3. (Configuration 5) The image degradation information is blur information. 4. The image processing device according to configuration 3. (Configuration 6) The control means When the amount of motion of the display device is smaller than a threshold, the information on the image degradation of the captured image at the combination position is acquired with a first amount of calculation; when the amount of motion of the display device is greater than the threshold, controlling the image processing means to acquire the information on the image degradation of the captured image at the synthesis position with a second amount of calculation smaller than the first amount of calculation. 6. The image processing device according to any one of configurations 1 to 5. (Configuration 7) When the amount of motion of the display device is greater than the threshold, the information on image degradation of the captured image at the synthesis position is information on image degradation at a position of the captured image corresponding to a specific position of the image of the virtual object. 7. The image processing device according to configuration 6, (Configuration 8) When the amount of motion of the display device is smaller than the threshold value, the information on image deterioration of the captured image at the synthesis position is information on image deterioration at a position of the captured image corresponding to a position of each pixel of the image of the virtual object. 8. The image processing device according to configuration 6 or 7. (Configuration 9) The control means when the amount of motion of the display device is greater than the threshold, acquiring first correction information corresponding to the information on the image deterioration acquired with the first amount of calculation, and controlling the image processing means to perform the image processing on the image of the virtual object based on the first correction information; when the amount of motion of the display device changes from a state in which it is greater than the threshold to a state in which it is smaller than the threshold, acquire second correction information corresponding to the information on image deterioration acquired with the second amount of calculation, and control the image processing means to perform the image processing on the image of the virtual object based on the correction information included between the first correction information and the second correction information. 9. The image processing device according to any one of configurations 6 to 8. (Configuration 10) The control means controls the threshold value in accordance with characteristics of an optical system of the imaging device. 10. The image processing device according to any one of configurations 6 to 9. (method) a display control step of controlling the display device so as to display a composite image obtained by combining an image of a virtual object with an image captured by the imaging device; acquiring a synthesis position for synthesizing an image of the virtual object in the captured image; an image processing step of performing image processing based on information of image deterioration of the captured image at the synthesis position on the image of the virtual object; a control step of controlling the image processing in accordance with an amount of movement of the display device; 13. An image processing method comprising: (program) A program for causing a computer to function as each of the means of the image processing device according to any one of configurations 1 to 10. [Explanation of symbols]

[0098] 110: HMD, 140: image processing device, 142: CG drawing unit, 302: image processing unit; 303: additional amount determining unit; 304: Composition section, 305: Drawing section

Claims

1. a display control means for controlling the display device to display a composite image obtained by combining an image of a virtual object with an image captured by the imaging device; an acquisition means for acquiring a synthesis position for synthesizing an image of the virtual object in the captured image; an image processing means for performing image processing based on information about image degradation of the captured image at the synthesis position on the image of the virtual object; a control means for controlling the image processing in accordance with the amount of movement of the display device; and The control means When the amount of motion of the display device is smaller than a threshold, the information on the image degradation of the captured image at the combining position is acquired with a first amount of calculation; and when the amount of motion of the display device is greater than the threshold, controlling the image processing means to acquire the information on the image degradation of the captured image at the combining position with a second amount of calculation that is smaller than the first amount of calculation.

1. An image processing device comprising:

2. the imaging device captures an image of a range corresponding to the movement of the display device; The synthesis position is a position corresponding to the range captured by the imaging device.

2. The image processing device according to claim 1, wherein:

3. the information on image degradation is information caused by characteristics of the optical system of the imaging device; 3. The image processing device according to claim 1, wherein the image processing device is a computer.

4. The information on image degradation is information on vignetting.

4. The image processing device according to claim 3.

5. The information on image degradation is blur information.

4. The image processing device according to claim 3.

6. When the amount of movement of the display device is greater than the threshold value, the image captured at the combining position is The information on image degradation of the captured image is information on image degradation at a position of the captured image corresponding to a specific position of the image of the virtual object.

2. The image processing device according to claim 1, wherein:

7. when the amount of motion of the display device is smaller than the threshold, the information on image degradation of the captured image at the synthesis position is information on image degradation at a position of the captured image corresponding to a position of each pixel of the image of the virtual object.

2. The image processing device according to claim 1, wherein:

8. The control means when the amount of motion of the display device changes from a state where it is smaller than the threshold value to a state where it is larger than the threshold value, acquires first correction information corresponding to the information on image deterioration acquired with the first amount of calculation, and controls the image processing means to perform the image processing on the image of the virtual object based on the first correction information; when the amount of motion of the display device changes from a state in which it is greater than the threshold to a state in which it is smaller than the threshold, acquire second correction information corresponding to the information on image deterioration acquired with the second amount of calculation, and control the image processing means to perform the image processing on the image of the virtual object based on correction information included between the first correction information and the second correction information.

2. The image processing device according to claim 1, wherein:

9. the control means controls the threshold value in accordance with characteristics of the optical system of the imaging device.

2. The image processing device according to claim 1, wherein:

10. a display control step of controlling the display device to display a composite image obtained by combining an image of a virtual object with an image captured by the imaging device; an acquisition step of acquiring a synthesis position at which an image of the virtual object is synthesized in the captured image; an image processing step of performing image processing based on information about image degradation of the captured image at the synthesis position on the image of the virtual object; a control step of controlling the image processing in accordance with an amount of movement of the display device; and In the control step, When the amount of motion of the display device is smaller than a threshold, the information on the image degradation of the captured image at the combining position is acquired with a first amount of calculation; When the amount of motion of the display device is greater than the threshold, the image processing is controlled so as to acquire the information on the image degradation of the captured image at the combining position with a second amount of calculation that is smaller than the first amount of calculation. An image processing method comprising:

11. A program for causing a computer to function as each of the means of the image processing apparatus according to claim 1 or 2.