Image processing apparatus, image processing method, and program

JP2024160556A5Pending Publication Date: 2026-05-11CANON KK
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
CANON KK
Filing Date
2023-05-01
Publication Date
2026-05-11

AI Technical Summary

Technical Problem

In MR technology, combining a captured image with a CG image often results in temporal mismatches and incorrect depth relationships, leading to unnatural and strange images for the user due to image correction techniques that fail to accurately align the CG image with the real-world image.

Method used

An image processing device that includes an acquisition unit for real-world images, a motion acquisition unit for device movement, and CG generation means, with pixel interpolation and correction units to align the CG image with the captured image based on motion information, ensuring correct depth and temporal alignment.

Benefits of technology

The solution provides a seamless and natural MR experience by accurately aligning CG and captured images, maintaining correct depth relationships and eliminating temporal discrepancies, resulting in a smooth and natural image presentation for the user.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To make it possible to provide an image that does not give a sense of discomfort to an HMD user.SOLUTION: An image processing apparatus acquires a captured image obtained by an imaging apparatus capturing a real world, and acquires motion information on the imaging apparatus in a space of the real world. Furthermore, the image processing apparatus generates a computer graphics image including combining information related to combining with the captured image, on the basis of the motion information on the imaging apparatus, separates the computer graphics image including the combining information into an image channel and a combining information channel, corrects the image channel by pixel interpolation according to the motion information on the imaging apparatus, and corrects the combining information channel by pixel replacement according to the motion information on the imaging apparatus. Then, the image processing device combines the captured image and the corrected computer graphics image.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an image processing technique for synthesizing a captured image and a CG image. [Background technology]

[0002] In recent years, so-called MR (Mixed Reality) technology has become known as a technology for seamlessly fusing the real world and the virtual world in real time. A video see-through HMD (Head Mounted Display) is known as one of the devices that realizes the MR technology. The video see-through HMD is configured to capture the real world that approximately coincides with the position of the HMD user's eyes using a video camera or the like, and to display an MR image that is a composite of the captured image and a CG (Computer Graphics) image.

[0003] The process of generating an MR image in which a CG image is synthesized with a captured image is often performed by an external image processing device or the like that is capable of communicating with the HMD. The image processing device receives a captured image captured by a camera of the HMD, calculates the position and orientation of the HMD (the position and orientation of the head of the HMD user) based on the captured image, generates a CG image based on the calculation result, and transmits the MR image synthesized with the captured image to the HMD. That is, the CG image in the MR image generated by the image processing device contains a time delay due to the calculation of the position and orientation of the HMD and the generation process based on the calculation result. Therefore, when an HMD user views an MR image in which a CG image including the time delay is synthesized with a captured image, the CG image may be delayed relative to the captured image, creating an unnatural image. In response to this, Patent Document 1 discloses a technique in which image correction is performed to cancel out the delay in a CG image in accordance with the movement of the head of an HMD user, and then the image is synthesized with a captured image. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] JP 2019-95916 A Summary of the Invention [Problem to be solved by the invention]

[0005] Incidentally, when generating an MR image by synthesizing a captured image of the real world with a CG image, it is necessary to ensure that the depth relationship between the objects and the CG image shown in the captured image is correct. However, when image correction is performed to cancel the delay of the CG image as in the technology disclosed in Patent Document 1, the depth relationship between the objects and the CG image in the captured image may not match, resulting in an image that feels unnatural to the HMD user.

[0006] Therefore, an object of the present invention is to provide an image that does not feel strange to the user of the HMD. [Means for solving the problem]

[0007] The image processing device of the present invention is characterized by having an acquisition means for acquiring an image captured by an imaging device of the real world, an operation acquisition means for acquiring motion information of the imaging device within the space of the real world, a CG generation means for generating a computer graphics image including synthesis information related to synthesis with the captured image based on the motion information of the imaging device, a CG correction means for separating the computer graphics image including the synthesis information into an image channel and a synthesis information channel, and performing correction by pixel interpolation on the image channel in accordance with the motion information of the imaging device and correction by pixel replacement on the synthesis information channel in accordance with the motion information of the imaging device, and a synthesis means for synthesizing the captured image and the computer graphics image corrected by the CG correction means. Effect of the Invention

[0008] According to the present invention, it is possible to provide an image that does not give an uncomfortable feeling to the user of the HMD. [Brief description of the drawings]

[0009] [Figure 1] FIG. 1 is a diagram illustrating an example of the overall configuration of an MR system. [Diagram 2] 1 is a diagram illustrating a detailed configuration example of an MR system according to a first embodiment. [Diagram 3] 4 is a flowchart of image processing according to the first embodiment. [Figure 4] FIG. 13 is a diagram illustrating an example of a configuration for correcting a CG image based on a time difference. [Diagram 5] 13 is a flowchart of a process for correcting a CG image based on a time difference. [Figure 6] 11A and 11B are diagrams used to briefly explain correction of a CG image based on a time difference. [Figure 7] 11A to 11C are diagrams used to explain in detail the correction of a CG image based on a time difference. [Figure 8] 4A to 4C are diagrams used for detailed explanation of correction of a CG image in the first embodiment. [Figure 9] FIG. 11 is a diagram illustrating an example of the configuration of an image processing device according to a second embodiment. [Figure 10] 10 is a flowchart of image processing according to the second embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0010] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. The following embodiments do not limit the present invention, and not all of the combinations of features described in the present embodiments are necessarily essential to the solution of the present invention. The configuration of the embodiment may be appropriately modified or changed depending on the specifications of the device to which the present invention is applied and various conditions (conditions of use, environment of use, etc.). In addition, the embodiment may be configured by appropriately combining parts of each embodiment described below. In the following embodiments, the same configurations are described with the same reference symbols.

[0011] <First embodiment> FIG. 1 is a diagram showing a schematic configuration example of an MR (Mixed Reality) system according to the first embodiment. 1, the MR system includes, for example, a head mounted display (HMD) 101 that is worn on the head of a user, and an image processing device 103 having a display device 102 and an operation unit 104. Hereinafter, a user wearing the HMD on his / her head will be referred to as an HMD user.

[0012] The HMD 101 is a video see-through type HMD. The HMD 101 transmits a captured image, captured by an imaging device such as a video camera, of the real world that is approximately the same as that observed from the pupil position of the HMD user to the image processing device 103. The image processing device 103 calculates the position and orientation of the HMD 101, that is, the position and orientation of the head of the HMD user, based on the captured image received from the HMD 101, and generates a computer graphics image (CG image) based on the calculation result. The image processing device 103 then generates an MR image by combining the CG image with the captured image, and transmits the MR image to the HMD 101. The HMD 101 presents the MR image received from the image processing device 103 so that the HMD user can view it. This allows the HMD user to experience an MR space.

[0013] Communication between the HMD 101 and the image processing device 103 is performed wirelessly or by wire. Wireless communication is performed by wireless connection using a small-scale network such as a wireless local area network (WLAN) or a wireless personal area network (WPAN). Note that, although the image processing device 103 and the HMD 101 have separate hardware configurations in the example of Fig. 1, it is also possible to integrate the image processing device 103 by implementing all of the functions of the image processing device 103 in the HMD 101.

[0014] Fig. 2 is a diagram showing main functional units of the HMD 101 and the image processing device 103 according to the first embodiment of the MR system shown in Fig. 1. Fig. 3 is a flowchart showing a processing flow in the image processing device 103 according to the first embodiment.

[0015] In the HMD 101, an imaging unit 201 captures an image of the real world (external world). The imaging unit 201 includes an objective optical system, an image sensor, and the like for capturing an image of the real world (external world). The display unit 202 is a presentation device for presenting an image to an HMD user, and includes an eyepiece optical system, a display, etc. The display unit 202 presents an image sent from the image processing device 103, which will be described later, to the HMD user. Note that the display unit 202 may be, for example, a presentation device of a retina scan type using MEMS (Micro Electro Mechanical Systems), etc. The sensing unit 203 senses the movement of the HMD 101 in a space of the real world. The sensing unit 203 has, for example, an IMU (Inertial Measurement Unit), an acceleration sensor, an angular velocity sensor, etc. as a sensor for sensing the movement of the HMD 101. In this embodiment, the sensor of the sensing unit 203 is an IMU.

[0016] In the image processing device 103, the imaging processing unit 211 executes imaging processing on a captured image of the real world (external world) captured by the imaging unit 201. Here, the imaging processing executed by the imaging processing unit 211 includes demosaic processing, shading correction, noise reduction, distortion correction, etc., and is processing for converting the captured image from the imaging unit 201 into an image corresponding to human visual characteristics. Then, the image that has been subjected to the imaging processing by the imaging processing unit 211 is sent to the synthesis unit 212 and the position and orientation calculation unit 215, respectively.

[0017] The motion calculation unit 214 is a motion acquisition unit that acquires motion information of the imaging unit 201 of the HMD 101 in the real world space, that is, motion information of the head of the HMD user (hereinafter referred to as HMD motion information), based on sensing information from the sensing unit 203 of the HMD 101. The motion calculation unit 214 calculates information such as the movement, inclination, and rotation of the HMD 101 as HMD motion information in the real world space, based on the sensing information sent from the IMU of the sensing unit 203. Then, the motion calculation unit 214 sends the calculated HMD motion information to the position and orientation calculation unit 215 and a correction unit 220 described later.

[0018] The position and orientation calculation unit 215 is a position and orientation acquisition unit that acquires the position and orientation of the HMD 101 in the real world space, that is, the position and orientation of the head of the HMD user, based on the image after imaging processing by the imaging processing unit 211 and HMD movement information from the motion calculation unit 214. First, the position and orientation calculation unit 215 calculates the relationship between a world coordinate system that indicates the real world and a camera coordinate system in an image captured by the imaging unit 201 of the HMD 101, based on the captured image after development processing and the HMD movement information. Then, the position and orientation calculation unit 215 calculates the position and orientation of the HMD 101 with respect to the real world, based on the relationship between the world coordinate system and the camera coordinate system.

[0019] As a method for calculating the position and orientation of the HMD, for example, a method can be used in which a reference marker or the like is placed in the real world, an image is acquired by the imaging unit 201 configured as a stereo camera, and the position and orientation of the HMD are calculated from the positional relationship of the marker in the captured image. Alternatively, a method can be used for calculating the position and orientation of the HMD 101 by using an external sensor or the like that constantly monitors where the HMD 101 is located in the world coordinate system, without using a stereo camera. Detailed configurations and explanations for implementing these position and orientation calculation methods are omitted. In this embodiment, any of these position and orientation calculation methods may be used, and there is no particular limitation. The position and orientation information of the HMD calculated by the position and orientation calculation unit 215 is sent to a CG generation unit 216 .

[0020] A content DB (database) 217 ​​stores data of CG content that is the basis of CG images. The CG generation unit 216 reads CG content from the content DB 217 based on the position and orientation information from the position and orientation calculation unit 215, and generates a CG image based on the CG content. First, the CG generation unit 216 calculates the position and orientation of the captured image at which the CG image should be superimposed, based on the position and orientation information calculated by the position and orientation calculation unit 215. Then, the CG generation unit 216 reads CG content for generating a CG image at that position and orientation from the content DB 217, and renders the CG image. The CG image generated by the CG generation unit 216 is sent to a separation unit 221 which constitutes a CG correction unit, which will be described later.

[0021] The separation unit 221 separates the CG image generated by the CG generation unit 216 into a color channel (hereinafter referred to as color CH) which is color information of the image, and channel data (hereinafter referred to as composite information CH) consisting of composite information, which will be described later. Details of the composite information and the channel separation process in the separation unit 221 will be described later. The color CH and composite information CH data separated by the separation unit 221 are then sent to the correction unit 220 which, together with the separation unit 221, constitutes a CG correction unit. The correction unit 220 corrects the data of the color CH and the combination information CH separated by the separation unit 221, based on the HMD movement information calculated by the movement calculation unit 214. Details of the correction process by the correction unit 220 will be described later. The data of the color CH and the combination information CH after correction by the correction unit 220 is then sent to the combination unit 212.

[0022] The synthesis unit 212 generates a CG image (corrected CG image) from the image after imaging processing by the imaging processing unit 211 and the color CH and synthesis information CH data after correction by the correction unit 220 described later. Furthermore, the synthesis unit 212 synthesizes the corrected CG image with the image after imaging processing by the imaging processing unit 211 to generate a synthesized image (MR image). The synthesized image (MR image) by the synthesis unit 212 is then sent to the display unit 202 of the HMD 101 and displayed. This allows the user of the HMD 101 to experience MR.

[0023] Before explaining the separation unit 221 and correction unit 220 according to this embodiment and the flowchart of FIG. 3, the time difference that occurs between the image captured by the imaging unit 201 of the HMD 101 and the CG image generated by the CG generation unit 216 will be explained. As described above, the CG generation unit 216 generates a CG image based on the image after development processing by the imaging processing unit 211 and the position and orientation information calculated by the position and orientation calculation unit 215 using the HMD movement information calculated by the movement calculation unit 214. However, the processing load in the CG generation unit 216 varies greatly depending on the CG image to be generated. For example, when the processing load is large and rendering of the CG image takes a long time, a time mismatch may occur between a captured image of the real world and the rendered CG image. Since the composite image (MR image) viewed by the HMD user is an image obtained by combining a captured image of the real world and a CG image, if there is a time mismatch between the captured image and the CG image, the image will be unnatural to the HMD user.

[0024] Hereinafter, a configuration and process as an example that can generate a synthetic image with less awkwardness by correcting the time difference between a captured image and a CG image will be described with reference to Figs. 4 to 7.

[0025] Fig. 4 is a diagram showing an example of the configuration of an MR system having an image processing device 403 as an example capable of generating a synthetic image with less awkwardness by considering the time difference between a captured image and a CG image. The image processing device 403 shown in Fig. 4 does not include the separation unit 221 and the correction unit 220 of the image processing device 103 of the first embodiment shown in Fig. 2, but instead includes an image correction unit 413. In the example of the configuration in Fig. 4, components that perform roughly the same processes as the components shown in Fig. 2 described above are given the same reference symbols as in Fig. 2, and detailed descriptions thereof will be omitted.

[0026] The image correction unit 413 in Figure 4 performs image correction processing on the CG image generated by the CG generation unit 216 based on the time difference between the captured image and the CG image, thereby generating a CG image that can eliminate the sense of incongruity in the composite image caused by the time mismatch between the captured image and the CG image. 5 is a flowchart showing the flow of image correction processing in image correction unit 413. In the following flowcharts, the symbol S represents a processing step.

[0027] First, as the process of S501, the image correction unit 413 receives HMD movement information such as the movement, inclination, and rotation of the HMD from the movement calculation unit 214. Next, the image correction unit 413 acquires the time difference between the captured image and the CG image as the process of S502. For example, the image correction unit 413 sets the average time required for rendering the CG image as a fixed delay value, and acquires the time difference between the fixed delay amount and the captured image. Alternatively, the image correction unit 413 sets the time required for rendering the CG image as a variable delay amount, and acquires the time difference between the variable delay amount and the captured image. Note that these time difference acquisition methods are merely examples and are not particularly limited.

[0028] Next, in processing S503, the image correction unit 413 calculates the position to which the CG image should be moved relative to the captured image based on the HMD movement information obtained in S501 and the time difference information obtained in S502, and calculates a homography matrix corresponding to the destination position. Then, in the process of S504, the image correction unit 413 performs image conversion by homography conversion using the homography matrix calculated in S503, and performs correction processing on the CG image. As a method of the correction processing at this time, a method such as bilinear interpolation can be considered. The image correction unit 413 makes the appearance of the CG image match the captured image by performing the process shown in the flowchart of Fig. 5. The CG image corrected by the image correction unit 413 is sent to the synthesis unit 212, and is synthesized with the captured image as described above to generate a synthesized image.

[0029] FIG. 6 is a diagram used to explain the sense of incongruity of a composite image caused by a temporal mismatch between a captured image and a CG image, and the effect of reducing the sense of incongruity in a composite image generated using a CG image after image correction processing by the image correction unit 413 described above. 6 shows a captured image of the real world, a CG image generated by the CG generating unit 216, a CG image after image correction by the image correcting unit 413 (hereinafter referred to as a corrected CG image), and a composite image obtained by combining the captured image and the corrected CG image. Fig. 6(a) shows an example of a captured image 601a, a CG image 602a, and a composite image 604a obtained by combining the captured image 601a and the CG image 602a when the HMD 101 is stationary. Fig. 6(b) shows an example of a captured image 601b, a CG image 602b, and a composite image 604b obtained by combining the captured image 601b and the CG image 602b when the HMD user, for example, shakes his / her head. 6(c) shows an example of a captured image 601c and a CG image 602c similar to those in FIG. 6(b), a corrected CG image 603c, and a composite image 604c obtained by combining the captured image 601c and the corrected CG image 603c. It is assumed that the captured image is, for example, an image of a sphere placed in the real world, the CG image is, for example, a rectangular image, and the composite image is an image in which the sphere in the captured image is partially covered by the CG image.

[0030] As shown in the example of Fig. 6(a), when the HMD 101 is stationary, even if there is a time mismatch between the captured image 601a and the CG image 602a, the rendering position of the CG image in the composite image 604a does not appear strange to the eye relative to the captured image of the real world. On the other hand, as shown in the example of Fig. 6(b), when the HMD user shakes his / her head, the CG image 602b is delayed from the captured image 601b by the rendering time, so that the rendering position of the CG image in the composite image 604b is shifted relative to the captured image of the real world. In other words, the HMD user feels strange because the position of the CG image in the actual composite image is shifted from the CG image that should be seen when the HMD user shakes his / her head.

[0031] In response to this, the image correction unit 413 performs image correction processing on the CG image 602c so that the CG image 602c coincides in time with the captured image 601c in accordance with the movement of the HMD 101 as described above. That is, in the image correction processing by the image correction unit 413, the CG image 602c is corrected to coincide with the change in the captured image 601c caused by the head swing of the HMD user, as shown in Fig. 6(c), to obtain a corrected CG image 603c. As a result, even if the captured image 601c and the CG image 602c do not coincide in time, the composite image 604 in which the corrected CG image 603c is composited becomes an image that gives little discomfort to the HMD user.

[0032] However, when a captured image of the real world is combined with a CG image after the above-mentioned image correction process, the depth relationship between the objects and the CG image in the captured image may not match. When the depth relationship between the objects and the CG image in the captured image does not match in this way, the image will feel unnatural to the HMD user.

[0033] Hereinafter, an example will be described in which the depth relationship between an object or the like shown in a captured image and a CG image becomes inconsistent as a result of image correction processing by the image correction unit 413. As described above, when generating a composite image by synthesizing a CG image with a captured image of the real world, the depth relationship between the objects and the CG image shown in the captured image must be correct. For this reason, the imaging processing unit 211 also acquires depth information in the real world corresponding to the captured image by the imaging unit 201. As a method for acquiring depth information corresponding to the captured image, for example, a method called LiDAR (Light Detection and Ranging) that acquires depth information from the time difference between irradiation of laser light and reflected light can be mentioned. In addition, there is also a method of acquiring depth information from a parallax image using a stereo camera. Detailed configurations and explanations for implementing these depth information acquisition methods are omitted. Any of these depth information acquisition methods may be used, and are not particularly limited.

[0034] Furthermore, the CG generating unit 216 generates a CG image including synthesis information based on the CG content read from the content DB 217. The synthesis information is information related to synthesis of a captured image and a CG image, and in this example includes alpha channel information and depth channel information, and is used as additional information of the CG image. The alpha channel information is transparency information that indicates transparency, and the depth channel information is depth information.

[0035] Fig. 7 is a diagram showing an example of a case where a CG image including an alpha channel and a depth channel is corrected by the image correction unit 413 of Fig. 4 described above, and then synthesized with a captured image to generate a synthesized image. Fig. 7(a) shows a captured image 701a, a CG image 702a, a corrected CG image 703a, and a synthesized image 704a obtained by synthesizing the captured image 701a and the corrected CG image 703a, similar to the example of Fig. 6(c). In Fig. 7(a) as in the example of Fig. 6(a), the captured image 701a is an image captured in a state where a sphere is placed in the real world, the CG image 702a is a rectangular image, and the synthesized image 704a is an image in which a part of the sphere in the captured image is covered with a CG image.

[0036] The difference between Fig. 7(a) and Fig. 6(c) is that the CG image 702a includes an alpha channel and a depth channel. As described above, it is assumed that the captured image 701a has depth information in the real world acquired. Fig. 7(b) shows enlarged images 711b-714b of the regions E in Fig. 7(a), and Fig. 7(c) shows examples of data 721c-723c of values ​​A and Z for each pixel in the enlarged images 711b-713b in Fig. 7(b).

[0037] The alpha channel value A for each pixel shown in Fig. 7(c) is a value that represents the transparency and is a value between 0% and 100%. For example, when A=0%, the CG image is in a transparent state (CG image is 0%, captured image is 100%), and when A=100%, the CG image is in a masked state (CG image is 100%, captured image is 0%). To represent a semi-transparent state, for example, by setting A=25%, it is possible to represent a semi-transparent CG image (CG image is 25%, captured image is 75%).

[0038] The value Z of the depth channel for each pixel is a value that represents depth information, and a value between 0 and 10 is used. For example, Z=0 indicates the boundary of the rear clip plane in the CG drawing space, and Z=10 indicates the boundary of the front clip plane in the CG drawing space. That is, Z=0 indicates that the depth information in the space represents the deepest part, and Z=10 indicates that the depth information in the space represents the foremost part. By changing the value Z of the depth channel from 0 to 10, it becomes possible to represent spatial depth information in a two-dimensional image. Note that the depth information acquired for the captured image is also represented as the value Z, similar to the depth channel. In this example, the alpha channel value A and the depth channel value Z have been described as described above, but the method of clipping information is not particularly limited.

[0039] Here, in the enlarged image 711b of Fig. 7(b) which is an enlargement of the region E of the boundary part of the sphere in the captured image 701a of Fig. 7(a), the depth information value Z corresponding to each pixel is assumed to be a value as shown in the data 721c of Fig. 7(c). That is, as shown in Fig. 7(c), the depth information of the boundary part of the sphere and the inside of the sphere in the captured image 701a is all assumed to be Z=5. For the sake of simplicity, the depth information of the region other than the sphere in the captured image 701a is omitted.

[0040] In addition, in an enlarged image 712b in FIG. 7(b) which is an enlargement of an area E of the boundary part of the rectangular CG image 702a in FIG. 7(a), the alpha channel value A and the depth channel value Z for each pixel are assumed to be values ​​as shown in FIG. 7(c). In the example of FIG. 7(c), the depth channel value Z of each pixel in the rectangular CG image 702a is all assumed to be Z=6. Meanwhile, the alpha channel value A of each pixel is assumed to be A=100% (CG image is 100% and the captured image is 0%) at the edge part which is the boundary with the captured image, whereas it is A=50% (CG image is 50% and the captured image is also 50%) inside the edge part. Note that there are no alpha channel and depth channel values ​​in areas other than the rectangular CG image.

[0041] In this way, when the value Z of the CG image 702a is 6 and the value Z of the sphere of the captured image 701a is 5, the rectangular CG image 702a is placed in front of the sphere of the captured image 701a. Furthermore, in the example of Fig. 7, the value A of the edge portion of the CG image 702a is 100%, so the captured image is 100% masked at the edge portion, while the value A of the inside of the edge portion is 50%, so the captured image is visible inside the edge portion. In other words, from the viewpoint of the HMD user, the CG image is placed in front of the sphere, and the sphere behind the CG image is visible through the CG image.

[0042] For example, if the HMD user shakes his / her head, the image correction unit 413 performs image correction processing such as performing homography transformation using the homography matrix as described above on the rectangular CG image 702a in accordance with the movement of the HMD 101. That is, the image correction unit 413 performs homography transformation on the CG image 702a in accordance with the movement of the HMD 101 to generate a corrected CG image 703a. However, when attention is paid to the value A of the alpha channel and the value Z of the depth channel at this time, the original value A (transparency information) and value Z (depth information) may be lost due to the influence of the homography transformation.

[0043] This will be described using CG image 702a, enlarged image 712b, and data 722c before homography conversion, and corrected CG image 703a, enlarged image 713b, and data 723c after homography conversion in FIG. For example, when homography transformation is performed on the rectangular CG image 702a in FIG. 7(a) in accordance with the movement of the HMD 101 when the HMD user shakes his / her head, the rectangular CG image 702a is transformed into a parallelogram CG image 703a. In other words, when homography transformation is performed, the edge portion of the rectangular CG image 702a is transformed from a straight line as in the enlarged image 712b to a smoothed diagonal line as in the enlarged image 713b. The smoothing process for smoothing such diagonal lines is realized by interpolating pixel data from the pixels surrounding the edge portion, and is a common process for image correction. However, at this time, in the pixels of the edge portion of the corrected CG image 703a, the values ​​A and Z of the surrounding pixels are also interpolated during the smoothing process when converting from a straight line to a diagonal line, and the information of the original values ​​A and Z may be lost.

[0044] For example, let us take pixels P700 and P701 at the edge of CG image 702a before homography transformation and the corresponding pixels Q700 and Q701 in CG image 703a after homography transformation. The alpha channel value A of pixels P700 and P701 is 100%, and the depth channel value Z is 6. After the edge is transformed from a straight line to a diagonal line by homography transformation, the positions of pixels P700 and P701 move to the positions of pixels Q700 and Q701.

[0045] Here, if we look at pixel P700, the pixel position has simply moved due to the homography transformation, so value A and value Z of pixel Q700 are maintained at 100% of value A of pixel P700 and value Z of 6, and no problem occurs.

[0046] On the other hand, when focusing on pixel P701, in addition to the movement of the pixel position by homography transformation, the value A of pixel Q701 changes to 50% and the value Z changes to 3 due to the interpolation process from surrounding pixels by smoothing. That is, in pixel Q701, the value Z of the depth channel in particular changes to 3, which is smaller than the value Z of 5 of the depth information of the captured image 701a. In this case, since the pixel of the captured image 701a is on the near side, when this corrected CG image 703a is synthesized with the captured image 701a, pixel O701 of the captured image 701a at a position corresponding to pixel Q701 of the corrected CG image 703a will be positioned on the near side. That is, in the corrected CG image 703a, the front-to-back relationship in the depth direction that should be originally maintained is reversed, pixel Q701 is buried, and the synthesized image becomes inappropriate.

[0047] As described above, if a correction process is performed on a CG image taking into account the time difference between the captured image and the CG image, the depth relationship between objects, etc. shown in the captured image and the CG image may not match, resulting in an inappropriate composite image.

[0048] <Configuration and Processing for Image Correction According to the First Embodiment> Therefore, the image processing device 103 according to this embodiment shown in Fig. 2 separates a CG image into a color channel and a composite information channel including the above-mentioned alpha channel and depth channel, and performs separate image correction processing for each channel. That is, the image processing device 103 according to the first embodiment shown in Fig. 2 includes a CG correction unit consisting of a separation unit 221 and a correction unit 220, instead of the image correction unit 413 as described in Fig. 4.

[0049] In the image processing device 103 of this embodiment, the CG generation unit 216 generates a CG image including synthesis information of the alpha channel and the depth channel as described above. Hereinafter, this CG image will be referred to as a CG image with synthesis information. Then, the CG image with synthesis information is sent to the separation unit 221.

[0050] The separation unit 221 separates the CG image with the composition information into a color channel (color CH) which is color information of the image, and a composition information channel (composition information CH) which includes an alpha channel and a depth channel. The color CH is input to a plane correction unit 222 of the correction unit 220, and the composition information CH is input to a space correction unit 223 of the correction unit 220.

[0051] Hereinafter, processing performed in plane correction section 222 and space correction section 223 of correction section 220 will be described with reference to the flowchart in Fig. 3. Note that the processing of S301 to S303 in the flowchart in Fig. 3 is executed commonly by plane correction section 222 and space correction section 223.

[0052] First, the plane correction unit 222 and the spatial correction unit 223 acquire the above-mentioned HMD movement information calculated by the movement calculation unit 214 as the process of S301. Next, in the process of S302, the plane correction unit 222 and the space correction unit 223 obtain the above-mentioned time difference between the captured image and the CG image.

[0053] Next, in the process of S303, the plane correction unit 222 and the space correction unit 223 calculate the position to which the CG image should be moved based on the HMD movement information acquired in S301 and the time difference information acquired in S302. Then, the plane correction unit 222 and the space correction unit 223 calculate a homography matrix according to the destination.

[0054] Next, in S304, the correction unit 220 branches the process depending on whether it is color CH or composite information CH. That is, in the case of color CH, the correction unit 220 advances the process to S305, and in the case of CG composite information, the correction unit 220 advances the process to S306.

[0055] When the process proceeds to S305, the plane correction unit 222 executes image correction (image conversion) processing on the color CH of the CG image using the homography matrix calculated in S303. At this time, the color CH of the CG image needs to be two-dimensionally matched with the captured image that follows the movement of the HMD 101. For this reason, the plane correction unit 222 calculates coordinate information below the unit coordinate as conversion coordinates from the homography matrix, and executes image correction by pixel interpolation such as bilinear interpolation with the conversion coordinates as the reference pixel position. That is, the image correction by pixel interpolation in the plane correction unit 222 is a correction process in which color information at the reference pixel position is referenced, and a value obtained by smoothing the color information corresponding to the reference pixel position is used as a correction value for pixel interpolation.

[0056] On the other hand, when the process proceeds to S306, the spatial correction unit 223 performs image correction (image conversion) on the synthesis information CH (alpha channel / depth channel) of the CG image using the homography matrix calculated in S303. At this time, the above-mentioned interpolation process is not performed on the alpha channel value A and the depth channel value Z, but pixel replacement is performed so as to spatially match the captured image. That is, the spatial correction unit 223 performs image correction without interpolation process based on the reference pixel position calculated from the homography matrix in the same manner as described above. In the image correction without interpolation process in the spatial correction unit 223, for example, synthesis information of the synthesis information CH corresponding to the reference pixel position is used as a correction value for pixel replacement. Note that the correction value for pixel replacement may be any one of the minimum value, median value, maximum value, and neighboring value of the synthesis information corresponding to the reference pixel position.

[0057] In this manner, the image processing device 103 of this embodiment separates the CG image into a color CH and a composite information CH including an alpha channel and a depth channel, and performs separate image correction processing for each channel.

[0058] The effects of the image processing device 103 of this embodiment will be described below with reference to FIG. FIG. 8 is a diagram showing an example in which a composite image is generated by performing image correction processing on a CG image divided into a color CH and composite information CH as described above, and then compositing the corrected CG image with a captured image. FIG. 8(a) shows a captured image 801a, a CG image 802a, a corrected CG image 803a according to this embodiment, and a composite image 804a obtained by compositing the captured image 801a and the corrected CG image 803a, as in the example of FIG. 7(a). Note that the captured image 801a and the CG image 802a in FIG. 8(a) are similar to the captured image 701a and the CG image 702a in FIG. 7(a). The enlarged images 811b and 812b in FIG. 8(b) are similar to the enlarged images 711b and 712b in FIG. 7(b), and the data 821c and 822c in FIG. 8(c) are similar to the data 721c and 722c in FIG. 7(c).

[0059] The corrected CG image 803a in Fig. 8(a) and the enlarged image 813b in Fig. 8(b) are CG images after image correction has been performed on only the color CH by the above-mentioned plane correction unit 222. Moreover, in the corrected CG image 803a and the enlarged image 813b, similar to the example of the corrected CG image 703a and the enlarged image 713b in Fig. 7, the edges have been smoothed by smoothing processing.

[0060] Data 823c in Fig. 8(c) indicates the values ​​A and Z of each pixel after image correction processing is performed on the synthesis information CH (alpha channel / depth channel) by the spatial correction unit 223. In the case of this embodiment, the spatial correction unit 223 performs image correction without interpolation processing from the reference pixel position calculated from the homography matrix, so the values ​​A and Z of each pixel after image correction on the synthesis information CH become as shown in data 823c.

[0061] In the data 723c in FIG. 7(c) described above, the value A of pixel Q701 changes to 50% and the value Z to 3 due to the smoothing interpolation process, and in particular the value Z changes from the original value, causing the front-to-back relationship between the captured image and the CG image at that pixel to be reversed. In contrast, in the configuration of FIG. 3, the color CH and the composite information CH (alpha channel, depth channel) are separated, and the composite information CH is not subjected to interpolation processing. Therefore, the values ​​A and Z of pixels R800 and R801 in data 823c in FIG. 8(c), which correspond to pixels Q700 and Q701 in data 723c in FIG. 7(c), are retained as original values. In particular, image correction is performed on pixel R801, retaining information that the value A is 100% and the value Z is 6.

[0062] Therefore, the composite image 804a (enlarged image 814b) obtained by combining the captured image 801a and the corrected CG image 803a in the combining unit 212 is an appropriate composite image in which the front-to-back relationship in the depth direction that should be maintained is correct. As described above, in the image processing device 103 of the first embodiment, when performing image correction, the CG image is separated into a color CH and a composite information CH, and by performing separate image correction processing on each, it becomes possible to appropriately correct the CG image and composite it with the captured image. Therefore, according to this embodiment, it is possible to provide an image that does not feel strange to the user of the HMD.

[0063] In the above description, an example is given in which the captured image is masked when the alpha channel value A is 100%, but a channel of mask information (mask channel) may be used in addition to the alpha channel. That is, in the present embodiment, an example is given in which the composite information CH includes two channels, an alpha channel and a depth channel, but the composite information CH may include three channels, an alpha channel, a depth channel, and a mask channel. In addition, the composite information CH may include at least one of the alpha channel, the depth channel, and the mask channel.

[0064] <Second embodiment> In the first embodiment described above, an example was given in which, in image correction to eliminate a temporal mismatch between a CG image including an alpha channel and a depth channel and a captured image, the CG image is separated into a color CH and a composite information CH, and each is corrected independently. In the second embodiment, an example will be described in which, in addition to image correction to a CG image similar to that in the first embodiment, image correction is also performed on a captured image captured by the imaging unit 201. Note that image correction to a CG image is performed in a CG correction unit composed of a separation unit 221 and a correction unit 220 similar to those in the first embodiment described above, and therefore a description thereof will be omitted.

[0065] An HMD user experiencing MR using the HMD 101 visually recognizes the captured image captured by the imaging unit 201 as an image of the real world. However, the captured image visually recognized by the HMD user is an image after photoelectric conversion and the like are performed by the imaging unit 201 and further imaging processing (demosaic processing, shading correction, noise reduction, distortion correction, etc.) is performed by the imaging processing unit 211. In other words, the captured image visually recognized by the HMD user is an image including delay time due to electrical conversion time such as photoelectric conversion and imaging processing, and the delay time is several tens of milliseconds. In other words, a time mismatch occurs between the real world and the captured image visually recognized by the HMD user. For this reason, for example, when the HMD user quickly shakes his head, the image visually recognized by the HMD user becomes an unnatural image delayed by the delay time compared to the appearance when directly viewing the real world without using the HMD.

[0066] For this reason, in the image processing device 903 of the second embodiment, it is also possible to perform image correction on the captured image to correct the temporal discrepancy between the real world and the captured image, based on the HMD movement information calculated by the movement calculation unit 214. However, when a captured image contains depth information, the depth information that should be held may change when image correction is performed on the captured image to correct the time discrepancy between the real world and the captured image. If the depth information that should be held changes, the front-to-back relationship in the depth direction with the CG image synthesized by the synthesis unit 212 may be reversed, resulting in an unnatural synthesized image as described above.

[0067] Therefore, the image processing device 903 of the second embodiment separates a captured image including depth information into a color channel (color CH) of the image and a channel of depth information (hereinafter referred to as depth CH), and performs independent image correction processing on each. Fig. 9 is a diagram showing main functional units of the HMD 101 and image processing device 903 of the MR system according to the second embodiment. In Fig. 9, the same functional units as those in Fig. 2 are given the same reference numerals as those in Fig. 2, and their explanations will be omitted as appropriate. Fig. 10 is a flowchart showing the flow of processing in the image processing device 903 of the second embodiment.

[0068] As shown in FIG. 9, the image processing device 903 according to the second embodiment further includes a separation unit 911 and a correction unit 920 as an image correction unit, in addition to the functional units of the image processing device 103 as shown in FIG. 2 described above. The separation unit 911 separates the captured image after development processing by the imaging processing unit 211 into a color CH and a depth CH. Then, the color CH separated by the separation unit 911 is input to a plane correction unit 922 of the correction unit 920, and the depth CH is input to a spatial correction unit 923 of the correction unit 920.

[0069] Hereinafter, processing performed by plane correction unit 922 and space correction unit 923 of correction unit 920 will be described with reference to the flowchart of Fig. 10. Note that the processing of S1001 to S1003 in the flowchart of Fig. 10 is executed commonly by plane correction unit 922 and space correction unit 923. First, the plane correction unit 922 and the spatial correction unit 923 acquire the above-mentioned HMD movement information calculated by the movement calculation unit 214 as the process of S1001.

[0070] Next, as processing in S1002, the plane correction unit 922 and the spatial correction unit 923 acquire the delay time included in the captured image, that is, the delay time due to the electrical conversion time due to photoelectric conversion in the imaging unit 201 and the imaging processing time in the imaging processing unit 211. Next, in the process of S1003, the plane correction unit 922 and the space correction unit 923 calculate a position to which the captured image should be moved based on the HMD movement information acquired in S1001 and the delay time information acquired in S1002. Then, the plane correction unit 922 and the space correction unit 923 calculate a homography matrix according to the destination.

[0071] Next, in S1004, the correction unit 920 branches the process between the color CH and the depth CH. That is, the correction unit 920 advances the process to S1005 in the case of the color CH, and advances the process to S1006 in the case of the depth CG.

[0072] When the process proceeds to S1005, the plane correction unit 922 executes image correction processing on the color CH of the captured image using the homography matrix calculated in S1003. At this time, since the color CH of the captured image needs to be moved by the delay time, the plane correction unit 922 executes image correction by pixel interpolation such as bilinear interpolation from the reference pixel position calculated from the homography matrix.

[0073] On the other hand, when the process proceeds to S1006, the spatial correction unit 923 performs correction on the depth CH of the captured image using the homography matrix calculated in S1003. At this time, the spatial correction unit 923 performs pixel replacement so as to spatially match the depth CH with the captured image without performing interpolation processing or the like. That is, the spatial correction unit 923 performs correction processing using depth information according to the reference pixel position calculated from the homography matrix as a correction value for pixel replacement. In other words, the spatial correction unit 923 performs image correction without interpolation processing according to the reference pixel position calculated from the homography matrix.

[0074] Thereafter, the composition unit 212 composes the captured image after image correction with the CG image after image correction in the same manner as described above. In the second embodiment, the CG image is also a CG image that has been image-corrected to eliminate the time discrepancy with the captured image as described in the first embodiment. A composite image obtained by combining the captured image and the CG image after image correction is sent to the HMD 101 and displayed on the display unit 202. As a result, the composite image displayed on the display unit 202 can be displayed as an image in which the time discrepancy with the real world is eliminated for both the captured image and the CG image.

[0075] As described above, in the second embodiment, in addition to the correction of the CG image as in the first embodiment, the captured image including the depth information is separated into the color CH and the depth CH, and separate image correction processing is performed for each. This makes it possible to provide the HMD user with a composite image obtained by combining the captured image and the CG image, each of which has been appropriately corrected for time mismatch, that is, to provide the HMD user with a natural MR experience.

[0076] In the second embodiment, an example is given in which both image correction for a captured image and image correction for a CG image similar to that in the first embodiment are performed. However, if no time difference occurs between the captured image and the CG image and image correction according to the time difference for the CG image as described above is not necessary, only image correction for the captured image may be performed.

[0077] The present invention can also be realized by supplying a program that realizes one or more functions of the above-mentioned embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions. The above-mentioned embodiments are merely examples of concrete implementations of the present invention, and the technical scope of the present invention should not be interpreted as being limited by them. In other words, the present invention can be implemented in various forms without departing from its technical concept or main features.

[0078] The disclosure of this embodiment includes the following configuration, method, and program. (Configuration 1) An acquisition means for acquiring an image captured by an imaging device of the real world; a motion acquisition means for acquiring motion information of the imaging device in the real world space; a CG generating means for generating a computer graphics image including synthesis information related to synthesis with the captured image based on motion information of the imaging device; a CG correction means for separating a computer graphics image including the composite information into an image channel and a composite information channel, and correcting the image channel by pixel interpolation in accordance with the motion information of the imaging device, and correcting the composite information channel by pixel replacement in accordance with the motion information of the imaging device; a synthesis means for synthesizing the captured image and the computer graphics image corrected by the CG correction means; 13. An image processing device comprising: (Configuration 2) The acquisition means also acquires depth information of the real world corresponding to the captured image, 2. The image processing device according to claim 1, wherein the synthesis information includes at least one of depth information and transparency information. (Configuration 3) 3. The image processing device according to configuration 1 or 2, wherein the CG correction means calculates a homography matrix based on motion information of the imaging device, and performs the correction based on transformation coordinates calculated from the homography matrix. (Configuration 4) The image processing device according to configuration 3, wherein the converted coordinates are coordinate information equal to or smaller than unit coordinates, and the correction by pixel interpolation is a correction in which a value obtained by smoothing color information corresponding to the coordinate information is used as a correction value for pixel interpolation. (Configuration 5) The image processing device according to configuration 3 or 4, wherein the converted coordinates are coordinate information equal to or smaller than unit coordinates, and the correction by pixel replacement is correction using synthesis information corresponding to the coordinate information as a correction value for pixel replacement. (Configuration 6) The image processing device according to configuration 3 or 4, wherein the converted coordinates are coordinate information equal to or smaller than unit coordinates, and the correction by pixel replacement is a correction using any one of a minimum value, a median value, a maximum value, and a neighboring value of synthesis information corresponding to the coordinate information as a correction value for pixel replacement. (Configuration 7) The acquisition means also acquires depth information of the real world corresponding to the captured image, a captured image correction means for separating the captured image into an image channel and a depth information channel, correcting the image channel by pixel interpolation according to the motion information of the imaging device, and correcting the depth information channel by pixel replacement according to the motion information of the imaging device; 7. The image processing device according to any one of configurations 1 to 6, wherein the combining means combines the captured image corrected by the captured image correction means and the computer graphics image corrected by the CG correction means. (Configuration 8) an acquisition means for acquiring a captured image of a real world captured by an imaging device and depth information of the real world corresponding to the captured image; a motion acquisition means for acquiring motion information of the imaging device in the real world space; a CG generating means for generating a computer graphics image based on the motion information of the imaging device; a captured image correction means for separating the captured image into an image channel and a depth information channel, and correcting the image channel by pixel interpolation in accordance with motion information of the imaging device, and correcting the depth information channel by pixel replacement in accordance with the motion information of the imaging device; a synthesis means for synthesizing the captured image corrected by the captured image correction means and the computer graphics image generated by the CG generation means; 13. An image processing device comprising: (Configuration 9) 9. The image processing device according to configuration 8, wherein the captured image correction means calculates a homography matrix based on motion information of the image capturing device, and performs the correction based on transformation coordinates calculated from the homography matrix. (Configuration 10) The image processing device according to configuration 9, characterized in that the converted coordinates are coordinate information equal to or smaller than unit coordinates, and the correction by pixel interpolation is a correction in which a value obtained by smoothing color information corresponding to the coordinate information is used as a correction value for pixel interpolation. (Configuration 11) The image processing device according to configuration 9 or 10, characterized in that the transformation coordinates are coordinate information equal to or smaller than unit coordinates, and the correction by pixel replacement is a correction in which the depth information corresponding to the coordinate information is used as a correction value for pixel replacement. (Configuration 12) the imaging device has a sensing unit that senses a movement of the imaging device in a space of the real world; the motion acquisition means acquires motion information of the imaging device based on sensing information by the sensing unit; a position and orientation acquisition means for acquiring a position and orientation of the image capture device in the real world space based on the motion information of the image capture device, 12. The image processing device according to any one of configurations 1 to 11, wherein the CG generating means generates the computer graphics image based on the position and orientation of the imaging device. (Configuration 13) 13. The image processing device according to configuration 12, wherein the sensing unit is an IMU (Inertial Measurement Unit). (Method 1) An acquisition step of acquiring an image captured by an imaging device of the real world; a motion acquisition step of acquiring motion information of the imaging device in the real world space; a CG generation step of generating a computer graphics image including synthesis information related to synthesis with the captured image based on the motion information of the imaging device; a CG correction step of separating a computer graphics image including the synthesis information into an image channel and a synthesis information channel, correcting the image channel by pixel interpolation according to the motion information of the imaging device, and correcting the synthesis information channel by pixel replacement according to the motion information of the imaging device; a synthesis step of synthesizing the captured image and the computer graphics image corrected by the CG correction step; 13. An image processing method comprising: (Method 2) an acquisition step of acquiring a captured image of a real world captured by an imaging device and depth information of the real world corresponding to the captured image; a motion acquisition step of acquiring motion information of the imaging device in the real world space; a CG generation step of generating a computer graphics image based on the motion information of the imaging device; a captured image correction step of separating the captured image into an image channel and a depth information channel, correcting the image channel by pixel interpolation according to the motion information of the imaging device, and correcting the depth information channel by pixel replacement according to the motion information of the imaging device; a synthesis step of synthesizing the captured image corrected by the captured image correction step and the computer graphics image generated by the CG generation step; 13. An image processing method comprising: (Program 1) A program for causing a computer to function as the image processing device according to any one of configurations 1 to 13. [Explanation of symbols]

[0079] 101: HMD, 103: image processing device, 201: imaging unit, 202: display unit, 203: sensing unit, 211: imaging processing unit, 212: synthesis unit, 214: motion calculation unit, 215: position and orientation calculation unit, 216: CG generation unit, 220: correction unit, 221: separation unit, 217: content DB

Claims

1. An acquisition means that acquires an image captured by an imaging device that captures the real world, An operation acquisition means for acquiring motion information of the imaging device within the space of the real world, A CG generation means generates a computer graphics image that includes composite information related to the synthesis with the captured image, based on the motion information of the imaging device. A CG correction means that separates a computer graphics image containing the composite information into image channels and composite information channels, performs correction on the image channels by pixel interpolation according to the motion information of the imaging device, and performs correction on the composite information channels by pixel replacement according to the motion information of the imaging device, A synthesis means for combining the captured image and the computer graphics image corrected by the CG correction means, An image processing apparatus characterized by having

2. The acquisition means also acquires depth information of the real world corresponding to the captured image, The image processing apparatus according to claim 1, characterized in that the composite information includes at least one of depth information and transparency information.

3. The image processing apparatus according to claim 1, characterized in that the CG correction means calculates a homography matrix based on motion information of the imaging device and performs the correction based on the transformed coordinates calculated from the homography matrix.

4. The image processing apparatus according to claim 3, characterized in that the transformed coordinates are coordinate information with decimal precision, and the correction by pixel interpolation is a correction in which a value obtained by smoothing color information corresponding to the coordinate information is used as the correction value for pixel interpolation.

5. The image processing apparatus according to claim 3, characterized in that the transformed coordinates are coordinate information with decimal precision, and the correction by pixel replacement is a correction that uses composite information corresponding to the coordinate information as the correction value for pixel replacement.

6. The image processing apparatus according to claim 3, characterized in that the transformed coordinates are coordinate information with decimal precision, and the correction by pixel replacement is a correction in which one of the minimum, intermediate, maximum, or nearest neighbor values ​​of the composite information corresponding to the coordinate information is used as the correction value for pixel replacement.

7. The acquisition means also acquires depth information of the real world corresponding to the captured image, The system further includes image correction means that separates the captured image into an image channel and a depth information channel, corrects the image channel by pixel interpolation according to the motion information of the imaging device, and corrects the depth information channel by pixel replacement according to the motion information of the imaging device. The image processing apparatus according to claim 1, characterized in that the synthesis means synthesizes the captured image corrected by the captured image correction means and the computer graphics image corrected by the CG correction means.

8. An acquisition means that acquires an image captured by an imaging device of the real world and depth information of the real world corresponding to the image captured, An operation acquisition means for acquiring motion information of the imaging device within the space of the real world, A CG generation means that generates a computer graphics image based on the motion information of the imaging device, Image correction means for separating the captured image into an image channel and a depth information channel, correcting the image channel by pixel interpolation according to the motion information of the imaging device, and correcting the depth information channel by pixel replacement according to the motion information of the imaging device, A synthesis means for combining the image image corrected by the image image correction means and the computer graphics image generated by the CG generation means, An image processing apparatus characterized by having

9. The image processing apparatus according to claim 8, characterized in that the image acquisition image correction means calculates a homography matrix based on motion information of the imaging device and performs the correction based on the transformed coordinates calculated from the homography matrix.

10. The image processing apparatus according to claim 9, characterized in that the transformed coordinates are coordinate information with decimal precision, and the correction by pixel interpolation is a correction in which a value obtained by smoothing color information corresponding to the coordinate information is used as the correction value for pixel interpolation.

11. The image processing apparatus according to claim 9, characterized in that the transformed coordinates are coordinate information with decimal precision, and the correction by pixel replacement is a correction that uses the depth information corresponding to the coordinate information as the correction value for pixel replacement.

12. The imaging device has a sensing unit that senses the movement of the imaging device within the space of the real world, The motion acquisition means acquires motion information of the imaging device based on the sensing information from the sensing unit. The system further includes position and orientation acquisition means for acquiring the position and orientation of the imaging device in the real world space based on the movement information of the imaging device, The image processing apparatus according to any one of claims 1 to 11, characterized in that the CG generation means generates the computer graphics image based on the position and orientation of the imaging device.

13. The image processing apparatus according to claim 12, characterized in that the sensing unit is an IMU (Internal Measurement Unit).

14. The acquisition process involves the imaging device capturing images of the real world, and A motion acquisition step for acquiring motion information of the imaging device within the space of the real world, A CG generation step that generates a computer graphics image including composite information related to the synthesis with the captured image based on the motion information of the imaging device, A CG correction step is performed which involves separating the computer graphics image containing the composite information into image channels and composite information channels, correcting the image channels by pixel interpolation according to the motion information of the imaging device, and correcting the composite information channels by pixel replacement according to the motion information of the imaging device. A synthesis step that combines the captured image and the computer graphics image corrected by the CG correction step, An image processing method characterized by having the following features.

15. An acquisition step in which an imaging device acquires an image of the real world captured by the imaging device and depth information of the real world corresponding to the image of the captured image, A motion acquisition step for acquiring motion information of the imaging device within the space of the real world, A CG generation step that generates a computer graphics image based on the motion information of the imaging device, Image correction step: Separating the captured image into an image channel and a depth information channel, correcting the image channel by pixel interpolation according to the motion information of the imaging device, and correcting the depth information channel by pixel replacement according to the motion information of the imaging device. A synthesis step that combines the image image corrected by the image image correction step and the computer graphics image generated by the CG generation step, An image processing method characterized by having the following features.

16. Computers, An acquisition means that acquires an image captured by an imaging device that captures the real world, An operation acquisition means for acquiring motion information of the imaging device within the space of the real world, A CG generation means generates a computer graphics image that includes composite information related to the synthesis with the captured image, based on the motion information of the imaging device. A CG correction means that separates a computer graphics image containing the composite information into image channels and composite information channels, performs correction on the image channels by pixel interpolation according to the motion information of the imaging device, and performs correction on the composite information channels by pixel replacement according to the motion information of the imaging device, A synthesis means for combining the captured image and the computer graphics image corrected by the CG correction means, A program that makes an image processing device function as having an image processing device.

17. Computers, An acquisition means that acquires an image captured by an imaging device of the real world and depth information of the real world corresponding to the image captured, An operation acquisition means for acquiring motion information of the imaging device within the space of the real world, A CG generation means that generates a computer graphics image based on the motion information of the imaging device, Image correction means for separating the captured image into an image channel and a depth information channel, correcting the image channel by pixel interpolation according to the motion information of the imaging device, and correcting the depth information channel by pixel replacement according to the motion information of the imaging device, A synthesis means for combining the image image corrected by the image image correction means and the computer graphics image generated by the CG generation means, A program that makes an image processing device function as having an image processing device.