Image processing device, electronic device, control method for image processing device, program

By generating and encoding mask images with reduced bit depth to correct transparency in MR systems, the system effectively reduces data transmission and maintains image quality in mixed reality environments.

JP2026065354APending Publication Date: 2026-04-15CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
CANON KK
Filing Date
2024-10-03
Publication Date
2026-04-15

AI Technical Summary

Technical Problem

Existing MR systems struggle to properly convey the transparency of semi-transparent regions in virtual images due to increased data transmission and mosquito noise from high compression rates, leading to deteriorated image quality.

Method used

The system generates mask images with a shallower bit depth than the alpha image to distinguish transparency regions, encodes and transmits these images, and applies masking to correct transparency values, reducing data amount while maintaining image quality.

Benefits of technology

This approach allows for accurate transparency determination of semi-transparent regions, reducing data transmission and preventing image degradation, thus enhancing image quality in mixed reality systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026065354000001_ABST
    Figure 2026065354000001_ABST
Patent Text Reader

Abstract

This method allows one device to more accurately perceive the transparency of a virtual image containing semi-transparent regions, while minimizing the amount of data transmitted. [Solution] The image processing apparatus includes: acquisition means for acquiring a virtual image and an alpha image representing the transparency of each region of the virtual image; generation means for distinguishing a plurality of regions in the virtual image with different transparency and generating at least one mask image having a shallower bit depth than the alpha image based on the alpha image; encoding means for encoding the data amount of each of the images, the at least one mask image, the virtual image, and the alpha image, by compressing it; and transmission means for transmitting each of the encoded images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image processing apparatus, an electronic device, a control method for an image processing apparatus, and a program.

Background Art

[0002] In recent years, as a technology for fusing the real world and the virtual world in real time and seamlessly, a composite reality technology, so-called MR (Mixed Reality) technology, is known. For example, there is an MR system that uses a video see-through type HMD (Head Mounted Display). In the MR system, an imaging unit built in the HMD captures an object that substantially coincides with the object observed from the pupil position of the user. Then, a composite image in which a CG (Computer Graphics) image is superimposed on the image obtained by the imaging is presented to the user, whereby the user can experience the MR space.

[0003] In the MR system, when synthesizing the image of the object and the CG image, in addition to the CG image, an alpha image having information on an alpha value representing the transparency is used to perform an alpha blending process. As a result, in the composite image, a semi-transparent portion where the object appears to be transparent over the CG can be expressed. However, when performing the synthesis process with a device different from the device that draws the CG image, it is necessary to transmit the alpha image in addition to the CG image from the device that draws the CG image to the device that performs the synthesis process. Therefore, the data transmission amount increases.

[0004] However, when non-reversible encoding processing with a high compression rate is performed on the compression of the CG image and the alpha image, mosquito noise occurs in the vicinity of the edge in the image. For this reason, the image quality of the composite image deteriorates. Patent Document 1 describes a technique for reducing noise in the vicinity of the edge of the composite image by using a mask image indicating a transparent color region.

Prior Art Documents

Patent Documents

[0005] [Patent Document 1] Japanese Patent Publication No. 2009-005290 [Overview of the project] [Problems that the invention aims to solve]

[0006] However, Patent Document 1 does not take into account the existence of semi-transparent regions in the character image, which is a virtual image. Therefore, in Patent Document 1, the device cannot properly convey the transparency of semi-transparent regions and other elements necessary for generating a composite image to other devices.

[0007] Therefore, the present invention aims to provide a technology that allows one device to more appropriately determine the transparency of a virtual image having a semi-transparent region, while suppressing the amount of data transmitted. [Means for solving the problem]

[0008] One aspect of the present invention is, Acquisition means for acquiring a virtual image and an alpha image representing the transparency of each region of the virtual image, A generation means for distinguishing multiple regions with different transparency in the virtual image and generating at least one mask image having a shallower bit depth than the alpha image, based on the alpha image, The images of the at least one mask image, the virtual image, and the alpha image A coding means that encodes data by compressing the amount of data, A transmission means for transmitting each of the encoded images, This is an image processing apparatus characterized by having [a certain feature]. [Effects of the Invention]

[0009] According to the present invention, the amount of data transmitted can be reduced while allowing one device to more appropriately determine the transparency of a virtual image having a semi-transparent region to be perceived by another device. [Brief explanation of the drawing]

[0010] [Figure 1] This is a diagram showing the configuration of the MR system according to Embodiment 1. [Figure 2] This is a diagram illustrating the generation of a composite image according to Embodiment 1. [Figure 3] This is a configuration diagram of the HMD and image processing device according to Embodiment 1. [Figure 4] This is a diagram illustrating mosquito noise according to Embodiment 1. [Figure 5] This figure illustrates the generation of a transparent mask image according to Embodiment 1. [Figure 6] This figure illustrates the generation of a background mask image according to Embodiment 1. [Figure 7] This is a diagram illustrating the masking process according to Embodiment 1. [Figure 8] This figure illustrates the generation of a semi-transparent outer mask image according to Modification 1. [Figure 9] This diagram illustrates the masking process related to Modification Example 1. [Figure 10] This figure illustrates the generation of a comprehensive mask image according to Embodiment 2. [Figure 11] This is a diagram illustrating the masking process according to Embodiment 2. [Modes for carrying out the invention]

[0011] The embodiments will be described in detail below with reference to the attached drawings. Note that the following embodiments do not limit the invention as defined in the claims. While the embodiments describe multiple features, not all of these features are essential to the invention, and the features may be combined in any way. Furthermore, in the attached drawings, identical or similar configurations are given the same reference numerals, and redundant descriptions are omitted.

[0012] <Embodiment 1> Using FIG. 1, a configuration example of the MR system (image processing system; information processing system) according to Embodiment 1 will be described. The wireless MR system according to Embodiment 1 includes an HMD 101, a control device 102, and a computer device 103.

[0013] The HMD 101 is an example of a display device (head-mounted display device) that can be worn on the user's head. The HMD 101 acquires a captured image of the real space. Further, the HMD 101 generates a composite image representing a composite reality space by synthesizing a CG image representing a virtual space and the captured image. The HMD 101 displays the composite image. Thereby, a composite image representing a composite reality space is presented in front of the eyes of the user wearing the HMD 101 on the head.

[0014] In FIG. 1, an interface separate from the head-mounted portion of the HMD 101 is depicted, but the interface may be integrated with the head-mounted portion. In the wireless MR system, a wireless communication method for forming a small-scale network, such as WLAN (Wireless Local Area Network) or WPAN (Wireless Personal Area Network), is used.

[0015] Thus, the HMD 101 and the control device 102 are connected wirelessly. However, the HMD 101 and the control device 102 may be connected by wire, or may be connected by a combination of wireless and wire. That is, the connection form between the HMD 101 and the control device 102 is not limited to a specific connection form.

[0016] The HMD 101 may operate on power supplied from a battery built into the HMD 101, or may operate on power supplied from the outside via a power cable. That is, the method of supplying power to the HMD 101 is not limited to a specific method.

[0017] The control device 102 transmits the captured image and / or position and orientation information transmitted from the HMD 101 to the computer device 103. The control device 102 also transmits the CG image transmitted from the computer device 103 to the HMD 101.

[0018] The computer device 103 determines the position and orientation of the HMD 101 (the position and orientation of the imaging unit of the HMD 101) based on the captured image and / or position and orientation information received from the control device 102. The computer device 103 generates a computer graphics (CG) image representing the virtual space as seen from the viewpoint with the determined position and orientation. Then, the computer device 103 transmits the CG image to the control device 102.

[0019] (Regarding the synthesis process) Using Figure 2, we will explain the process of generating a composite image based on the captured image, the CG image, and the alpha image (an alpha image containing information on alpha values ​​representing the transparency of each region of the CG image).

[0020] The captured image 201 is represented in RGB, YUV, or YCbCr color format. Each pixel of the captured image 201 contains multiple bits of information. The captured image 201 contains markers 202 that are artificially placed in real space. In Figure 2, for simplicity of explanation, the number of markers in the captured image 201 is set to 1, but in reality, the captured image 201 may contain multiple markers. Furthermore, if SLAM (Simultaneous Localization and Mapping) using natural feature points (feature points not artificially placed) in the image is used for position and orientation estimation, the captured image 201 does not need to contain markers.

[0021] The computer device 103 extracts "markers 202 or natural feature points" contained in the captured image 201. Then, the computer device 103 determines the position and orientation of the HMD 101 based on "markers 202 or natural feature points" and / or "position and orientation information received from the control device 102". Subsequently, the computer device 103 generates a CG image 203 representing the virtual space as seen from a viewpoint with the determined position and orientation.

[0022] CG image 203 is represented in RGB, YUV, or YCbCr color format. Each pixel in CG image 203 has multiple bits of information. CG image 203 contains object 204. In CG image 203, areas other than object 204 are filled with black, which has a specific pixel value (e.g., 0).

[0023] The computer device 103 generates an alpha image 205 containing alpha value information representing the transparency of each region of the CG image 203, along with the CG image 203. The alpha image 205 is represented in grayscale. Each pixel in the alpha image 205 has multiple bits of information. The background region 206 represents the region of the alpha image 205 where the transparency of the CG image 203 is 100% (the region where the CG image 203 is not displayed) when compositing is performed. The background region 206 is filled with black pixels having a specific pixel value (for example, 0). On the other hand, the CG region 207 represents the region where the transparency of the CG image 203 is 0% (the region where only the CG image 203 is displayed). The CG region 207 is filled with white pixels having a specific pixel value (for example, 255). The semi-transparent region 208 represents the region where the transparency of the CG image 203 is neither 0% nor 100% (the region where the transparency corresponds to the alpha value). The pixel values ​​of the semitransparent region 208 are, for example, other than "black and white". The pixel values ​​(intermediate gradations from 1 to 254) are as follows. The background region 206 may represent an area where the transparency of the CG image 203 is between 95% and 100%, and the CG region 207 may represent an area where the transparency is between 0% and 5%. The semi-transparent region 208 may represent an area where the transparency of the CG image 203 is between 5% and 95%.

[0024] Then, based on the alpha image 205, the HMD 101 generates a composite image 209 of the mixed reality space by combining the captured image 201 and the CG image 203. The HMD 101 displays the generated composite image 209. In this way, by performing the synthesis process using the alpha image 205, the area corresponding to the semi-transparent region 208 in the composite image 209 (the car window in Figure 2) is semi-transparent, and the area through which the captured image 201 is transmitted is displayed.

[0025] In Figure 1, the computer device 103 and the control device 102 are shown as separate devices, but the computer device 103 and the control device 102 may be integrated. Below, we will describe a configuration in which the computer device 103 and the control device 102 are integrated. Hereafter, the "device in which the computer device 103 and the control device 102 are integrated" will be referred to as the "image processing device 104".

[0026] (Regarding the functional configuration of HMDs and image processing units) Using the block diagram in Figure 3, examples of the functional configurations of the HMD 101 and the image processing device 104 will be explained. The HMD 101 includes an imaging unit 301, an attitude sensor 302, a display unit 303, an image processing unit 304, an encoding unit 305, an interface 306, a decoding unit 307, a mask processing unit 308, a synthesis unit 309, and a display image processing unit 310. In addition, some of the components of the HMD 101 may be included in the control device (electronic device) that controls the HMD 101.

[0027] The imaging unit 301 acquires captured images by capturing images of the real world. The captured image of one frame is used as an image to be combined with a virtual image (hereinafter sometimes referred to as a "background image") and also as an image used to generate position and orientation information (hereinafter sometimes referred to as a "position image"). The imaging unit 301 has an image sensor for the left eye and an image sensor for the right eye. The image sensor for the left eye captures a moving image of the real world corresponding to the left eye of the HMD 101 user. The image sensor for the left eye outputs an image of each frame in the captured moving image (captured image). The image sensor for the right eye captures a moving image of the real world corresponding to the right eye of the HMD 101 user. The image sensor for the right eye outputs an image of each frame in the captured moving image (captured image). In other words, the imaging unit 301 acquires a stereo image with parallax (parallax that approximately coincides with the positions of the left and right eyes of the HMD 101 user) as the captured image. Furthermore, in an HMD for an MR system, it is preferable that the central optical axis of the imaging range of the imaging unit 301 is positioned so as to substantially coincide with the user's line of sight.

[0028] Each of the image sensors, one for the left eye and one for the right eye, has an optical system and an imaging device. Light incident from the outside world passes through the optical system to the imaging device. The imaging device outputs an image corresponding to the incident light as an image.

[0029] The image sensor used in the imaging unit 301 of the HMD101 can be either a rolling shutter type or a global shutter type, taking into consideration various factors such as pixel count, image quality, noise, sensor size, power consumption, and cost. Furthermore, the imaging unit 301 can use a combination of rolling shutter type and global shutter type image sensors depending on the application.

[0030] For example, in the imaging unit 301, a rolling shutter type image sensor capable of acquiring higher quality images is used as the image sensor for acquiring the background image. The system uses a global shutter image sensor that eliminates image blur. Image blur is a phenomenon that occurs due to the operating principle of the rolling shutter system, in which exposure processing is started sequentially for each line in the scanning direction. Specifically, image blur is known as a phenomenon in which the subject is deformed and recorded as blurred when the image sensor or subject moves during the exposure time due to a time lag in the exposure timing of each line. In the global shutter system, exposure processing is performed simultaneously for all lines, so there is no time lag in the exposure timing of each line, and image blur does not occur.

[0031] The attitude sensor 302 measures various data necessary to determine the position and orientation of the HMD 101. Based on the measured data, the attitude sensor 302 detects the orientation of the HMD 101 and outputs orientation information. The attitude sensor 302 is implemented using magnetic sensors (including geomagnetic sensors), ultrasonic sensors, acceleration sensors, angular velocity sensors, etc.

[0032] The display unit 303 has a display unit for the right eye and a display unit for the left eye. The image for the left eye, representing the mixed reality space, is displayed on the left eye display unit. The image for the right eye, representing the mixed reality space, is displayed on the right eye display unit. Both the left eye display unit and the right eye display unit have a display optical system and a display element. The display optical system may be an eccentric optical system such as a free-form prism, or a normal coaxial optical system or an optical system with a zoom mechanism. For the display element, for example, a small liquid crystal display, an organic EL display, or a retinal scanning type device using MEMS (Micro Electro Mechanical Systems) can be used. Light from the image displayed on the display element enters the user's eye of the HMD 101 via the display optical system.

[0033] The image processing unit 304 performs various image processing operations on the image captured by the imaging unit 301. Here, the image processing performed on the background image and the image processing performed on the position image may be different operations.

[0034] The coding unit 305 performs encoding processing to compress the amount of data in the captured image after image processing has been performed by the image processing unit 304.

[0035] Interface 306 is a transmitting unit that transmits the encoded captured image and attitude information output from the attitude sensor 302 to the image processing unit 104. Interface 306 is also a receiving unit that receives the encoded image (CG image, alpha image, non-transparent mask image, and background mask image) from the image processing unit 104. Interface 306 transmits and receives control signals including setting information for each device. The non-transparent mask image is an image that distinguishes between areas that have no transparency at all (areas with 0% transparency) and areas that have transparency (areas with transparency that are not 0%). The background mask image is an image that distinguishes between areas of objects in the CG image (areas with transparency that are not 100%) and areas other than objects (areas with transparency that are 100%). Details of the non-transparent mask image and the background mask image will be described later using Figures 5 and 6.

[0036] The decoding unit 307 performs decoding (decoding) each encoded image (CG image, alpha image, transparency mask image, background mask image). The decoding unit 307 performs decoding processing corresponding to the encoding scheme used in the encoding performed by the coding unit 317 of the image processing device 104. The coding unit 317 may perform a common encoding process for each image, or it may perform different encoding processes for each image depending on its intended use.

[0037] The mask processing unit 308 uses the decoded (decoded) transparency mask image and background mask image to perform masking (correction) on the decoded alpha image. Details of the masking process will be described later with reference to Figure 7.

[0038] The synthesis unit 309 generates a composite image by combining the captured image and the decoded CG image based on the alpha image after mask processing.

[0039] The display image processing unit 310 performs various image processing operations on the composite image generated by the synthesis unit 309. The image processing performed here may include, for example, processing to correct the influence on the composite image caused by individual variations in the display devices or display optical systems that constitute the display unit 303. Specifically, the image processing may include offset or gain adjustment processing, pixel defect correction, or distortion correction processing of the display optical system.

[0040] The image processing device 104 will now be described. The image processing device 104 includes an interface 311, a decoding unit 312, a position and orientation detection unit 313, a content DB (database) 314, a drawing unit 315, a mask generation unit 316, and a coding unit 317.

[0041] Interface 311 is a receiving unit that receives encoded captured images and attitude information output from the attitude sensor 302 from the HMD 101. Interface 311 is also a transmitting unit that transmits each encoded image (CG image, alpha image, transparent mask image, background mask image) to the HMD 101. Interface 311 also transmits and receives control signals, including setting information for each device.

[0042] Interface 311 may transmit the encoded CG image, alpha image, and mask image (transparency mask image and background mask image) by wireless communication using different frequency bands. For example, interface 311 may transmit the alpha image using the 5GHz band and the two mask images using the 2.4GHz band. By using multiple frequency bands in this way, image transmission can be made faster, and the transmission of images with large amounts of data can also be made possible.

[0043] The decoding unit 312 performs the process of decoding the encoded captured image.

[0044] The position and orientation detection unit 313 determines the position and orientation of the left eye imaging unit and the right eye imaging unit based on the decoded captured image and the orientation information received from the HMD 101 via the interface 311. Since the process for determining the position and orientation of the imaging unit based on the captured image and the orientation information measured by the sensor is well known, a description of this technology will be omitted.

[0045] Content DB314 stores various types of data (virtual space data) necessary for rendering images in the virtual space. This virtual space data includes, for example, data that defines each object that makes up the virtual space (for example, data that defines the geometric shape, color, texture, position, and orientation of the object). It also includes, for example, data that defines the light sources placed in the virtual space (for example, data that defines the type, position, and orientation of the light sources).

[0046] The drawing unit 315 constructs (acquires) a virtual image using the virtual space data stored in the content DB 314. The drawing unit 315 draws (acquires) an image (CG image and alpha image) of the virtual space as seen from a viewpoint having the position and orientation of the left eye imaging unit determined by the position and orientation detection unit 313. The drawing unit 315 also draws an image (CG image and alpha image) of the virtual space as seen from a viewpoint having the position and orientation of the right eye imaging unit determined by the position and orientation detection unit 313.

[0047] The mask generation unit 316 generates a mask based on the alpha image drawn (acquired) by the drawing unit 315. It generates a transparent outer mask image and a background mask image.

[0048] The coding unit 317 performs encoding processing to compress the data size of the CG image, alpha image, transparent outer mask image, and background mask image.

[0049] Here, the number of pixels in the CG image, alpha image, transparent outer mask image, and background mask image may be the same or may differ depending on the application. For example, the number of pixels in the alpha image may be less than the number of pixels in each mask image. The encoding process performed on the CG image, alpha image, transparent outer mask image, and background mask image may be the same or may be different.

[0050] For example, since CG images contain objects that the user observes, high resolution and image quality are required. Therefore, a high number of pixels in a CG image is desirable. For encoding CG images with these characteristics, it is advisable to use, for example, H.265 / HEVC (High Efficiency Video Coding), which, although having a higher processing load among lossy compression methods, offers good compression efficiency and high image quality. On the other hand, an alpha image is information used to determine the blending ratio between the background image and the CG image in a composite image, and is not an image that the user observes. For this reason, the number of pixels in the alpha image can be less than that of the CG image. For encoding alpha images with these characteristics, it is also acceptable to use JPEG (Joint Photographic Experts Group), which, although not highly efficient among lossy compression methods, can keep the processing load low.

[0051] The transparency mask image and background mask image are not images that the user observes, just like the alpha image. However, these images are used for masking to remove "mosquito noise generated by lossy compression near the edge between the transparent and non-transparent regions of the CG image within the alpha image." In other words, the transparency mask image and background mask image affect the perceived resolution and noise level near the edge between the captured image and the transparent region of the CG image in the composite image. For this reason, the number of pixels in the transparency mask image and background mask image should be equivalent to the number of pixels in the CG image. These images can be represented with 1 bit per pixel (i.e., a shallower bit depth than the alpha image), and high compression efficiency can be expected even when using lossless compression. For this reason, "PNG (Portable Network Graphics), a lossless compression method," can be used to encode these images.

[0052] Furthermore, the code unit 317 may reduce the number of pixels (resolution) of the CG image or alpha image if the communication conditions (radio wave conditions in the case of wireless communication), such as the communication speed, are worse than certain conditions. This reduces the total amount of data transmitted, making it easier to transmit each image even when communication conditions are poor. In addition, if the code unit 317 reduces the number of pixels of the CG image, it may also reduce the number of pixels of the alpha image. This prevents the alpha image from becoming an unnecessarily large amount of data, as the alpha image indicates the transparency of the CG image.

[0053] In addition to the components mentioned above, the MR system may also include, for example, a control unit for configuring each device, or an acquisition unit for obtaining configuration information from an external source, in either the HMD 101 or the image processing device 104.

[0054] Refer to Figures 4A to 4E to explain the mosquito noise near the edges of the alpha image due to the effects of lossy compression.

[0055] Figure 4A shows the alpha image 205 before irreversible compression. Figure 4B shows the alpha image 401 after decoding following irreversible compression. Figure 4C is a magnified view of region 402 of the decoded alpha image 401. In Figure 4C, edge 403 indicates the boundary between the semi-transparent region and the CG region, and edge 404 indicates the boundary between the background region and the CG region.

[0056] Figure 4D shows the region 402 with the background area (i.e., pixel value 0), which represents complete black, replaced with gray. In Figure 4D, the black pixels that remain and have not been replaced with gray are pixels whose original pixel value was 0, but which have changed to a different pixel value due to the effects of lossy compression.

[0057] Figure 4E shows the CG region (i.e., pixel value 255) within region 402 replaced with gray. In Figure 4D, the white pixels that remain without being replaced with gray are pixels whose original pixel value was 255, but which have changed to a different pixel value due to the effects of lossy compression.

[0058] Pixels in the background area shown in Figure 4D with a value other than 0, and pixels in the CG area shown in Figure 4E with a value other than 255, will be made semi-transparent when the background image and CG image are combined. This results in a decrease in the image quality of the combined image.

[0059] (Generating a transparent mask image) Referring to Figure 5, we will explain how to generate an out-of-transparency mask image from an alpha image.

[0060] In the alpha image 205, the pixel value of the background region 206 is 0, and the pixel value of the CG region 207 is 255. The pixel value of the semitransparent region 208 is one of the values ​​from 1 to 254. The mask generation unit 316 of the image processing device 104 converts pixels in the alpha image 205 with a pixel value of 254 or less to pixels with a pixel value of 0. The mask generation unit 316 also converts pixels in the alpha image 205 with a pixel value of 255 to pixels with a pixel value of 1. As a result, the mask generation unit 316 generates a non-transparent mask image 501. In the non-transparent mask image 501, the region with a pixel value of 0 is the transparent region 502, and the region with a pixel value of 1 is the non-transparent region 503.

[0061] (Generating a background mask image) Referring to Figure 6, we will explain how to generate a background mask image from an alpha image.

[0062] The alpha image 205 shown in Figure 6 is identical to the alpha image 205 shown in Figure 5. The mask generation unit 316 of the image processing device 104 maintains pixels with a pixel value of 0 in the alpha image 205 as pixels with a pixel value of 0. The mask generation unit 316 also converts pixels with a pixel value of 1 or more in the alpha image 205 to pixels with a pixel value of 1. As a result, the mask generation unit 316 generates a background mask image 601. In the background mask image 601, the area with a pixel value of 0 is the area outside the object region 602, and the area with a pixel value of 1 is the object region 603.

[0063] (Masking the alpha image) Referring to Figure 7, a method for applying a mask to the decoded alpha image (a method for correcting the alpha image) using the transparent outer mask image 501 and the background mask image 601 will be explained.

[0064] First, the mask processing unit 308 sequentially references the pixel values ​​of each pixel in the transparent outer mask image 501 and the background mask image 601, and determines the pixel values ​​of each referenced mask image. Then, according to Table 701, the mask processing unit 308 determines the corresponding pixel positions of each referenced mask image. The pixel values ​​of pixels in alpha image 401 are changed (corrected).

[0065] For example, if a pixel in the alpha image 401 has a pixel value of 1 in both the transparency mask image 501 and the background mask image 601, the mask processing unit 308 replaces the pixel value of that pixel in the alpha image 401 with 255. If a pixel has a pixel value of 0 in both the transparency mask image 501 and the background mask image 601 with 1, the mask processing unit 308 maintains the pixel value of that pixel in the alpha image 401. Furthermore, if a pixel has a pixel value of 0 in both the transparency mask image 501 and the background mask image 601 with 0, the mask processing unit 308 replaces the pixel value of that pixel in the alpha image 401 with 0. The mask processing unit 308 completes the masking (correction) of the decoded alpha image by performing this process for all pixels.

[0066] In this way, the decoded alpha image is masked using a non-transparent mask image and a background mask image. For example, if the pixel value of a pixel that was 255 before encoding changes to a different pixel value due to the effects of lossy compression, the pixel value of that pixel can be restored to its original value of 255. Similarly, if the pixel value of a pixel that was 0 before encoding changes to a different pixel value due to the effects of lossy compression, the pixel value of that pixel can be restored to its original value of 0. Therefore, even if the virtual image has semi-transparent regions, the image processing device 104 can transmit the mask image and alpha image to the HMD 101 to more appropriately determine the transparency of the virtual image.

[0067] As described above, according to Embodiment 1, the HMD 101 uses the alpha image after masking to synthesize the captured image and the decoded CG image. This prevents the background image and CG image from becoming semi-transparent in areas that should not be semi-transparent due to the effects of lossy compression. Therefore, it is possible to improve the image quality of the synthesized image while suppressing the increase in the total amount of data when transmitting the alpha image to the HMD 101.

[0068] In Embodiment 1, the alpha image after masking was used to combine the background image and the CG image, but it may be used for other purposes. For example, the alpha image after masking may be used to correct or transform the CG image. For instance, pixels with a pixel value of 1 or more in the alpha image after masking are regions in the CG image where virtual objects are depicted. By cutting out these regions from the CG image, an image in which only virtual objects are depicted can be appropriately extracted.

[0069] <Example 1> In Embodiment 1, masking was performed using a transparent outer mask image 501 and a background mask image 601 as mask images. However, alternatively, a semi-transparent outer mask image that masks the area excluding the area corresponding to the semi-transparent region 208 in the alpha image may also be used.

[0070] Using Figure 8, we will explain how to generate a semitransparent outer mask image 801 from an alpha image.

[0071] The alpha image 205 shown in Figure 8 is identical to the alpha image 205 shown in Figure 5. The mask generation unit 316 of the image processing device 104 converts the pixel values ​​of pixels with pixel values ​​between 1 and 254 in the alpha image 205 to 1. The mask generation unit 316 also converts the pixel values ​​of pixels with pixel values ​​of 255 or 0 in the alpha image 205 to 0. As a result, the mask generation unit 316 generates a semitransparent outer mask image 801. In the semitransparent outer mask image 801, the area with a pixel value of 0 is the semitransparent outer region 802, and the area with a pixel value of 1 is the semitransparent region 803.

[0072] Figure 9 shows a table illustrating the combinations of mask images used for decoding and the method for masking the alpha image after decoding.

[0073] As shown in Figure 9, by combining two of the following mask images—the transparent outer mask image 501, the background mask image 601, and the semi-transparent outer mask image 801—masking similar to that in Embodiment 1 is possible.

[0074] Referring to Table 902, the method for masking the decoded alpha image using the transparent out-of-mask image 501 and the semi-transparent out-of-mask image 801 will be explained. First, the mask processing unit 308 sequentially references the pixel values ​​of each pixel in the transparent out-of-mask image 501 and the semi-transparent out-of-mask image 801, and determines the pixel values ​​of each referenced mask image. Then, the mask processing unit 308 replaces the pixel values ​​of the pixels in the decoded alpha image 401 corresponding to the pixel positions of each referenced mask image with the values ​​listed in Table 902. The mask processing unit 308 completes the masking of the decoded alpha image by processing all pixels.

[0075] Referring to Table 903, the method for masking the decoded alpha image using the background mask image 601 and the semi-transparent outer mask image 801 will be described. First, the mask processing unit 308 sequentially references the pixel values ​​of each pixel in the background mask image 601 and the semi-transparent outer mask image 801, and determines the pixel values ​​of each referenced mask image. Then, the mask processing unit 308 replaces the pixel values ​​of the pixels in the decoded alpha image 401 corresponding to the pixel positions of each referenced mask image with the values ​​listed in Table 903. The mask processing unit 308 completes the masking of the decoded alpha image by processing all pixels.

[0076] <Embodiment 2> Embodiment 1 described a method for performing masking on the decoded alpha image using two types of mask images. Embodiment 2 describes a method for performing masking using one mask image (hereinafter referred to as the "inclusive mask image").

[0077] Using Figure 10, we will explain how to generate a comprehensive mask image 1001 from an alpha image.

[0078] The alpha image 205 shown in Figure 10 is identical to the alpha image 205 shown in Figure 5. The mask generation unit 316 of the image processing device 104 converts the pixel values ​​of pixels with pixel values ​​between 1 and 254 in the alpha image 205 to 1. The mask generation unit 316 maintains the pixel values ​​of pixels with pixel values ​​of 0 in the alpha image 205 at 0. The mask generation unit 316 also converts the pixel values ​​of pixels with pixel values ​​of 255 in the alpha image 205 to 2. As a result, the mask generation unit 316 generates the comprehensive mask image 1001.

[0079] In the comprehensive mask image 1001, the areas with a pixel value of 0 are the background image area 1002. The areas with a pixel value of 1 are the semi-transparent image area 1003, and the areas with a pixel value of 2 are the CG image area 1004.

[0080] The mask processing unit 308 first sequentially references the pixel values ​​of each pixel in the comprehensive mask image 1001 and determines the pixel values ​​of each referenced mask image. Then, the mask processing unit 308 replaces the pixel values ​​of the pixels in the decoded alpha image 401 corresponding to the pixel positions of each referenced mask image with the values ​​listed in Table 1101 shown in Figure 11.

[0081] For example, if the pixel value of a certain pixel in the decoded alpha image 401 is 2, the mask processing unit 308 replaces the pixel value of that pixel in the alpha image 401 with 255. Also, the mask processing unit 308 applies a comprehensive mask to a certain pixel. If the pixel value of image 1001 is 1, the pixel value of that pixel in alpha image 401 is maintained as is. Also, if the pixel value of a certain pixel in the comprehensive mask image 1001 is 0, the mask processing unit 308 replaces the pixel value of that pixel in alpha image 401 with 0. The mask processing unit 308 completes the masking of the decoded alpha image by performing this process for all pixels.

[0082] Furthermore, the comprehensive mask image 1001 itself can be represented with 2 bits per pixel (a bit depth shallower than that of the alpha image), and high compression efficiency can be expected even when using lossless compression. For this reason, PNG, a lossless compression method, can be used to encode the comprehensive mask image 1001. The number of pixels and compression method for each image described above are given as examples to provide a specific explanation of Embodiment 2, and are not intended to limit the scope to these examples.

[0083] As described above, according to Embodiment 2, the HMD 101 uses the alpha image after masking to synthesize the captured image and the decoded CG image. This prevents the background image and CG image from becoming semi-transparent and transparent in areas that should not be semi-transparent due to the effects of lossy compression. Therefore, it is possible to improve the image quality of the synthesized image while suppressing the increase in the amount of data transmitted to the HMD 101 for the alpha image.

[0084] Furthermore, in the above, "If A is greater than or equal to B, proceed to step S1; if A is less than (lower than) B, proceed to step S2" may be rephrased as "If A is greater than (higher than) B, proceed to step S1; if A is less than or equal to B, proceed to step S2." Conversely, "If A is greater than (higher than) B, proceed to step S1; if A is less than or equal to B, proceed to step S2" may be rephrased as "If A is greater than or equal to B, proceed to step S1; if A is less than (lower than) B, proceed to step S2." Therefore, as long as no contradiction arises, "greater than or equal to A" may be rephrased as "greater than (higher; longer; more) than A," and "less than or equal to A" may be rephrased as "less than (lower; shorter; fewer) than A." And "greater than (higher; longer; more) than A" may be rephrased as "greater than or equal to A," and "less than (lower; shorter; fewer) than A" may be rephrased as "less than or equal to A."

[0085] The various controls described above may or may not be performed by a single piece of hardware (e.g., a processor or circuit). Multiple pieces of hardware (e.g., multiple processors, multiple circuits, or a combination of one or more processors and one or more circuits) may share the processing to control the entire device.

[0086] Furthermore, the above-mentioned processors are processors in a broad sense, including general-purpose processors and specialized processors. General-purpose processors include, for example, CPUs (Central Processing Units), MPUs (Micro Processing Units), and DSPs (Digital Signal Processors). Specialized processors include, for example, GPUs (Graphics Processing Units), ASICs (Application Specific Integrated Circuits), and PLDs (Programmable Logic Devices). Programmable logic devices include, for example, FPGAs (Field Programmable Gate Arrays) and CPLDs (Complex Programmable Logic Devices).

[0087] Furthermore, although embodiments of the present invention have been described in detail, the present invention is not limited to these specific embodiments, and various forms that do not depart from the spirit of the invention are also included in the present invention. Moreover, each of the embodiments described above is merely one embodiment of the present invention, and it is possible to combine each embodiment as appropriate.

[0088] <Other Embodiments> The present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit that implements one or more functions.

[0089] The above-disclosed embodiments include the following configurations, methods, and programs. (Composition 1) Acquisition means for acquiring a virtual image and an alpha image representing the transparency of each region of the virtual image, A generation means for distinguishing multiple regions with different transparency in the virtual image and generating at least one mask image having a shallower bit depth than the alpha image, based on the alpha image, A coding means that encodes the data amount of each of the images, the at least one mask image, the virtual image, and the alpha image, by compressing the data amount of each image. A transmission means for transmitting each of the encoded images, An image processing apparatus characterized by having (Configuration 2) The at least one mask image comprises a mask image that distinguishes a region having a first range of transparency from a region not having the first range of transparency in the virtual image, and a mask image that distinguishes a region having a second range of transparency different from the first range from a region not having the second range of transparency in the virtual image. The image processing apparatus according to configuration 1, characterized in that... (Composition 3) The at least one mask image is a mask image that distinguishes a region having a first range of transparency in the virtual image from a region having a second range of transparency different from the first range, and a region that does not have transparency in either the first or second range. The image processing apparatus according to configuration 1, characterized in that... (Composition 4) The first range and the second range are, respectively, a range where the transparency of the virtual image is 0%, a range where the transparency of the virtual image is 100%, and a range where the transparency of the virtual image is neither 0% nor 100%. The image processing apparatus according to configuration 2 or 3, characterized by the above. (Composition 5) The encoding means is capable of encoding the virtual image, the alpha image, and the mask image in different ways. An image processing apparatus according to any one of configurations 1 to 4, characterized by the above. (Composition 6) The aforementioned symbolic means is The virtual image and the alpha image are encoded using a lossy compression method. The aforementioned mask image is encoded using a lossless compression method. An image processing apparatus according to any one of configurations 1 to 5, characterized by the above. (Composition 7) The number of pixels in the alpha image is less than the number of pixels in the mask image. An image processing apparatus according to any one of configurations 1 to 6, characterized by the above. (Composition 8) The transmission means is capable of transmitting each of the encoded images by wireless communication using different frequency bands. An image processing apparatus according to any one of configurations 1 to 7, characterized by the above. (Composition 9) A receiving means for receiving each of the encoded images from an image processing device described in any of configurations 1 to 8, Decoding means for decoding each of the encoded images, Correction means for correcting the decoded alpha image based on the at least one decoded mask image, An electronic device characterized by having the following features. (Composition 10) The system includes a synthesis means for generating a composite image by combining the first image and the decoded virtual image based on the corrected alpha image. The electronic device according to configuration 9, characterized by the features described therein. (Composition 11) The electronic device is a display device that can be worn on the user's head and displays the composite image. The electronic device according to configuration 10, characterized by the above. (method) An acquisition step of acquiring a virtual image and an alpha image representing the transparency of each region of the virtual image, A generation step of distinguishing multiple regions with different transparency in the virtual image and generating at least one mask image having a shallower bit depth than the alpha image, based on the alpha image, A coding step of encoding by compressing the data amount of each of the images, the at least one mask image, the virtual image, and the alpha image, A transmission step of transmitting each of the encoded images, A control method for an image processing apparatus, characterized by having the following features. (program) A program for causing a computer to function as one of the means of an image processing apparatus described in any of configurations 1 to 8. [Explanation of symbols]

[0090] 104: Image processing device, 315: Drawing unit, 316: Mask generation unit, 317: Sign unit, 311: Interface

Claims

1. Acquisition means for acquiring a virtual image and an alpha image representing the transparency of each region of the virtual image, A generation means for distinguishing multiple regions with different transparency in the virtual image and generating at least one mask image having a shallower bit depth than the alpha image, based on the alpha image, A coding means that encodes the data amount of each of the images, the at least one mask image, the virtual image, and the alpha image, by compressing the data amount of each image. A transmission means for transmitting each of the encoded images, An image processing apparatus characterized by having

2. The at least one mask image includes a mask image that distinguishes a region having a first range of transparency from a region not having the first range of transparency in the virtual image, and a mask image that distinguishes a region having a second range of transparency different from the first range from a region not having the second range of transparency in the virtual image. The image processing apparatus according to feature 1.

3. The at least one mask image is a mask image that distinguishes a region having a first range of transparency in the virtual image from a region having a second range of transparency different from the first range, and a region that does not have transparency in either the first or second range. The image processing apparatus according to feature 1.

4. The first range and the second range are, respectively, a range where the transparency of the virtual image is 0%, a range where the transparency of the virtual image is 100%, and a range where the transparency of the virtual image is neither 0% nor 100%. The image processing apparatus according to claim 2.

5. The encoding means is capable of encoding the virtual image, the alpha image, and the mask image in different ways. The image processing apparatus according to feature 1.

6. The aforementioned symbolic means is The virtual image and the alpha image are encoded using a lossy compression method. The aforementioned mask image is encoded using a lossless compression method. The image processing apparatus according to feature 1.

7. The number of pixels in the alpha image is less than the number of pixels in the mask image. The image processing apparatus according to feature 1.

8. The transmission means is capable of transmitting each of the encoded images by wireless communication using different frequency bands. The image processing apparatus according to feature 1.

9. Receiving means for receiving each of the encoded images from the image processing apparatus according to any one of claims 1 to 8, Decoding means for decoding each of the encoded images, Correction means for correcting the decoded alpha image based on the at least one decoded mask image, An electronic device characterized by having the following features.

10. The system includes a synthesis means for generating a composite image by combining the first image and the decoded virtual image based on the corrected alpha image. The electronic device according to feature 9.

11. The electronic device is a display device that can be worn on the user's head and displays the composite image. The electronic device according to feature 10.

12. An acquisition step of acquiring a virtual image and an alpha image representing the transparency of each region of the virtual image, A generation step of distinguishing multiple regions with different transparency in the virtual image and generating at least one mask image having a shallower bit depth than the alpha image, based on the alpha image, A coding step of encoding by compressing the data amount of each of the images, the at least one mask image, the virtual image, and the alpha image, A transmission step of transmitting each of the encoded images, A control method for an image processing apparatus, characterized by having the following features.

13. A program for causing a computer to function as one of the means of an image processing apparatus according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Image processing device and program

    JP2009005290A