Image processing apparatus, image processing method, and storage medium

The image processing device enhances mixed reality systems by setting a reference depth to selectively display real objects of interest, improving user concentration and immersion by hiding distracting real-world elements.

JP2026004879APending Publication Date: 2026-01-15CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024102924
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-26
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

Existing mixed reality systems using see-through HMDs struggle to maintain user concentration and immersion due to real-world objects within the wearer's field of view, such as people or objects in hallways, which can disrupt the experience.

Method used

An image processing device sets a reference depth based on user input or object detection, generating mixed reality images where real objects closer to this depth are visible while objects further back are hidden, using a combination of real and virtual images.

Benefits of technology

Improves user concentration and sense of immersion by selectively displaying real objects of interest while hiding distracting real-world elements, enhancing the mixed reality experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026004879000001_ABST
    Figure 2026004879000001_ABST
Patent Text Reader

Abstract

To improve concentration and immersion of a wearer in a mixed reality image using a see-through type HMD.SOLUTION: The image processing apparatus includes a distance setting unit configured to set a reference depth indicating a distance at which a real object reflected in the real image is allowed to be visually recognized in the mixed reality image, and a generation unit configured to generate the mixed reality image in which the real image and the virtual reality image are combined according to the reference depth.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to image processing techniques for mixed reality. [Background technology]

[0002] In recent years, a technology called mixed reality (MR), which blends real and virtual spaces, has been gaining popularity. One method for achieving mixed reality is to use a see-through head-mounted display (HMD). In an MR system using a see-through HMD, the wearer observes a composite image in which a virtual reality image created using computer graphics (CG) is superimposed on a real image captured by a camera built into the HMD. The wearer can simultaneously view both real and virtual objects through the composite image (hereinafter referred to as a "mixed reality image"). One application of such an MR system using a see-through HMD is office work using a virtual display. In this case, the wearer can perform office work by viewing a virtual display outputting a PC screen while viewing a real keyboard and documents. This eliminates the need for a desktop display in the real space, offering the advantage of allowing office work on a large-screen virtual display regardless of location. Here, the relationship between real and virtual objects is important in the mixed reality image viewed by the wearer of the see-through HMD. In this regard, a technology has been proposed that solves the problem of the wearer coming into contact with a real object because the real object is occluded by the virtual object by displaying the virtual object within a predetermined distance from the wearer transparently (see Patent Document 1). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2016-4493 Summary of the Invention [Problem to be solved by the invention]

[0004] However, when experiencing mixed reality images using a see-through HMD, people or objects in the real world within the wearer's field of view can sometimes reduce the wearer's concentration and sense of immersion. For example, in the aforementioned office work example, suppose there is a hallway ahead of the wearer's line of sight (an extension of the virtual display). In such an environment, people passing through the hallway can disrupt the wearer's concentration and sense of immersion. To avoid this, real objects that the wearer wants to see, such as a keyboard used for operation, must be displayed visibly in the mixed reality image, while other real objects are preferably hidden from view. The technology of Patent Document 1, which allows virtual objects within a certain distance to be transparent, was unable to address this issue. [Means for solving the problem]

[0005] According to the technology disclosed herein, an image processing device generates a mixed reality image that combines a real image and a virtual reality image and is displayed on a see-through imaging display device worn by a person on the head, and is characterized by having a distance setting means that sets a reference depth that represents the distance that allows real objects reflected in the real image to be visible in the mixed reality image, and a generation means that generates the mixed reality image that combines the real image and the virtual reality image in accordance with the reference depth. [Effects of the Invention]

[0006] According to the present disclosure, it is possible to improve the wearer's concentration and sense of immersion in mixed reality images using a see-through HMD. [Brief explanation of the drawings]

[0007] [Figure 1] (a) is a diagram showing the appearance of the MR system and an example of how it is worn, and (b) is a diagram showing the hardware configuration of the image processing device. [Figure 2] FIG. 1 is a diagram showing an example of the hardware configuration of a video see-through HMD. [Figure 3] FIG. 2 is a functional block diagram showing the software configuration of the image processing device according to the first embodiment. [Figure 4] (a) and (b) are diagrams showing the coordinate system. [Figure 5] 5 is a flowchart showing the flow of processing for generating a mixed reality image in the image processing device according to the first embodiment. [Figure 6] 10(a) and 10(b) are explanatory diagrams of the reference depth. [Figure 7] 10(a) to 10(c) are diagrams illustrating the process of generating a mixed reality image. [Figure 8] 10 is a flowchart showing the flow of processing for generating a mixed reality image according to a first modification of the first embodiment. [Figure 9] 10(a) and 10(b) are explanatory diagrams of the reference depth. [Figure 10] 10(a) and 10(b) are explanatory diagrams of the reference depth. [Figure 11] FIG. 10 is a functional block diagram showing the software configuration of an image processing apparatus according to a second embodiment. [Figure 12] 10 is a flowchart showing the flow of processing for generating a mixed reality image in an image processing device according to a second embodiment. [Figure 13] 10 is a flowchart showing the flow of processing for generating a mixed reality image according to a second modification of the second embodiment. [Figure 14] FIG. 11 is a functional block diagram showing the software configuration of an image processing device according to a third embodiment. [Figure 15] 11 is a flowchart showing the flow of processing for generating a mixed reality image in an image processing device according to a third embodiment. [Figure 16] FIG. 10A is a diagram showing an example of an object detection result image, and FIG. 10B is a diagram showing an example of an object detection result table. [Figure 17] FIG. 11 is a diagram showing an example of a mixed reality image according to the third embodiment. [Figure 18] FIG. 10 is a functional block diagram showing the software configuration of an image processing device according to a fourth embodiment. [Figure 19]10 is a flowchart showing the flow of processing for generating a mixed reality image in an image processing device according to a fourth embodiment. [Figure 20] FIG. 10 is a diagram showing an example of a display flag table. DETAILED DESCRIPTION OF THE INVENTION

[0008] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Note that the following embodiments do not limit the present invention, and not all of the combinations of features described in the embodiments are necessarily essential to the solution of the present invention. Note that the same components will be described with the same reference numerals.

[0009] [Embodiment 1] In embodiment 1, a reference depth is set based on user input, and a mixed reality image is generated in which real objects closer to the reference depth are visible, but real objects further back than the reference depth are not visible.

[0010] <System configuration> 1(a) is a diagram showing the appearance of an MR system that reproduces mixed reality images and an example of how it is worn. The MR system 1 is composed of an HMD 10, which is a see-through imaging and display device worn on the head of a person, and an image processing device 20 that generates mixed reality images that combine real space and virtual space and provides the images to the HMD 10. The HMD 10 and the image processing device 20 are connected via a cable 30. Note that the connection between the HMD 10 and the image processing device 20 is not limited to a wired connection and may be a wireless connection.

[0011] <Hardware configuration of image processing device> FIG. 1(b) is a diagram showing an example of the hardware configuration of the image processing apparatus 20. In FIG. 1(b), the CPU 101 uses the RAM 102 as a work memory, executes programs stored in the ROM 103 and the hard disk drive (HDD) 105, and controls the operations of each block described later via the system bus 110. The HDD interface (hereinafter, the interface is denoted as "I / F") 104 connects secondary storage devices such as the HDD 105 and an optical disk drive. The HDD I / F 104 is an I / F such as Serial ATA (SATA), for example. The CPU 101 can read data from the HDD 105 and write data to the HDD 105 via the HDD I / F 104. Further, the CPU 101 can expand the data stored in the HDD 105 into the RAM 102, and conversely, can also save the data expanded in the RAM 102 to the HDD 105. And the CPU 101 can execute the data expanded in the RAM 102 as a program. The input I / F 106 connects an input device 107 such as a keyboard and a mouse. The input I / F 106 is a serial bus I / F such as USB or IEEE 1394, for example. The CPU 101 can read operation signals of the input device 107 via the input I / F 106. The output I / F 108 connects an output device 109 such as the HMD 10 and a liquid crystal display device. The output I / F 108 is a video output I / F such as DVI or HDMI (registered trademark), for example. The CPU 101 can send video data via the output I / F 108 and display a predetermined video on the HMD 10 or the liquid crystal display device.

[0012] <Hardware Configuration of HMD> FIG. 2 illustrates an example of the hardware configuration of a video see-through HMD 10. The HMD 10 has multiple RGB cameras 201 and an inertial measurement unit (IMU) (not shown) to achieve inside-out position tracking. The IMU detects three-dimensional inertial motion (translational and rotational motion in three orthogonal axes) and is composed of a gyro sensor for detecting rotational motion and an acceleration sensor for detecting translational motion. The HMD 10 also has a distance sensor 202, such as a LiDAR (Light Detection and Ranging) sensor, for acquiring depth information. The HMD 10 also has a left-eye display 203 and a right-eye display 205, each of which is configured using a liquid crystal panel or an organic electroluminescence (EL) panel, for displaying images for the left and right eyes. Furthermore, left-eye eyepieces 204 and right-eye eyepieces 206 are disposed in front of the displays 203 and 205, respectively. The wearer observes enlarged virtual images of the images displayed for the left and right eyes through the lenses 204 and 206. The HMD 10 generates a left-eye image and a right-eye image based on a mixed reality image provided by the image processing device 20, and displays the left-eye image on the left-eye display 203 and the right-eye image on the right-eye display 206. By providing an appropriate parallax between the left-eye image and the right-eye image, the wearer can perceive an image with a sense of depth. The HMD 10 also includes a dial 207 that accepts operation instructions from the wearer. Note that the HMD 10 has other components in addition to those described above, but these are not the focus of the present invention and will not be described here.

[0013] <Software configuration of image processing device> Fig. 3 is a functional block diagram showing the software configuration (logical configuration) of an image processing device 20 according to this embodiment. In Fig. 3, the image processing device 20 has an input receiving unit 11, a reference distance setting unit 12, a data acquisition unit 13, a VR image generation unit 14, and an MR image generation unit 15. The MR image generation unit 15 also has an inside / outside determination unit 16 and a synthesis unit 17. Each unit will be described below.

[0014] The input receiving unit 11 receives various input operations (user inputs) by the wearer of the HMD 10. Information on the received user inputs is output to the reference distance setting unit 12.

[0015] The reference distance setting unit 12 sets a reference distance (hereinafter referred to as "reference depth") that allows a real object to be visible in a mixed reality image. The set reference depth is output to the inside / outside determination unit 16.

[0016] The data acquisition unit 13 acquires real image data from the HMD 10. As shown in FIG. 4(a), the position of each pixel in the real image is specified by a uv coordinate system in which the horizontal and vertical directions are represented by u and v, respectively. The number of pixels in the u direction (width) is w, and the number of pixels in the v direction (height) is h. Each pixel in the real image has color value data for three channels: red, green, and blue (8-bit RGB values ​​for each channel). The data acquisition unit 13 also acquires depth information from the HMD 10, which represents the distance in real space from the wearer (≈HMD 10) to a real object. In this embodiment, the depth information acquired is data for an image with the same number of pixels as the real image (hereinafter referred to as a "real depth image"). Each pixel in the real depth image stores a read value from the distance sensor 202, which is the distance value (in meters) to the real object displayed at the same pixel position in the real image. FIG. 4(b) shows a coordinate system for the distance indicated by each pixel in the real depth image. In this embodiment, a so-called camera coordinate system is used, in which the wearer of the HMD 10 is centered, with the z-axis pointing forward, the y-axis pointing vertically, and the x-axis pointing horizontally. In other words, the distance represented by the real depth image is the distance (unit: m) from the origin, where the wearer of the HMD 10 is the origin. Instead of using the distance sensor 202, a real depth image may be generated and acquired by applying a known stereo matching technique to the acquired real image. Stereo matching is a technique for determining the three-dimensional position of a subject based on the principle of triangulation from the deviation of the subject depicted in image data acquired from two imaging devices located at different positions. Stereo matching determines the three-dimensional coordinates of the real object depicted in each pixel of the real image. The coordinate system used here is a camera coordinate system in which the origin of the three-dimensional coordinates is the center of the head of the HMD wearer. In stereo matching, when the three-dimensional coordinates of a point on a real object reflected at a pixel (u', v') are (x', y', z'), the Euclidean distance expressed by the following equation (1) is stored at the position of pixel (u', v') in the real depth image. Then, by doing this for all pixels, a depth image is obtained.

[0017] TIFF2026004879000002.tif6150 The real image data thus acquired is output to the synthesis unit 17, and the real depth image data, which is depth information, is output to the inside / outside determination unit 16.

[0018] The VR image generation unit 14 generates a virtual reality image that represents a virtual object, such as a virtual display. The VR image generation unit 14 also generates a depth image (hereinafter referred to as a "virtual reality depth image") in which distance values ​​are stored in pixels that correspond one-to-one to pixels of the virtual reality image as depth information that represents the distance to the virtual object represented by the virtual reality image. The width w and height h of the virtual reality depth image are the same as the width w and height h of the real image, just like the real depth image. The generated virtual reality image and data of the virtual reality depth image are output to the synthesis unit 17.

[0019] The MR image generator 15 is composed of an inside / outside determination unit 16 and a composition unit 17, and generates a mixed reality image that is the result of combining a real image and a virtual reality image. The width w and height h of the mixed reality image are the same as the width w and height h of the real image.

[0020] The inside / outside determination unit 16 determines whether a real object shown in the real image is located in front of the reference depth (whether the real object is located inside or outside the boundary indicated by the reference depth) based on the real depth image. The determination result is output to the synthesis unit 17.

[0021] The synthesis unit 17 synthesizes the real image and the virtual reality image to generate a mixed reality image so that real objects determined by the inside / outside determination unit 16 to be in front of the reference depth are visible, and real objects determined to be behind the reference depth are invisible. Data of the generated mixed reality image is transmitted to the HMD 10.

[0022] The software configuration (logical configuration) of the image processing device 20 has been described above.

[0023] <Operation flow of image processing device> Fig. 5 is a flowchart showing the flow of processing for generating a mixed reality image in the image processing device 20 according to this embodiment. The series of processes shown in the flowchart of Fig. 5 are executed on a frame-by-frame basis. A detailed explanation will be given below with reference to the flowchart of Fig. 5. In the following explanation, the symbol "S" means step.

[0024] In S501, the data acquisition unit 13 acquires from the HMD 10 data of a real image obtained by capturing an image of a real space with the RGB camera 201 and a real depth image obtained by measuring the same real space with the distance sensor 202.

[0025] In S502, the next process to be executed is determined depending on whether a reference depth D_ref for the mixed reality image to be generated has been set. If it has not yet been set, the process proceeds to S503, where the real image acquired in S501 is displayed on the two displays 203 and 205 in the HMD 10. On the other hand, if the reference depth D_ref has already been set, S505 is executed next. Note that even if the reference depth D_ref has already been set, if a new user input is detected as described in S504 below, the reference depth D_ref may be set again.

[0026] In S504, the reference distance setting unit 12 sets a reference depth D_ref, which is a distance at which a real object is allowed to be displayed in the mixed reality image to be generated, based on a user input. In this embodiment, the reference depth D_ref is set according to a user instruction received via the input receiving unit 11, specifically, according to the value (input value) of the rotation amount of the dial 207 when the wearer turns the dial 207 mounted on the HMD 10. The wearer operates the dial 207 while viewing real images displayed on the two displays 203 and 205 in the HMD 10. For example, the dial 207 can be set to any distance within a range of 0 m to 10 m (the minimum input value corresponds to a distance of 0 m, and the maximum rotation amount corresponds to a distance of 10 m). In this case, the distance according to the input value of the dial 207 by the wearer is set as the reference depth D_ref. Furthermore, the distances corresponding to the minimum and maximum rotation amounts of the dial 207 may be variable depending on the size of the real space. In this case, the distance M from the HMD 10 to the wall is first obtained using a wall detection technique using known machine learning or the like. Then, the distance M divided by the maximum rotation amount multiplied by the dial input value is calculated, converted into a distance equivalent to the dial input value, and the reference depth D_ref is set. FIGS. 6A and 6B are explanatory diagrams of the reference depth D_ref according to this embodiment, showing a range equidistant from the center of the wearer's head. FIG. 6A is a plan view, and FIG. 6B is a side view. In both figures, a dashed line 601 indicates the reference depth D_ref. Upon completion of the above processing, the MR image generator 15 executes the subsequent processing of S505 to S510. While the wearer is performing an operation to set the reference depth, the user may be able to confirm the range of the reference depth corresponding to the operation. Specifically, a line representing the reference depth (see dashed line 701 in FIG. 7A, described later) may be superimposed on the real image, or the real image may be color-coded to indicate areas in front of and behind the reference depth.

[0027] In S505, the VR image generation unit 14 generates depth information (virtual reality depth image) that represents the distance to a virtual reality image and a virtual object reflected in the virtual reality image. Specifically, information such as the shape, texture, position, and orientation of a virtual object prepared in advance is read from HDD 105 or the like, and a virtual reality image and a corresponding virtual reality depth image are generated using this information through a known rendering technique. Here, the position and orientation information of the virtual object is set in advance, for example, to position the virtual object 2 m in front of HMD 10 so as to face the wearer directly. Note that the virtual reality image and virtual reality depth image may be generated and stored in advance and then read from HDD 105.

[0028] In S506, a pixel position of interest in the real image acquired in S501 is determined. Here, the image coordinates of the pixel position of interest are represented as (ui, vi). Prior to determining the pixel position of interest, a buffer for temporarily storing data of the mixed reality image being generated is also allocated in, for example, RAM 102.

[0029] In S507, the inside / outside determination unit 16 determines whether the real object reflected at the pixel position of interest (ui, vi) is closer or farther than the reference depth D_ref set in S504. Specifically, the pixel value (distance value in m) stored at the pixel position of interest (ui, vi) in the real depth image acquired in S501 is compared with the set reference depth D_ref to determine whether the reference depth D_ref is greater. In the aforementioned FIGS. 6(a) and 6(b), the area inside the dashed line 601 as seen from the wearer of the HMD 10 is the range that is closer than the reference depth D_ref. If the result of the determination is that the pixel value stored at the pixel position of interest (ui, vi) in the real depth image is equal to or less than the reference depth D_ref, S508 is executed next. If the reference depth D_ref is greater, S509 is executed next.

[0030] In S508, the synthesis unit 17 determines the color value of the pixel position (ui, vi) of interest in the mixed reality image by performing a process (depth-aware synthesis process) that synthesizes the real image and the virtual reality image while taking into account the front-to-back relationship between the real object and the virtual object in the wearer's line of sight. In the depth-aware synthesis process, the object that is closer to the wearer of the HMD 10 is rendered between the real object displayed at the pixel position (ui, vi) of interest in the real image and the virtual object displayed at the pixel position (ui, vi) of interest in the virtual reality image. The specific process is as follows: First, the distance value of the pixel position (ui, vi) of interest in the real depth image is compared with the distance value of the pixel position (ui, vi) of interest in the virtual reality depth image. If the distance value of the pixel position (ui, vi) of interest in the real depth image is smaller, the color value of the pixel position (ui, vi) of the real image is stored as the color value of the pixel position (ui, vi) of the mixed reality image being generated in the buffer. On the other hand, if the distance value of the pixel position (ui, vi) of interest in the virtual reality depth image is smaller, the color value of the pixel position (ui, vi) of interest in the virtual reality image is stored as the color value of the pixel position (ui, vi) of interest in the mixed reality image being generated in the buffer. Note that if the distance values ​​of both are equal, it is possible to determine in advance whether to use the color value of the real image or the color value of the virtual reality image and process accordingly.

[0031] In S509, the composition unit 17 determines the color value of the pixel position (ui, vi) of interest in the mixed reality image by a process of combining the real image and the virtual reality image so that the real object is not visible and only the virtual object is visible. In this composition process, the virtual object reflected at the pixel position (ui, vi) of interest in the virtual reality image is always drawn. Specifically, the color value of the pixel position (ui, vi) of interest in the virtual reality image is stored as the color value of the pixel position (ui, vi) of interest in the mixed reality image being generated in the buffer.

[0032] In S510, it is determined whether all pixels of the real image acquired in S501 have been processed. If there are unprocessed pixels, the process returns to S506, where the next pixel of interest (ui, vi) is determined and similar processing continues. On the other hand, if all pixels have been processed, S511 is executed next. Figures 7(a) to 7(c) are diagrams illustrating a mixed reality image obtained by this embodiment. Figure 7(a) shows a real image, Figure 7(b) shows a virtual reality image, and Figure 7(c) shows a mixed reality image obtained by combining the real image of Figure 7(a) and the virtual reality image of Figure 7(b). The real image of Figure 7(a) shows a desk 703 with a keyboard 702 on it, and a passerby 704 is visible behind it. The dashed line 701 indicates, for illustrative purposes, a reference depth D_ref that is not actually visible in the real image. The virtual reality image of Figure 7(b) shows (draws) a virtual display 705 against a virtual background of natural scenery. In the flow of FIG. 5 described above, a synthesis process is performed in which, for areas closer to the reference depth D_ref, the closer of the real object and the virtual object is rendered based on depth information, and for areas further back than the reference depth D_ref, the virtual object is rendered regardless of depth information. In other words, among the real objects shown in the real image of FIG. 7(a), real objects that are further back than the reference depth D_ref indicated by the dashed line 701 are not rendered. As a result, in the image area of ​​the mixed reality image of FIG. 7(c) corresponding to a distance greater than the reference depth D_ref, the virtual display 705 and a virtual natural landscape are rendered, and the passerby 704 in the real space is not rendered. On the other hand, in the image area of ​​the mixed reality image of FIG. 7(c) corresponding to a distance closer than the reference depth D_ref, the real object or the virtual object with the smaller distance value is rendered. Therefore, the keyboard 702 and part of the desk 703, which are closer, and the base of the virtual display 705, which is closer than the desk 703, are rendered (the virtual background, which is farther away, is not rendered).

[0033] In S511, the data of the mixed reality image obtained by the processing up to this point is transmitted and output to the HMD 10. Then, in the HMD 10, an image for the left eye and an image for the right eye are generated based on the mixed reality image, and are displayed on the two displays 203 and 205, respectively.

[0034] In S512, it is determined whether to end the generation of the mixed reality image. For example, if it is detected that an end button (not shown) provided on the HMD 10 is pressed or that the wearer has removed the HMD 10, the generation of the mixed reality image ends. If it is determined that the generation has ended, this flow is exited. On the other hand, if the generation is to continue, the process returns to S501 and continues.

[0035] The above is a series of processing flows for generating a mixed reality image in the image processing device 20 according to this embodiment. Through these processing steps, the wearer of the HMD 10 can view a mixed reality video in real time.

[0036] <Variation 1> In the above-described embodiment, the user directly sets the reference depth. However, the user may also indirectly set the reference depth by identifying a physical object. This automatically resets the reference depth according to the distance to the physical object even if the position of the physical object changes, thereby reducing the effort required to reset the reference depth. Figure 8 is a flowchart showing the flow of processing for generating a mixed reality image according to this modification. This flowchart differs from the flowchart in Figure 5 in that step S801 is executed instead of step S504; the other common steps are numbered the same. The following describes step S801, which is the difference.

[0037] In S801, the reference distance setting unit 12 estimates the three-dimensional position of a specific physical object based on user input, and sets the distance to the obtained three-dimensional position as a reference depth D_ref. Specifically, first, the input receiving unit 11 receives user input specifying a physical object. In this case, the specifying method may be, for example, the wearer pointing at the specific physical object with their finger or a stick, or using a dedicated controller, or any other method. Next, a known machine learning technique is applied to the physical image acquired in S501 to obtain a three-dimensional direction vector in which the wearer is pointing with their finger or the like. Then, a known ray tracing technique is used to cast a ray along the estimated direction vector and estimate the three-dimensional position of the colliding physical object. Then, the distance from the wearer to the estimated three-dimensional position of the physical object is obtained, and the obtained distance is set as the reference depth. Each step from S505 onwards is executed using the reference depth D_ref thus set.

[0038] The position of a physical object required for ray tracing may be calculated, for example, using the stereo matching technique described above. Here, immediately after the user specifies a physical object, the distance determined from the three-dimensional position of the physical object estimated based on the specification can be used as the reference depth. On the other hand, the relative positional relationship between the wearer and the physical object may change if the wearer moves or if the wearer moves the physical object. In this case, the three-dimensional position of the physical object may be estimated using an object tracking technique such as a known particle filter. Even if a physical object has already been specified, step S801 may be re-executed if the wearer inputs a new physical object specification.

[0039] Alternatively, a method other than pointing to a physical object may be used, such as inputting the name of the target physical object using a keyboard or mouse, and the physical object may be detected using a known object detection technique. In this case, the distance to the physical object detected by object detection is used as the reference depth. For example, by performing object detection for each frame, the distance to the target physical object can be updated, allowing for movement of the wearer or the physical object to be accommodated. Furthermore, if the physical object has a wireless communication function such as Bluetooth, the distance to the physical object may be calculated using a known distance calculation method based on wireless communication (e.g., radio wave intensity) and used as depth information. In this case, the user may select a specific physical object from a list of Bluetooth-connected physical objects, and the distance to the selected physical object may be calculated. Note that even without a user selection, the device may automatically select the farthest physical object from among devices currently connected via Bluetooth, and use the distance to the selected physical object as the reference depth.

[0040] Alternatively, a plurality of physical objects may be designated, and the distance to the farthest physical object among the designated plurality of physical objects may be set as the reference depth.

[0041] <Variation 2> In the above-described embodiment, the distance indicated by the reference depth was the distance along the x-axis, y-axis, and z-axis. However, real objects that reduce the wearer's concentration and sense of immersion often exist in the horizontal direction, and objects in the vertical direction (y-axis direction) are often not very noticeable. For example, in the aforementioned office work example, passersby and other distracting objects are primarily a problem in the xz plane, while the only objects in the vertical direction are the floor and ceiling, and such real objects are not very noticeable. Therefore, the distance indicated by the reference depth may be the distance along the x-axis and z-axis (distance on the xz plane). In this case, the reference depth D_ref is indicated by the dashed line 901 in the plan view of FIG. 9(a) and the side view of FIG. 9(b). In this case, when the three-dimensional coordinates of a real object reflected at a certain pixel position (u', v') are (x', y', z'), the Euclidean distance expressed by the following equation (2) is calculated. Then, by repeating the process of storing the obtained Euclidean distance at the pixel position (u', v') for all pixels, a real depth image and a virtual reality depth image are obtained.

[0042] TIFF2026004879000003.tif6150 Furthermore, the distance indicated by the reference depth may be the distance on the z-axis (depth component only). In this case, the reference depth D_ref is indicated by the dashed line 1001 in the plan view of Figure 10(a) and the side view of Figure 10(b). In this case, when the three-dimensional coordinates of a real object reflected at a pixel position (u', v') are (x', y', z'), the Euclidean distance expressed by the following equation (3) is calculated. Then, the obtained Euclidean distance is stored at the pixel position (u', v') and this process is repeated for all pixels to obtain a real depth image and a virtual reality depth image.

[0043] TIFF2026004879000004.tif6150<Variation 3> In the above-described embodiment, the reference depth is set based on the wearer's dial operation on the HMD 10. However, instead of dial operation, the reference depth may be set based on the operation of other hardware, such as a button or touch panel, or based on the wearer's hand gestures. Furthermore, a numerical value (distance value) in meters may be directly input from a keyboard or the like. In this case, the input numerical value is stored in RAM 102 or the like and read and used so that the wearer does not need to input the numerical value for each frame. Note that an arbitrary numerical value, such as 1.0 (m), is set as the initial value so that a mixed reality image can be generated without waiting for input by the wearer. Furthermore, the processing of S503 may be performed in a separate thread using multithreading technology so that the mixed reality image can be displayed even while the wearer is inputting.

[0044] <Variation 4> In S507, the depth information is considered on a pixel-by-pixel basis and the composite is performed so that the real object and the virtual object that is closer to the viewer is displayed. However, for example, it is also possible to render all real objects that are closer than the reference depth. In this case, to prevent a virtual object, such as a virtual display, from being occluded by a real object and not being rendered, the virtual object needs to be placed further back than the reference depth.

[0045] As described above, according to this embodiment, a reference depth is set based on user input, and a mixed reality image is generated in which real objects closer to the reference depth are visible but real objects further back are not visible, thereby improving the sense of immersion and concentration of the HMD wearer.

[0046] [Embodiment 2] In recent years, systems that allow an HMD wearer to arbitrarily move the position of a virtual object in a mixed reality image have become popular. For example, in the aforementioned office work example, a situation can be considered in which the user adjusts the position of the virtual display so that it appears as if it were actually present on the desk. If the method of the aforementioned first embodiment is applied to such a system, the user must separately set the reference depth and the position of the virtual object, which is time-consuming. Therefore, a mode in which the position of a virtual object is set and the reference depth is automatically set based on the set virtual object position will be described as a second embodiment. Note that the hardware configurations of the HMD 10 and the image processing device 20 are similar to those of the first embodiment, and therefore a description thereof will be omitted. The following description will focus on the differences.

[0047] <Software configuration of image processing device> 11 is a functional block diagram showing the software configuration (logical configuration) of the image processing device 20 according to this embodiment. The major difference from the functional block diagram of FIG. 3 of the first embodiment is that a VR position setting unit 1101 is added. This VR position setting unit 1101 sets the position of a virtual object to be rendered in a virtual reality image and a mixed reality image based on a user input. Position information of the set virtual object is output to the VR image generation unit 14 and the reference distance setting unit 12.

[0048] <Operation flow of image processing device> Fig. 12 is a flowchart showing the flow of processing for generating a mixed reality image in the image processing device 20 according to this embodiment. The difference from the flowchart of Fig. 5 of the first embodiment is that steps S1201 to S1203 are executed instead of steps S503 and S504, and other common steps are numbered the same. The following describes the differences, steps S1201 to S1203.

[0049] In S1201, the VR image generation unit 14 reads data stored in the custom depth buffer and generates a provisional virtual reality image. Here, the custom depth buffer is an image buffer referenced during rendering. The custom depth buffer holds depth information indicating the distance to a virtual object (pixels in the virtual reality image where a virtual object is displayed store the distance value of the displayed portion, and other pixels store "NULL," meaning zero). The VR image generation unit 14 reads the initial values ​​of the custom depth buffer prepared in advance from the HDD 105 or the like, performs rendering, and generates a provisional virtual reality image in which a virtual object, such as a virtual display, is drawn at a default position. Data of the generated provisional virtual reality image is sent to the HMD 10 and displayed on the two displays 203 and 205 within the HMD 10.

[0050] TIFF2026004879000005.tif53170

[0051] TIFF2026004879000006.tif21150Note that if the wearer does not perform a hand gesture or the like and the input receiving unit 11 cannot detect user input, the change in the hand position will be zero and the current position of the virtual object (default position) will be maintained as is.

[0052] In S1203, the reference distance setting unit 12 sets a reference depth D_ref based on the position of the virtual object set in S1202. Specifically, the distance from the wearer to the virtual object is determined by referring to the data stored in the above-mentioned custom depth buffer, and the determined distance is set and derived as the reference depth D_ref. When determining the distance, it is sufficient to find the minimum value, average value, median value, mode value, representative value, etc., within the data stored in the custom depth buffer. For example, to find the shortest distance, it is sufficient to scan the values ​​stored in the custom depth buffer and find the smallest value. The distance to the virtual object thus obtained is set as the reference depth D_ref.

[0053] Using the reference depth D_ref thus set, each step from S505 onwards is executed.

[0054] In S1203 described above, the distance derived from the value stored in the custom depth buffer is directly set as the reference depth D_ref, but a buffer may also be provided. The derived distance may also be adjustable by the wearer. For example, the determined distance may be temporarily displayed on the two displays 203 and 205 of the HMD 10, and a value corresponding to an input operation by the wearer using the dial 207 or the like may be added to or subtracted from the determined distance to set the reference depth D_ref. In this way, the wearer can adjust the reference depth value automatically derived according to the position of the virtual object to suit their own intentions.

[0055] <Variation 1> In this embodiment, an example has been described in which a virtual object is set at an arbitrary position based on an instruction from the wearer, and the distance to the set position is set as the reference depth D_ref. Alternatively, for example, a line or mark representing the reference depth may be drawn as a virtual object (see the dashed line 701 in FIG. 7(a) described above). In this case, the distance to the line or mark may be derived based on a hand gesture or the like with respect to the line or mark, and the reference depth D_ref may be set.

[0056] <Variation 2> In the above-described embodiment, the reference depth is automatically set based on the position of a virtual object specified by the user. However, the position of the virtual object for automatically setting the reference depth may be automatically set based on the position of a real object that the user wants to view in the mixed reality image, and the reference depth may be automatically set based on the position of the virtual object.

[0057] Fig. 13 is a flowchart showing the flow of processing for generating a mixed reality image in the image processing device 20 according to this modification. The difference from the flowchart in Fig. 12 described above is that instead of steps S1201 to S1203, S503 is executed first, followed by steps S1301 and S1302; other common steps are numbered the same. Only the differences will be described below.

[0058] In S503, the real image acquired in S501 is displayed on the two displays 203 and 205 in the HMD 10.

[0059] In S1301, the reference distance setting unit 12 estimates the three-dimensional position of a specific physical object based on a user input, and sets the position of a virtual object based on the obtained three-dimensional position. Specifically, first, the input receiving unit 11 receives a user input specifying a physical object. Here, the method of specifying the physical object and the method of estimating its three-dimensional position are as described in S801 of Modification 1 of the first embodiment. Then, a predetermined position behind the estimated three-dimensional position of the physical object is set as the position of the virtual object. In this case, the predetermined position is determined in advance, for example, to be 0.1 m away from the three-dimensional position of the physical object in the z-axis direction.

[0060] In S1302, the reference distance setting unit 12 sets a reference depth according to the position of the virtual object set in S1301. Specifically, the distance from the wearer to the estimated three-dimensional position of the virtual object is calculated, and the calculated distance is set as the reference depth. Each step from S505 onwards is executed using the reference depth D_ref thus set. This completes the contents of this modified example.

[0061] As described above, according to this embodiment, the reference depth is automatically set according to the position of a virtual object set based on user input, thereby eliminating the need to separately set the position of a virtual object and the reference depth.

[0062] [Embodiment 3] In the mixed reality images generated by the methods of the first and second embodiments, real or virtual objects are displayed in front of the reference depth using depth-aware synthesis, and virtual objects are displayed behind the reference depth. Therefore, real objects that straddle the boundary of the reference depth are displayed cut off midway (for example, in the aforementioned FIG. 7(c) , the far corner of the desk 703 is cut off and replaced with a virtual landscape), resulting in an unnatural mixed reality image. Therefore, a mode in which real objects partially within the reference depth are displayed in their entirety will be described as a third embodiment. Since the hardware configurations of the HMD 10 and the image processing device 20 are similar to those of the first and second embodiments, their description will be omitted, and the following description will focus on the differences. Although the following description focuses on the differences when based on the second embodiment, it is also possible to apply the first embodiment (including its modified examples) as a base.

[0063] <Software configuration of image processing device> 14 is a functional block diagram showing the software configuration (logical configuration) of the image processing device 20 according to this embodiment. The major difference from the functional block diagram of FIG. 11 of the second embodiment is that an object detection unit 1401 is added. This object detection unit 1401 performs processing to detect a real object from an input real image. Position information of the detected real object is output to the inside / outside determination unit 16 of the MR image generation unit 15.

[0064] <Operation flow of image processing device> Fig. 15 is a flowchart showing the flow of processing for generating a mixed reality image in the image processing device 20 according to this embodiment. The difference from the flowchart of Fig. 12 of the second embodiment is that steps S1501 to S1506 are executed instead of steps S506 to S510, and other common steps are numbered the same. The following describes the differences, steps S1501 to S1506.

[0065] In S1501, the object detection unit 1401 applies a known object detection process to the real image acquired in S501 to detect real objects reflected in the real image, and generates an image representing the detection result (hereinafter referred to as the "object detection result image"). FIG. 16(a) is a diagram showing an example of the object detection result image. The size of the object detection result image is the same as that of the real image, having a width w and a height h. The object detection result image also holds an ID that uniquely represents each detected real object at each pixel position corresponding to each pixel of the real image. In the example of FIG. 16(a), two black pixel blocks represent detected real objects. The black pixel block with ID=1 represents a chair, the black pixel block with ID=2 represents a desk, and white pixels represent non-detection areas. Furthermore, the object detection unit 1401 generates a table as shown in FIG. 16(b) (hereinafter referred to as the "object detection result table") in which the type (class) and likelihood of each real object are associated with the ID assigned to each detected real object. It is desirable to exclude ceilings, walls, floors, ground, etc. from object detection. If such large real objects were included in the detection targets, even a portion of them would occupy most of the screen, and almost all objects in the real space would be displayed in the mixed reality image. Furthermore, after including the real objects in the detection targets, three-dimensional shape information of the detected real objects may be calculated, and real objects larger than a predetermined size may be excluded from the mixed reality image and object detection result table based on geometric information such as volume, area, and length. The generated object detection result image and object detection result table data are stored in RAM 102. For ease of explanation, the following steps S1502 to S1504 are described in units of real objects, but the actual processing is performed in units of pixels.

[0066] In S1502, a target real object is determined from all detected real objects. Then, in S1503, the inside / outside determination unit 16 determines whether the target real object is closer or farther than the reference depth D_ref based on the object detection result image and the real depth image. Specifically, a pixel region in which the ID of the target real object is stored in the object detection result image is first identified. Then, for the identified pixel region, the pixel values ​​in the real depth image are compared with the reference depth D_ref to determine whether the reference depth D_ref is greater. If the result of the determination is that the pixel values ​​in the real depth image are less than or equal to the reference depth D_ref, S1504 is executed next. If the reference depth D_ref is greater, S1505 is executed next.

[0067] In S1504, the synthesis unit 17 determines the color value of the pixel region of the target real object in the mixed reality image by a synthesis process that takes depth into consideration. Specifically, for a pixel region in the object detection result image in which the ID of the target real object is stored, the pixel value (distance value) of the real depth image is compared with the pixel value (distance value) of the virtual reality depth image. If the pixel value of the real depth image is smaller, the color value of the real image is stored as the color value of the corresponding pixel position of the mixed reality image being generated in the buffer. On the other hand, if the pixel value of the virtual reality depth image is smaller, the color value of the virtual reality image is stored as the color value of the corresponding pixel position of the mixed reality image being generated in the buffer. Note that if both distance values ​​are equal, it is possible to determine in advance whether to adopt the color value of the real image or the color value of the virtual reality image, and to process accordingly.

[0068] In S1505, the composition unit 17 determines the color value of the pixel area of ​​the target real object in the mixed reality image by a composition process that does not take depth into account. In the composition process that does not take depth into account, the virtual object in the virtual reality image is always drawn. Specifically, the color value of the virtual reality image is stored as the color value of the corresponding pixel area of ​​the mixed reality image being generated in the buffer.

[0069] In S1506, it is determined whether all physical objects detected in S1501 have been processed. If the determination shows that there are unprocessed physical objects, the process returns to S1502, where the next target physical object is determined, and similar processing continues. On the other hand, if all physical objects have been processed, S511 is then executed. FIG. 17 is a diagram showing a mixed reality image obtained by the method of this embodiment, and is a mixed reality image obtained by combining the real image of FIG. 7(a) and the virtual reality image of FIG. 7(b) described above. Compared with the mixed reality image of FIG. 7(c) described above, it can be seen that the entire desk 703 is depicted in the mixed reality image of FIG. 17. The mixed reality image generated in this manner is displayed in S511. This completes the content of this embodiment.

[0070] <Variation 1> For example, in the situation shown in FIG. 7( a) above, a portion of the desk 703 may be located closer to the user than the reference depth, while the keyboard 702 and documents (not shown) placed on the desk 703 are located further back than the reference depth. When this embodiment is applied to such a case, a different ID from that of the desk 703 may be assigned to the keyboard 702 and documents. Depending on the reference depth, the entire desk 703 may be displayed, but the documents and keyboard 702 on the desk may not be displayed. In such a case, it may be preferable to also display the objects placed on the desk 703. To avoid the need for the wearer to perform adjustments for this purpose, when a specific physical object located closer to the user than the reference depth is displayed in its entirety, other physical objects located above and below the specific physical object may also be displayed. Furthermore, to obtain such a mixed reality image, the specific physical object located closer to the user than the reference depth and the physical objects located above and below it may be combined using, for example, a known clustering technique and treated as a single physical object. Furthermore, for real objects that are treated as exceptions, such as ceilings, walls, floors, and the ground, even if a part of them is above or below, rather than displaying the entire object, the method of embodiment 1 or 2 may be applied to display only the part that is in front of the reference depth.

[0071] <Variation 2> Furthermore, in the above-described embodiment, an example was described in which if a part of a physical object is closer than the reference depth, the entire physical object can be displayed. However, if the entire physical object is not closer than the reference depth, the physical object may not be displayed even partially. This also makes it possible to prevent the physical object from being cut off midway. In this case, it is sufficient to determine whether all pixel values ​​in the real depth image for a pixel region in the object detection result image where the ID of the target physical object is stored are equal to or greater than the reference depth value. Alternatively, this determination may be made based on the pixel value at a representative point in the pixel region where the ID of the target physical object is stored.

[0072] As described above, according to this embodiment, it is possible to prevent a real object that exists across a boundary between reference depths from being displayed in a disconnected manner in the mixed reality image.

[0073] [Embodiment 4] Because the distance to a real object changes when the wearer moves, in the methods of the first and second embodiments, a portion of a real object located at the very edge of the reference depth is switched between displayed and hidden in the mixed reality image even with a slight movement of the wearer. In the method of the third embodiment, in which even a portion of a real object is displayed in its entirety if it is within the reference depth, the entire real object is switched between displayed and hidden in accordance with the wearer's movements. If a real object flickers in the mixed reality image in this way, the wearer's sense of immersion and concentration will be significantly reduced. Therefore, a fourth embodiment will be described in which, once a reference depth is set, the display / hidden state of the real object is not switched until a new reference depth is set. Since the hardware configurations of the HMD 10 and the image processing device 20 are similar to those of the first to third embodiments, a description thereof will be omitted, and the following description will focus on the differences. Although the following description focuses on the differences when based on the third embodiment, it is also possible to apply the first and second embodiments (including their variations) as bases.

[0074] <Software configuration of image processing device> 18 is a functional block diagram showing the software configuration (logical configuration) of the image processing device 20 according to this embodiment. A major difference from the functional block diagram of FIG. 14 of the third embodiment is that a flag setting unit 1801 is added. This flag setting unit 1801 performs processing to set a display flag as information indicating that a real object determined to be in front of the reference depth is to be displayed in a mixed reality image. The set flag information is output to the synthesis unit 17.

[0075] <Operation flow of image processing device> Fig. 19 is a flowchart showing the flow of processing for generating a mixed reality image in the image processing device 20 according to this embodiment. The difference from the flowchart of Fig. 15 of embodiment 3 is that steps S1901 to S1912 are executed instead of steps S1501 to S1506, and other common steps are numbered the same. The following describes steps S1901 to S1912, which are the differences.

[0076] In S1901, similar to S1501 described above, the object detection unit 1401 applies a known object detection process to the real image acquired in S501 to detect real objects appearing in the real image, and generates an object detection result image and an object detection result table.

[0077] In S1902, the inside / outside determination unit 16 determines whether the reference depth has changed since the previous processing. For example, it determines whether the reference depth value of the previous frame stored in the RAM 102 or the like is equal to the reference depth value of the current frame. If the reference depth has changed, S1903 is executed next. On the other hand, if it has not changed, S1908 is executed next. Note that since there is no previous frame immediately after the start of this flow, S1902 is skipped and S1908 is executed immediately.

[0078] In S1903, a target real object is determined from all the real objects detected in S1901. Then, in S1904, the inside / outside determination unit 16 determines whether the target real object is closer or farther than the reference depth D_ref based on the object detection result image and the real depth image. The result of the determination is sent to the flag setting unit 1801. If the target real object is closer than the reference depth D_ref, S1905 is executed, and if it is farther, S1906 is executed.

[0079] In S1905, the flag setting unit 1801 sets the value of the display flag for the target physical object to "TRUE", which means that the target physical object is to be made visible. In addition, in S1906, the flag setting unit 1801 sets the value of the display flag for the target physical object to "FALSE", which means that the target physical object is not to be made visible. The display flag set in this way is stored by adding a "display flag" column to the object detection result table, as shown in FIG. 20, for example. The initial value of the display flag is "FALSE".

[0080] In S1907, it is determined whether all physical objects detected in S1901 have been processed. If the result of the determination is that there are unprocessed physical objects, the process returns to S1903, where the next target physical object is determined, and similar processing is continued. On the other hand, if all physical objects have been processed, S1908 is executed next.

[0081] In S1908, a physical object to be focused on is determined from all the physical objects detected in S1901. Then, in S1909, the next process to be executed is determined depending on the value of the display flag of the focused physical object. If the value of the display flag is "TRUE", S1910 is executed next, and if the value of the display flag is "FALSE", S1911 is executed next.

[0082] In S1910, similar to S1504 described above, the composition unit 17 determines the color value of the pixel region of the target physical object in the mixed reality image by composition processing that takes depth into consideration. Also, in S1911, similar to S1505 described above, the composition unit 17 determines the color value of the pixel region of the target physical object in the mixed reality image by composition processing that does not take depth into consideration.

[0083] In S1912, similarly to S1506 described above, it is determined whether or not all physical objects detected in S1901 have been processed. If the result of the determination indicates that there are unprocessed physical objects, the process returns to S1908, where the next target physical object is determined, and similar processing is continued. On the other hand, if all physical objects have been processed, S511 is executed next. This completes the content of this embodiment.

[0084] As described above, according to this embodiment, it is possible to prevent the display / non-display of a real object from being frequently switched in accordance with the movement of the wearer.

[0085] <Other embodiments> The present disclosure can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.

[0086] The present disclosure also includes the following configurations and methods.

[0087] [Configuration 1] An image processing device that generates a mixed reality image by combining a real image and a virtual reality image, and that is displayed on a see-through imaging display device that is worn on a person's head, a distance setting means for setting a reference depth representing a distance at which a real object shown in the real image is allowed to be visible in the mixed reality image; a generating means for generating the mixed reality image by combining the real image and the virtual reality image in accordance with the reference depth; 1. An image processing device comprising:

[0088] [Configuration 2] an acquisition means for acquiring depth information representing a distance between a user wearing the imaging and display device and a real object shown in the real image; the generating means determines whether or not to make the real object shown in the real image visible in the mixed reality image based on a result of comparison between the distance indicated by the reference depth and the distance indicated by the depth information, and generates the mixed reality image. 2. The image processing device according to configuration 1,

[0089] [Configuration 3] The generating means If the distance indicated by the depth information is smaller than the distance indicated by the reference depth, the mixed reality image is generated in such a way that the real object reflected in the real image is visible; If the distance indicated by the depth information is greater than the distance indicated by the reference depth, the mixed reality image is generated in such a way that the real object reflected in the real image is not visible. 3. The image processing device according to configuration 2.

[0090] [Configuration 4] The generating means For a real object whose distance indicated by the depth information is shorter than the distance indicated by the reference depth, a composite image is generated by combining the real object in the line of sight of the wearer and the virtual object shown in the virtual reality image, whichever is closer to the wearer, so that the object is visible; and For a real object whose distance indicated by the depth information is greater than the distance indicated by the reference depth, the real object is not made visible in the line of sight of the wearer, but a synthesis is performed to make the virtual object reflected in the virtual reality image visible, thereby generating the mixed reality image. 3. The image processing device according to configuration 2.

[0091] [Configuration 5] Further, the device has a receiving unit for receiving an instruction from the wearer. the distance setting means sets the reference depth based on an instruction from the wearer. 5. The image processing device according to any one of configurations 2 to 4,

[0092] [Configuration 6] a receiving means for receiving an instruction from the wearer; a position setting means for setting a position of the virtual object based on an instruction from the wearer; and The distance setting means determining the position of the virtual object set by the position setting means; setting the reference depth based on the distance from the wearer to the determined position; 5. The image processing device according to configuration 4.

[0093] [Configuration 7] The position setting means Identifying the position of the real object based on an instruction from the wearer; setting a position of the virtual object based on the identified position of the real object; The distance setting means determining the position of the virtual object set by the position setting means; setting the reference depth based on the distance from the wearer to the determined position; 7. The image processing device according to configuration 6,

[0094] [Configuration 8] The image processing device according to configuration 6 or 7, characterized in that the distance setting means sets the reference depth by adding or subtracting a distance according to an instruction from the wearer to the distance from the wearer to the determined position.

[0095] [Configuration 9] the receiving means receives designation of a real object that the wearer desires to be visible in the mixed reality image from among the real objects shown in the real image; the distance setting means sets the distance from the wearer to the physical object designated by the wearer as the reference depth. 6. The image processing device according to configuration 5.

[0096] [Configuration 10] a detection unit configured to detect a real object from the real image acquired by the acquisition unit; The image processing device described in any one of configurations 2 to 4, characterized in that when the distance indicated by the depth information for a part of a real object of interest among the real objects detected by the detection means is smaller than the distance indicated by the reference depth, the generation means generates the mixed reality image in which the entire real object of interest reflected in the real image is visible.

[0097] [Configuration 11] a detection unit configured to detect a real object from the real image acquired by the acquisition unit; The image processing device described in any one of configurations 2 to 4, characterized in that when the distance indicated by the depth information for a part of a real object of interest among the real objects detected by the detection means is smaller than the distance indicated by the reference depth, the generation means generates the mixed reality image by performing a synthesis so that the entire real object shown in the real image and the virtual object shown in the virtual reality image, whichever is closer to the wearer, is visible.

[0098] [Configuration 12] a detection unit configured to detect a real object from the real image acquired by the acquisition unit; The image processing device described in any one of configurations 2 to 4, characterized in that the generation means generates the mixed reality image in which the entire target real object reflected in the real image is visible when the distance indicated by the depth information for the entire target real object among the real objects detected by the detection means is smaller than the distance indicated by the reference depth.

[0099] [Configuration 13] The image processing device according to configuration 11 or 12, characterized in that the generation means generates the mixed reality image in which other real objects placed above and below the real object of interest, the entirety of which is made visible in the synthesis, are also made visible.

[0100] [Configuration 14] The image processing device described in any one of configurations 1 to 4, characterized in that once the reference depth is set by the distance setting means, the generation means generates the mixed reality image without changing the real objects that are made visible until a new reference depth is set again.

[0101] [Configuration 15] The image processing device according to any one of configurations 2 to 14, wherein the distance indicated by the reference depth is the distance on each of the x-axis, y-axis, and z-axis, which are expressed in a camera coordinate system with the z-axis directed forward of the wearer, the y-axis directed vertically, and the x-axis directed horizontally.

[0102] [Configuration 16] The image processing device according to any one of configurations 2 to 14, wherein the distance indicated by the reference depth is the distance between the x-axis and the z-axis expressed in a camera coordinate system in which the z-axis is in a forward direction of the wearer, the y-axis is in a vertical direction, and the x-axis is in a horizontal direction.

[0103] [Configuration 17] The image processing device according to any one of configurations 2 to 14, wherein the distance indicated by the reference depth is a distance on the z-axis expressed in a camera coordinate system having the z-axis in a forward direction of the wearer, the y-axis in a vertical direction, and the x-axis in a horizontal direction.

[0104] [Method 1] An image processing method for generating a mixed reality image by combining a real image and a virtual reality image, the mixed reality image being displayed on a see-through imaging display device worn by a person on the head, comprising: a distance setting step of setting a reference depth representing a distance that allows a real object shown in the real image to be visible in the mixed reality image; a generating step of generating a mixed reality image by combining a real image and a virtual reality image according to the reference depth; An image processing method comprising:

[0105] [Configuration 18] 18. A program for causing a computer to function as the image processing device according to any one of configurations 1 to 17.

Claims

1. An image processing device that generates a mixed reality image by combining a real image and a virtual reality image, and that is displayed on a see-through imaging display device that is worn on a person's head, a distance setting means for setting a reference depth representing a distance at which a real object shown in the real image is allowed to be visible in the mixed reality image; a generating means for generating the mixed reality image by combining the real image and the virtual reality image in accordance with the reference depth; 1. An image processing device comprising:

2. an acquisition means for acquiring depth information representing a distance between a user wearing the imaging and display device and a real object shown in the real image; the generating means determines whether or not to make the real object shown in the real image visible in the mixed reality image based on a result of comparison between the distance indicated by the reference depth and the distance indicated by the depth information, and generates the mixed reality image.

2. The image processing device according to claim 1, wherein:

3. The generating means If the distance indicated by the depth information is smaller than the distance indicated by the reference depth, the mixed reality image is generated in such a way that the real object reflected in the real image is visible; If the distance indicated by the depth information is greater than the distance indicated by the reference depth, the mixed reality image is generated in such a way that the real object reflected in the real image is not visible.

3. The image processing device according to claim 2.

4. The generating means For a real object whose distance indicated by the depth information is shorter than the distance indicated by the reference depth, a composite image is generated by combining the real object in the line of sight of the wearer and the virtual object shown in the virtual reality image, whichever is closer to the wearer, so that the object is visible; and For a real object whose distance indicated by the depth information is greater than the distance indicated by the reference depth, the real object is not made visible in the line of sight of the wearer, but a synthesis is performed to make the virtual object reflected in the virtual reality image visible, thereby generating the mixed reality image.

3. The image processing device according to claim 2.

5. Further, the device has a receiving unit for receiving an instruction from the wearer. the distance setting means sets the reference depth based on an instruction from the wearer.

3. The image processing device according to claim 2.

6. a receiving means for receiving an instruction from the wearer; a position setting means for setting a position of the virtual object based on an instruction from the wearer; and The distance setting means determining the position of the virtual object set by the position setting means; setting the reference depth based on the distance from the wearer to the determined position; 5. The image processing device according to claim 4.

7. The position setting means Identifying the position of the real object based on an instruction from the wearer; setting a position of the virtual object based on the identified position of the real object; The distance setting means determining the position of the virtual object set by the position setting means; setting the reference depth based on the distance from the wearer to the determined position; 7. The image processing device according to claim 6,

8. 8. The image processing device according to claim 6, wherein the distance setting means sets the reference depth by adding or subtracting a distance according to an instruction from the wearer to the distance from the wearer to the determined position.

9. the receiving means receives designation of a real object that the wearer desires to be visible in the mixed reality image from among the real objects shown in the real image; the distance setting means sets the distance from the wearer to the physical object designated by the wearer as the reference depth.

6. The image processing device according to claim 5,

10. a detection unit configured to detect a real object from the real image acquired by the acquisition unit; The image processing device described in claim 2, characterized in that, when the distance indicated by the depth information for a part of a real object of interest among the real objects detected by the detection means is smaller than the distance indicated by the reference depth, the generation means generates the mixed reality image in which the entire real object of interest reflected in the real image is visible.

11. a detection unit configured to detect a real object from the real image acquired by the acquisition unit; The image processing device described in claim 2, characterized in that, when the distance indicated by the depth information for a part of a real object of interest among the real objects detected by the detection means is smaller than the distance indicated by the reference depth, the generation means generates the mixed reality image by performing a synthesis so that the entire object of the real object shown in the real image and the virtual object shown in the virtual reality image, whichever is closer to the wearer, is visible.

12. a detection unit configured to detect a real object from the real image acquired by the acquisition unit; The image processing device described in claim 2, characterized in that the generation means generates the mixed reality image in which the entire target real object reflected in the real image is visible when the distance indicated by the depth information for the entire target real object among the real objects detected by the detection means is smaller than the distance indicated by the reference depth.

13. The image processing device according to claim 11 or 12, characterized in that the generating means generates the mixed reality image in which other real objects placed above and below the real object of interest, the entirety of which is made visible in the synthesis, are also made visible.

14. The image processing device described in claim 1, characterized in that once the reference depth is set by the distance setting means, the generation means generates the mixed reality image without changing the real objects that are made visible until a new reference depth is set again.

15. 3. The image processing device according to claim 2, wherein the distance indicated by the reference depth is the distance on each of the x-axis, y-axis, and z-axis, which are expressed in a camera coordinate system having the z-axis in a forward direction of the wearer, the y-axis in a vertical direction, and the x-axis in a horizontal direction.

16. 3. The image processing device according to claim 2, wherein the distance indicated by the reference depth is the distance between the x-axis and the z-axis expressed in a camera coordinate system in which the z-axis is in the forward direction of the wearer, the y-axis is in the vertical direction, and the x-axis is in the horizontal direction.

17. 3. The image processing device according to claim 2, wherein the distance indicated by the reference depth is a distance on the z-axis expressed in a camera coordinate system having the z-axis in a forward direction of the wearer, the y-axis in a vertical direction, and the x-axis in a horizontal direction.

18. An image processing method for generating a mixed reality image that combines a real image and a virtual reality image and is displayed on a see-through imaging display device that is worn on a person's head, comprising: a distance setting step of setting a reference depth representing a distance that allows a real object shown in the real image to be visible in the mixed reality image; a generating step of generating the mixed reality image by combining the real image and the virtual reality image according to the reference depth; An image processing method comprising:

19. A program for causing a computer to execute the image processing method according to claim 18.

Citation Information

Patent Citations

  • Image processor and control method thereof

    JP2016004493A