Head-mounted display and stereoscopic image display method

WO2026105510A1PCT designated stage Publication Date: 2026-05-21JAPAN DISPLAY INC
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
JAPAN DISPLAY INC
Filing Date
2025-10-08
Publication Date
2026-05-21

Smart Images

  • Figure JP2025035751_21052026_PF_FP_ABST
    Figure JP2025035751_21052026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention addresses the problem of reducing the weight of a head-mounted display to lower its cost. A head-mounted display system (100) includes: a lens; a coded aperture that narrows external light passing through the lens, to a predetermined coded pattern through a liquid crystal panel; a camera (14) having an image sensor; a depth calculation unit (20) that obtains the depth distribution of an image, which is captured by the camera (14), from the image and a known blur function related to the coded pattern; a stereoscopic image generation unit (21) that generates, on the basis of the image and the depth distribution, an image for a right eye and an image for a left eye; and a stereoscopic display (13) that displays the image for the right eye and the image for the left eye such that the images correspond to the wearer's right and left eyes, respectively.
Need to check novelty before this filing date? Find Prior Art

Description

Head-Mounted Display and Stereoscopic Image Display Method

[0008] ,

[0001] The present invention relates to a head-mounted display and a stereoscopic image display method.

[0002] Patent Document 1 describes a head-mounted display. In a non-transmissive type, that is, a type of head-mounted display that blocks external light rays from the user, in order to allow the user to view the surrounding situation, it is necessary to display an image captured by a camera mounted on the head-mounted display on an internal display.

[0003] At this time, in the head-mounted display, in order to display stereoscopic images corresponding to the user's left and right eyes, the head-mounted display must be provided with a pair of two stereoscopic cameras corresponding to the left and right images respectively.

[0004] International Publication No. 2018 / 101162

[0005] Since the head-mounted display is worn on the user's head, its weight is closely related to the burden on the user. Also, for the widespread use of the head-mounted display, its cost should be low.

[0006] The present invention has been made in view of such circumstances, and its object is to reduce the weight of the head-mounted display and reduce its cost.

[0007] To solve the above problems, the head-mounted display system according to the present application includes a lens, an encoding aperture that narrows external light passing through the lens into a predetermined encoding pattern by a liquid crystal panel, a camera having an image sensor, an image captured by the camera, a depth calculation unit that obtains a depth distribution of the image from a known blur function related to the encoding pattern, a stereoscopic image generation unit that generates a right-eye image and a left-eye image based on the image and the depth distribution, and a stereoscopic display that displays the right-eye image and the left-eye image corresponding to the left and right eyes of the wearer respectively.

[0008] Furthermore, in order to solve the above problems, the stereoscopic image display method according to this application comprises a lens, an encoding aperture that focuses the ambient light passing through the lens to a predetermined encoding pattern using a liquid crystal panel, a step of capturing an image with a camera having an image sensor, a step of determining the depth distribution of the image from the image captured by the camera and a known blur function related to the encoding pattern, a step of generating a right-eye image and a left-eye image based on the image and the depth distribution, and a step of displaying the right-eye image and the left-eye image on a head-mounted display corresponding to the left and right eyes of the wearer, respectively.

[0009] This is an external view of a head-mounted display system according to a preferred embodiment of the present invention. This is a schematic diagram showing the structure of a camera. This is a diagram showing the principle of depth estimation using an encoded aperture. This is a functional block diagram of the head-mounted display system. This is a diagram illustrating the process of generating a stereo image using an example. This is a diagram illustrating the process of generating a stereo image using an example. This is a diagram illustrating how a three-dimensional object is placed in a three-dimensional model space by a three-dimensional object placement unit. This is a diagram showing an example in which missing parts of a stereo image are filled in. This is a flowchart illustrating the procedure of a stereoscopic image display method according to a preferred embodiment of the present invention. This is a flowchart illustrating the procedure of a stereoscopic image display method according to another example of an embodiment of the present invention.

[0010] In this application, the drawings may schematically represent the width, thickness, shape, etc., of each part in order to clarify the explanation, but these are merely examples and do not limit the interpretation of the present invention. In this specification and in each drawing, elements having the same function as those described with respect to previously shown drawings are denoted by the same reference numerals, and redundant explanations may be omitted.

[0011] Furthermore, in the detailed description of the present invention, when defining the positional relationship between one component and another component, "above" and "below" include not only cases where the component is located directly above or directly below another component, but also cases where other components are interposed between them, unless otherwise specified.

[0012] Figure 1 is an external view of a head-mounted display system 100 according to a preferred embodiment of the present invention. The head-mounted display 100 includes a head-mounted display 1 worn on the user's head and a control unit 2 that controls the operation of the head-mounted display 1 and supplies power to it. The control unit 2 may be an information processing device having the configuration of a general-purpose computer, and it is responsible for externally handling information processing that cannot be handled by the head-mounted display 1 alone. Therefore, the control unit 2 and the head-mounted display 1 are connected via wired or wireless means to enable information communication.

[0013] If the head-mounted display 1 does not have an internal power source such as a battery, it needs to be powered from an external source. Power to the head-mounted display 1 may be supplied via the controller 2 or from a separate power source. Power may be supplied via wired or wireless power. If the head-mounted display 1 has an internal power source, power supply to the head-mounted display 1 is not required while it is in use.

[0014] Furthermore, if the information processing functions that the controller 2 should perform can be done by the information processing circuit built into the head-mounted display 1 itself, then, as shown in the figure, there is no need to prepare the controller 2 separately from the head-mounted display 1, and it is acceptable to build it into the head-mounted display 1. The head-mounted display 1 is preferably lightweight for the purpose of being worn on the user's head, and a configuration in which it receives power from an external source and uses a separate controller 2 as shown in the figure is conceivable. However, since a wired connection would restrict the freedom of movement of the user wearing the head-mounted display 1, it is also possible to configure the head-mounted display system 100 with the head-mounted display 1 alone. How the head-mounted display system 100 is configured may be designed appropriately depending on the weight of the head-mounted display 1 and the intended use.

[0015] The head-mounted display 1 includes a goggle section 10 that covers both of the wearer's eyes, a headphone section 11 that covers the wearer's ears and has speakers that output sound, and a fixing belt 12 that is fixed to the wearer's head and supports the head-mounted display 1. Inside the goggle section 10, there is a stereo display 13 that displays a stereoscopic image to the wearer by displaying a right eye image and a left eye image corresponding to the wearer's left and right eyes, respectively, and a camera 14 that captures the scenery in front of the head-mounted display 1.

[0016] The structure of the stereo display 13 can be anything as long as it is capable of presenting stereoscopic images to the wearer. This could be a glasses-free stereoscopic display with lenticular lenses or the like on the surface of a flat panel display such as a liquid crystal display, or a display that projects images from separate displays for the left and right fields of view onto the wearer's eyes via appropriate optical systems such as mirrors or lenses, or any other type of stereo display 13 known to be used.

[0017] In the example shown in Figure 1, the camera 14 is located at the upper center of the goggles 10, but it can be located at any position that allows it to capture the scenery directly in front of the head-mounted display 1. In addition to any position on the goggles 10, it may also be located, for example, at the top of the wearer's head where the fixing belt 12 is attached.

[0018] In order to present the external, frontal view of the head-mounted display 1 to the wearer as a stereoscopic image using the stereo display 13, images for the right eye and left eye, captured from different viewpoints, are required. For this reason, generally, two cameras corresponding to the left and right viewpoints must be provided, or a so-called stereo camera equipped with an optical system for capturing stereo images must be used. However, in the head-mounted display system 100 according to this embodiment, only one monocular camera 14 is used.

[0019] Figure 2 is a schematic diagram showing the structure of the camera 14. In this embodiment, the camera 14 has the configuration of a so-called digital camera, and consists of a lens group 140 including at least one lens, an encoding aperture 141 that narrows the ambient light passing through the lens group according to a predetermined encoding pattern, a shutter 142, and an image sensor 143, all housed in a housing 144.

[0020] The lens group 140 may be a set of imaging lenses commonly used in cameras, and may be capable of adjusting the focal length, depth of field, and zoom magnification as appropriate. The material, coating, number of groups, and number of lenses constituting the lens group 140 are not particularly limited, and the lens group 140 may be a fixed-focus single lens. The adjustment of the lens group 140 may be performed automatically or manually.

[0021] The encoding aperture 141 partially blocks external light passing through the lens group 140 using a specific mask pattern. Specific examples of the encoding aperture 141 include a black plate with an opening of a specific pattern shape, or a transparent plate such as glass with a specific black pattern printed on its surface. However, in this embodiment, the encoding aperture 141 can be a liquid crystal shutter using a liquid crystal panel, allowing the aperture pattern to be changed. For example, the presence or absence of the encoding aperture can be switched by switching the display of a specific mask pattern on or off. Alternatively, the type of encoding aperture can be changed by switching between multiple types of mask patterns. Furthermore, a dot matrix type liquid crystal display may be used as the liquid crystal shutter for the encoding aperture 141, allowing for the display of an arbitrary mask pattern, as well as the display of a normal aperture pattern with a circular opening.

[0022] The shutter 142 functions as a shutter in a normal camera and adjusts the amount of ambient light exposed to the image sensor 143. A general mechanical shutter may be used for the shutter 142, but in this embodiment, a liquid crystal shutter is used. Note that if the exposure amount is adjusted by controlling the timing of charge sweep from each element constituting the image sensor 143, the shutter 142 may be omitted.

[0023] The image sensor 143 is a two-dimensional optical sensor. There are no limitations on the type of image sensor 13; it may be a general CMOS (complementary metal-oxide-semiconductor) sensor or a CCD (charge-coupled device). The image detected by the image sensor 13 may be a color image or a monochrome image, but in this embodiment, a color image is output.

[0024] In this embodiment, the depth calculation unit 20 built into the control unit 2 can determine the depth distribution, which is the distribution of depth within an image, from the image obtained from the image sensor 143 of the camera 14, which is the distance from the camera 14 to the surface of the subject.

[0025] Figure 3 shows the principle of depth estimation using the coded aperture 141. The figure schematically shows how ambient light passing through the coded aperture 141 is refracted by the lens group 140 (shown here as a single lens for simplicity) and strikes the image sensor 143. In the figure, (a) shows the optical path of light rays from the subject surface that is closer than the focal length of the lens group 140, and (b) shows the optical path of light rays from the subject surface that is further than the focal length of the lens group 140, both shown by dashed lines.

[0026] In the examples of Figure 3(a) and (b), light rays from the subject surface do not form an image on the image sensor 143 and are captured as blurred. In this case, if the PSF (General Blur Function) on the surface of the image sensor 13 of the light rays that have passed through the coding aperture 141 is known, then restoring the blurred image by deconvolution using such PSF is equivalent to estimating the depth of each part of the image, and thus the depth to the subject surface in the image can be estimated by calculation.

[0027] In this case, as shown in (a), the light rays encoded by passing through the encoding aperture 141 enter the image sensor 13 in the direction of encoding shown in 3a, without changing the geometric positional relationship of the encoding.

[0028] For convenience, Figure 3 shows the coding patterns of the coding aperture 141 projected onto the surface of the image sensor 13 in coding directions 3a and 3b. However, this is not actually the case. The figure shows that the spatial frequency characteristics of the blur of light rays entering the image sensor 143 follow the PSF corresponding to the coding pattern shown in coding direction 3a. Hereafter, this geometric positional relationship will be referred to as the forward direction, and the PSF in this case will be referred to as the forward PSF.

[0029] In contrast, in the case shown in (b), the light rays encoded by passing through the encoding aperture 141 enter the image sensor 143 with the encoding direction shown in 3b, with the geometric positional relationship of the encoding inverted vertically and horizontally. The encoding direction 3b in Figure 3, like the encoding direction 3a, indicates that the spatial frequency characteristics of the blur of the light rays entering the image sensor 143 follow the PSF corresponding to the encoding pattern shown in encoding direction 3b. Hereafter, this geometric positional relationship will be referred to as the reverse direction, and the PSF in this case will be referred to as the reverse direction PSF.

[0030] In this case, if the forward PSF and the reverse PSF are the same, it is not possible to determine from the blur of the image captured by the image sensor 143 whether the position of the subject surface is far or near relative to the focal length of the lens group 140. On the other hand, since the coding direction 3a and the coding direction 3b are in a positional relationship rotated 180 degrees with respect to the optical axis, if the coding pattern in the coding aperture 141 is a figure that can distinguish between these two, that is, a double-asymmetric figure as exemplified in Figure 3, then the forward PSF and the reverse PSF will be different, and the distance of the subject surface relative to the focal length of the lens group 140 can also be determined using each PSF.

[0031] Figure 4 is a functional block diagram of the head-mounted display system 100. The head-mounted display system 100 is configured to input an image from the camera 14 to the depth calculation unit 20 to determine the depth distribution of the image, generate a stereo image from the image and the obtained depth distribution in the stereo image generation unit 21, and output it to the stereo display 13. The depth distribution and image obtained by the depth calculation unit 20 are further sent to the three-dimensional object placement unit 22 and used for placement in the three-dimensional model space held by the three-dimensional model space holding unit 23. In this embodiment, the depth calculation unit 20, stereo image generation unit 21, three-dimensional object placement unit 22, and three-dimensional model space holding unit 23 are implemented by software executed in the controller 2 or by circuits provided in the controller 2.

[0032] The depth calculation unit 20 determines the depth distribution of an image from the image captured by the camera 14 and a known blur function related to the encoding pattern of the camera 14's encoded aperture 141. Then, using this image and depth distribution, the stereo image generation unit 21 can generate a stereo image, that is, a right-eye image and a left-eye image to be displayed on the stereo display 13.

[0033] Figures 5 and 6 illustrate the process of generating a stereo image using an example. Here, we assume that the subject is a cube with letters written on each face, as shown in Figure 5(a). The imaging direction is the arrow shown in Figure 5(a), that is, an angle looking down from slightly above the front of the subject (the face marked "A").

[0034] At this time, the image captured by the camera 14 will be as shown in Figure 5(b). That is, the front and top surfaces of the subject are captured. From the blur of each part of the image in (b), the depth calculation unit 20 can calculate the depth of each part of the image, and therefore the distance from the camera 14 to each part of the subject can be determined.

[0035] If the distance to each part of the subject is known, the projection position of each part of the subject when the viewpoint position is changed can be easily calculated. Therefore, the stereo image generation unit 21 can generate the right-eye image and left-eye image shown in Figure 6 by deforming the image using the depth distribution obtained by the depth calculation unit 20. By displaying these right-eye and left-eye images on a stereo display, the image captured by the monocular camera 14 can be displayed to the wearer as a stereo image.

[0036] However, in the right-eye and left-eye images shown in Figure 6, as indicated by the diagonal lines, when generating a stereo image by changing the viewpoint position, there may be missing parts that are not included in the original image because they are blind spots, and therefore cannot be reconstructed as a stereo image. As long as these missing parts are based solely on the images captured by the camera 14, it is impossible to construct a stereo image from them. Therefore, these missing parts may be filled in with pixels of a predetermined specific color, such as black or white.

[0037] However, the head-mounted display system 100 according to this embodiment may have a configuration to further compensate for such missing parts. This corresponds to the three-dimensional object placement unit 22 and the three-dimensional model space holding unit 23 shown in Figure 4.

[0038] The three-dimensional object placement unit 22 places three-dimensional objects in the three-dimensional model space based on the image obtained by the camera 14 and the depth distribution obtained by the depth calculation unit 20. Here, the three-dimensional model space is a virtual three-dimensional space associated with the space in which the head-mounted display 1 exists.

[0039] Figure 7 illustrates how a three-dimensional object is placed in the three-dimensional model space by the three-dimensional object placement unit 22. Since the relative distance between the subject surface captured by the camera 14 and the camera 14 is known from the depth distribution, the position of the subject surface in the three-dimensional model space is determined if the position and orientation of the camera 14, i.e., the head-mounted display 1, are known. Therefore, the three-dimensional object placement unit 22 can place the captured subject surface as an object in the three-dimensional model space.

[0040] If the image shown in Figure 5(b) is captured by camera 14, then the faces of the three-dimensional object corresponding to the subject shown in Figure 7, with the letters "A" and "B" drawn on them, will be added as objects. Furthermore, if the side of the subject is captured from a different angle, then, as shown in Figure 7, a face with the letter "C" drawn on it will be added as a three-dimensional object. In this way, the wider the range captured by camera 14, the more the three-dimensional model space becomes a mapping of real-world objects as three-dimensional models within the three-dimensional model space.

[0041] The position and orientation of the head-mounted display 1 may be constantly tracked by equipping the head-mounted display 1 with various motion sensors, such as an accelerometer and an angular velocity sensor. Alternatively, feature points may be extracted from the external image captured by the camera 14 using a known algorithm, and the position and orientation of the head-mounted display 1 may be determined based on the change in the position of the feature points in the image, or both methods may be used in combination.

[0042] The three-dimensional objects placed by the three-dimensional object placement unit 22 are held together with the three-dimensional model space in the three-dimensional model space holding unit 23, and are referenced by the stereo image generation unit 21 as needed and used to generate stereo images.

[0043] That is, as shown in FIG Figure FIG Figure 8, the stereo image generation unit 21 complements the missing portions not included in the image captured by the camera 14 based on the three-dimensional object. When the missing portions are included in the stereo image to be generated, the stereo image generation unit 21 complements the missing portions of the stereo image on the assumption that the three-dimensional object arranged in the three-dimensional model space can be seen based on the position and orientation of the head-mounted display 1. When the three-dimensional object corresponding to the missing portion has not been arranged yet, as before, it is complemented by filling it with pixels of a specific color determined in advance.

[0044] In the embodiment described above, the stereo image generation unit 21 generates a stereo image based on the image captured by the camera 14 and its depth distribution, and has been described as complementing only the missing portions included in the stereo image based on the three-dimensional object arranged in the three-dimensional model space. However, the present invention is not limited to this, and the stereo image generation unit 21 may generate a stereo image exclusively based on the three-dimensional object arranged in the three-dimensional model space.

[0045] Figure 9 is a flowchart for explaining the procedure of the stereoscopic image display method according to the present embodiment.

[0046] First, in step ST1, an image of the subject is captured by the camera 14 having the lens 140, the encoding diaphragm 141 that narrows the external light passing through the lens 140 with a predetermined encoding pattern, and the image sensor 143. In the subsequent step ST2, the depth calculation unit 3 obtains the depth distribution of the image from the image captured by the camera 14 and the known blur function related to the encoding pattern.

[0047] Furthermore, in step ST3, the stereo image generation unit 21 generates a right-eye image and a left-eye image based on the image and the depth distribution. At this time, in step ST4, the three-dimensional object arrangement unit 22 arranges a three-dimensional object in the three-dimensional model space based on the image and the depth distribution.

[0048] In step ST5, the stereo image generation unit 21 determines whether there is a missing portion in the generated stereo image. If there is a missing portion, the process proceeds to step ST6, and the missing portion is complemented based on the three-dimensional object arranged in the three-dimensional model space.

[0049] If there is no missing portion in the stereo image, and after the missing portion is complemented, finally in step ST7, the stereo image, that is, the right-eye image and the left-eye image, is displayed on the stereo display 13 of the head-mounted display 1.

[0050] Note that the above flow described with reference to FIG. 9 is a case where the stereo image generation unit 21 generates a stereo image based on the image captured by the camera 14 and its depth distribution, and only for the missing portion included in the stereo image, it is complemented based on the three-dimensional object arranged in the three-dimensional model space.

[0051] On the other hand, when the stereo image generation unit 21 generates a stereo image exclusively based on the three-dimensional object arranged in the three-dimensional model space, as shown in FIG. 10, following steps ST1, ST2, and ST4 already described with reference to FIG. 9, in step ST8, the stereo image generation unit 21 generates a stereo image based on the three-dimensional object arranged in the three-dimensional model space, and in step ST7, it may be displayed on the stereo display 13 of the head-mounted display 1.

[0052] 1. Head-mounted display 2. Control unit 3a, 3b Encoding directions 10. Goggle unit 11. Headphone unit 12. Fixed belt 13. Stereo display 14. Camera 20. Depth calculation unit 21. Stereo image generation unit 22. Three-dimensional object arrangement unit 23. Three-dimensional model space holding unit 140. Lens group 141. Encoding aperture 142. Shutter 143. Image sensor 144. Housing 100. Head-mounted display system

Claims

1. A head-mounted display system comprising: a lens; an encoding aperture that focuses ambient light passing through the lens to a predetermined encoding pattern using a liquid crystal panel; a camera having an image sensor; an image captured by the camera; a depth calculation unit that determines the depth distribution of the image from a known blur function related to the encoding pattern; a stereo image generation unit that generates a right-eye image and a left-eye image based on the image and the depth distribution; and a stereo display that displays the right-eye image and the left-eye image corresponding to the wearer's left and right eyes, respectively.

2. The head-mounted display system according to claim 1, comprising: a three-dimensional model space holding unit that holds a three-dimensional model space; and a three-dimensional object placement unit that places three-dimensional objects in the three-dimensional model space based on the image and the depth distribution, wherein the stereo image generation unit generates the right-eye image and the left-eye image based on the three-dimensional objects.

3. The head-mounted display system according to claim 2, wherein the stereo image generation unit generates the right-eye image and the left-eye image based on the image, and fills in any missing portions not included in the image based on the three-dimensional object.

4. A stereoscopic image display method comprising: a step of capturing an image with a camera having a lens, an encoding aperture that focuses ambient light passing through the lens into an encoded pattern using a predetermined liquid crystal panel; a step of determining the depth distribution of the image from the image captured by the camera and a known blur function related to the encoded pattern; a step of generating a right-eye image and a left-eye image based on the image and the depth distribution; and a step of displaying the right-eye image and the left-eye image on a head-mounted display corresponding to the wearer's left and right eyes, respectively.

5. A stereoscopic image display method according to claim 4, comprising the step of arranging three-dimensional objects in the three-dimensional model space based on the image and the depth distribution, wherein the right-eye image and the left-eye image are generated based on the three-dimensional objects.

6. The stereoscopic image display method according to claim 5, wherein when generating the right-eye image and the left-eye image based on the aforementioned image, missing portions not included in the aforementioned image are filled in based on the three-dimensional object.