Head-mounted display terminal, control method, and information display system

By dynamically switching between binocular and monocular modes based on external detection, the head-mounted display terminal addresses convergence-accommodation inconsistencies, improving visibility and reducing double images.

WO2026105326A1PCT designated stage Publication Date: 2026-05-21MAXELL LTD
View PDF 12 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
MAXELL LTD
Filing Date
2024-11-18
Publication Date
2026-05-21

AI Technical Summary

Technical Problem

Head-mounted display terminals, particularly in binocular mode, suffer from double images due to convergence-accommodation inconsistency when users attempt to focus on both real objects and displayed images simultaneously, especially when the real object is closer than the display surface.

Method used

The head-mounted display terminal switches between binocular and monocular modes based on the detection of surrounding objects, using external detectors to determine the state and adjust display modes accordingly.

Benefits of technology

This approach improves visibility by ensuring that images from both eyes fuse correctly, reducing double images and enhancing user comfort by aligning convergence and accommodation distances.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024040789_21052026_PF_FP_ABST
    Figure JP2024040789_21052026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention provides technology capable of improving visibility in a head-mounted display terminal having a binocular mode and a monocular mode. A head-mounted display terminal 1 has a processor 101 as well as a right display and left display that display video. The head-mounted display terminal 1 has an external detector for acquiring external information. The processor 101 involves: processing for detecting an object in the surroundings on the basis of data acquired by the external detector; processing for determining the status of at least one among the object and video; and processing for switching, on the basis of the determination, between a binocular mode for displaying video on both the right display and left display and a monocular mode for displaying video on one among the right display and left display.
Need to check novelty before this filing date? Find Prior Art

Description

Head-mounted display terminal, control method, and information display system

[0001] The present invention relates to a head-mounted display terminal, a control method, and an information display system.

[0002] In a head-mounted display terminal such as an AR glass, a VR glass, or a head-mounted display that can display an image on displays corresponding to the left and right eyes, there is known a head-mounted display terminal having a binocular mode in which an image is displayed on the displays of both eyes and a monocular mode in which an image is displayed only on one of the displays corresponding to the left and right eyes. The binocular mode has the advantage that depth perception is easy due to binocular parallax and it is excellent in three-dimensional display.

[0003] Patent Document 1 describes a technique in which when the power supply voltage of a device is lower than a predetermined value, the driving of at least one of a plurality of display means is stopped and the remaining display means are driven, so that when the power supply capacity becomes low, an image or the like can be displayed for a longer time.

[0004] Patent Document 2 describes a technique in which, by alternately switching between a first state in which the emission of image light is stopped on one of a right-eye image light generation unit and a left-eye image light generation unit and the emission of image light is executed on the other, and a second state in which the emission of image light is executed on one and the emission of image light is stopped on the other, low power consumption is realized without degrading the image quality of the virtual image recognized by the user.

[0005] Patent Document 3 describes a technique in which when an image display is being executed, the fatigue degree of each of the left and right eyes of an observer is independently detected, and based on the detection result, the image display by either the left-eye display unit or the right-eye display unit is substantially stopped, thereby preventing eye fatigue while being able to view necessary content information without interruption.

[0006] Japanese Patent Application Laid-Open No. 8-179275, Japanese Patent Application Laid-Open No. 2012-168221, Japanese Patent Application Laid-Open No. 2012-211959

[0007] In recent years, there has been an increase in use cases where images are displayed on a portion of the transparent display of a head-mounted display terminal, allowing users to view the displayed images while simultaneously viewing a real object through the transparent display. In such use cases, it is difficult to focus on both the real object and the image simultaneously, especially when the real object is closer than the display surface. In such cases, particularly in binocular mode, the images from both eyes do not fuse, resulting in a double image and significantly impairing the user's visibility.

[0008] Patent documents 1 to 3 describe a technology for automatically stopping the display on one of multiple displays, but they do not describe switching between binocular and monocular modes to address the issue of double images occurring in binocular mode, which is relevant to user visibility.

[0009] Therefore, the present invention aims to provide a technology that can improve visibility in a head-mounted display terminal equipped with binocular mode and monocular mode.

[0010] To solve the above problems, one representative head-mounted display terminal of the present invention is a head-mounted display terminal comprising a processor and a right display and a left display for displaying images, wherein the head-mounted display terminal is equipped with an external detector for acquiring external information, and the processor has the processes of detecting surrounding objects based on the data acquired by the external detector, determining at least one state of an object or image, and switching between a binocular mode in which images are displayed on both the right display and the left display and a monocular mode in which images are displayed on either the right display or the left display based on the determination.

[0011] According to the present invention, visibility can be improved in a head-mounted display terminal equipped with binocular mode and monocular mode.

[0012] Other issues, configurations, and effects not mentioned above will be clarified by the following description of the embodiments.

[0013] This is a block diagram showing an example of the configuration of a head-mounted display terminal. This is a diagram showing an example of the appearance of a head-mounted display terminal. This is a diagram showing an example of the appearance of a head-mounted display terminal. This is a diagram showing an example of the overall configuration when a head-mounted display terminal is linked with an external information terminal. This is a diagram showing an example of the overall configuration when a head-mounted display terminal is linked with an external information terminal. This is a block diagram showing an example of the configuration of an external information terminal. This is a diagram illustrating the relationship between the position of an object and the convergence angle. This is a diagram illustrating the relationship between the distance between the user and the object and the convergence angle at that time. This is a diagram for explaining convergence-accommodation inconsistency. This is a diagram showing the occurrence of a double image in a head-mounted display terminal due to convergence-accommodation inconsistency. This is a diagram showing the occurrence of a double image in a head-mounted display terminal due to convergence-accommodation inconsistency. This is a diagram illustrating the relationship between distance and VAC amount. This is a diagram illustrating the relationship between distance and VAC amount. This is a diagram illustrating an example of mode change from binocular mode to monocular mode in this embodiment. This is a diagram illustrating an example of mode change from binocular mode to monocular mode in this embodiment. This is a diagram illustrating another example of mode change from binocular mode to monocular mode in this embodiment. This figure illustrates another example of mode change from binocular mode to monocular mode in this embodiment. This flowchart shows an example of mode change processing performed on a head-mounted display terminal. This flowchart shows a first example of processing for mode transition condition determination (S1). This flowchart shows a second example of processing for mode transition condition determination (S1). This flowchart shows a third example of processing for mode transition condition determination (S1). This table shows an example of processing for setting confirmation (S10). This table shows an example of processing for actual object detection (S11). This table shows an example of processing for actual object determination (S12). This table shows an example of processing for display mode confirmation (S13). This table shows an example of processing for mode change (S2). This flowchart shows a first example of processing for mode transition release condition determination (S3). This flowchart shows a second example of processing for mode transition release condition determination (S3). This flowchart shows a third example of processing for mode transition release condition determination (S3). This table shows an example of processing for actual object detection (S30). This table shows an example of processing for actual object determination (S31). This table shows an example of processing for display mode confirmation (S32). This table shows an example of the process for changing modes (S4).

[0014] Embodiments of the present invention will be described below with reference to the drawings. The embodiments are illustrative examples for explaining the present invention, and have been omitted and simplified as appropriate for clarity of explanation. The present invention can also be implemented in various other forms. Unless otherwise specified, each component may be singular or plural.

[0015] The positions, sizes, shapes, and ranges of the components shown in the drawings may not represent their actual positions, sizes, shapes, and ranges in order to facilitate understanding of the invention. Therefore, the present invention is not necessarily limited to the positions, sizes, shapes, and ranges disclosed in the drawings.

[0016] When there are multiple components with the same or similar function, they may be described using the same symbol but with different subscripts. Furthermore, when it is not necessary to distinguish between these multiple components, the subscripts may be omitted in the description.

[0017] In embodiments, processing performed by executing a program may be described. Here, the computer executes the program using a processor (e.g., CPU (Central Processing Unit), GPU (Graphics Processing Unit)) and performs processing defined by the program using memory resources (e.g., memory) and interface devices (e.g., communication ports). Therefore, the main entity performing the processing by executing the program may be the processor. The processor includes transistors and other circuits and is considered circuitry or processing circuitry. Similarly, the main entity performing the processing by executing the program may be a controller, device, system, computer, or node having a processor. The main entity performing the processing by executing the program may be an arithmetic unit and may include dedicated circuits that perform specific processing. Here, dedicated circuits include, for example, FPGAs (Field Programmable Gate Arrays), ASICs (Application Specific Integrated Circuits), CPLDs (Complex Programmable Logic Devices), etc.

[0018] The program may be installed on the computer from the program source. The program source may be, for example, a program distribution server or a storage medium readable by the computer. If the program source is a program distribution server, the program distribution server includes a processor and storage resources for storing the program to be distributed, and the processor of the program distribution server may distribute the program to other computers. In addition, in some embodiments, two or more programs may be implemented as a single program, or one program may be implemented as two or more programs.

[0019] Head-mounted display terminals are devices that display video information within the user's field of view, such as head-mounted displays and AR glasses.

[0020] First, an example of the configuration of a head-mounted display terminal will be explained with reference to Figure 1.

[0021] Figure 1 is a block diagram showing an example of the configuration of a head-mounted display terminal.

[0022] The head-mounted display terminal 1 is a device capable of displaying information based on AR (Augmented Reality). As shown in Figure 1, this head-mounted display terminal 1 includes a processor 101, a storage device 110, an input I / F 120, a video input / output device 130, an audio input / output device 140, a sensor group 150, a communication I / F 160, an expansion I / F 171, a timer 172, and an actuator 173.

[0023] The processor 101 is configured using a CPU and the like, and is connected to various configurations via the bus 102.

[0024] The storage device 110 includes a volatile memory 111 and a non-volatile memory 112. The volatile memory 111 is the main memory and is configured using, for example, RAM (Random Access Memory). The processor 101 temporarily stores data such as programs in the volatile memory 111 and performs data processing. The non-volatile memory 112 is an auxiliary storage device that stores data nonvolatilely. The non-volatile memory 112 is configured using a non-volatile storage medium and stores data such as programs. The non-volatile memory 112 stores, for example, a basic operation PRG 112a, an AR viewing PRG 112b, and a proximity transmission PRG 112c. The basic operation PRG 112a is, for example, a program related to the OS (Operating System). The AR viewing PRG 112b is a program used for viewing AR content. Here, an example with an AR viewing PGR is shown, but it may also include a VR viewing PGR, which is a program used for viewing VR (Virtual Reality) content, or an MR viewing PGR, which is a program used for viewing MR (Mixed Reality) content. The proximity transmission PGR 112c is a program for transmitting proximity information to the outside using proximity communication. These programs may each be separate programs. However, it is not limited to this, and a single program may include and provide the basic operation function, AR viewing function, and proximity transmission function. Alternatively, different divided programs may provide the aforementioned functions. Furthermore, the programs and data stored in these storage devices 110 may be modified or newly added using the communication I / F 160 described later.

[0025] The input I / F 120 is an interface used by the user for operation input, and information that the user wishes to input is input via the input I / F 120. The input I / F 120 may be configured to accept user operation by having the user operate a predetermined button switch 121, for example, a power button, volume buttons, etc. These button switches 121 may be provided on the head-mounted display terminal 1 itself, or they may be provided on a controller or another terminal connected to the head-mounted display terminal 1 via the communication I / F 160. Furthermore, the input I / F 120 may be configured to accept user operation based on the detection of the user's gaze, the detection of the user's hands, the detection of the user's gestures, the detection of a predetermined sound, etc. Also, the input I / F 120 may be configured to accept user operation by having the user operate a pointer on the display. Moreover, the input I / F 120 may be configured to accept user operation by having the user operate an operating device connected via the expansion I / F 171, which will be described later.

[0026] The video input / output device 130 comprises a display 131, an image signal processing unit 132, and an out-camera 133. The display 131 is configured to output video. For example, the display 131 is a transparent display, and a user wearing the head-mounted display terminal 1 can view the outside world through the display 131. The display 131 can display, for example, an image superimposed on the outside world (AR image), a pointer used for user operation, etc. The image signal processing unit 132 is configured for processing image signals and can be configured using, for example, VRAM (video RAM), a dedicated image processing circuit (LSI), an image (video) signal processor, software, etc. The display 131 is not limited to a transparent display and may be an opaque display. The out-camera 133 is configured to capture the outside world and is a camera unit that inputs image data of subjects in the outside world by converting light input from the lens into an electrical signal using an electronic device such as a CCD (Charge Coupled Device) or CMOS (Complementary Metal Oxide Semiconductor) sensor. The rear camera 133 may capture images of the user's hands, gestures, etc., and the head-mounted display terminal 1 may acquire the image data captured by the rear camera 133 as instruction information for input operations, etc., and perform predetermined processing. The display 131 may be an opaque display and may be configured to allow the user to see the outside world by displaying images from cameras that capture the outside world, such as the rear camera 133. Although not shown in the figure, it may also be equipped with an in-camera that captures the user's eyes.

[0027] The audio input / output device 140 includes a speaker 141, an audio signal processing unit 142, and a microphone 143. The speaker 141 outputs, for example, the audio (sound) of the content played during the experience of the head-mounted display terminal 1. The speaker 141 can also inform the user of various notification information by voice. The audio signal processing unit 142 is a configuration used for processing audio signals and can be configured using, for example, a circuit board, an audio signal processor, software, etc. The microphone 143 collects the user's own voice, external sounds, etc., and converts them into audio data. The user may speak voice instructions such as input operations, and the head-mounted display terminal 1 may acquire the audio data collected by the microphone as instruction information such as input operations and execute predetermined processing.

[0028] The sensor group 150 may include, for example, a positioning sensor 151, a geomagnetic sensor 152, a distance measuring sensor 153, an acceleration sensor 154, a gyroscope sensor 155, a gaze detection sensor 156, and the like. The sensor group 150 may also include different types of sensors than those listed above. Furthermore, the sensor group 150 may be configured in which the above sensors are appropriately omitted. In addition, the sensor group 150 may be omitted from the head-mounted display terminal 1, and sensors used for processing may be appropriately provided.

[0029] The positioning sensor 151 is a GNSS sensor. The positioning sensor 151 is a device that receives signals from GNSS (Global Navigation Satellite System) satellites in the sky and is used to detect the current position of the head-mounted display terminal 1. The head-mounted display terminal 1 can use the positioning sensor 151 to detect its own position (in other words, the position of the user wearing the head-mounted display terminal 1).

[0030] The geomagnetic sensor 152 is a sensor that detects the Earth's magnetic field and the direction in which the head-mounted display terminal 1 is facing. By using a three-axis type sensor that detects the geomagnetic field in the vertical direction in addition to the front-back and left-right directions, it is also possible to detect the movement of the head-mounted display terminal 1 by capturing the changes in the geomagnetic field in response to the movement of the head-mounted display terminal 1. This makes it possible to detect the posture of the user wearing the head-mounted display terminal 1.

[0031] The distance measuring sensor 153 is a sensor that measures the distance from the head-mounted display terminal 1 to an object, measures the position of the object, and can capture the shape of an object as a three-dimensional object. Examples of distance measuring sensors 153 include LiDAR (Light Detection and Ranging), which irradiates an object with laser light such as infrared light and measures the scattered light that reflects back; TOF (Time Of Flight) sensors, which measure the reflection time of pulsed light irradiated onto a subject for each pixel; and millimeter-wave radar, which emits millimeter-wave radio waves and captures the reflected waves. Furthermore, the distance measuring sensor 153 may be a sensor that measures based on the angle at which reflected light from an object is received. In other words, the distance measuring sensor 153 may be a triangulation-type sensor. Also, in the head-mounted display terminal 1, the distance measuring sensor 153 may be configured as a stereo camera that performs measurements based on a parallax image.

[0032] Furthermore, measuring instruments that detect the external environment around the user, such as the rear camera 133 and the distance measuring sensor 153 mentioned above, are called external environment detectors. External environment detectors acquire two-dimensional or three-dimensional information including depth, and acquire information about the external environment by simultaneously or in time-divided, receiving information on infrared and visible light emitted and reflected from external objects.

[0033] The acceleration sensor 154 is a sensor that detects acceleration, which is the change in velocity per unit time, and can capture movement, vibration, shock, etc. The acceleration sensor 154 can detect the tilt and direction of the head-mounted display terminal 1 worn by the user. The gyro sensor 155 is a sensor that detects angular velocity in the rotational direction, and can capture the vertical, horizontal, and diagonal orientation. Therefore, the orientation, such as the tilt and direction of the head-mounted display terminal 1, can be detected using the acceleration sensor 154 and the gyro sensor 155.

[0034] The gaze detection sensor 156 may include a left-eye gaze sensor for detecting the gaze of the left eye and a right-eye gaze sensor for detecting the gaze of the right eye, and can detect the movement and direction of the left and right eyes to capture the user's viewpoint, which is the target of their gaze. The head-mounted display terminal 1 may perform eye tracking using the gaze detection sensor 156 and, for example, acquire the detection result as input information. The head-mounted display terminal 1 may then execute predetermined processing based on this input information. For example, the head-mounted display terminal 1 may detect the user's gaze toward a predetermined display on the display and acquire the detection result as input information to execute predetermined processing corresponding to that display. Alternatively, the gaze may be detected from an image of the user's eyes acquired by an in-camera. The in-camera may be a camera that captures images in the visible light range or a camera that captures images in the infrared light range. If it is a camera that captures images in the infrared light range, an infrared LED may be provided as illumination.

[0035] The communication interface 160 is an interface used for communication and includes, for example, a LAN communication interface 161, a short-range wireless communication interface 162, and a telephone network communication interface 163.

[0036] The LAN communication interface 161 is an interface used for communication over a LAN (Local Area Network). The LAN communication interface 161 is connected to the network via an access point (AP) device, for example, by wireless connection such as Wi-Fi (registered trademark), and transmits and receives data with other devices on the network.

[0037] The short-range wireless communication interface 162 is a communication interface for short-range wireless communication with devices within short-range wireless communication range. Short-range wireless communication is performed, for example, using electronic tags, but is not limited to this. If the head-mounted display terminal 1 is near a device and this device is at least capable of wireless communication, short-range wireless communication may be performed using Bluetooth®, IrDA (Infrared Data Association®), Zigbee®, HomeRF (Home Radio Frequency®), or wireless LAN (IEEE 802.11a, IEEE 802.11b, IEEE 802.11g).

[0038] Telephone network communication I / F163 is used for third-generation mobile communication systems (hereinafter referred to as "3G") such as GSM (Registered Trademark) (Global System for Mobile Communications), W-CDMA (Wideband Code Division Multiple Access), CDMA2000, and UMTS (Universal Mobile Telecommunications System), or communication methods called LTE (Long Term Evolution), fourth generation (4G), and fifth generation (5G). It connects to the communication network through base stations using the mobile communication network and transmits and receives information with servers on the communication network.

[0039] The expansion I / F 171 is an interface for extending the functionality of the head-mounted display terminal 1, and is, for example, a USB device connection terminal. Various devices can be connected to the expansion I / F 171, and it is used for purposes such as transmitting and receiving data and charging. For example, an operating device separate from the head-mounted display terminal 1, such as a keyboard, key buttons, or touch keys, may be connected to the expansion I / F 171. The head-mounted display terminal 1 may then acquire the content input to the operating device as instruction information such as input operations and execute predetermined processing. The expansion I / F 171 can be configured to allow connection of devices via wired or wireless connections.

[0040] Timer 172 is a timer that holds the current time in the real world, and holds a time such as Coordinated Universal Time (UTC) as the current time. The timer may be configured as software or may be configured using an RTC (Real Time Clock). Actuator 173 is for conveying physical movement such as vibration to the user. The actuator may operate electrically as long as it generates physical force, or may utilize a piezoelectric element, magnetic force, or hydraulic pressure. Although not shown in FIG. 1, the head-mounted display terminal 1 may be provided with a battery.

[0041] Next, an example of the configuration of the head-mounted display terminal will be described while referring to FIGS. 2A and 2B.

[0042] FIGS. 2A and 2B are diagrams showing an example of the appearance of the head-mounted display terminal.

[0043] In this example, the head-mounted display terminal 1 is a glasses-type HMD (Head-Mounted Display) and has the same configuration as the configuration described above.

[0044] FIG. 2A shows an example of a head-mounted display terminal using a transmissive display. As shown in FIG. 2A, an out-camera is disposed at the front of the glasses-type frame 193. That is, an out-camera 133L is disposed on the left front side of the frame 193, and an out-camera 133R is disposed on the right front side of the frame 193. Further, an out-camera 133F and a distance measuring sensor 153 are disposed on the front center side of the frame. Furthermore, a left display 131L, which is a left transmissive display, and a right display 131R, which is a right transmissive display, are disposed at the front of the frame 193.

[0045] On the left side of the frame 193, a microphone 143 and a left speaker 141L are arranged. On the right side of the frame 193, a right speaker 141R, a sensor group 150, a communication I / F 160, and a control device 192 are arranged. Here, the control device 192 is a device that implements a processor 101, a storage device 110, an input I / F 120, an expansion I / F 171, a timer 172, and an actuator 173. Also, a line-of-sight detection sensor 156 is arranged inside the front part of the frame 193.

[0046] Regarding the head-mounted display terminal 1 shown in FIG. 2, the out cameras (133L, 133R, 133F) are configurations related to the out camera 133. The left display 131L and the right display 131R are configurations related to the display 131. The speakers (141L, 141R) are configurations related to the speaker 141. Although not shown, a battery may be built into the frame 193 or the like, or may be connected to an external battery. Also, it may be connected to other housings or mobile terminals to supply power.

[0047] FIG. 2B is an example of a head-mounted display terminal using a non-transmissive display by a virtual image method. As in this example, a display 131 may be built into the upper part of the glass within the frame 193, and the image of the display 131 may be reflected by the glass so that the user can visually recognize the image as a virtual image. In the example of FIG. 2B, the left display 131L and the right display 131R themselves may be non-transmissive displays. Also, the display 131 is not limited to the upper part of the frame 193, and may be built into other places within the frame 193 as long as the reflected light from the display 131 can be visually recognized. The generation of the virtual image in this example may be performed by providing a prism between the glass and the pupil and controlling the polarization of the light emitted from the display 131, or the light emitted from the display 131 may be reflected or totally reflected within the glass and guided to the pupil.

[0048] In this embodiment, an example is described in which the processor 101 of the head-mounted display terminal 1 performs the processing described later. However, as shown in Figures 3A and 3B, the head-mounted display terminal 1 may communicate with an external information terminal via the communication I / F 160, and the head-mounted display terminal 1 and the external information terminal may work together to perform the processing.

[0049] Figures 3A and 3B show an example of the overall configuration when a head-mounted display terminal is linked with an external information terminal.

[0050] Figure 3A shows an example in which a head-mounted display terminal 1 is connected wirelessly to an external information terminal 2 via a communication interface 160. The external information terminal 2 is, for example, a smartphone, a smartwatch, or a PC.

[0051] Figure 3B also shows an example in which the head-mounted display terminal 1 is connected to an external information terminal 2 via a wired connection through a communication interface 160.

[0052] In this embodiment, the head-mounted display terminal 1 is connected to an external information terminal 2 wirelessly or via a wired connection, and for example, the external information terminal 2 may perform some or all of the processing described later. Furthermore, the external information terminal 2 may supply power to the head-mounted display terminal 1.

[0053] Furthermore, while Figures 3A and 3B illustrate smartphones, smartwatches, and PCs as examples of external information terminals 2, the method is not limited to these. It can be applied to any wearable device worn by the user, such as a ring or bracelet, that has similar functionality.

[0054] Figure 4 is a block diagram showing an example of the configuration of the external information terminal 2.

[0055] An example of the configuration of the external information terminal 2 is the same as the configuration of the head-mounted display terminal 1 in Figure 1; therefore, the same reference numerals are used for the same components, and their explanation is omitted.

[0056] Furthermore, the external information terminal 2 does not need to have all the configurations shown in Figure 4; it only needs to have the configurations necessary to perform some or all of the processing described later. The external information terminal 2 may also have configurations other than those shown in Figure 4. The external information terminal 2 uses these functions to complement the function of the head-mounted display terminal 1 in displaying images to the user.

[0057] Figure 5 illustrates the relationship between the position of an object and its convergence angle when a human perceives a physical object in real space.

[0058] The upper and lower diagrams show the difference in distance to the object and the convergence angle when a user fixates on the object 501. Convergence is the eye movement in which the eyeballs rotate inward when both eyes fixate on an object. This causes the lines of sight from both eyes to intersect, allowing the eye to focus on the object. The angle between the object and the lines of sight of both eyes at this time is called the convergence angle.

[0059] In the upper and lower diagrams, the distance from the user's eye 500 to the object 501 is different. The distance L1 from the user's eye 500 to the object 501, shown in the upper diagram, is larger than the distance L2 from the user's eye 500 to the object 501, shown in the lower diagram. At the same time, the convergence angle θ1, which is the angle at which the lines of sight of both eyes intersect when the user fixates on the object 501, shown in the upper diagram, is smaller than the convergence angle θ2 when the user fixates on the object 501, shown in the lower diagram. Thus, the convergence angle when the user fixates on the object 501 becomes smaller as the object being fixed on is farther from the eye, and larger as it is closer to the eye.

[0060] Furthermore, when a user fixates on an object, the eye adjusts the thickness of its lens to focus on the object. For example, when viewing distant objects, the lens thins to increase the accommodative distance, and when viewing nearby objects, the lens thickens to shorten the accommodative distance, thereby forming an image on the retina.

[0061] Thus, we perceive the depth of an object using convergence, which is the depth perception in binocular vision, and accommodation, which is the depth perception in monocular vision. On the other hand, in binocular vision using a display where the image surface is always fixed, the focus is always fixed on the image surface, so accommodation does not occur, and depth is perceived using only convergence.

[0062] Figure 6 shows the relationship between the distance between the user and the object, and the convergence angle at that distance.

[0063] In Figure 6, the horizontal axis represents the distance from the user's eyes to the object, and the vertical axis represents the convergence angle between the object and the user's binocular lines of sight.

[0064] Plot 601 shows the convergence angle at each distance, and plot 602 shows the relative convergence angle based on the convergence angle when the object is located 2m from the user's eye.

[0065] As shown in Figure 6, the further the object is from the user's eye, the more the convergence angle asymptotically approaches 0 degrees. When the object is far away, the change in the convergence angle with respect to distance is small, while when the object is near the user, the change in the convergence angle with respect to distance is large. As shown in plot 602, with 2m as the reference distance, the change in the convergence angle when the object moves further away is only about 1 degree, whereas when the object moves closer, the change in the convergence angle is several degrees to tens of degrees.

[0066] Furthermore, for distant objects, the angle of light reaching the retina is limited, resulting in a deeper depth of field, similar to having a narrow aperture. Conversely, for objects close to the user, the angle of light rays reaching the retina is wider, resulting in a shallower depth of field, similar to having a wide aperture.

[0067] However, in head-mounted display devices such as AR glasses, VR glasses, and head-mounted displays that can display images on separate displays for the left and right eyes, the position of the image the user sees depends on the position of the display, and the distance between the user's eyes and the image surface is often fixed. In such devices, the sense of depth of the image is represented only by changes in convergence, without changing the focal adjustment distance, which can sometimes feel unnatural.

[0068] Figure 7 is a diagram illustrating the vergence-accommodation contradiction.

[0069] Vergence-accommodation conflict (VAC) is a condition in which the convergence distance (the distance from the eyeball to the point where the lines of sight from both eyes intersect) does not match the accommodation distance (the distance to the object that is focused on by accommodation).

[0070] The upper diagram shows an example where the focal length L3 and image distance L4 coincide, while the lower diagram shows an example where the focal length L3 and image distance L4 do not coincide.

[0071] In the head-mounted display terminal 1, the image is always displayed at a constant distance from the eye 700. In the example in Figure 7, the image surface of the head-mounted display terminal 1 is always displayed at layer 701, which is L4 away from the user's eye.

[0072] In the diagram above, the object that the user is looking at, such as another device or an external object, is located in layer 702 at the same distance as layer 701 where the image surface exists. In this case, when the user looks at the object, the focal length L3 is guided to the distance corresponding to layer 702 where the object is located. On the other hand, the image distance L4 is always fixed at the display position of the image. Therefore, the focal length L4 and the image distance L3 coincide. As a result, both layer 701 where the image surface exists and layer 702 where the object being looked at exists are present at the focal length L3, and both are in focus.

[0073] On the other hand, the diagram below shows the case where an object the user is looking at, such as another device or an external object, is located at a distance closer than the layer 701 where the image surface exists. In this case, the focal length L3 is guided to a distance that matches the layer 702 where the object is located. Meanwhile, the image distance L4 is always fixed at the display position of the image. Therefore, the focal length L4 and the image distance L3 do not coincide.

[0074] When the object the user is focusing on, such as another device or an external object, is far from the image plane, the depth of field increases, allowing both to be in focus. However, when the object the user is focusing on is close to the image plane, the depth of field decreases, making it possible to focus on either the image or the object, but not both. As a result, the object that is not in focus appears blurred.

[0075] Furthermore, the convergence angle θ1 with respect to layer 701, where the image surface exists, is different from the convergence angle θ2 with respect to layer 702, where the object being viewed exists. When the head-mounted display terminal 1 displays images that take the convergence angle θ1 into account on the left and right displays, if the user rotates their eyeballs to a state where the convergence angle θ2 is reached, the images on the left and right displays will not fuse, and will be perceived by the user as two separate images (double images).

[0076] While it is common in the real world for multiple objects to be out of focus simultaneously, the phenomenon where information perceived by the left and right eyes is not fused and is instead seen as two separate images is a problem that occurs with head-mounted display devices that utilize binocular vision.

[0077] Furthermore, as shown in Figure 6, the change in convergence angle becomes larger in areas close to the user, resulting in a larger difference between the convergence angle relative to the image and the convergence angle relative to the object being gazed upon.

[0078] Due to this vergence adjustment inconsistency, the left and right images on the head-mounted display terminal 1 are not fused, resulting in the phenomenon of double images. This phenomenon becomes more pronounced as the difference in vergence angles increases.

[0079] Figures 8A and 8B illustrate the occurrence of double images in a head-mounted display terminal due to vergence accommodation discrepancies.

[0080] Figure 8A shows an example where the focal plane 801 and the image surface 802 of the head-mounted display terminal coincide, while Figure 8B shows an example where the focal plane 801 and the image surface 802 do not coincide.

[0081] The focal plane 801 is the plane on which the object the user is fixated on, for example, the screen of a mobile information terminal within the field of view, is located, and it changes depending on the distance to the object the user is fixated on.

[0082] The image surface 802 is the surface on which the image from the head-mounted display terminal 1 is displayed. The image surface 802 is designed for each device and is displayed at a certain distance from the user's eyes.

[0083] In Figure 8A, since the focal plane 801 and the image plane 802 coincide, the user can focus on and view both the object and the image simultaneously.

[0084] Figure 8B shows the case where the object moves near the user. In this case, the user focuses and converges on the nearby object. However, the adjustment distance and convergence angle when the user fixates on the object no longer match the convergence and accommodation presented by the head-mounted display terminal's image. As a result, the image displayed on the head-mounted display terminal 1 becomes blurry due to lack of focus, and a double image is generated due to poor fusion of the left and right displays caused by the mismatch in convergence angles.

[0085] Figures 9A and 9B illustrate the relationship between distance and VAC quantity.

[0086] The VAC (Vergence-Accommodation Conflict) can be quantified by comparing the optical power required to focus on an object at the convergence distance with the optical power required to focus on an object at the accommodation distance.

[0087] Here, optical power is defined as the reciprocal of the adjustment distance of the optical system, measured in meters, and expressed as diopters (D or m). -1 It is expressed as follows: Therefore, the VAC quantity is expressed as the difference between the reciprocal of the congestion distance and the reciprocal of the adjustment distance.

[0088] In Figure 9A, the depth presentation distance due to convergence is shown on the horizontal axis, and the VAC amount is shown on the vertical axis.

[0089] Plots 901, 902, and 903 show the change in VAC quantity when the depth presentation distance due to congestion is changed, with image planes formed at distances of 1 m, 2 m, and 3 m, respectively.

[0090] When the VAC (Vital Acceleration) level increases, users experience discomfort and fatigue due to misalignment between convergence and accommodation, making fusion difficult. When the misalignment between convergence and accommodation is small, the images presented to the left and right eyes can be fused. However, when the misalignment is large, the images from the left and right eyes are not fused and are perceived as double images.

[0091] Generally, a VAC level of approximately 0.25D is considered a comfortable viewing range, and a VAC level of approximately 0.5D is considered an acceptable viewing range. However, as is clear from Figure 9A, regardless of the distance at which the image surface is set, the difference between the position of the image surface and the depth presentation distance due to convergence becomes large, resulting in a high VAC level. Therefore, when the image surface is fixed, it is not possible to keep the VAC level low in all areas. Consequently, fusion defects occur in head-mounted display terminals that use a binocular viewing system.

[0092] In Figure 9B, the distance to the image plane is shown on the horizontal axis, and the adjustment distance is shown on the vertical axis.

[0093] Plots 904 and 905 show the change in accommodation distance when the distance to the image plane is changed, when the VAC amount is 0.25D and 0.5D, respectively. As shown in this figure, by detecting the distance of the object being gazed at in a timely manner and adjusting the convergence angle when presenting the image to the accommodation distance of the object being gazed at, the VAC amount can be kept low at all times. A specific example is that by detecting the distance of the object being gazed at using the out-camera 133 of the head-mounted display terminal 1, the distance measuring sensor 155, the gaze detection sensor 156, etc., the convergence angle of the displayed image can be controlled and corrected so that the images of the left and right eyes are always fused.

[0094] As shown in Figures 9A and 9B, the amount of VAC increases particularly in the vicinity of the user. The area from 50 cm to 1 m from the user roughly corresponds to the distance at which one performs tasks on a PC, while the area below 50 cm corresponds to the distance at which one performs tasks using a device held by the user, such as a tablet or smartphone. Many general AR glasses are designed to display images several meters away, assuming that the user will see the image and the surrounding environment simultaneously. Therefore, when viewing close-range devices such as PCs and smartphones, it is difficult to view the image on the AR glasses at the same time. In particular, when fixating on a smartphone, the image on the AR glasses not only fails to come into focus, but the convergence angle presented by the image and the convergence angle when viewing the object being fixed on are significantly different, resulting in a phenomenon where the images presented to both eyes do not fuse and appear double.

[0095] Figures 10A and 10B illustrate an example of mode change from binocular mode to monocular mode in this embodiment.

[0096] As shown in Figure 10A, when a user is viewing an object 1002 that is closer than the image 1001 on the head-mounted display terminal 1, in binocular mode, the focal adjustment distance for the image 1001 and the convergence angle presented by the image do not match the focal adjustment distance and convergence angle for the object 1002, resulting in a blurred double image of the image 1001.

[0097] Therefore, as shown in Figure 10B, the head-mounted display terminal 1 changes mode from a binocular mode, which displays images on both the left display 131L and the right display 131R, to a monocular mode, which displays images on only one of the left display 131L or the right display 131R. This allows the user to view the image 1001 without experiencing double images. Double images occur when the images displayed on the left display 131L and the right display 131R do not fuse together; therefore, an image displayed on only one of the displays is not perceived as double. It should be noted that while it is difficult for the user to simultaneously focus on the image 1001 and the object 1002, which are at different distances, the user can focus on either the image 1001 or the object 1002 and view them. The phenomenon of being unable to simultaneously focus on objects at different distances is a phenomenon that occurs daily in the real world and is not a problem unique to head-mounted display terminals.

[0098] Whether the user is viewing the object 1002 can be determined by the head-mounted display terminal 1 detecting that the object 1002 is closer than the image 1001, or by recognizing that it is close based on the size of the object 1002 displayed on the screen. Specifically, the distance between the user and the object can be detected using the head-mounted display terminal 1's rear camera 133 or distance measuring sensor 155, or the object that the user is fixated on can be detected using a gaze detection sensor 156.

[0099] Figures 11A and 11B illustrate another example of mode change from binocular mode to monocular mode in this embodiment.

[0100] Figure 11A shows an example of image display that, like the side-view mode, moves the image 1101 to the periphery of the field of view in advance, making the surrounding environment, such as the central region of the field of view, easier to see (making the surrounding environment the primary focus of viewing). If the user is viewing an object 1102 that is closer than the image 1101 on the head-mounted display terminal 1, in binocular mode, the convergence distance presented by the accommodation distance and convergence angle between the image 1101 and the object 1102 does not match, resulting in a blurred double image of the image 1101.

[0101] When the device is set to side-view mode, it is assumed that the user will see the image and the actual object simultaneously. Therefore, as shown in Figure 11B, the head-mounted display terminal 1 changes its mode from binocular mode, which displays the image on both the left display 131L and the right display 131R, to monocular mode, which displays the image on only one of the left display 131L or the right display 131R. This allows the user to see the image 1101 without experiencing a double image. While it is difficult for the user to focus on both the image 1101 and the object 1102 simultaneously, they can focus on either the image 1101 or the object 1102. In this way, by pre-setting whether to use binocular mode or monocular mode for each display mode of the head-mounted display terminal, the occurrence of double images can be prevented regardless of the object detection result.

[0102] Furthermore, monocular mode is not limited to a mode in which no image is displayed on one of the displays, but also includes a mode in which the same image object is not displayed on both the left and right displays. A double image is a phenomenon in which images that should be displayed in the same position are displayed with a gap, and images that are displayed in positions that are already far apart do not necessarily overlap. Therefore, in monocular mode, different image objects may be displayed on the left and right displays, as long as the same image object is not displayed on both displays.

[0103] Figure 12 is a flowchart showing an example of a mode change process performed by the head-mounted display terminal 1. This process is achieved by the processor 101 of the head-mounted display terminal 1 executing a program stored in the non-volatile memory 112. In this embodiment, an example is described in which the head-mounted display terminal 1 performs this process, but the external information terminal 2 may also perform part or all of this process.

[0104] When the head-mounted display terminal 1 starts processing (S0), it performs a mode transition condition determination (S1) to determine whether to transition from a mode in which images are displayed on both the left and right displays (binocular mode) to a mode in which images are displayed on only one of the left or right displays (monocular mode).

[0105] In the mode transition condition determination (S1), if there is a setting to select the use of monocular vision, it is first checked whether this setting is enabled.

[0106] Next, the system checks whether there are any real-world objects that should be focused on, other than the image displayed within the field of view. It also determines whether those objects are unsuitable for simultaneous viewing with the image.

[0107] The processing example for determining the mode transition condition (S1) will be explained in detail later with reference to Figures 13 to 19.

[0108] If the mode transition condition determination in S1 determines that the video is not suitable for simultaneous viewing with an object in real space, the mode is changed and the video is adjusted (S2).

[0109] When a user is fixated on an object in real space, the eye's accommodation and convergence distances may differ from those of the displayed image. A mismatch in accommodation distance manifests as a focus error, while a mismatch in convergence distance manifests as a double image due to fusion difficulty. As an example of adjustment, the mode is changed from binocular mode to monocular mode, and the image that was displayed on both displays is shown on only one display.

[0110] The process of changing modes (S2) will be explained in detail later with reference to Figure 20.

[0111] Next, a mode transition cancellation condition determination is performed (S3) to determine whether the mode transition conditions in S2 are no longer met.

[0112] The process for determining the mode transition release condition (S3) will be explained in detail later with reference to Figures 21-26.

[0113] If the mode transition cancellation condition check in S3 determines that the mode transition condition is no longer met, the mode is changed (S4) and the process is terminated (S5).

[0114] The process of changing modes (S4) will be explained in detail later with reference to Figure 27.

[0115] Figure 13 is a flowchart showing a first example of the mode transition condition determination (S1).

[0116] If the process transitions from the previous step (S0) to the mode transition condition determination (S1), for example, as shown in Figure 13, the head-mounted display terminal 1 performs setting confirmation (S10) and actual object detection (S11). Then the process transitions to the next step (S2).

[0117] Figure 14 is a flowchart showing a second example of the mode transition condition determination (S1).

[0118] If the process transitions from the previous step (S0) to the mode transition condition determination (S1), for example, as shown in Figure 14, the head-mounted display terminal 1 performs setting confirmation (S10), actual object detection (S11), and actual object determination (S12). Then the process transitions to the next step (S2).

[0119] Figure 15 is a flowchart showing a third example of the mode transition condition determination (S1).

[0120] If the process transitions from the previous step (S0) to the mode transition condition determination (S1), for example, as shown in Figure 15, the head-mounted display terminal 1 performs setting confirmation (S10) and display mode confirmation (S13). Then the process transitions to the next step (S2).

[0121] The following shows examples of each process, but this embodiment may apply one example or a combination of multiple examples.

[0122] Figure 16 is a table showing an example of the setting confirmation process (S10).

[0123] In Example 1, the head-mounted display terminal 1 checks its settings and proceeds to the next step if the following condition (1) is met: (1) The head-mounted display terminal 1 is capable of switching between monocular and binocular modes. If the device does not support switching between monocular and binocular modes, or if switching is prohibited, the mode cannot be switched at all. In that case, the process is terminated.

[0124] In Example 2, the head-mounted display terminal 1 checks the settings conditions of the head-mounted display terminal 1, and if the following condition (2) is met, it proceeds to the next step. (2) The head-mounted display terminal 1 is set to enable monocular mode. Switching between monocular mode and binocular mode is possible, and the user may be able to choose whether to use mode switching through the settings screen or a message to the user. In such cases, it is checked whether the user has selected to switch modes, and a decision is made whether to continue or terminate the process.

[0125] In Example 3, the head-mounted display terminal 1 checks its settings and proceeds to the next step if the following condition (3) is met. (3) The head-mounted display terminal 1 is not already started in monocular mode. Also, if monocular mode is selected and started from the beginning instead of binocular mode, further mode switching is not possible. The system refers to the currently running display mode and decides whether to continue or terminate the process.

[0126] In Example 4, the head-mounted display terminal 1 checks the settings conditions or operating mode of the head-mounted display terminal 1, and if none of the above conditions (1) to (3) are met, it proceeds to S5 and terminates the process.

[0127] Figure 17 is a table showing an example of the process for detecting a real object (S11).

[0128] In Example 1, the head-mounted display terminal 1 proceeds to the next step if the following condition (1) is met: (1) The head-mounted display terminal 1's rear camera 133 detects a real-world object within its field of view. If a real-world object is captured in the image taken by the head-mounted display terminal 1's rear camera 133, that object may be impairing the visibility of the image. Therefore, if an object is detected, the system proceeds to the next step to determine the detected object. Furthermore, if the head-mounted display terminal 1 is equipped with multiple cameras, it is possible to acquire information including the distance to the object.

[0129] In Example 2, the head-mounted display terminal 1 proceeds to the next step if the following condition (2) is met: (2) When the distance sensor 153 of the head-mounted display terminal 1 detects an object in real space within its field of view. If the distance sensor 153 of the head-mounted display terminal 1 detects an object in the area scanned by the user using infrared light or the like, that object may be impairing the visibility of the image. Therefore, if an object is detected, the system proceeds to the next step to make a judgment regarding the detected object. Furthermore, if the head-mounted display terminal 1 is also equipped with a camera, it is possible to acquire information about the object along with the camera's image.

[0130] In Example 3, the head-mounted display terminal 1 proceeds to the next step if the following condition (3) is met: (3) The user confirms an object in real space within the field of view of the head-mounted display terminal 1, inputs that information via the input I / F 120, and the head-mounted display terminal 1 accepts the user's input. In Examples 1 and 2, the head-mounted display terminal 1 detects an object that may be impairing the user's visibility based on the information it has detected, but the user may also confirm the object that is impairing their visibility. Processing based on detection information from cameras and sensors is performed assuming the user's state, but it does not necessarily reflect the state of each individual user. Therefore, by allowing the user to specify the object that is impairing their visibility, it becomes possible to perform processing that is more tailored to the state of each individual user.

[0131] In Example 4, if none of the above conditions (1) to (3) are met, the head-mounted display terminal 1 continues S11 until the user inputs a termination command. If the user inputs a termination command, the process proceeds to S5 and terminates.

[0132] Figure 18 is a table showing an example of the process for determining the actual object (S12). This process determines whether the detected object is a factor that causes VAC. In order to avoid frequent switching of display modes, the process may be configured to proceed only when the conditions described in the example below occur continuously for a certain period of time, and the system is deemed to have met the conditions.

[0133] In Example 1, the head-mounted display terminal 1 proceeds to the next step if the following condition (1) is met: (1) The distance between the head-mounted display terminal 1 and the detected object is smaller than a specified value. When the distance between the head-mounted display terminal 1 and the object is close, fusion defects may occur. This phenomenon is particularly noticeable when the object is located close to the head-mounted display terminal 1 relative to the display distance of the image. By setting a specified distance to the object when switching display modes, the occurrence of double images due to fusion defects can be predicted.

[0134] In Example 2, the head-mounted display terminal 1 proceeds to the next step if the following condition (2) is met: (2) The size of the actual object detected by the head-mounted display terminal 1 within its field of view is larger than a specified value. For example, the size of the actual object detected by the head-mounted display terminal 1 on the left display 131L and the right display 131R is larger than the size of the image displayed on the left display 131L and the right display 131R. If the size of the object within the user's field of view is large, it attracts the user's attention, and the user will stare at the object. By defining the size of the object within the field of view, it is possible to predict the occurrence of double images due to fusion failure and switch the display mode. This defined size of the object may be changed depending on the size of the displayed image. If the image size is large, the defined size may be increased, and if the image size is small, the defined size may be decreased. Also, if the object and the image overlap within the field of view, and the image is displayed in front of the object, the area of ​​the object that overlaps with the image does not need to be considered in determining the size.

[0135] In Example 3, the head-mounted display terminal 1 proceeds to the next step if the following condition (3) is met: (3) The total area of ​​the actual objects detected by the head-mounted display terminal 1 within the field of view is greater than a specified value. When there are multiple objects detected by the head-mounted display terminal 1, even if each object is small, the multiple objects may cover a large portion of the field of view. By defining the total area of ​​multiple objects, the occurrence of double images due to fusion defects can be predicted, and the display mode can be switched. In this case, if the objects overlap within the field of view, the area of ​​the overlapping objects does not need to be considered when determining the size.

[0136] In Example 4, the head-mounted display terminal 1 proceeds to the next step if the following condition (4) is met: (4) The actual object detected by the head-mounted display terminal 1 is located within a defined area in the video display area of ​​the head-mounted display terminal 1, for example, in the center of the video display area. The human eye has a central field of vision and a peripheral field of vision. The central field of vision has high resolution and can distinguish fine differences, but the peripheral field of vision has lower resolution and can only distinguish rough movements. Therefore, fixation on an object is often done with the central field of vision. The center of the video display area of ​​the head-mounted display terminal 1 corresponds to the central field of vision of the user wearing the head-mounted display terminal 1. Therefore, since the center of the video display area is the area where the object being fixed on exists, by determining whether the object has entered that area, it is possible to predict the occurrence of double images due to fusion failure and switch the display mode. Alternatively, the size of the object's intrusion within the defined area may be defined and used for the determination.

[0137] In Example 5, the head-mounted display terminal 1 proceeds to the next step if the following condition (5) is met: (5) The actual object detected by the head-mounted display terminal 1 is located closer to the image display position of the head-mounted display terminal 1. When the distance between the head-mounted display terminal 1 and the actual object differs from the distance between the head-mounted display terminal 1 and the image display distance, fusion defects may occur. This phenomenon is particularly pronounced when the actual object is located close to the head-mounted display terminal 1 relative to the image display distance. By defining a specified distance of the object relative to the image display position when switching display modes, the occurrence of double images due to fusion defects can be predicted.

[0138] In Example 6, the head-mounted display terminal 1 proceeds to the next step if the following condition (6) is met: (6) The difference between the position of the actual object detected by the head-mounted display terminal 1 and the image display position of the head-mounted display terminal 1 exceeds 0.25D. As described in the explanation of Figure 9A, a VAC amount of up to approximately 0.25D is considered to be within the comfortable viewing range. Therefore, by calculating the VAC amount between the detected object and the image display surface and comparing it with a VAC amount of 0.25D, it is possible to predict the occurrence of double images due to fusion defects.

[0139] In Example 7, the head-mounted display terminal 1 proceeds to the next step if the following condition (7) is met: (7) The difference between the position of the actual object detected by the head-mounted display terminal 1 and the image display position of the head-mounted display terminal 1 exceeds 0.5D. Similar to Example 6, a VAC amount of approximately 0.5D is considered to be within the acceptable viewing range. Therefore, by calculating the VAC amount between the detected object and the image display surface and comparing it with a VAC amount of 0.5D, it is possible to predict the occurrence of double images due to fusion defects.

[0140] In Example 8, the head-mounted display terminal 1 proceeds to the next step if the following condition (8) is met: (8) The gaze detection sensor 156 detects that the user is gazing at the real object detected by the head-mounted display terminal 1. Even if an object is in the field of view of the user wearing the head-mounted display terminal 1, detecting the object alone does not necessarily mean that the user is gazing at it. If the gaze detection sensor 156 confirms that the user's gaze direction is towards the detected object, it can be determined that the user is gazing at the object, and the occurrence of double images due to fusion defects can be predicted.

[0141] In Example 9, the head-mounted display terminal 1 proceeds to the next step if the following condition (9) is met: (9) The user selects to activate monocular display mode, and the head-mounted display terminal 1 accepts the user's selection. In this case, since the user has selected monocular mode, monocular mode is implemented regardless of whether an object is detected or not. The same applies if monocular mode is already activated.

[0142] In Example 10, if none of the above conditions (1) to (9) are met, the head-mounted display terminal 1 continues S11 until the user inputs a termination command. If the user inputs a termination command, the process proceeds to S5 and terminates.

[0143] Furthermore, the processes described in Examples 1 through 10 may be performed individually, or multiple examples may be combined to perform more detailed processing.

[0144] Figure 19 is a table showing an example of the display mode confirmation process (S13). This process confirms which display mode the head-mounted display terminal 1 is using. In particular, it is used when a display restriction mode is being used, which displays video only in a portion of the video display area, allowing the surrounding environment to be seen in the other areas, rather than a mode that displays video across the entire video display area.

[0145] In Example 1, the head-mounted display terminal 1 proceeds to the next step if the following condition (1) is met: (1) The head-mounted display terminal 1 has multiple display modes, and the mode that reduces the size of the display is activated.

[0146] In Example 2, the head-mounted display terminal 1 proceeds to the next step if the following condition (2) is met: (2) The head-mounted display terminal 1 has multiple display modes, and a mode that limits the display area, for example, a mode that displays only in the four corners of the display area, is activated.

[0147] In Example 3, the head-mounted display terminal 1 proceeds to the next step if the following condition (3) is met: (3) The head-mounted display terminal 1 has multiple display modes, and a mode that restricts display to an area including the center of the field of view is activated.

[0148] In Example 4, the head-mounted display terminal 1 proceeds to the next step if the following condition (4) is met: (4) The head-mounted display terminal 1 has multiple display modes, and a mode that restricts the display to an area of ​​50% or more of the field of view is activated. When the display modes (1) to (4) above are activated, the display is designed to allow viewing of objects other than the image. In such displays, it is possible to gaze at objects other than the image. When displaying in these modes, it is possible to switch to monocular mode in advance, regardless of whether an object has been detected, in order to suppress the occurrence of fusion defects.

[0149] In Example 5, if none of the above conditions (1) to (4) are met, the head-mounted display terminal 1 continues S13 until the user inputs a termination command. If the user inputs a termination command, the process proceeds to S5 and terminates.

[0150] The above describes an example of determining the display mode, but you may also determine whether to switch between monocular and binocular modes based on the displayed content. For example, use monocular mode when displaying 2D content and binocular mode when displaying 3D content. 3D content may display different images on the left and right displays to utilize the user's parallax. Therefore, a sufficient stereoscopic effect cannot be obtained in monocular mode. On the other hand, 2D content displays the same image on both the left and right displays, so using monocular mode does not affect the user.

[0151] Figure 20 is a table showing an example of the mode change process (S2). In this process, the display mode is changed based on the result of the mode transition condition determination (S1) in the preceding stage.

[0152] In Example 1, the head-mounted display terminal 1 displays an image only on the predetermined display (either the left display 131L or the right display 131R), and does not display an image on the other display. This example includes cases where the user selects a predetermined display in the settings screen, or where the system displays the image on a predetermined display.

[0153] In Example 2, the head-mounted display terminal 1 stores past usage history in the storage device 110 and displays an image only on the display that was used during the previous monocular mode display (left display 131L and right display 131R), while not displaying an image on the other display. This example records the display information used during monocular mode and describes the process for using the same display in monocular mode the next time.

[0154] In Example 3, the head-mounted display terminal 1 stores past usage history in the storage device 110 and displays an image only on the left display 131L and the right display 131R that was not used during the previous monocular mode display, while not displaying an image on the other display. In this example, by using the left and right displays alternately as the displays used in monocular mode, it is possible to balance the fatigue of both eyes. It is also possible to equalize the usage time of the displays.

[0155] In Example 4, the head-mounted display terminal 1 stores past usage history in the storage device 110, and when a difference of more than a specified value occurs in the usage time of the left display 131L and the right display 131R, it switches the display used for monocular display. In this example, by recording the device usage time of the left and right displays, it is possible to manage the degradation of the left and right displays so that it does not occur concentrated on only one side. If a significant difference occurs in the operating time of the left and right displays, the display used is switched in order to average them out.

[0156] In Example 5, the head-mounted display terminal 1 displays an image only on the display selected by the user from the left display 131L and the right display 131R, and does not display an image on the other display. In this example, only the display preferred by the user can be used in monocular mode.

[0157] In Example 6, the head-mounted display terminal 1 displays a message to switch to monocular mode on the left display 131L or the right display 131R, or on both displays, and switches to monocular mode. This allows the user to know that they are switching to monocular mode before the image on one of the displays goes blank.

[0158] In Example 7, the head-mounted display terminal 1 displays a message to switch to monocular mode on the left display 131L or the right display 131R, and switches to monocular mode after obtaining user approval. This allows the user to know that the device is switching to monocular mode before the image on one of the displays goes blank, and also allows them to pay attention to the switch to monocular mode.

[0159] In Example 8, the head-mounted display terminal 1 displays an image only on the display on the user's dominant eye side, either the left display 131L or the right display 131R, and does not display an image on the other display side. The user's dominant eye is set in advance, for example, via the input I / F 120, and stored in the storage device 110. Even when a user uses binocular vision, they primarily use the eye whose brain preferentially processes information to see objects. This preferentially processing eye is called the dominant eye. Since the user primarily sees objects based on information from their dominant eye, image information may also be displayed on the display corresponding to the dominant eye side.

[0160] In Example 9, the head-mounted display terminal 1 displays an image only on the display on the side that is not the user's dominant eye, either the left display 131L or the right display 131R, and does not display an image on the other display. Continuously using the dominant eye frequently can place excessive strain on that eye, potentially causing physical ailments such as headaches and eye strain. Therefore, to alleviate eye strain, it is acceptable to display an image only on the display on the side that is not the dominant eye.

[0161] In Example 10, the head-mounted display terminal 1 does not display any images on either the left display 131L or the right display 131R. Furthermore, if there is concern about VAC generation, the image display may be temporarily suspended. This makes it possible to prevent deterioration of visibility associated with images.

[0162] Figure 21 is a flowchart showing a first example of the mode transition release condition determination (S3).

[0163] If the process transitions from the previous step (S2) to the mode transition release condition determination (S3), for example, as shown in Figure 21, the head-mounted display terminal 1 performs actual object detection (S30). Then the process transitions to the next step (S4).

[0164] Figure 22 is a flowchart showing a second example of the mode transition release condition determination (S3).

[0165] If the process transitions from the previous step (S2) to the mode transition release condition determination (S3), for example, as shown in Figure 22, the head-mounted display terminal 1 performs actual object detection (S30) and actual object determination (S31). Then the process transitions to the next step (S4).

[0166] Figure 23 is a flowchart showing a third example of the mode transition release condition determination (S3).

[0167] If the process transitions from the previous step (S2) to the mode transition cancellation condition determination (S3), for example, as shown in Figure 23, the head-mounted display terminal 1 performs a display mode confirmation (S32). Then the process transitions to the next step (S4).

[0168] Figure 24 is a table showing an example of the process for detecting a real object (S30).

[0169] In Example 1, the head-mounted display terminal 1 proceeds to S4 and changes to binocular mode when the following condition (1) is met: (1) The head-mounted display terminal 1's rear camera 133 does not detect any real-world objects within its field of view. If a real-world object is no longer visible in the image captured by the head-mounted display terminal 1's rear camera 133, it may be determined that the visibility of the image is not impaired by the object, and the mode may be changed to binocular mode. This process may be performed not only when the object disappears from the camera's image, but also when the object is located more than a certain distance away from the user in the image, when its size in the image becomes smaller than a certain size, or when it is detected that the user is not looking at an object in the image.

[0170] In Example 2, the head-mounted display terminal 1 proceeds to S4 and changes to binocular mode when the following condition (2) is met. (2) When the distance sensor 153 of the head-mounted display terminal 1 does not detect an object in real space within the field of view. When the distance sensor 153 of the head-mounted display terminal 1 no longer detects an object in the area scanned by the user using infrared light or the like, it may be determined that the visibility of the image is not impaired by an object and the mode may be changed to binocular mode. This process may be performed not only when an object is no longer detected within the measurement area of ​​the distance sensor, but also when the user moves more than a certain distance away from the user within the measurement area, when the size of the object within the measurement area becomes smaller than a certain size, or when it is detected that the user is not looking at an object detected by the distance sensor.

[0171] In Example 3, the head-mounted display terminal 1 proceeds to S4 and changes to binocular mode when the following condition (3) is met: (3) The user inputs a monocular mode termination instruction via the input I / F 120, and the head-mounted display terminal 1 accepts the user's input. If the user determines that their visibility is not impaired, it means that they are not fixating on objects other than the image, or the VAC amount with objects other than the image is small, so the mode may be changed to binocular mode.

[0172] In Example 4, if none of the above conditions (1) to (3) are met, the head-mounted display terminal 1 repeatedly executes S3 and continues in monocular mode (first processing example in Figure 21). In this case, it may be determined that there is an object that is being continuously viewed and is interfering with the fusion of the image, and the monocular mode may be continued.

[0173] In Example 5, the head-mounted display terminal 1 proceeds to S31 if none of the above conditions (1) to (3) are met (second processing example in Figure 22). In this case, there is a continuous object within the user's field of view that causes fixation, and it is necessary to determine whether the object is causing fixation.

[0174] Figure 25 is a table showing an example of the process for determining the actual object (S31). For these determinations, in order to avoid frequent switching of display modes, the process may be configured to proceed only when the conditions described in the example below occur continuously for a certain period of time, in which case the condition is deemed to be met.

[0175] In Example 1, the head-mounted display terminal 1 proceeds to S4 and changes to binocular mode when the following condition (1) is met: (1) The distance between the head-mounted display terminal 1 and the detected object is greater than or equal to a specified value. When the distance between the head-mounted display terminal 1 and the object is far, the occurrence of fusion defects is suppressed. For this reason, if it is determined that the distance between the object and the user is greater than or equal to a certain specified value, the system may switch to binocular mode.

[0176] In Example 2, the head-mounted display terminal 1 proceeds to S4 and changes to binocular mode when the following condition (2) is met: (2) The size of the actual object detected by the head-mounted display terminal 1 within the field of view is less than or equal to a specified value. For example, the size of the actual object detected by the head-mounted display terminal 1 on the left display 131L and the right display 131R is less than or equal to the size of the image displayed on the left display 131L and the right display 131R. When the size of the object within the user's field of view is small, the effect of attracting the user's attention is small, and the user finds it difficult to focus on the object. The size of the object within the field of view is defined, and if it is determined that the size of the object is smaller than or equal to the defined value, the system may switch to binocular mode. This defined size of the object may be changed depending on the size of the displayed image. If the image size is large, the defined size may be increased, and if the image size is small, the defined size may be decreased. Furthermore, if the object and the image overlap within the field of view, and the image is displayed in front of the object, the area of ​​the object that overlaps with the image does not need to be considered when determining its size.

[0177] In Example 3, the head-mounted display terminal 1 will proceed to S4 and change to binocular mode if the following condition (3) is met: (3) The total area of ​​the actual objects detected by the head-mounted display terminal 1 within the field of view is less than or equal to a specified value. When there are multiple objects detected by the head-mounted display terminal 1, even if each individual object is small, multiple objects may cover a large portion of the field of view. The total area of ​​multiple objects is defined, and if it is determined that the total size of the objects is less than or equal to a specified value, the system may switch to binocular mode. In this case, if the objects overlap within the field of view, the overlapping areas do not need to be considered in determining the size.

[0178] In Example 4, the head-mounted display terminal 1 proceeds to S4 and changes to binocular mode if the following condition (4) is met: (4) The actual object detected by the head-mounted display terminal 1 is not located within a defined area in the image display area of ​​the head-mounted display terminal 1, for example, in the center of the display area. Since the center of the image display area is the area where the object being viewed is located, the intrusion of the object into that area is determined, and if it is determined that the object is not located within the defined area, the system may switch to binocular mode.

[0179] In Example 5, the head-mounted display terminal 1 proceeds to S4 and changes to binocular mode when the following condition (5) is met: (5) The actual object detected by the head-mounted display terminal 1 is located farther away from the image display position of the head-mounted display terminal 1. When the actual object is located farther away from the head-mounted display terminal 1 and the display distance of the image, the occurrence of fusion defects is suppressed. Therefore, a specified distance of the object relative to the image display position can be set, and if it is determined that the object is located farther away from the display distance of the image, the system may switch to binocular mode.

[0180] In Example 6, the head-mounted display terminal 1 proceeds to S4 and changes to binocular mode when the following condition (6) is met: (6) The difference between the position of the actual object detected by the head-mounted display terminal 1 and the image display position of the head-mounted display terminal 1 is within 0.25D. When the VAC amount is within approximately 0.25D, it is considered a comfortable viewing range without fusion defects. Therefore, the VAC amount between the detected object and the image display surface is calculated, and if it is determined to be less than 0.25D, the system may switch to binocular mode.

[0181] In Example 7, the head-mounted display terminal 1 proceeds to S4 and changes to binocular mode if the following condition (7) is met: (7) The difference between the position of the actual object detected by the head-mounted display terminal 1 and the image display position of the head-mounted display terminal 1 is within 0.5D. When the VAC amount is within approximately 0.5D, it is considered an acceptable viewing range even if fusion defects occur. Therefore, the VAC amount between the detected object and the image display surface is calculated, and if it is determined to be less than 0.5D, the system may switch to binocular mode.

[0182] In Example 8, the head-mounted display terminal 1 proceeds to S4 and changes to binocular mode if the following condition (8) is met: (8) The gaze detection sensor 156 does not detect that the user is gazing at the real object detected by the head-mounted display terminal 1. If the gaze detection sensor 156 does not confirm that the user's gaze direction is looking at the detected object, it may be determined that the user is not gazing at the object and the system may switch to binocular mode. In Example 9, the head-mounted display terminal 1 proceeds to S4 and changes to binocular mode if the following condition (9) is met: (9) The size of the real object detected by the head-mounted display terminal 1 is less than or equal to the size of the image displayed on the left display 131L or the right display 131R. In this case, the presence of the image becomes greater, so it is considered that the user will gaze at the image, and the system changes to binocular mode. If the size of the image is larger than the detected object, the attention-grabbing effect on the detected object decreases. If the detected object is determined to be smaller than the displayed image, the system may switch to binocular mode. In Example 10, the head-mounted display terminal 1 proceeds to S4 and changes to binocular mode when the following condition (10) is met: (10) The user selects to activate binocular mode, and the head-mounted display terminal 1 accepts the user's selection. In this case, since the user has selected binocular mode, binocular mode is implemented regardless of whether an object is detected or not. The same applies if binocular mode is already activated.

[0183] In Example 11, if the head-mounted display terminal 1 does not meet any of the above conditions (1) to (10), it repeatedly executes S3 and continues in monocular mode.

[0184] Figure 26 is a table showing an example of the display mode confirmation process (S32). Here, the system switches to binocular mode according to the display mode setting of the head-mounted display terminal 1.

[0185] In Example 1, the head-mounted display terminal 1 proceeds to S4 and changes to binocular mode when the following condition (1) is met: (1) When there are multiple display modes for the head-mounted display terminal 1, and the mode for displaying a reduced size is deactivated.

[0186] In Example 2, the head-mounted display terminal 1 proceeds to S4 and changes to binocular mode when the following condition (2) is met: (2) When there are multiple display modes for the head-mounted display terminal 1, and a mode that limits the display area, for example, a mode that displays only in the four corners of the display area, is deactivated.

[0187] In Example 3, the head-mounted display terminal 1 will proceed to S4 and change to binocular mode if the following condition (3) is met: (3) If there are multiple display modes for the head-mounted display terminal 1, and the mode that restricts display in the area including the center of the field of view is deactivated.

[0188] In Example 4, the head-mounted display terminal 1 will proceed to S4 and change to binocular mode if the following condition (4) is met: (4) If there are multiple display modes for the head-mounted display terminal 1, and a mode that restricts display in an area of ​​50% or more of the field of view is deactivated.

[0189] In Example 5, the head-mounted display terminal 1 proceeds to S4 and changes to binocular mode when the following condition (5) is met: (5) When there are multiple display modes for the head-mounted display terminal 1 and a mode that allows video display in the entire display area is selected. When the display modes (1) to (5) above are activated, the display is intended to be viewed with video as the main content. In such displays, it is not intended to focus on objects other than the video. When displaying in these modes, it is possible to switch to binocular mode regardless of whether an object has been detected in order to comfortably view the video content.

[0190] In Example 6, if none of the above conditions (1) to (5) are met, the head-mounted display terminal 1 continues S3 until the user inputs a termination command. If the user inputs a termination command, the system proceeds to S4 and changes to binocular mode.

[0191] Figure 27 is a table showing an example of the mode change (S4) process.

[0192] In Example 1, the head-mounted display terminal 1 displays images on the left display 131L and the right display 131R4.

[0193] In Example 2, the head-mounted display terminal 1 displays a message to switch to binocular mode on the left display 131L or the right display 131R, and switches to binocular mode.

[0194] In Example 7, the head-mounted display terminal 1 displays a message to switch to binocular mode on the left display 131L or the right display 131R, and switches to binocular mode after obtaining user approval.

[0195] Although embodiments have been described above, the present invention is not limited to the embodiments described above, and includes various modifications and equivalent configurations within the spirit of the attached claims. For example, the embodiments described above are described in detail for the purpose of explaining the present invention in an easy-to-understand manner, and the present invention is not necessarily limited to having all the configurations described. Also, for example, some of the configurations of the embodiments may be added, deleted, or replaced with other configurations.

[0196] 1: Head-mounted display terminal 101: Processor 110: Storage device 120: Input I / F 130: Video input / output device 131: Display 140: Audio input / output device 150: Sensor group 160: Communication I / F

Claims

1. A head-mounted display terminal comprising a processor and a right display and a left display for displaying images, wherein the head-mounted display terminal is equipped with an external detector for acquiring external information, and the processor has the following processes: detecting surrounding objects based on data acquired by the external detector; determining at least one state of the object or the image; and switching between a binocular mode in which the image is displayed on both the right display and the left display and a monocular mode in which the image is displayed on one of the right display and the left display based on the determination.

2. A head-mounted display terminal according to claim 1, wherein the processor includes a process for detecting the position of an object and determining the distance of the object from the head-mounted display terminal, and a process for switching between the monocular mode and the binocular mode based on the determination.

3. The head-mounted display terminal according to claim 2, wherein the processor determines the distance from the head-mounted display terminal to the object, and the head-mounted display terminal uses the monocular mode when the distance to the object is less than a specified value.

4. The head-mounted display terminal according to claim 3, wherein the processor determines the distance from the head-mounted display terminal to the object, and the head-mounted display terminal uses the monocular mode if the period during which the distance to the object is less than a predetermined value is longer than a predetermined time.

5. A head-mounted display terminal according to claim 1, wherein the processor includes a process for detecting the position of an object and determining the distance of the object to the display position of the image, and a process for switching between the monocular mode and the binocular mode based on the determination.

6. A head-mounted display terminal according to claim 1, wherein the processor has a process for determining the size of the object, and a process for switching between the monocular mode and the binocular mode based on the determination.

7. A head-mounted display terminal according to claim 1, wherein the processor has the processes of determining the total area of ​​the object and switching between the monocular mode and the binocular mode based on the determination.

8. A head-mounted display terminal according to claim 1, wherein the processor includes a process for determining whether the object is located within a defined area of ​​the display area, and a process for switching between the monocular mode and the binocular mode based on the determination.

9. A head-mounted display terminal according to claim 1, wherein the processor includes a process for detecting the user's gaze on an object and determining whether the user is fixating on the object, and a process for switching between the monocular mode and the binocular mode based on the determination.

10. A head-mounted display terminal according to claim 1, wherein the processor switches to the binocular mode when it does not detect an object within the field of view of the head-mounted display terminal.

11. A head-mounted display terminal according to claim 1, wherein the processor has a process for confirming permission to use the switch between monocular mode and binocular mode in the settings of the head-mounted display terminal, and a process for switching between the monocular mode and the binocular mode based on the result of the confirmation.

12. A head-mounted display terminal according to claim 1, wherein the processor has a process for notifying the user of switching between the binocular mode and the monocular mode.

13. A head-mounted display terminal according to claim 1, wherein the processor switches the display that displays the image when the difference in usage time between the left display and the right display is greater than or equal to a specified value in the monocular mode.

14. A control method for a head-mounted display terminal comprising a processor and a right display and a left display for displaying images, wherein the processor detects an object within the field of view of the head-mounted display terminal, and the control method switches between a binocular mode in which the image is displayed on both the right display and the left display and a monocular mode in which the image is displayed on one of the right display and the left display, based on the state of the object or the image.

15. An information display system comprising a head-mounted display terminal having a right display and a left display for displaying images, and an external information terminal for communicating with the head-mounted display terminal, wherein the head-mounted display terminal is equipped with an external detector for acquiring external information, the external information terminal is equipped with a processor, and the processor has the following processes: a process for detecting surrounding objects based on data acquired by the external detector; a process for determining at least one state of the object or the image; and a process for switching between a binocular mode in which the image is displayed on both the right display and the left display and a monocular mode in which the image is displayed on one of the right display and the left display based on the determination.