Head-mounted display device for changing disparity of stereo image, and operating method thereof

WO2026177390A1PCT designated stage Publication Date: 2026-08-27SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2026/001100
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-24
Filing Date
2026-01-19
Publication Date
2026-08-27

Smart Images

  • Figure KR2026001100_27082026_PF_FP_ABST
    Figure KR2026001100_27082026_PF_FP_ABST
Patent Text Reader

Abstract

A head-mounted display device for adjusting a disparity of a stereo image, and an operating method thereof are provided. The head-mounted display device can display a stereo image including a left-eye image and a right-eye image, receive a user's zoom-in input enlarging the size of the stereo image, determine the position and size of a region of interest from the stereo image on the basis of the received zoom-in input, shift the position of the region of interest by a shift value in the left-eye image and / or the right-eye image to primarily adjust the disparity between the region of interest in the left-eye image and the region of interest in the right-eye image, and enlarge the size of the region of interest to the size of the stereo image so as to secondarily adjust the disparity.
Need to check novelty before this filing date? Find Prior Art

Description

Head-mounted display device for changing the parallax of a stereo image and method of operation thereof

[0001] The present disclosure relates to a head-mounted display (HMD) device for changing the disparity of a stereo image and a method of operating the same. Specifically, the present disclosure discloses a head-mounted display device and a method of operating the same for adjusting the disparity of a stereo image when user input is received to zoom in and enlarge a stereo image being displayed.

[0002] Virtual reality (VR) refers to a specific environment or situation created by artificial technology that is similar to reality but not actually real, or the technology itself. Devices that enable the experience of virtual reality include, for example, head-mounted displays (HMDs). Among head-mounted display devices, see-through HMDs allow users to see the real world and can enhance the user experience by augmenting virtual objects onto the real world.

[0003] Augmented reality (AR) is a technology that overlays virtual images onto the physical environment or real-world objects to display them together. Examples of augmented reality devices utilizing this technology include augmented reality glasses. Augmented reality glasses are useful in daily life for tasks such as information retrieval, navigation, and photography. In particular, among augmented reality glasses, smart glasses are worn as fashion items and are primarily used for outdoor activities.

[0004] Users can view image content, such as images or videos captured using mobile devices like smartphones, by displaying them through virtual reality devices or augmented reality devices, such as head-mounted display devices. In particular, with the recent proliferation of head-mounted display devices, there is an increasing demand to experience vivid realism and three-dimensionality by viewing stereo image content captured using stereo cameras through head-mounted display devices. Stereo image content provides binocular images captured from the left and right using multiple cameras to the user's eyes, thereby enabling the user to perceive a sense of depth based on the disparity between the binocular images while viewing the image content.

[0005] Viewers of stereo image content according to conventional technology do not provide a zoom function. That is, when a user views stereo image content, they can only perceive depth from the position where the stereo image content was captured, and there is no change in depth related to the user's forward and backward movements or gestures. In the case of conventional technology, since the user cannot perceive a change in depth even when inputting movements or gestures, there is a technical limitation in that the user's immersiveness is reduced.

[0006] One aspect of the present disclosure provides a method for a head-mounted display device to change the disparity of a stereo image. A method of operation of a head-mounted display device according to one embodiment of the present disclosure may include receiving a zoom-in input from a user to enlarge the size of a stereo image while displaying a stereo image including a left-eye image and a right-eye image. A method of operation of a head-mounted display device according to one embodiment of the present disclosure may include determining the location and size of a region of interest within the entire area of ​​the stereo image based on the received zoom-in input. A method of operation of a head-mounted display device according to one embodiment of the present disclosure may include a step of adjusting the disparity between the region of interest in the left-eye image and the right-eye image by shifting the location of the region of interest determined in at least one of the left-eye image and the right-eye image by a shift value calculated to have an inverse relationship with the size of the region of interest. A method of operation of a head-mounted display device according to one embodiment of the present disclosure may include a step of increasing the size of the region of interest to the size of the entire area of ​​the stereo image to increase the adjusted disparity.

[0007] One aspect of the present disclosure provides a head-mounted display device for adjusting the parallax of a stereo image. In one embodiment of the present disclosure, the head-mounted display device may include a sensor comprising at least one of a position sensor and an Inertial Measurement Unit (IMU) sensor; a left-eye display for displaying a left-eye image included in the stereo image and a right-eye display for displaying a right-eye image included in the stereo image; at least one processor comprising processing circuitry; and a memory for storing one or more instructions. By executing the one or more instructions individually or collectively by at least one processor, the head-mounted display device may receive a user's zoom-in input to enlarge the size of the stereo image and determine the location and size of a region of interest within the entire area of ​​the stereo image based on the received zoom-in input. By executing the above one or more instructions individually or collectively by at least one processor, the head-mounted display device can adjust the disparity between the region of interest in the left-eye image and the right-eye image by shifting the position of the region of interest determined in at least one of the left-eye image and the right-eye image by a shift value calculated to have an inverse relationship with the size of the region of interest. By executing the above one or more instructions individually or collectively by at least one processor, the head-mounted display device can increase the adjusted disparity by expanding the size of the region of interest to the size of the entire area of ​​the stereo image.

[0008] One aspect of the present disclosure provides a method for a head-mounted display device to change the disparity of a stereo image. In one embodiment of the present disclosure, a method of operation of a head-mounted display device may include receiving a user's zoom-out input to reduce the size of the stereo image while displaying a stereo image including a left-eye image and a right-eye image. A method of operation of a head-mounted display device may include determining the location and size of a region of interest within the entire area of ​​the stereo image based on the received zoom-out input. A method of operation of a head-mounted display device may include reducing the size of the region of interest based on the zoom-out input. A method of operation of a head-mounted display device may include adjusting the disparity between the region of interest in each of the left-eye image and the right-eye image as the size of the region of interest is reduced. The size of the region of interest may be determined to be inversely proportional to the amount of movement of the received zoom-out input. The minimum value of the size of the region of interest may be determined based on a preset minimum value of the resolution of the stereo image.

[0009] The present disclosure can be easily understood from the combination of the following detailed description and the accompanying drawings, where reference numerals denote structural elements.

[0010] FIG. 1 is a conceptual diagram illustrating the operation of a head-mounted display device according to one embodiment of the present disclosure to adjust the disparity of a stereo image based on a user's zoom-in input.

[0011] FIG. 2 is a flowchart illustrating a method for a head-mounted display device according to one embodiment of the present disclosure to adjust the disparity of a stereo image based on a user's zoom-in input.

[0012] FIG. 3 is a block diagram illustrating the components of a head-mounted display device according to one embodiment of the present disclosure.

[0013] FIG. 4 is a flowchart illustrating a method for a head-mounted display device according to one embodiment of the present disclosure to determine the location and size of a region of interest on a stereo image based on the distance traveled due to a user's movement.

[0014] FIG. 5a is a diagram illustrating the operation of a head-mounted display device according to one embodiment of the present disclosure determining the location of a region of interest based on the line of sight of the user's two eyes.

[0015] FIG. 5b is a diagram illustrating the operation of a head-mounted display device according to one embodiment of the present disclosure determining the size of a region of interest based on the distance traveled due to a user's movement.

[0016] FIG. 5c is a graph illustrating the relationship between the user's movement distance and the size of the region of interest according to one embodiment of the present disclosure.

[0017] FIG. 5d is a graph illustrating the relationship between the user's travel distance and the size of the region of interest according to one embodiment of the present disclosure.

[0018] FIG. 6 is a flowchart illustrating a method for a head-mounted display device according to one embodiment of the present disclosure to determine the location and size of a region of interest on a stereo image based on a user's pinch zoom input.

[0019] FIG. 7a is a diagram illustrating the operation of a head-mounted display device according to one embodiment of the present disclosure determining the location and size of a region of interest on a stereo image based on a user's pinch zoom input.

[0020] FIG. 7b is a graph illustrating the relationship between the drag distance of a pinch zoom input and the size of the region of interest according to one embodiment of the present disclosure.

[0021] FIG. 7c is a graph illustrating the relationship between the drag distance of a pinch zoom input and the size of the region of interest according to one embodiment of the present disclosure.

[0022] FIG. 8 is a diagram illustrating the operation of a head-mounted display device according to one embodiment of the present disclosure determining the position and size of a region of interest on a stereo image based on the user's two-eyed gaze information and pinch zoom input.

[0023] FIG. 9a is a diagram illustrating the operation of a head-mounted display device according to one embodiment of the present disclosure to change the position of a region of interest when the position of the region of interest is outside the frame of a stereo image.

[0024] FIG. 9b is a diagram illustrating the operation of a head-mounted display device according to one embodiment of the present disclosure to change the position of a region of interest when the position of the region of interest is outside the frame of a stereo image.

[0025] FIG. 10 is a flowchart illustrating a method for a head-mounted display device to adjust the parallax of a region of interest according to one embodiment of the present disclosure.

[0026] FIG. 11a is a diagram illustrating the operation of a head-mounted display device according to one embodiment of the present disclosure adjusting the parallax of a region of interest.

[0027] FIG. 11b is a diagram illustrating the operation of a head-mounted display device according to one embodiment of the present disclosure adjusting the parallax of a region of interest.

[0028] FIG. 11c is a diagram illustrating the operation of a head-mounted display device according to one embodiment of the present disclosure adjusting the parallax of a region of interest.

[0029] FIG. 12 is a flowchart illustrating a method for a head-mounted display device according to one embodiment of the present disclosure to determine a shift value for parallax adjustment of a region of interest.

[0030] FIG. 13 is a diagram illustrating a disparity map to explain the operation of a head-mounted display device determining a shift value according to one embodiment of the present disclosure.

[0031] FIG. 14a is a graph illustrating the relationship between the size of the region of interest and the shift value according to one embodiment of the present disclosure.

[0032] FIG. 14b is a graph illustrating the relationship between the size of the region of interest and the shift value according to one embodiment of the present disclosure.

[0033] FIG. 15 is a diagram illustrating the operation of a head-mounted display device according to one embodiment of the present disclosure that increases the parallax of a region of interest by expanding the region of interest based on a zoom-in input.

[0034] FIG. 16 is a flowchart illustrating a method for a head-mounted display device according to one embodiment of the present disclosure to adjust the parallax of a stereo image based on a user's zoom-in input.

[0035] FIG. 17 is a conceptual diagram illustrating the operation of a head-mounted display device according to one embodiment of the present disclosure to adjust the disparity of a stereo image based on a user's zoom-out input.

[0036] FIG. 18 is a flowchart illustrating a method for a head-mounted display device according to one embodiment of the present disclosure to adjust the disparity of a stereo image based on a user's zoom-out input.

[0037] FIG. 19 is a graph illustrating the relationship between the drag distance of a user’s movement distance or pinch zoom input and the size of the region of interest according to one embodiment of the present disclosure.

[0038] In describing the present disclosure, technical details that are well known in the technical field to which the present disclosure belongs and are not directly related to the present disclosure are omitted. This is intended to convey the essence of the present disclosure more clearly without obscuring it by omitting unnecessary explanations. Furthermore, the terms described below are defined considering their functions within the present disclosure, and these definitions may vary depending on the intentions or practices of the user or operator. Therefore, their definitions should be based on the content throughout this specification.

[0039] For the same reason, some components in the attached drawings have been exaggerated, omitted, or schematically depicted. Additionally, the dimensions of each component do not entirely reflect their actual dimensions. Identical or corresponding components in each drawing have been assigned the same reference numbers.

[0040] The advantages and features of the present disclosure, and the methods for achieving them, will become clear by referring to the embodiments described below in detail together with the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below but may be implemented in various different forms. The disclosed embodiments are provided to ensure that the disclosure of the present disclosure is complete and to fully inform those skilled in the art of the scope of the disclosure. An embodiment of the present disclosure may be defined according to the claims. Throughout the specification, the same reference numerals indicate the same components. Furthermore, in describing an embodiment of the present disclosure, if it is determined that a detailed description of a related function or configuration might unnecessarily obscure the essence of the present disclosure, such detailed description is omitted. Additionally, terms described below are defined considering their functions in the present disclosure, and these may vary depending on the intentions or conventions of the user or operator. Therefore, their definitions should be based on the content throughout the specification.

[0041] In one embodiment of the present disclosure, each block of the flowcharts and combinations of the flowcharts may be executed by computer program instructions. Computer program instructions may be loaded onto a processor of a general-purpose computer, a computer for special purposes, or other programmable data processing equipment, and the instructions executed through the processor of the computer or other programmable data processing equipment may produce means for performing the functions described in the flowchart block(s). Computer program instructions may also be stored in computer-available or computer-readable memory that may be directed toward the computer or other programmable data processing equipment to implement the function in a specific manner, and instructions stored in computer-available or computer-readable memory may produce a manufactured item containing instruction means for performing the function described in the flowchart block(s). Computer program instructions may also be loaded onto a computer or other programmable data processing equipment.

[0042] Additionally, each block of the flowchart may represent a module, segment, or part of code containing one or more executable instructions for executing a specified logical function(s). In one embodiment, the functions mentioned in the blocks may occur out of order. For example, two blocks shown in succession may be executed substantially simultaneously or in reverse order depending on the function.

[0043] In one embodiment of the present disclosure, the term “part” used may refer to software or hardware components such as a Field Programmable Gate Array (FPGA) or an Application Specific Integrated Circuit (ASIC), and the “part” may perform a specific role. Meanwhile, the “part” is not limited to software or hardware. The “part” may be configured to reside in an addressable storage medium or may be configured to run one or more processors. In one embodiment of the present disclosure, the “part” may include components such as software components, object-oriented software components, class components, and task components, as well as processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables. Functions provided through a specific component or a specific “part” may be combined or separated into additional components to reduce their number. Additionally, in one embodiment of the present disclosure, the “part” may include one or more processors.

[0044] As used in this disclosure, the expression “configured to” may be replaced, depending on the context, with, for example, “suitable for,” “having the capacity to,” “designed to,” “adapted to,” “made to,” or “capable of.” The term “configured to” may not necessarily mean only “specifically designed to” in hardware. Instead, in some situations, the expression “system configured to” may mean that the system is “capable of” in conjunction with other devices or components. For example, the phrase “processor configured to perform A, B, and C” may mean a dedicated processor for performing the said operations (e.g., an embedded processor), or a generic-purpose processor (e.g., a CPU or an application processor) capable of performing said operations by executing one or more software programs stored in memory.

[0045] In addition, when a component is described in the present disclosure as being "connected" or "connected" to another component, it should be understood that the component may be directly connected to or directly connected to the other component, but unless otherwise specifically stated, it may also be connected or connected through another component in between.

[0046] Additionally, the description "at least one of A, B, and C" means that it may be any one of 'A', 'B', 'C', 'A and B', 'A and C', 'B and C', and 'A, B, and C'.

[0047] Functions related to artificial intelligence according to the present disclosure are operated through a processor and memory. The processor may be composed of one or more processors. In this case, the one or more processors may be general-purpose processors such as CPUs, APs, and DSPs (Digital Signal Processors), graphics-dedicated processors such as GPUs and VPUs (Vision Processing Units), or artificial intelligence-dedicated processors such as NPUs. The one or more processors control the processing of input data according to predefined operation rules or artificial intelligence models stored in memory. Alternatively, if the one or more processors are artificial intelligence-dedicated processors, the artificial intelligence-dedicated processors may be designed with a hardware structure specialized for processing a specific artificial intelligence model.

[0048] The predefined rules of operation or artificial intelligence models are characterized by being created through learning. Here, being created through learning means that a predefined rules of operation or artificial intelligence models configured to perform a desired characteristic (or objective) are created by a basic artificial intelligence model being trained using multiple learning data by a learning algorithm. Such learning may be performed on the device itself where the artificial intelligence according to the present disclosure is executed, or it may be performed through a separate server and / or system. Examples of learning algorithms include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but are not limited to the examples described above.

[0049] An artificial intelligence model may be composed of multiple neural network layers. Each of the multiple neural network layers has multiple weight values ​​and performs neural network operations through operations between the results of previous layers and the multiple weights. The multiple weights possessed by the multiple neural network layers can be optimized based on the learning results of the artificial intelligence model. For example, the multiple weights may be updated so that the loss value or cost value obtained from the artificial intelligence model during the learning process is reduced or minimized. Artificial neural networks may include deep neural networks (DNNs), such as Convolutional Neural Networks (CNNs), Deep Neural Networks (DNNs), Recurrent Neural Networks (RNNs), Restricted Boltzmann Machines (RBMs), Deep Belief Networks (DBNs), Bidirectional Recurrent Deep Neural Networks (BRDNNs), or Deep Q-Networks, but are not limited to the examples mentioned above.

[0050] In the present disclosure, "Augmented Reality" means displaying a virtual image together within a physical environment space of the real world or displaying a virtual image together with a real object.

[0051] In the present disclosure, a ‘Head Mounted Display (HMD) device’ is a device worn on a user’s head and provides an augmented reality experience to the user. In one embodiment of the present disclosure, the head mounted display device may include, for example, an augmented reality helmet, a face mounted display (FMD) device worn on a user’s face, or augmented reality glasses in the form of glasses.

[0052] In the present disclosure, a 'stereo image' refers to a plurality of images captured using a plurality of cameras spaced apart by a baseline, and represents an image capable of providing a sense of depth to the user through the disparity between the plurality of images.

[0053] Embodiments of the present disclosure are described below with reference to the attached drawings so that those skilled in the art can easily implement them. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein.

[0054] Embodiments of the present disclosure will be described in detail below with reference to the drawings.

[0055] FIG. 1 is a conceptual diagram illustrating the operation of a head-mounted display device (100) according to one embodiment of the present disclosure adjusting the disparity of a stereo image based on a zoom-in input from a user (10).

[0056] A head-mounted display (HMD) device (100) is a device capable of displaying virtual images together within a physical environment space of the real world, or displaying real objects and virtual images together, thereby expressing augmented reality. In the embodiment illustrated in FIG. 1, the head-mounted display device (100) may be implemented as a device worn on the head of a user (10). However, it is not limited thereto, and the head-mounted display device (100) may include, for example, an augmented reality helmet, a face-mounted display (FMD) device worn on the user's face, or augmented reality glasses in the shape of glasses.

[0057] Referring to FIG. 1, the head-mounted display device (100) has a left eye image (i L ) and right eye image(i R A stereo image including ) can be displayed (Operation ①).

[0058] The head-mounted display device (100) receives a zoom-in input from the user to enlarge the stereo image while the stereo image is being displayed, and can determine a region of interest (ROI) based on the received zoom-in input (operation ②).

[0059] The head-mounted display device (100) has a left eye image (i L ) and right eye image(i R The disparity between the regions of interest (ROI1, ROI2) in each can be adjusted (Operation ③).

[0060] The head-mounted display device (100) enlarges the size of the region of interest based on the zoom-in input, and the left eye image (i L) and right eye image(i R ) The disparity between the regions of interest (ROI1, ROI2) in each can be increased (Operation ④).

[0061] FIG. 2 is a flowchart illustrating a method for a head-mounted display device (100) according to one embodiment of the present disclosure to adjust the disparity of a stereo image based on a user's zoom-in input.

[0062] Hereinafter, the function and / or operation of the head-mounted display device (100) according to the embodiment illustrated in FIG. 1 will be described in detail with reference to FIG. 1 and FIG. 2 together.

[0063] In step S210 of FIG. 2, the head-mounted display device (100) receives a zoom-in input from a user to enlarge the size of the stereo image while displaying the stereo image. In the present disclosure, the 'stereo image' is a plurality of images captured using a plurality of cameras spaced apart by a baseline, and may include a left-eye image viewed through the user's left eye and a right-eye image viewed through the user's right eye. The stereo image may provide a sense of depth to the user through the user's binocular disparity. The stereo image may consist of a still image or a video. For example, the stereo image may be a video in which disparity correction is performed on a binocular video and the corrected disparity is stored in a side-by-side (SBS) or MV-HEVC (Multiview High Efficiency Video Coding) format. For example, the stereo image may be a still image in which disparity correction is performed on a binocular image and the corrected disparity is stored in SBS or Samsung Extended Format (SEF).

[0064] In one embodiment of the present disclosure, a stereo image may be captured and acquired by a plurality of cameras included in a head-mounted display device (100). However, it is not limited thereto, and in one embodiment of the present disclosure, the stereo image may be captured by a plurality of cameras included in an external device, for example, a mobile device, and the head-mounted display device (100) may acquire the stereo image from the external device through a wired or wireless communication network. For example, the head-mounted display device (100) may be paired with an external device through a short-range wireless communication network such as Bluetooth or WiFi Direct, and may receive data of the stereo image captured by the external device through the short-range wireless communication network.

[0065] The head-mounted display device (100) can display a stereo image. Referring together to operation 1 of FIG. 1, the head-mounted display device (100) may include a left-eye display (150L) and a right-eye display (150R). The head-mounted display device (100) displays the left-eye image (i) among the stereo images through the left-eye display (150L). L Displays ) and through the right eye display (150R) the right eye image (i R It can display ).

[0066] The head-mounted display device (100) has a left eye image (i L ) and right eye image(i R A user's zoom-in input that enlarges the size of ) can be received. In one embodiment of the present disclosure, the user's zoom-in input may include at least one of a distance traveled due to the user's movement and a pinch zoom input by a hand gesture.

[0067] Referring again to FIG. 2, in step S220, the head-mounted display device (100) determines the location and size of the region of interest within the entire area of ​​the stereo image based on the user's zoom-in input. In one embodiment of the present disclosure, the head-mounted display device (100) receives a user's zoom-in input to enlarge the size of the stereo image while the stereo image including the left eye image and the right eye image is being displayed, and can determine the location and size of the region of interest based on the zoom-in input.

[0068] Referring to operation ② of FIG. 1, when a zoom-in input is received due to a position change caused by the movement of a user (10) wearing a head-mounted display device (100), the head-mounted display device (100) can recognize the direction of gaze of both eyes of the user (10) using an eye tracking sensor and detect a gaze point where the direction of gaze of both eyes converges. Based on the location of the detected gaze point, the head-mounted display device (100) can determine the position coordinate values ​​of the center points (Pc1, Pc2) of the regions of interest (ROI1, ROI2). The head-mounted display device (100) can acquire a change in position including the position, rotation, and translation of the head-mounted display device (100) caused by the movement of the user (10) using at least one of a GPS sensor and an IMU sensor (inertial measurement unit), and thereby measure the distance traveled (d) caused by the movement of the user (10). A head-mounted display device (100) can determine the size of a region of interest (ROI1, ROI2) centered on the position coordinate values ​​of a center point (Pc1, Pc2) based on a measured travel distance (d). In one embodiment of the present disclosure, the head-mounted display device (100) can determine the area size of the region of interest (ROI1, ROI2) as a value inversely proportional to the measured travel distance (d). For example, the area size of the region of interest (ROI1, ROI2) can be calculated as a value having an inverse relationship with respect to the travel distance (d) according to linear, logarithmic, or exponential, etc.

[0069] When a pinch zoom input is received by a hand gesture of a user (10), the head-mounted display device (100) can determine the location and size of the region of interest (ROI1, ROI2) based on the pinch zoom input. Referring to operation ② of FIG. 1, the head-mounted display device (100) can obtain a hand image by taking a picture of the user's (10) hand (20) using a front camera, and can obtain information regarding the position of the fingers and the drag distance (d) according to the movement of the fingers by analyzing the hand image using an artificial intelligence model trained to recognize hand joint positions or hand gestures or known image processing. The head-mounted display device (100) can determine the middle coordinate value among the position coordinate values ​​of the fingers as the center point (Pc1, Pc2) of the region of interest (ROI1, ROI2) and determine the size of the region of interest (ROI1, ROI2) based on the drag distance (d) by the fingers. The head-mounted display device (100) can determine the size of the region of interest (ROI1, ROI2) as a value inversely proportional to the drag distance (d). For example, the area size of the region of interest (ROI1, ROI2) can be calculated as a value having an inverse relationship with the drag distance (d) according to linear, logarithmic, or exponential, etc.

[0070] Although not illustrated in the drawings, in one embodiment of the present disclosure, the head-mounted display device (100) may receive zoom input from an external controller. Since the specific method for determining the location and size of the region of interest based on the zoom input from the external controller is identical to pinch zoom input except that drag input is performed using a controller instead of a user's finger, a redundant description is omitted.

[0071] A head-mounted display device (100) includes a left eye image (i) in a stereo image based on a zoom-in input. L ) and right eye image(iR Regions of interest (ROI1, ROI2) can be set for each of them. Referring to the embodiment illustrated in FIG. 1, the head-mounted display device (100) has a left eye image (i L Determine the location and size of the first region of interest (ROI1) within ), and the right eye image (i R The location and size of the second region of interest (ROI2) can be determined within ).

[0072] Referring again to FIG. 2, in step S230, the head-mounted display device (100) adjusts the disparity between the regions of interest in the left-eye image and the right-eye image by shifting the position of the region of interest determined in at least one of the left-eye image and the right-eye image by a shift value. In one embodiment of the present disclosure, the head-mounted display device (100) can adjust the disparity of the region of interest by shifting the position of the region of interest in the left-eye image to the left along the horizontal direction (X-axis direction) by a shift value, and by shifting the position of the region of interest in the right-eye image to the right along the horizontal direction by a shift value. Referring together to operation ③ illustrated in FIG. 1, the head-mounted display device (100) [addresses] the left-eye image (i L The position of the first region of interest (ROI1) within ) is shifted to the left along the X-axis by a shift value, and the right eye image (i R The position of the second region of interest (ROI2) within the image can be shifted to the right along the X-axis by a shift value. As the position of the region of interest is shifted, the parallax (δ1) between the regions of interest in the stereo image can be increased by a value corresponding to twice the shift value compared to the parallax (δ0) before shifting.

[0073] In the present disclosure, the 'shift value' is a left eye image (i) to adjust the parallax between regions of interest set in a stereo image. L The first region of interest (ROI1) and right eye image (i) in )R It is an indicator value representing the amount of positional shift that moves the position of the second region of interest (ROI2) in the X-axis direction. In one embodiment of the present disclosure, the shift value is a left eye image (i L ) and right eye image(i R It can be calculated based on the maximum disparity of the region of interest obtained from a disparity map containing information regarding the disparity between ). For example, the shift value can be calculated as a value inversely proportional to the maximum disparity of the region of interest. However, it is not limited to this, and the shift value can be calculated based on the ratio of the size of the region of interest to the size of the entire region of the stereo image.

[0074] However, the present disclosure is not limited thereto, and a head-mounted display device (100) according to one embodiment of the present disclosure may move only one of the first region of interest (ROI1) and the second region of interest (ROI2) by a shift value along the X-axis direction. In one embodiment of the present disclosure, the head-mounted display device (100) may move one region of interest by twice the shift value.

[0075] Referring to FIG. 2, in step S240, the head-mounted display device (100) expands the size of the region of interest to the size of the entire stereo image area, thereby increasing the adjusted parallax. The head-mounted display device (100) can expand the size of the region of interest to the size of the entire stereo image area based on the user's zoom-in input. As the size of the region of interest is expanded, the parallax adjusted in step S230 can be increased proportionally to the expanded size of the region of interest. Referring together to operation ④ of FIG. 1, the head-mounted display device (100) expands the size of the first region of interest (ROI1') to the left eye image (i LEnlarge to the same size as the entire size of ), and the size of the second region of interest (ROI2') is the right eye image (i R It can be enlarged to be equal to the total size of the image. In this case, the disparity (δ2) between the enlarged first region of interest (ROI1') and the enlarged second region of interest (ROI2') can be increased by the value obtained by dividing the disparity (δ1) adjusted in step S230 by the region of interest ratio value. In the present disclosure, the 'region of interest ratio value' may be a value representing the ratio between the size of the region of interest before enlargement and the total size of the stereo image. The region of interest ratio value may be calculated through an operation of dividing the size of the region of interest before enlargement by the size of the total area of ​​the stereo image.

[0076] In one embodiment of the present disclosure, a head-mounted display device (100) can perform post-processing to improve the quality, such as the resolution of an image, through methods such as pixel interpolation after enlarging the size of the region of interest. In one embodiment of the present disclosure, the image quality improvement may be achieved by using a deep learning model or by using conventional image processing methods such as image interpolation.

[0077] The method of using deep learning models involves utilizing super-resolution networks such as DRCT, SwinIR, or SRCNN. DRCT (Dense-residual-connected Transformer) is an artificial intelligence model based on SwinTransformer and Convolutional Neural Networks (CNN) trained through supervised learning, where low-resolution images are used as input and high-resolution images are used as output ground truths. SwinIR is an artificial intelligence model based on SwinTransformer and Convolutional Neural Networks trained through supervised learning, where low-resolution images are used as input and high-resolution images are used as ground truths. Here, 'SwinTransformer' refers to a Transformer that performs shifted-window self-attention. SRCNN is an artificial intelligence model based on a Convolutional Neural Network (CNN) that is trained through a supervised learning method in which a low-resolution image is applied as input and a high-resolution image is applied as output ground truth. A head-mounted display device (100) can input an image of an enlarged region of interest (ROI1', ROI2') into a super-resolution model such as DRCT, SwinIR, or SRCNN, and obtain a high-quality image of the region of interest with improved resolution as an output value based on the inference result of the super-resolution model.

[0078] As a technique to improve image quality, the head-mounted display device (100) may use an image interpolation algorithm such as Bilinear or Bicubic.

[0079] The head-mounted display device (100) can display high-quality images of the region of interest with improved resolution through the left eye display (150L) and the right eye display (150R).

[0080] Recently, with the widespread adoption of virtual reality / augmented reality devices such as head-mounted display devices, there is an increasing demand to experience vivid realism and three-dimensionality by viewing stereo image content captured using stereo cameras through these devices. However, stereo image content viewers based on conventional technology have the problem of not providing a zoom function. That is, when a user views stereo image content, they can only perceive depth at the location where the stereo image content was captured, and there is no change in depth related to the user's forward and backward movements or gestures. In the case of conventional technology, since the user cannot perceive a change in depth even when inputting movements or gestures, there is a problem of reduced immersion.

[0081] The present disclosure aims to provide a head-mounted display device (100) and a method of operation thereof, which, when a zoom-in input is received from a user to enlarge the size of a stereo image of an area of ​​interest while displaying a stereo image, adjusts the disparity of the stereo image based on the zoom-in input to provide the user with a stereoscopic effect by not only enlarging the image but also changing the depth value of the stereo image.

[0082] A head-mounted display device (100) according to the embodiment illustrated in FIGS. 1 and 2 determines the position and size of a region of interest within a stereo image based on positional movement of the head-mounted display device (100) due to user movement or pinch zoom input, and a left eye image (i L ) and right eye image(iR By shifting the position of the region of interest in each of the X-axis directions in opposite directions, the left eye image (i L ) and right eye image(i R The parallax between regions of interest in the first stage is adjusted, and the size of the region of interest is expanded to the size of the entire stereo image area based on the zoom-in input, thereby increasing the first-stage adjusted parallax between regions of interest a second time. Through this, the head-mounted display device (100) according to one embodiment of the present disclosure not only expands the stereo image according to the user's movement or pinch zoom input, but also provides a technical effect of providing a three-dimensional sense to the user by adjusting the sense of depth and enhancing the user's immersion. In addition, the head-mounted display device (100) according to one embodiment of the present disclosure can enable the user to edit spatial content by adjusting the parallax of the stereo image.

[0083] FIG. 3 is a block diagram illustrating the components of a head-mounted display device (100) according to one embodiment of the present disclosure.

[0084] Referring to FIG. 3, a head-mounted display device (100) may include a sensor (110), a camera (120), a processor (130), a memory (140), and a display (150). The sensor (110), camera (120), processor (130), memory (140), and display (150) may each be electrically and / or physically connected to one another. FIG. 3 illustrates only essential components for explaining the function and / or operation of the head-mounted display device (100), and the components included in the head-mounted display device (100) are not limited to those illustrated in FIG. 3. In one embodiment of the present disclosure, the head-mounted display device (100) may further include a battery that supplies driving power to the sensor (110), camera (120), processor (130), and display (150). In one embodiment of the present disclosure, the head-mounted display device (100) may further include a communication interface that pairs with an external device (e.g., a mobile device such as a smartphone or tablet PC) via a short-range wireless communication network (e.g., Bluetooth or Wi-Fi Direct) and receives a stereo image from the external device.The communication interface may be composed of a hardware device that performs data communication with an external device or server using at least one of, for example, Wireless LAN, Wi-Fi, Wi-Fi Direct, Bluetooth, BLE (Bluetooth Low Energy), infrared communication (IrDA, infrared Data Association), NFC (Near Field Communication), Wibro (Wireless Broadband Internet), WiMAX (World Interoperability for Microwave Access), SWAP (Shared Wireless Access Protocol), WiGig (Wireless Gigabit Alliance), or RF communication.

[0085] The sensor (110) is configured to acquire a measurement value by sensing the movement, positional shift, rotation, speed change, etc. of the head-mounted display device (100). In one embodiment of the present disclosure, the sensor (110) may include a position sensor (112), an IMU sensor (114), and an eye-tracking sensor (116).

[0086] The position sensor (112) is a sensor configured to acquire position information of the head-mounted display device (100), and can be implemented, for example, as a Global Positioning System (GPS) sensor.

[0087] The IMU (Inertial Measurement Unit) sensor (114) is a sensor configured to measure the movement speed, direction, angle, and gravitational acceleration of the head-mounted display device (100) through a combination of an accelerometer, a gyroscope, and a magnetometer. In particular, the gyroscope can measure the amount of positional movement including rotation and translation of the head-mounted display device (100). The processor (130) can use the IMU sensor to obtain 6 DoF (6 Degree of Freedom) measurements including 3-dimensional positional coordinate values ​​(x-axis, y-axis, and z-axis coordinate values) and 3-axis angular velocity values ​​(roll, yaw, and pitch) of the head-mounted display device (100).

[0088] An eye tracking sensor (116) is a device configured to detect the direction of a user's gaze and to track the direction of the gaze. The eye tracking sensor (116) can detect the direction of the user's gaze by detecting an image of a person's pupil or iris, or by detecting the direction or amount of reflected light, such as near-infrared light, reflected from the cornea. In one embodiment of the present disclosure, the eye tracking sensor (116) may include an infrared light source and an infrared camera. However, it is not limited thereto, and the eye tracking sensor (116) of the present disclosure may be composed of a sensor utilizing any known gaze detection and / or tracking technology.

[0089] The eye tracking sensor (116) includes a left eye tracking sensor and a right eye tracking sensor, and can detect the direction of gaze of the user's left eye and the direction of gaze of the user's right eye, respectively. Detecting the direction of gaze of the user may include the operation of obtaining gaze information related to the user's gaze. In one embodiment of the present disclosure, a gaze point where the direction of gaze of the user's left eye detected by the eye tracking sensor (116) and the direction of gaze of the user's right eye detected by the right eye tracking sensor converge is detected, and location coordinate value information of the gaze point can be obtained. The eye tracking sensor (110) can provide the location coordinate value information of the gaze point to the processor (130).

[0090] A camera (120) is configured to capture an object and acquire an image. The camera (120) may include a lens module, an image sensor, and an image processing module. The camera (120) may acquire a still image or video of an object by means of an image sensor (e.g., CMOS or CCD). The video may include multiple image frames that are continuously acquired by capturing an object through the camera (120). The image processing module may store a still image consisting of a single image frame acquired through the image sensor or video data consisting of multiple image frames in an image data storage within a memory (140).

[0091] The camera (120) may include two or more cameras. The camera (120) may be implemented as a plurality of stereo cameras. The plurality of stereo cameras may be spaced apart by a baseline. In one embodiment of the present disclosure, the plurality of stereo cameras (120) may capture a real-world object to obtain a left-eye image and a right-eye image.

[0092] In one embodiment of the present disclosure, the camera (120) can capture a hand image by photographing the user's hand. For example, the camera (120) can capture a hand gesture of the user performing pinch zoom.

[0093] The processor (130) can execute one or more instructions of a program stored in memory (140). The processor (130) may be composed of hardware components that perform arithmetic, logic, and input / output operations and image processing. Although the processor (130) is depicted as a single element in FIG. 3, it is not limited thereto. In one embodiment of the present disclosure, the processor (130) may be composed of one or more elements.

[0094] The processor (130) may include various processing circuits and / or multiple processors. For example, the term "processor" as used in the present disclosure, including in the claims, may include at least one processor and various processing circuits. In the at least one processor, one or more processors may be configured to perform the various functions described in the present disclosure in a distributed manner, individually and / or collectively. In the present disclosure, "processor," "at least one processor," and "one or more processors" may be configured to perform various functions. However, these terms cover, without limitation, situations where one processor performs some of the functions and other processor(s) perform other parts of the functions, and situations where a single processor can perform all functions. Additionally, the at least one processor may include a combination of processors performing various functions of the disclosed functions in a distributed manner. The at least one processor may execute program instructions to achieve or perform various functions.

[0095] The processor (130) may be implemented as a general-purpose processor such as a CPU (Central Processing Unit), AP (Application Processor), DSP (Digital Signal Processor), a graphics-dedicated processor such as a GPU (Graphic Processing Unit) or VPU (Vision Processing Unit), or an artificial intelligence-dedicated processor such as an NPU (Neural Processing Unit). The processor (130) may be controlled to process input data according to predefined operation rules or an artificial intelligence model. Alternatively, if the processor (130) is an artificial intelligence-dedicated processor, the artificial intelligence-dedicated processor may be designed with a hardware structure specialized for processing a specific artificial intelligence model.

[0096] The memory (140) may be composed of at least one type of storage medium, such as a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., SD or XD memory), RAM (Random Access Memory), SRAM (Static Random Access Memory), ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), PROM (Programmable Read-Only Memory), or an optical disk.

[0097] The memory (140) may store instructions related to functions and / or operations for the head-mounted display device (100) to adjust the disparity of a stereo image based on a zoom-in input. In one embodiment of the present disclosure, the memory (140) may store at least one of instructions, an algorithm, a data structure, program code, and an application program that can be read by the processor (130). The instructions, algorithm, data structure, and program code stored in the memory (140) may be implemented in a programming or scripting language such as, for example, C, C++, Java, assembler, etc.

[0098] The memory (140) may store instructions, algorithms, data structures, or program codes related to the region of interest determination module (142) and the parallax adjustment module (144). A 'module' included in the memory (140) refers to a unit that processes a function or operation performed by the processor (130), and this may be implemented as software such as instructions, algorithms, data structures, or program code. Although not illustrated in the drawings, the memory (140) according to one embodiment of the present disclosure may further include a data storage that stores a previously acquired stereo image.

[0099] The processor (130) can be implemented by executing instructions or program codes stored in memory (140). Hereinafter, the functions and / or operations performed by the processor (130) by executing the instructions or program codes of each of the plurality of modules stored in memory (140) will be described in detail.

[0100] The processor (130) can control the display (150) to display a stereo image. In one embodiment of the present disclosure, the camera (120) is composed of a plurality of stereo cameras, and the processor (130) can display the stereo image obtained by capturing a real object through the plurality of stereo cameras through the display (150). However, it is not limited thereto, and the head-mounted display device (100) may also obtain image data of a stereo image captured by a plurality of cameras included in an external device, for example, a mobile device, through a wired or wireless communication network. For example, the head-mounted display device (100) includes a communication interface that pairs with an external device through a short-range wireless communication network such as Bluetooth or WiFi Direct, and can receive image data of a stereo image captured by the paired external device through the short-range wireless communication network.

[0101] The stereo image may include a left eye image and a right eye image. The processor (130) may control the display (150) to display the left eye image through the left eye display (150L) and to display the right eye image through the right eye display (150R). The stereo image may be, for example, a video in which parallax correction is performed on a binocular video and the corrected parallax is saved in a side-by-side (SBS) or MV-HEVC (Multiview High Efficiency Video Coding) format. For example, the stereo image may be a still image in which parallax correction is performed on a binocular image and saved in SBS or Samsung Extended Format (SEF).

[0102] The region of interest determination module (142) is composed of instructions or program code for executing a function and / or operation to set a region of interest within a stereo image based on a user's zoom-in input and to determine the location and size of the region of interest. In one embodiment of the present disclosure, the zoom-in input may include at least one of a distance traveled by a user's movement and a pinch zoom input by a hand gesture. By executing the instructions or program code of the region of interest determination module (142), the processor (130) can set a region of interest within a stereo image based on a zoom-in input including at least one of a distance traveled by a user's movement and a pinch zoom input, and determine the location and size of the region of interest. The processor (130) can set a region of interest for each of the left eye image and the right eye image included in the stereo image based on the zoom-in input.

[0103] When the zoom-in input is a positional shift caused by the movement of a user wearing a head-mounted display device (100), the processor (130) can determine the position of the center point of the area of ​​interest based on the positional coordinate values ​​regarding the user's gaze point obtained from the eye-tracking sensor (116). Additionally, the processor (130) can obtain information on the distance of movement caused by the user's movement based on the amount of positional change, including the position, rotation, and translation of the head-mounted display device (100), measured by the position sensor (112) and the IMU sensor (114). The processor (130) can determine the size of the area of ​​interest centered on the positional coordinate values ​​of the center point based on the distance of movement. In one embodiment of the present disclosure, the processor (130) can determine the area size of the area of ​​interest as a value inversely proportional to the distance of movement. For example, the area size of the area of ​​interest can be calculated as a value having an inverse relationship with the distance of movement according to linear, logarithmic, or exponential, etc. A specific embodiment in which the processor (130) determines the location and size of the region of interest based on the user's gaze direction and the distance traveled due to the user's movement will be described in detail with reference to FIGS. 4, 5a, 5b, 5c, and 5d.

[0104] When the zoom-in input is a pinch zoom input by a user's hand gesture, the processor (130) can determine the location and size of the region of interest based on the pinch zoom input. The processor (130) can obtain a hand image by photographing the user's hand performing the pinch zoom gesture through the camera (120) and obtain information regarding the position of the fingers and the drag distance according to the movement of the fingers by analyzing the hand image. In one embodiment of the present disclosure, the region of interest determination module (142) includes an artificial intelligence model trained to recognize hand joint positions or hand gestures, and the processor (130) can obtain information regarding the drag distance caused by the pinch zoom input by analyzing the hand image using the artificial intelligence model. However, it is not limited thereto, and the processor (130) can obtain information regarding the drag distance caused by the pinch zoom input by analyzing the hand image using known general image processing techniques. The processor (130) can obtain an intermediate coordinate value based on the position of the fingers, for example, the position coordinate values ​​of the index finger and thumb, and determine the intermediate coordinate value as the center point of the region of interest. Additionally, the processor (130) may determine the size of the region of interest based on the drag distance. In one embodiment of the present disclosure, the processor (130) may determine the size of the region of interest as a value inversely proportional to the drag distance. For example, the area size of the region of interest may be calculated as a value having an inverse relationship with the drag distance (d) according to linear, logarithmic, or exponential, etc. Specific embodiments in which the processor (130) determines the location and size of the region of interest based on the user's pinch zoom input will be described in detail with reference to FIGS. 6, 7a, and 7b.

[0105] In one embodiment of the present disclosure, when a pinch zoom input is received, the processor (130) determines the center point of the area of ​​interest based on the position coordinate values ​​of the gaze point recognized by the eye tracking sensor (116), and determines the size of the area of ​​interest centered on the center point based on the drag distance of the pinch zoom input. A specific embodiment in which the processor (130) determines the position of the center point of the area of ​​interest based on the gaze direction of the user's two eyes and determines the size of the area of ​​interest based on the drag distance of the pinch zoom input will be described in detail with reference to FIG. 8.

[0106] The disparity adjustment module (144) is composed of instructions or program code for executing a function and / or operation to adjust the disparity between regions of interest set in each of the left-eye image and the right-eye image based on a zoom-in input. The processor (130) can adjust the disparity between regions of interest set in each of the left-eye image and the right-eye image based on a zoom-in input by executing the instructions or program code of the disparity adjustment module (144). In one embodiment of the present disclosure, the processor (130) can perform a first adjustment of the disparity between regions of interest by shifting the region of interest set in the left-eye image and the region of interest set in the right-eye image in opposite directions along the X-axis direction, and then perform a second adjustment by increasing the size of the region of interest based on a zoom-in input to increase the first adjusted disparity by the zoom ratio.

[0107] The first adjustment of the disparity by the processor (130) includes the operation of adjusting the disparity by moving the position of at least one region of interest among the regions of interest set in each of the left eye image and the right eye image by a shift value. In the present disclosure, the 'shift value' is an indicator value representing the amount of positional shift that moves the position in the X-axis direction of the first region of interest set in the left eye image and the second region of interest set in the right eye image to adjust the disparity between the regions of interest set in the stereo image. In one embodiment of the present disclosure, the processor (130) may obtain information regarding the maximum value of the disparity within the region of interest from a disparity map containing information regarding the disparity between the left eye image and the right eye image, and may calculate the shift value based on the maximum value of the disparity within the region of interest. For example, the shift value may be calculated as a value inversely proportional to the maximum value of the disparity of the region of interest. However, it is not limited thereto, and the processor (130) may calculate the shift value based on the ratio of the size of the region of interest to the size of the entire area of ​​the stereo image.

[0108] The processor (130) can adjust the disparity of the regions of interest by shifting the position of the first region of interest set in the left-eye image by a shift value in the first direction along the X-axis direction, and shifting the position of the second region of interest set in the right-eye image by a shift value in the second direction along the X-axis direction, which is opposite to the first direction. For example, the first direction may be the left direction, and the second direction may be the right direction. As the position of the first region of interest set in the left-eye image is shifted by a shift value in the left direction and the position of the second region of interest set in the right-eye image is shifted by a shift value in the right direction, the disparity between the regions of interest in the stereo image may be increased by a value corresponding to twice the shift value compared to the disparity before shifting. An embodiment in which the processor (130) increases the disparity by shifting the first region of interest and the second region of interest in the left direction and the right direction, respectively, will be described in detail with reference to FIG. 11a.

[0109] However, the present disclosure is not limited thereto, and in one embodiment of the present disclosure, the processor (130) may shift only one of the first region of interest set in the left eye image and the second region of interest set in the right eye image by a shift value along the X-axis direction. The processor (130) may shift one region of interest by twice the shift value. For example, the processor (130) may shift the first region of interest to the left along the X-axis by twice the shift value. For example, the processor (130) may shift the second region of interest to the left along the X-axis by twice the shift value. An embodiment in which the processor (130) increases the parallax by shifting only one of the first region of interest and the second region of interest will be described in detail with reference to FIG. 11b and FIG. 11c.

[0110] The processor (130) can enlarge the size of the region of interest to the size of the entire stereo image based on a zoom-in input. The processor (130) can enlarge the size of the region of interest to increase the first-adjusted disparity and thus make a second adjustment to the disparity. As the size of the region of interest is enlarged, the first-adjusted disparity may increase in proportion to the enlarged size of the region of interest. In one embodiment of the present disclosure, the processor (130) can enlarge the size of the first region of interest to be equal to the entire size of the left eye image and enlarge the size of the second region of interest to be equal to the entire size of the right eye image based on a zoom-in input. In this case, the disparity between the enlarged first region of interest and the enlarged second region of interest may be increased by the value obtained by dividing the first-adjusted disparity by the ratio of the size of the region of interest before enlargement to the entire size of the stereo image. A specific embodiment of the processor (130) making a second adjustment to the disparity as it enlarges the size of the region of interest will be described in detail with reference to FIG. 15.

[0111] In one embodiment of the present disclosure, the processor (130) may perform post-processing to improve the quality, such as the resolution of an image, through methods such as pixel interpolation after enlarging the size of the region of interest. Image quality improvement may be achieved using a method utilizing a deep learning model or a method utilizing conventional image processing such as image interpolation. The processor (130) may, for example, use a super-resolution network such as DRCT, SwinIR, or SRCNN to increase the resolution of the enlarged region of interest and improve image quality. Since deep learning models such as DRCT, SwinIR, or SRCNN are identical to those described in FIG. 2, a redundant description is omitted. For example, the processor (130) may improve the image quality of the enlarged region of interest using an image interpolation algorithm such as Bilinear or Bicubic.

[0112] The processor (130) can display high-quality images of the region of interest with increased resolution through the left eye display (150L) and the right eye display (150R).

[0113] The display (150) is configured to display a stereo image under the control of the processor (130). In one embodiment of the present disclosure, the display (150) may include a left eye display (150L) that is positioned adjacent to the user's left eye to display a left eye image while the user is wearing the head-mounted display device (100), and a right eye display (150R) that is positioned adjacent to the user's right eye to display a right eye image.

[0114] The display (150) may be composed of at least one of, for example, a liquid crystal display (LCD), a thin film transistor-liquid crystal display, an organic light-emitting diode (OLED), a flexible display, a 3D display, and an electrophoretic display.

[0115] When the head-mounted display device (100) is implemented as an augmented reality device such as augmented reality glasses or an augmented reality helmet, the display (150) may be composed of a lens optical system and may include a waveguide and an optical engine. The optical engine may be composed of a projector that generates light of a virtual object composed of a virtual image and projects the light onto the waveguide. The optical engine may include, for example, an image panel, a lighting optical system, a projection optical system, etc. In one embodiment of the present disclosure, the optical engine may be placed in the frame or temples of the augmented reality device (e.g., augmented reality glasses).

[0116] FIG. 4 is a flowchart illustrating a method in which a head-mounted display device (100) according to one embodiment of the present disclosure determines the location and size of a region of interest on a stereo image based on the distance traveled due to a user's movement.

[0117] Step S410 illustrated in FIG. 4 represents an operation that embodies the operation of step S210 of FIG. 2. Steps S420 to S450 illustrated in FIG. 4 are steps that embodies the operation of step S220 of FIG. 2. After the operation of step S450 of FIG. 4 is performed, the operation of step S230 of FIG. 2 may be performed.

[0118] In step S410, the head-mounted display device (100) receives a zoom-in input resulting from the user's movement. In one embodiment of the present disclosure, a user wearing the head-mounted display device (100) may move toward a real-world object or change the position and orientation of the head-mounted display device (100) through head movement. The head-mounted display device (100) may measure the amount of change in position and orientation resulting from the user's movement in real time using a position sensor (e.g., a GPS sensor), an IMU sensor, or an eye-tracking sensor.

[0119] In step S420, the head-mounted display device (100) detects a gaze point where the gaze directions of both eyes converge by acquiring information regarding the gaze direction of the user's two eyes using an eye-tracking sensor. FIG. 5a shows a head-mounted display device (100) according to an embodiment of the present disclosure in which the user's two eyes (E L , E R This is a diagram illustrating the operation of determining the location of a region of interest (ROI) based on the gaze direction of the user. Referring to step S420 of FIG. 4 in conjunction with FIG. 5a, the head-mounted display device (100) uses an eye-tracking sensor (116, see FIG. 3) to [determine] the user's binoculars (E L , E R The gaze direction of the ) can be detected. The head-mounted display device (100) uses an eye tracking sensor (116) to detect the left eye (E L Detects the gaze direction of ), and the right eye (E R Detects the gaze direction of ), and the left eye (E L The direction of gaze of ) and the right eye (E R A gaze point (G), which is a point where the gaze direction of the ) converges, can be detected. The processor (130, see FIG. 3) of the head-mounted display device (100) can obtain position coordinate values ​​of the gaze point (G) on the stereo image (i).

[0120] Referring again to FIG. 4, in step S430, the head-mounted display device (100) determines the position coordinates of the center point of the region of interest based on the position of the gaze point. Referring together to the embodiment illustrated in FIG. 5a, the processor (130) of the head-mounted display device (100) determines the center point (P) of the region of interest (ROI) on the stereo image (i) based on the position coordinates of the gaze point. C ) can be determined. In one embodiment of the present disclosure, the center point (P) of the region of interest (ROI) C The position coordinate value of ) can be determined as the same coordinate value as the position coordinate value of the gaze point (G).

[0121] In step S440 of FIG. 4, the head-mounted display device (100) obtains a change in position of the head-mounted display device using at least one sensor and measures the distance traveled due to the user's movement. In one embodiment of the present disclosure, the head-mounted display device (100) may obtain position information using a position sensor (e.g., a GPS sensor) and obtain a change in position including rotation and translation of the head-mounted display device (100) due to the user's movement using an IMU sensor. The processor (130) of the head-mounted display device (100) may measure the distance traveled due to the user's movement based on the change in position.

[0122] In step S450, the head-mounted display device (100) determines the size of the region of interest centered on the position coordinates of the center point based on the measured travel distance. In one embodiment of the present disclosure, the processor (130) of the head-mounted display device (100) determines the center point (P) determined in step S430 based on the measured travel distance. C The size of the region of interest centered on the position coordinate values ​​(see Fig. 5a) can be determined.

[0123] FIG. 5b is a diagram illustrating the operation of a head-mounted display device (100) according to one embodiment of the present disclosure determining the size of a region of interest (ROI) based on the distance traveled (d) caused by the movement of a user. Referring to step S450 of FIG. 4 together with FIG. 5b, a user (10) wearing the head-mounted display device (100) can move toward a real object (50). For example, a stereo image (i) can be obtained by photographing the real object (50) using the camera (120, see FIG. 3) of the head-mounted display device (100), and the stereo image (i) can be displayed through a display (150, see FIG. 3). While the stereo image (i) is being displayed, the user (10) can walk or run toward the real object (50) while wearing the head-mounted display device (100). The processor (130) of the head-mounted display device (100) measures the distance traveled (d) caused by the movement of the user (10) using a position sensor or an IMU sensor, and based on the distance traveled (d), the center point (P) within the stereo image (i) CThe size of the region of interest (ROI) centered on the location coordinate values ​​of the image can be determined. For example, if the area of ​​the stereo image (i) has a width of w in the X-axis direction and a height of h in the Y-axis direction, the processor (130) can determine the size of the region of interest (ROI) by multiplying the width (w) and height (h) of the stereo image (i) by the region of interest ratio value (α). Since the minimum value of the region of interest (ROI) size cannot be 0 and the maximum value of the region of interest (ROI) size is the total size of the stereo image (i), the 'region of interest ratio value (α)' can be determined as a value within the range of greater than 0 and less than or equal to 1. In this case, the region of interest (ROI) can be determined as having a width and height of h, where α is the result of multiplying the region of interest ratio value (α) by the width (w) of the stereo image (i) in the X-axis direction and the region of interest ratio value (α) by the height (h) of the stereo image (i) in the Y-axis direction.

[0124] In one embodiment of the present disclosure, the processor (130) may determine the area size of the region of interest (ROI) to a value inversely proportional to the measured travel distance. In one embodiment of the present disclosure, the processor (130) may calculate a region of interest ratio value (α) for determining the size of the region of interest (ROI) based on the following Equation 1.

[0125]

[0126] FIG. 5c is a graph (500) illustrating the relationship between a user's travel distance (d) and the size of the region of interest according to an embodiment of the present disclosure. The graph (500) illustrated in FIG. 5c illustrates the relationship between the region of interest ratio value (α) and the travel distance (d) calculated by Equation 1. Referring to Equation 1 and the graph (500) illustrated in FIG. 5c, the longer the travel distance (d), the smaller the region of interest ratio value (α), i.e., the size of the region of interest (ROI). In other words, the travel distance (d) and the size of the region of interest (ROI) may have an inverse relationship. Since the size of the region of interest (ROI) in the graph (500) must be 0 or greater, the travel distance (d) reaches a maximum value (d max Even if it exceeds ), the region of interest ratio value (α) is the minimum value (α min ) greater than or equal to, and the minimum value (α) of the region of interest ratio value (α) min ) can be a value greater than 0.

[0127] However, the region of interest ratio value (α) is not limited to the value calculated by the aforementioned Equation 1 or the value shown in the graph (500) of FIG. 5c. FIG. 5d is a graph (510) showing the relationship between the user's travel distance (d) and the size of the region of interest according to one embodiment of the present disclosure. Referring to the graph (510) shown in FIG. 5d, the travel distance (d) and the size of the region of interest (ROI) may have an inverse relationship. For example, the region of interest ratio value (α) may be calculated as a value having an inverse relationship with respect to the travel distance (d) according to linear, logarithmic, or exponential, etc.

[0128] FIG. 6 is a flowchart illustrating a method in which a head-mounted display device (100) according to one embodiment of the present disclosure determines the location and size of a region of interest on a stereo image based on a user's pinch zoom input.

[0129] Step S610 illustrated in FIG. 6 represents an operation that embodies the operation of Step S210 of FIG. 2. In Step S610, the head-mounted display device (100) receives a pinch zoom input from a user that magnifies a stereo image using a hand gesture. The user may perform a pinch zoom gesture by closing and opening their fingers. The head-mounted display device (100) may acquire a hand image by photographing the user's hand performing the hand gesture using a camera (120, see FIG. 3), and may acquire information regarding the position of the fingers and the drag distance according to the movement of the fingers by analyzing the hand image. In one embodiment of the present disclosure, the processor (130, see FIG. 3) of the head-mounted display device (100) may recognize the pinch zoom input by inputting the hand image into an artificial intelligence model trained to recognize hand joint positions or hand gestures, and by analyzing the hand image through inference using the artificial intelligence model.

[0130] Step S620 illustrated in FIG. 6 represents an operation that embodies the operation of Step S220 of FIG. 2. After the operation of Step S620 of FIG. 6 is performed, the operation of Step S230 of FIG. 2 may be performed.

[0131] In step S620, the head-mounted display device (100) determines the size of the region of interest based on a value inversely proportional to the drag distance of the received pinch zoom input. The head-mounted display device (100) can obtain information regarding the drag distance caused by the user's finger movement by analyzing the pinch zoom input, and determine the size of the region of interest based on the obtained drag distance.

[0132] FIG. 7a is a diagram illustrating the operation of a head-mounted display device (100) according to one embodiment of the present disclosure determining the location and size of a region of interest on a stereo image based on a user's pinch zoom input. Referring to step S620 of FIG. 6 together with FIG. 7a, the processor (130) of the head-mounted display device (100) can obtain three-dimensional position coordinate values ​​of the finger positions, for example, the index finger (21) and the thumb (22), based on the analysis result of the hand image, and obtain two-dimensional position coordinate values ​​of points (P1, P2) corresponding to the three-dimensional position coordinate values ​​of the index finger (21) and the thumb (22), respectively, on the stereo image (i). The processor (130) can determine the size of a region of interest (ROI) within a stereo image (i) based on the two-dimensional position coordinate values ​​of a first point (P1) corresponding to the three-dimensional position of the index finger (21) and a second point (P2) corresponding to the three-dimensional position of the thumb (22). In one embodiment of the present disclosure, the processor (130) obtains the two-dimensional position coordinate value of the midpoint between the two-dimensional position coordinate value of the first point (P1) and the two-dimensional position coordinate value of the second point (P2), and the obtained two-dimensional position coordinate value is the center point (P) of the region of interest (ROI). C It can be determined as ). As the drag distance (d) changes due to pinch zoom input, the center point (P C The first distance (d1) toward the first point (P1) from ) and the center point (P CThe second distance (d2) toward the second point (P2) from the first distance (d1) and the second distance (d2) may be changed, and the size of the region of interest (ROI) may be changed as the first distance (d1) and the second distance (d2) are changed. For example, when the drag distance (d) is increased by the user's pinch zoom input, the size of the first distance (d1) and the second distance (d2) is increased, and accordingly, the size of the region of interest (ROI) may be increased. In one embodiment of the present disclosure, the size of the first distance (d1) and the second distance (d2) may be the same.

[0133] As in the embodiment illustrated in FIG. 7a, when the area of ​​the stereo image (i) has a width of w in the X-axis direction and a height of h in the Y-axis direction, the processor (130) can calculate an area of ​​interest ratio value (α) that determines the size of the area of ​​interest (ROI) based on the drag distance from the pinch zoom input. The processor (130) can determine the size of the area of ​​interest (ROI) by multiplying the width (w) and height (h) of the stereo image (i) by the area of ​​interest ratio value (α). Since the minimum value of the area of ​​interest (ROI) size cannot be 0 and the maximum value of the area of ​​interest (ROI) size is the total size of the stereo image (i), the 'area of ​​interest ratio value (α)' can be determined as a value within the range of greater than 0 and less than or equal to 1. In this case, the region of interest (ROI) can be determined with a width and height of α·w, which is the product of the region of interest ratio value (α) and the width (w) of the stereo image (i) in the X-axis direction, and α·h, which is the product of the region of interest ratio value (α) and the height (h) of the stereo image (i) in the Y-axis direction.

[0134] The processor (130) can determine the area size of the region of interest (ROI) to a value inversely proportional to the drag distance (d) by the pinch zoom input. In one embodiment of the present disclosure, the processor (130) can calculate the region of interest ratio value (α) for determining the size of the region of interest (ROI) based on the following Equation 2.

[0135]

[0136] FIG. 7b is a graph (700) illustrating the relationship between the drag distance (d) of a pinch zoom input and the size of the region of interest according to an embodiment of the present disclosure. The graph (700) illustrated in FIG. 7b illustrates the relationship between the region of interest ratio value (α) and the drag distance (d) calculated by Equation 2. Referring to Equation 2 and the graph (700) illustrated in FIG. 7b, the longer the drag distance (d) caused by the pinch zoom input, the smaller the region of interest ratio value (α), i.e., the size of the region of interest (ROI). In other words, the drag distance (d) and the size of the region of interest (ROI) may have an inverse relationship. Since the size of the region of interest (ROI) in the graph (700) must be 0 or greater, the drag distance (d) is at a maximum value (d max Even if it exceeds ), the region of interest ratio value (α) is the minimum value (α min ) greater than or equal to, and the minimum value (α) of the region of interest ratio value (α) min ) can be a value greater than 0.

[0137] However, the region of interest ratio value (α) is not limited to the value calculated by the aforementioned Equation 2 or the value shown in the graph (700) of FIG. 7b. FIG. 7c is a graph (710) showing the relationship between the drag distance (d) of a pinch zoom input and the size of the region of interest according to one embodiment of the present disclosure. Referring to the graph (710) shown in FIG. 7c, the drag distance (d) and the size of the region of interest (ROI) may have an inverse relationship. For example, the region of interest ratio value (α) may be calculated as a value having an inverse relationship with respect to the drag distance (d) according to linear, logarithmic, or exponential, etc.

[0138] In FIGS. 6, FIGS. 7a, and FIGS. 7b, the size of the region of interest (ROI) is shown and described as being determined based on the drag distance (d) of the pinch zoom input, but the present disclosure is not limited thereto. In one embodiment of the present disclosure, the head-mounted display device (100) may receive a zoom input from an external controller.

[0139] An external controller refers to an input controller that is mounted on a part of the user's body or carried by the user. In one embodiment of the present disclosure, the external controller may be paired with a head-mounted display device (100) using a short-range wireless communication network, such as Bluetooth or Wi-Fi Direct. The external controller includes an IMU sensor capable of tracking relative and absolute positions with respect to the head-mounted display device (100), and can acquire three-dimensional position coordinate values ​​of the external controller when the user operates the external controller to move its position. The head-mounted display device (100) can acquire information on three-dimensional position coordinate values ​​from the external controller in real time using a short-range wireless communication network, and can recognize the drag distance of a zoom input using the external controller based on the acquired three-dimensional position coordinate information.

[0140] The head-mounted display device (100) can determine the location and size of an area of ​​interest based on the drag distance caused by the user's external controller operation. Since the method by which the head-mounted display device (100) determines the location and size of an area of ​​interest based on the drag distance is the same as the method described in FIG. 6, FIG. 7a, and FIG. 7b, a redundant description is omitted.

[0141] FIG. 8 shows a head-mounted display device (100) according to one embodiment of the present disclosure, in which a user's two eyes (E L , E RThis is a diagram illustrating the operation of determining the location and size of a region of interest (ROI) on a stereo image (i) based on eye-tracking information and pinch zoom input by hand gesture.

[0142] Referring to FIG. 8, the head-mounted display device (100) uses an eye-tracking sensor (116, see FIG. 3) to track the user's two eyes (E L , E R The gaze direction of the ) can be detected. The head-mounted display device (100) uses an eye tracking sensor (116) to detect the left eye (E L Detects the gaze direction of ), and the right eye (E R Detects the gaze direction of ), and the left eye (E L The direction of gaze of ) and the right eye (E R A gaze point (G), which is the point where the gaze direction of ) converges, can be detected. A processor (130, see FIG. 3) of a head-mounted display device (100) can obtain two-dimensional position coordinate values ​​of the gaze point (G) on a stereo image (i). Based on the two-dimensional position coordinate values ​​of the gaze point (G), the processor (130) [determines] the center point (P) of the region of interest (ROI). C The location of ) can be determined. In one embodiment of the present disclosure, the center point (P) of the region of interest (ROI) C The position coordinate value of ) can be the same as the 2D position coordinate value of the gaze point (G).

[0143] A head-mounted display device (100) can determine the size of a region of interest (ROI) based on the drag distance (d) of a pinch zoom input by a user's hand gesture. In one embodiment of the present disclosure, the head-mounted display device (100) can acquire a hand image by photographing the user's hand (20) using a camera (120, see FIG. 3) and can recognize the positions of fingers (21, 22) by analyzing the hand image. A processor (130) of the head-mounted display device (100) can acquire three-dimensional position coordinate values ​​of the acquired finger positions, for example, the index finger (21) and the thumb (22), based on the analysis result of the hand image, and can determine the size of a region of interest (ROI) based on points (P1, P2) corresponding to the three-dimensional position coordinate values ​​of the index finger (21) and the thumb (22), respectively, on a stereo image (i). The processor (130) can determine the size of the region of interest (ROI) based on the position coordinate values ​​of the points (P1, P2) that change as the drag distance (d) from the pinch zoom input increases or decreases. Since the method by which the processor (130) determines the size of the region of interest (ROI) based on the drag distance (d) from the user's pinch zoom input is the same as described in FIG. 6, FIG. 7a, and FIG. 7b, a redundant description is omitted.

[0144] FIG. 9a is a diagram illustrating the operation of a head-mounted display device (100) according to one embodiment of the present disclosure to change the location of a region of interest (ROI) when the location of the region of interest (ROI) is outside the frame of a stereo image (i).

[0145] Referring to FIG. 9a, when a user views an area of ​​an external virtual background image as well as an area of ​​a stereo image (i), the region of interest (ROI) can be set to extend beyond the frame of the stereo image (i). The head-mounted display device (100) can change the position of the region of interest (ROI) so that the region of interest (ROI) does not extend beyond the frame of the stereo image (i).

[0146] In one embodiment of the present disclosure, a processor (130, see FIG. 3) of a head-mounted display device (100) determines that when the region of interest (ROI) extends beyond the width (w) of a frame of a stereo image (i) in the X-axis direction, the Y-axis direction edge (E) of the region of interest (ROI) ROI The position coordinate values ​​of ) and the Y-axis direction border (E) of the stereo image (i) i The distance between the position coordinate values ​​of ) offset(d off It can be determined as ). For example, the processor (130) can determine the offset (d) through the following Equation 3. off ) can be produced.

[0147]

[0148] Referring to Formula 3 above, offset (d off ) is the center point (P) of the region of interest (ROI). Cx It can be calculated by adding the size (cx) of the X-axis coordinate value of the 2D position coordinate of ) and the value (dx) corresponding to half of the width (α·w) of the region of interest (ROI), and subtracting the width (w) of the stereo image (i) from the sum.

[0149] The processor (130) offsets the location of the region of interest (ROI) by (d off The position of the region of interest (ROI') can be moved into the frame of the stereo image (i) by moving it by ). For example, the center point (P) of the region of interest (ROI') Cx The X-axis coordinate value (cx') of the 2D position coordinates of ') is the center point (P) of the region of interest (ROI) before translation.Cx Offset (d) from the magnitude (cx) of the X-axis coordinate value of ) off It can be determined by the value excluding ).

[0150] FIG. 9b is a diagram illustrating the operation of a head-mounted display device (100) according to one embodiment of the present disclosure to change the location of a region of interest (ROI) when the location of the region of interest (ROI) is outside the frame of a stereo image (i).

[0151] The embodiment illustrated in FIG. 9b is identical to the embodiment illustrated in FIG. 9a except that the position of the region of interest (ROI) is outside the frame of the stereo image (i) in the left direction, so a redundant description will be omitted.

[0152] As in the embodiment illustrated in FIG. 9b, when the position of the region of interest (ROI) extends beyond the width (w) of the frame of the stereo image (i) in the left direction, the processor (130, see FIG. 3) of the head-mounted display device (100) [explains] the left edge (E) in the Y-axis direction of the region of interest (ROI). ROI The position coordinates of ) and the left edge (E) in the Y-axis direction of the stereo image (i) i The distance between the position coordinate values ​​of ) offset(d off It can be determined as ). For example, the processor (130) can determine the offset (d) through the following Equation 4. off ) can be produced.

[0153]

[0154] Referring to Equation 4 above, offset (d off ) is the center point (P) of the region of interest (ROI) at the value (dx) corresponding to half the width (α·w) of the region of interest (ROI). Cx It can be calculated as the value obtained by subtracting the magnitude (cx) of the X-axis coordinate value of the 2D position coordinate of ).

[0155] The processor (130) offsets the location of the region of interest (ROI) by (d offThe position of the region of interest (ROI') can be moved into the frame of the stereo image (i) by moving it by ). For example, the center point (P) of the region of interest (ROI') Cx The X-axis coordinate value (cx') of the 2D position coordinates of ') is the center point (P) of the region of interest (ROI) before translation. Cx The magnitude (cx) and offset (d) of the X-axis coordinate value of ) off It can be determined by the sum of ).

[0156] In the embodiments illustrated in FIGS. 9a and 9b, the case where the location of the region of interest (ROI) extends out of the frame of the stereo image (i) along the X-axis direction is illustrated and described, but the present disclosure is not limited to the case illustrated in the drawings. In one embodiment of the present disclosure, even when the location of the region of interest (ROI) extends out of the frame of the stereo image (i) along the Y-axis direction, the head-mounted display device (100) can change the location of the region of interest (ROI) by applying the same method described in FIGS. 9a and 9b but with a different axis direction.

[0157] FIG. 10 is a flowchart illustrating a method for a head-mounted display device (100) according to one embodiment of the present disclosure to adjust the parallax of a region of interest.

[0158] Steps S1010 and S1020 illustrated in FIG. 10 are steps that embody the operation of step S230 of FIG. 2. Step S1010 illustrated in FIG. 10 can be performed after the operation of step S220 of FIG. 2 has been performed. Step S240 of FIG. 2 can be performed after the operation of step S1020 illustrated in FIG. 10 has been performed.

[0159] In step S1010, the head-mounted display device (100) may determine a shift value for shifting the position of the region of interest based on the size of the region of interest. In one embodiment of the present disclosure, the shift value may be calculated based on a region of interest ratio value representing the size of the region of interest relative to the size of the entire area of ​​the stereo image. However, it is not limited thereto, and the shift value according to one embodiment of the present disclosure may be calculated based on the maximum value of the disparity of the region of interest. Specific embodiments in which the head-mounted display device (100) determines the shift value will be described in detail with reference to FIGS. 10 to 14.

[0160] In step S1020, the head-mounted display device (100) adjusts the disparity between the first region of interest and the second region of interest by shifting at least one of the first region of interest in the left eye image and the second region of interest in the right eye image by a shift value along the X-axis.

[0161] FIG. 11a is a diagram illustrating the operation of a head-mounted display device (100) according to one embodiment of the present disclosure adjusting the parallax (δ) of a region of interest (ROI1, ROI2). Referring to step S1020 of FIG. 10 together with FIG. 11a, the head-mounted display device (100) through a left-eye display (150L) a left-eye image (i L Displays ) and through the right eye display (150R) the right eye image (i R ) can be displayed. The head-mounted display device (100) can set a region of interest based on the user's zoom-in input. The head-mounted display device (100) can display a left-eye image (i L Set the first region of interest (ROI1) in ), and the right eye image (i RA second region of interest (ROI2) can be set in the first region of interest (ROI1) and the second region of interest (ROI2). The position coordinate values ​​of feature points (e.g., pixel values ​​representing mountain peaks) in the first region of interest (ROI1) and the second region of interest (ROI2) may be separated by a specific parallax (δ).

[0162] The processor (130, see FIG. 3) of the head-mounted display device (100) captures the left eye image (i L The position of the first region of interest (ROI1) set in ) is shifted to the left along the X-axis by a shift value (s), and the right eye image (i R Image processing can be performed to move the position of the second region of interest (ROI2) set in ) to the right along the X-axis by a shift value (s). Through image processing that moves the region of interest, the position coordinate value of the center point of the first region of interest (ROI1') is the (x before change c , y c ) shifted along the X-axis in the negative direction (-) by a shift value (-s) (x c -s, y c It can be changed to ). The location coordinates of the center point of the second region of interest (ROI2') are the (x before change c , y c ) is shifted along the X-axis in the positive direction (+) by a shift value (+s) from (x c +s, y c It can be changed to ). As the position of the region of interest (ROI1, ROI2) moves in the opposite direction, the left eye image (i L The first region of interest (ROI1) and right eye image (i) in ) R The disparity (δ) between the second region of interest (ROI2) in ) can be increased by a value corresponding to twice the shift value and changed to δ+2s.

[0163] FIG. 11b is a diagram illustrating the operation of a head-mounted display device (100) according to one embodiment of the present disclosure adjusting the parallax (δ) of a region of interest (ROI1, ROI2).

[0164] Referring to step S1020 of FIG. 10 in conjunction with the embodiment illustrated in FIG. 11b, the processor (130, see FIG. 3) of the head-mounted display device (100) has a left eye image (i L The first region of interest (ROI1) and right eye image (i) set in ) R Only the position of the first region of interest (ROI1) among the second region of interest (ROI2) set in ). In one embodiment of the present disclosure, the processor (130) may perform image processing to move the position of the first region of interest (ROI1) to the left along the X-axis by twice the shift value (s) (2s). The processor (130) may not move the position of the second region of interest (ROI2) and may maintain it as is. Through image processing to move the region of interest, the position coordinate value (x) of the center point of the first region of interest (ROI1) c , y c ) can be shifted along the X-axis in the negative direction (-) by twice the shift value (-2s). The position coordinates of the center point of the modified first region of interest (ROI1') are (x c -2s, y c It can be changed to ). The location coordinates of the center point of the second region of interest (ROI2) (x c , y c ) may not change. As the position of the first region of interest (ROI1) moves, the left eye image (i L The first region of interest (ROI1) and right eye image (i) in ) R The disparity (δ) between the second region of interest (ROI2) in ) can be increased by a value corresponding to twice the shift value and changed to δ+2s.

[0165] In the embodiment illustrated in FIG. 11b, the processor (130) is illustrated and described as having moved the position of the first region of interest (ROI1) by twice the shift value, but the present disclosure is not limited thereto. In one embodiment of the present disclosure, the processor (130) may move the position of the first region of interest (ROI1) by the shift value. In this case, the time difference (δ) between the first region of interest (ROI1) and the second region of interest (ROI2) may be changed to δ+s.

[0166] FIG. 11c is a diagram illustrating the operation of a head-mounted display device (100) according to one embodiment of the present disclosure adjusting the parallax (δ) of a region of interest (ROI1, ROI2).

[0167] The embodiment illustrated in FIG. 11c is identical to the embodiment illustrated in FIG. 11b except that the position of the second region of interest (ROI2) is moved without moving the position of the first region of interest (ROI1), so redundant description is omitted.

[0168] Referring to step S1020 of FIG. 10 in conjunction with the embodiment illustrated in FIG. 11c, the processor (130, see FIG. 3) of the head-mounted display device (100) captures a right eye image (i R Image processing can be performed to move the position of the second region of interest (ROI2) set in ) along the X-axis to the right by twice the shift value (s) (2s). The processor (130) can maintain the position of the first region of interest (ROI1) without moving it. Through image processing that moves the region of interest, the position coordinate value (x) of the center point of the second region of interest (ROI2) c , y c ) can be moved along the X-axis in the positive direction (+) by twice the shift value (+2s). The position coordinates of the center point of the modified second region of interest (ROI2') are (x c +2s, y cIt can be changed to ). The location coordinate value of the center point of the first region of interest (ROI1) (x c , y c ) may not change. As the position of the second region of interest (ROI2) moves, the left eye image (i L The first region of interest (ROI1) and right eye image (i) in ) R The disparity (δ) between the second region of interest (ROI2) in ) can be increased by a value corresponding to twice the shift value and changed to δ+2s.

[0169] In the embodiment illustrated in FIG. 11c, the processor (130) is illustrated and described as having moved the position of the second region of interest (ROI2) by twice the shift value, but the present disclosure is not limited thereto. In one embodiment of the present disclosure, the processor (130) may move the position of the second region of interest (ROI2) by the shift value. In this case, the time difference (δ) between the first region of interest (ROI1) and the second region of interest (ROI2) may be changed to δ+s.

[0170] FIG. 12 is a flowchart illustrating a method for a head-mounted display device (100) according to one embodiment of the present disclosure to determine a shift value for parallax adjustment of a region of interest.

[0171] Steps S1210 to S1250 illustrated in FIG. 12 are steps that embody the operation of step S1010 of FIG. 10. Step S1210 illustrated in FIG. 12 can be performed after the operation of step S220 of FIG. 2 has been performed.

[0172] In step S1210, the head-mounted display device (100) identifies whether a disparity map is stored. The disparity map is a map containing information regarding the disparity of a stereo image, which may be acquired in advance and stored in the storage space of the memory (140, see FIG. 3) of the head-mounted display device (100). However, it is not limited thereto, and the head-mounted display device (100) may not store a disparity map.

[0173] When a parallax map is stored in the memory (140) of the head-mounted display device (100) (step S1220), the head-mounted display device (100) obtains the maximum value of the parallax of the region of interest from the parallax map. In one embodiment of the present disclosure, the processor (130, see FIG. 3) of the head-mounted display device (100) obtains parallax information of the region corresponding to the region of interest from the parallax map and can identify the maximum value of the parallax from the obtained parallax information.

[0174] FIG. 13 is a disparity map (D) for explaining the operation of a head-mounted display device (100) according to one embodiment of the present disclosure determining a shift value. i ) is a drawing illustrating. Referring to step S1220 of FIG. 12 together with FIG. 13, the head-mounted display device (100) has a parallax map (D i Obtain parallax information regarding the entire area of ​​the stereo image from ), and the region of interest (D ROI Information regarding the maximum value of the time difference within ) can be obtained.

[0175] Referring again to FIG. 12, in step S1230, the head-mounted display device (100) calculates a shift value based on the maximum value of the parallax of the region of interest. In one embodiment of the present disclosure, the processor (130) of the head-mounted display device (100) may calculate a shift value based on a value inversely proportional to the maximum value of the parallax of the region of interest. For example, the shift value may be calculated by the following Equation 5.

[0176]

[0177] Referring to Equation 5, the shift value(s) is the maximum value among the parallax of the entire stereo image (max(D i The maximum value among the disparities of the region of interest in )) (max(D ROI It can be calculated as a value corresponding to half of the value obtained by subtracting )). In Equation 5, the maximum value of the shift value(s) is the maximum value of the parallax of the entire stereo image (max(D i It may be calculated as a value corresponding to 1 / 2 of )), but the present disclosure is not limited thereto. In Formula 5, max(D i Instead of ), the maximum value (s) is pre-set to an arbitrary value. max ) may also be applied.

[0178] However, if a previously stored parallax map exists, the shift value(s) is not limited to being calculated by the above Equation 5. In one embodiment of the present disclosure, the shift value(s) is the maximum parallax value of the region of interest (max(D ROI It can be calculated as a value having an inverse relationship with respect to )) linear, logarithmic, or exponential, etc.

[0179] If the parallax map is not stored in the memory (140) of the head-mounted display device (100) (step S1240), the head-mounted display device (100) calculates a region of interest ratio value, which is the ratio of the size of the region of interest to the size of the entire area of ​​the stereo image. In one embodiment of the present disclosure, the processor (130) of the head-mounted display device (100) can calculate the region of interest ratio value through an operation of dividing the size of the region of interest by the entire size of the stereo image.

[0180] In step S1250, the head-mounted display device (100) calculates a shift value based on the region of interest ratio value. In one embodiment of the present disclosure, the processor (130) may calculate a shift value inversely proportional to the region of interest ratio value. For example, the shift value(s) may be calculated by the following Equation 6.

[0181]

[0182] FIG. 14a is a graph (1400) illustrating the relationship between the size of the region of interest and the shift value (s) according to an embodiment of the present disclosure. The graph (1400) illustrated in FIG. 14a illustrates the relationship between the region of interest ratio value (α) and the shift value (s). Referring to Equation 6 and the graph (1400) illustrated in FIG. 14a, the larger the region of interest ratio value (α), the smaller the size of the shift value (s). That is, the region of interest ratio value (α) and the size of the shift value (s) may have an inverse relationship. Since the size of the region of interest in the graph (1400) must be 0 or greater, the minimum value (α) of the region of interest ratio value (α) min ) is a value greater than 0, and the region of interest ratio value (α) is the minimum value (α min When ), the shift value(s) is the maximum value(s max It can be.

[0183] However, the shift value(s) is not limited to the value calculated by the aforementioned formula 6 or the value shown in the graph (1400) of FIG. 14a.

[0184] FIG. 14b is a graph (1410) illustrating the relationship between the size of the region of interest and the shift value (s) according to one embodiment of the present disclosure. Referring to the graph (1410) shown in FIG. 14b, the shift value (s) can be calculated as a value having an inverse relationship with respect to the region of interest ratio value (α) in a linear manner. However, it is not limited thereto, and the shift value (s) can also be calculated as a value having an inverse relationship with respect to the region of interest ratio value (α) in a logarithmic or exponential manner.

[0185] FIG. 15 is a diagram illustrating the operation of a head-mounted display device (100) according to one embodiment of the present disclosure, which increases the disparity of a region of interest (ROI1, ROI2) by expanding the region of interest based on a zoom-in input.

[0186] Referring to FIG. 15, the head-mounted display device (100) displays the left eye image (i) of the stereo image through the left eye display (150L). L Displays ) and through the right eye display (150R) the right eye image (i R Can display ). Left eye image (i L The first region of interest (ROI1) and right eye image (i) set in ) R The second region of interest (ROI2) set in ) can be set to a size determined by the width of α·w and the height of α·h. The parallax between the first region of interest (ROI1) and the second region of interest (ROI2) can be δ+2s, which is the original parallax (δ) before the shift of the region of interest increased by twice the shift value (s).

[0187] The head-mounted display device (100) determines the size of the region of interest (ROI1, ROI2) based on the user's zoom-in input using a stereo image (i L , i R ) can be enlarged to the full size. For example, the processor (130, see FIG. 3) of the head-mounted display device (100) can enlarge the size of a region of interest (ROI1, ROI2) having a width of α·w and a height of α·h to a stereo image (i) having a width of w and a height of h. L , i R ) It can be enlarged to the full size. Depending on the size of the region of interest (ROI1', ROI2') enlarged by the zoom-in input, the disparity may be increased. In one embodiment of the present disclosure, the disparity between the enlarged first region of interest (ROI1') and the enlarged second region of interest (ROI2') may be increased from δ+2s to (δ+2s) / α, which is the value obtained by dividing δ+2s by the region of interest ratio value (α). Since the region of interest ratio value (α) is determined to be a value within the range of greater than 0 and less than or equal to 1, the disparity ((δ+2s) / α) that increases due to the enlargement of the size of the region of interest (ROI1', ROI2') may be a value greater than the disparity (δ+2s) of the region of interest (ROI1, ROI2) before enlargement.

[0188] A head-mounted display device (100) according to the embodiment illustrated in FIGS. 10 to 15 receives a user's zoom-in input, and the left eye image (i L The first region of interest (ROI1) set in ) is directed to the left, and the right eye image (i RBy moving the second region of interest (ROI2) set in the ) to the right, a first adjustment regarding the parallax between regions of interest is performed, and by enlarging the size of the regions of interest (ROI1, ROI2), a second adjustment regarding the parallax is performed, thereby providing a technical effect that provides a sense of depth to the user and enhances the user's immersion. In particular, the head-mounted display device (100) according to one embodiment of the present disclosure can provide a sense of depth even to areas (e.g., regions of interest) where the user did not perceive a sense of depth in existing captured spatial image content. Furthermore, the head-mounted display device (100) according to one embodiment of the present disclosure can provide intuitive results using the same user experience (UX) even within various spatial image content.

[0189] FIG. 16 is a flowchart illustrating a method in which a head-mounted display device (100) according to one embodiment of the present disclosure adjusts the parallax of a stereo image based on a user's zoom-in input.

[0190] In step S1610, the head-mounted display device (100) receives a user's zoom-in input to enlarge the size of the stereo image while displaying the stereo image.

[0191] In step S1620, the head-mounted display device (100) determines the location and size of the region of interest within the entire area of ​​the stereo image based on the zoom-in input.

[0192] Steps S1610 and S1620 are identical to steps S210 and S220 shown in FIG. 2, respectively, so redundant descriptions are omitted.

[0193] In step S1630, the head-mounted display device (100) increases the size of the region of interest to the size of the entire stereo image area based on the zoom-in input, thereby increasing the disparity of the region of interest. In one embodiment of the present disclosure, the processor (130, see FIG. 3) of the head-mounted display device (100) increases the size of the region of interest based on the zoom-in input and can increase the disparity between the first region of interest set in the left-eye image and the second region of interest set in the right-eye image by the value obtained by dividing the original disparity before the expansion of the region of interest size by the region of interest ratio value (α). The region of interest ratio value (α) can be determined in step S1620. The region of interest ratio value (α) represents a ratio value indicating the size of the region of interest relative to the size of the entire stereo image area.

[0194] In step S1640, the head-mounted display device (100) adjusts the disparity of the region of interest by shifting the position of the enlarged region of interest in each of the left-eye image and the right-eye image by a shift value. In one embodiment of the present disclosure, the processor (130) may shift the position of the enlarged first region of interest in a first direction along the X-axis by a shift value, and shift the position of the enlarged second region of interest in a second direction along the X-axis by a shift value, which is opposite to the first direction. For example, the first direction may be the left direction and the second direction may be the right direction. The disparity between the enlarged first region of interest and the enlarged second region of interest may be increased by a value corresponding to twice the shift value relative to the disparity before the shift.

[0195] FIG. 17 is a conceptual diagram illustrating the operation of a head-mounted display device (100) according to one embodiment of the present disclosure adjusting the disparity of a stereo image (i) based on a user's zoom-out input.

[0196] Referring to FIG. 17, the head-mounted display device (100) has a left eye image (i L ) and right eye image(i R A stereo image including ) can be displayed (Operation ①).

[0197] The head-mounted display device (100) receives a zoom-out input from the user to reduce the stereo image while the stereo image is being displayed, and can determine a region of interest (ROI) based on the received zoom-out input (operation ②).

[0198] The head-mounted display device (100) reduces the size of the region of interest (ROI1, ROI2) based on the zoom-out input, and the left eye image (i L ) and right eye image(i R The parallax (δ) between the first region of interest (ROI1') and the second region of interest (ROI2'), each set in ), can be adjusted (Operation ③).

[0199] FIG. 18 is a flowchart illustrating a method in which a head-mounted display device (100) according to one embodiment of the present disclosure adjusts the disparity of a stereo image based on a user's zoom-out input.

[0200] Hereinafter, with reference to FIG. 17 and FIG. 18 together, the function and / or operation of the head-mounted display device (100) adjusting the parallax of a stereo image as a zoom-out input is received will be described in detail.

[0201] In step S1810 of FIG. 18, the head-mounted display device (100) receives a zoom-out input from a user to reduce the size of the stereo image while displaying the stereo image. The head-mounted display device (100) may display a stereo image including a left-eye image and a right-eye image. Referring together to operation 1 of FIG. 17, the head-mounted display device (100) includes a left-eye display (150L) and a right-eye display (150R), and through the left-eye display (150L), the left-eye image (i) of the stereo image L Displays ) and through the right eye display (150R) the right eye image (i R It can display ).

[0202] The head-mounted display device (100) has a left eye image (i L ) and right eye image(i R The device can receive a user's zoom-out input that reduces the size of the object. In one embodiment of the present disclosure, the user's zoom-out input may include at least one of a positional displacement distance caused by the user's movement and a pinch zoom input caused by a hand gesture. Referring together to operation ① illustrated in FIG. 17, the head-mounted display device (100) can receive a zoom-out input, for example, when a user (10) wearing the head-mounted display device (100) moves in a direction opposite to the real object, that is, a direction moving away from the real object. For example, the head-mounted display device (100) can receive a zoom-out input by capturing the hand (20) of the user (10) who performs a hand gesture of curling using the index finger and thumb of the hand (20), thereby acquiring a hand image, and by analyzing the hand image to recognize the pinch zoom input.

[0203] Referring again to FIG. 18, in step S1820, the head-mounted display device (100) determines the location and size of the region of interest within the entire area of ​​the stereo image based on the user's zoom-out input. Referring together to operation ② of FIG. 17, the head-mounted display device (100) [describes] the left eye image (i L ) and right eye image(i R While a stereo image including ) is being displayed, a user's zoom-out input to reduce the size of the stereo image is received, and the location and size of the region of interest (ROI1, ROI2) can be determined based on the zoom-out input.

[0204] When a zoom-out input is received due to a position change caused by the movement of a user (10) wearing a head-mounted display device (100), the head-mounted display device (100) can recognize the direction of gaze of both eyes of the user (10) using an eye-tracking sensor and detect a gaze point where the direction of gaze of both eyes converges. Based on the location of the detected gaze point, the head-mounted display device (100) can determine the position coordinate values ​​of the center points (Pc1, Pc2) of the region of interest. The head-mounted display device (100) can acquire a change in position including the position, rotation, and translation of the head-mounted display device (100) caused by the movement of the user (10) using at least one of a GPS sensor and an IMU sensor (inertial measurement unit), and thereby measure the distance traveled (d) caused by the movement of the user (10). The head-mounted display device (100) can determine the size of the region of interest centered on the position coordinate values ​​of the center points (Pc1, Pc2) based on the measured travel distance (d). In one embodiment of the present disclosure, the head-mounted display device (100) can determine the area size of the region of interest to a value inversely proportional to the measured travel distance (d).

[0205] When a pinch zoom input is received via a hand gesture of the user (10), the head-mounted display device (100) can determine the location and size of the region of interest based on the pinch zoom input. Referring to the embodiment illustrated in FIG. 17, the head-mounted display device (100) can acquire a hand image by photographing the user's (10) hand (20) using a front camera (120, see FIG. 3), and can acquire information regarding the position of the fingers and the drag distance (d) according to the movement of the fingers by analyzing the hand image using an artificial intelligence model trained to recognize hand joint positions or hand gestures or known image processing. The head-mounted display device (100) can determine the median coordinate value among the position coordinate values ​​of the fingers as the center point (Pc1, Pc2) of the region of interest and determine the size of the region of interest based on the drag distance (d) by the fingers. The head-mounted display device (100) can determine the size of the region of interest as a value inversely proportional to the drag distance (d).

[0206] Although not illustrated in the drawings, in one embodiment of the present disclosure, the head-mounted display device (100) may receive a zoom-out input from an external controller. Since the specific method for determining the location and size of the region of interest based on the zoom-out input from the external controller is identical to the pinch-zoom input except that drag input is performed using the controller instead of the user's finger, a redundant description is omitted.

[0207] A head-mounted display device (100) includes a left eye image (i) in a stereo image based on a zoom-out input. L ) and right eye image(i R A region of interest can be set for each of the following. Referring to the embodiment illustrated in FIG. 17, the head-mounted display device (100) has a left eye image (i LDetermine the location and size of the first region of interest (ROI1) within ), and the right eye image (i R The location and size of the second region of interest (ROI2) can be determined within ).

[0208] FIG. 19 is a graph (1900) illustrating the relationship between the size of a region of interest and the user's movement distance (d) or the drag distance (d) caused by pinch zoom input according to one embodiment of the present disclosure. Referring to step S1820 of FIG. 18 together with the graph (1900) shown in FIG. 19, the processor (130, see FIG. 3) of the head-mounted display device (100) can determine the size of the region of interest ratio value (α), which is proportional to the area size of the region of interest, as a value inversely proportional to the user's movement distance (d) or the drag distance (d) caused by pinch zoom input. The 'region of interest ratio value (α)' represents a value obtained by dividing the size of the region of interest by the size of the entire area of ​​the stereo image. In one embodiment of the present disclosure, the processor (130) can calculate the region of interest ratio value (α) for determining the size of the region of interest (ROI) based on the following Equation 7.

[0209]

[0210] Referring to Equation 7 and the graph (1900) illustrated in FIG. 19, the longer the travel distance (d) or drag distance (d), the smaller the region of interest ratio value (α), i.e., the size of the region of interest. That is, the travel distance (d) or drag distance (d) and the size of the region of interest may have an inverse relationship. However, the region of interest ratio value (α) is not limited to the value calculated by the aforementioned Equation 7 or the value illustrated in the graph (1900) of FIG. 19. For example, the region of interest ratio value (α) may be calculated as a value having an inverse relationship with respect to the travel distance (d) according to linear, logarithmic, or exponential, etc.

[0211] In the graph (1900), the movement distance (d) or drag distance (d) is at its maximum value (d max Even if it becomes longer with a value exceeding ), the magnitude of the region of interest ratio value (α) is the minimum value (α min It does not become smaller than ). In one embodiment of the present disclosure, the minimum value (α of the region of interest ratio) min ) can be determined based on a predetermined minimum value of the resolution of the stereo image content. That is, when the user moves away from a real-world object or zooms out the image via pinch zoom input, the resolution of the stereo image content decreases, but the movement distance (d) or the drag distance (d) of the pinch zoom input exceeds a certain level (e.g., the maximum value of d (d) max When the size becomes )), there may be no further change in the stereo image content. In this case, the size of the region of interest ratio value (α), which represents the size of the region of interest, becomes the minimum value (α min It can be determined as ). Afterwards, even if the movement distance (d) or drag distance (d) increases, the magnitude of the region of interest ratio value (α) becomes the minimum value (α min It does not decrease below )

[0212] Referring again to FIG. 18, in step S1830, the head-mounted display device (100) reduces the size of the region of interest based on the zoom-out input. Referring together to operation ③ of FIG. 17, the head-mounted display device (100) based on the zoom-out input, the left eye image (i L The first region of interest (ROI1) and right eye image (i) set in ) R The size of the second region of interest (ROI2) set in ) can be reduced. In one embodiment of the present disclosure, the processor (130) of the head-mounted display device (100) can determine the reduction ratio of the regions of interest (ROI1, ROI2) based on a value proportional to the user's movement distance (d) or the drag distance (d) by pinch zoom input.

[0213] In step S1840 of FIG. 18, the head-mounted display device (100) adjusts the disparity between the regions of interest in the left-eye image and the right-eye image, respectively, as the size of the region of interest is reduced. Referring together to the embodiment illustrated in FIG. 17, the disparity (δ') between the first region of interest (ROI1') and the second region of interest (ROI2') reduced by the zoom-out input may be smaller than the disparity (δ) between the first region of interest (ROI1) and the second region of interest (ROI2) before reduction. In one embodiment of the present disclosure, even if the size of the regions of interest (ROI1', ROI2') is reduced by the zoom-out input, the reduced disparity (δ') may have a value greater than or equal to 0. That is, the minimum value of the disparity (δ') may be a value greater than or equal to 0.

[0214] The head-mounted display device (100) according to the embodiment illustrated in FIGS. 17 to 19 can provide the user with a sense of depth in the opposite direction to zoom-in by reducing the size of the region of interest (ROI1, ROI2) and reducing the parallax between the regions of interest (ROI1, ROI2) when the user moves away from a real object or reduces the size of the stereo image through pinch zoom input, thereby improving the user's sense of immersion.

[0215] One aspect of the present disclosure provides a method for a head-mounted display device (100) to change the disparity of a stereo image. A method of operation of a head-mounted display device (100) according to one embodiment of the present disclosure may include a step (S210) of receiving a user's zoom-in input to enlarge the size of a stereo image while displaying a stereo image including a left-eye image and a right-eye image. A method of operation of a head-mounted display device (100) according to one embodiment of the present disclosure may include a step (S220) of determining the location and size of a region of interest within the entire area of ​​the stereo image based on the received zoom-in input. A method of operation of a head-mounted display device (100) according to one embodiment of the present disclosure may include a step (S230) of adjusting the disparity between the region of interest in the left-eye image and the right-eye image by shifting the location of the region of interest determined in at least one of the left-eye image and the right-eye image by a shift value calculated to have an inverse relationship with the size of the region of interest. A method of operation of a head-mounted display device (100) according to one embodiment of the present disclosure may include a step (S240) of increasing the adjusted parallax by expanding the size of the region of interest to the size of the entire region of the stereo image.

[0216] In one embodiment of the present disclosure, the user’s zoom-in input may include a positional movement of the head-mounted display device (100) caused by the movement of the user wearing the head-mounted display device (100). The step of determining the position and size of the area of ​​interest (S220) may include a step of recognizing a gaze point where the gaze directions of both eyes converge by obtaining information regarding the gaze directions of the user’s two eyes using an eye-tracking sensor (116) (S420), and a step of determining the positional coordinates of the center point of the area of ​​interest based on the position of the recognized gaze point (S430). The step of determining the location and size of the area of ​​interest (S220) may include the step of measuring the distance traveled due to the user's movement by obtaining a change in location including the position, rotation, and translation of the head-mounted display device (100) using at least one of a GPS sensor and an IMU sensor (inertial measurement unit) (S440), and the step of determining the size of the area of ​​interest centered on the location coordinate value of the center point determined based on the measured distance traveled (S450).

[0217] In one embodiment of the present disclosure, the size of the region of interest may be determined to be a value inversely proportional to the user's travel distance.

[0218] In one embodiment of the present disclosure, the step of receiving the user's zoom-in input (S210) may include the step of receiving the user's pinch zoom input (S610) for enlarging a stereo image using a hand gesture. The step of determining the location and size of the region of interest (S220) may include the step of determining the size of the region of interest (S620) based on a value inversely proportional to the drag distance of the received pinch zoom input.

[0219] In one embodiment of the present disclosure, the step (S220) of determining the location and size of the region of interest may include the step of determining the location coordinates of the center point of the region of interest based on the direction of the user's eyes of sight obtained using an eye-tracking sensor (116).

[0220] In one embodiment of the present disclosure, the step (S230) of adjusting the parallax of the stereo images may include the step (S1020) of shifting the position of a first region of interest corresponding to a region of interest in the left eye image by a shift value in a first direction along the X-axis, and shifting the position of a second region of interest corresponding to a region of interest in the right eye image by a shift value in a second direction opposite to the first direction along the X-axis.

[0221] In one embodiment of the present disclosure, the step of adjusting the disparity of the stereo image (S230) may include the step of obtaining a maximum value of the disparity of a region of interest from a disparity map containing information regarding the disparity between a left eye image and a right eye image (S1220), and the step of calculating a shift value based on the obtained maximum value of the disparity (S1230). The shift value may be calculated as a value inversely proportional to the maximum value of the disparity within the region of interest.

[0222] In one embodiment of the present disclosure, the step of adjusting the parallax of the stereo image (S230) may include the step (S1240, S1250) of calculating the shift value based on an interest region ratio value representing the ratio of the size of the interest region to the size of the entire area of ​​the stereo image. The shift value may be calculated as a value inversely proportional to the interest region ratio value.

[0223] In one embodiment of the present disclosure, the step of increasing the adjusted disparity (S240) may include increasing the adjusted disparity as the size of the region of interest is enlarged by a value obtained by dividing the adjusted disparity by the ratio between the size of the region of interest before enlargement and the total size of the stereo image.

[0224] One aspect of the present disclosure provides a head-mounted display device (100) for adjusting the parallax of a stereo image. In one embodiment of the present disclosure, the head-mounted display device (100) may include a sensor (110) comprising at least one of a position sensor (112) and an Inertial Measurement Unit (IMU) sensor (114); a left-eye display (150L) for displaying a left-eye image included in the stereo image, and a right-eye display (150R) for displaying a right-eye image included in the stereo image; at least one processor (130) comprising processing circuitry; and a memory (140) for storing one or more instructions. By executing one or more of the above commands individually or collectively by at least one processor (130), the head-mounted display device (100) can receive a user's zoom-in input to enlarge the size of the stereo image and determine the location and size of the region of interest within the entire area of ​​the stereo image based on the received zoom-in input. By executing one or more of the above commands individually or collectively by at least one processor (130), the head-mounted display device (100) can adjust the disparity between the region of interest in the left-eye image and the right-eye image by shifting the location of the region of interest determined in at least one of the left-eye image and the right-eye image by a shift value calculated to have an inverse relationship with the size of the region of interest.By executing one or more of the above instructions individually or collectively by at least one processor (130), the head-mounted display device (100) can increase the size of the region of interest to the size of the entire region of the stereo image, thereby increasing the adjusted parallax.

[0225] In one embodiment of the present disclosure, the user’s zoom-in input may include a positional movement of the head-mounted display device (100) caused by the movement of the user wearing the head-mounted display device (100). By executing one or more of the commands individually or collectively by at least one processor (130), the head-mounted display device (100) can acquire information regarding the direction of the user’s gaze of both eyes using an eye-tracking sensor (116), thereby recognizing a gaze point where the directions of the gaze of both eyes converge, and determine the positional coordinates of the center point of the region of interest based on the location of the recognized gaze point. By executing one or more of the above commands individually or collectively by at least one processor (130), the head-mounted display device (100) can obtain a change in position including the position, rotation, and translation of the head-mounted display device (100) using a sensor (110), thereby measuring the distance traveled due to the user's movement and determining the size of a region of interest centered on the position coordinates of a center point determined based on the measured distance traveled.

[0226] In one embodiment of the present disclosure, the size of the region of interest may be determined to be a value inversely proportional to the user's travel distance.

[0227] In one embodiment of the present disclosure, by executing one or more of the instructions individually or collectively by at least one processor (130), the head-mounted display device (100) receives a pinch zoom input from a user that enlarges a stereo image using a hand gesture, and can determine the size of the region of interest based on a value inversely proportional to the drag distance of the received pinch zoom input.

[0228] In one embodiment of the present disclosure, by executing one or more of the instructions individually or collectively by at least one processor (130), the head-mounted display device (100) can determine the position coordinates of the center point of the region of interest based on the direction of the user's eyes of sight obtained using the eye tracking sensor (116).

[0229] In one embodiment of the present disclosure, by executing one or more instructions individually or collectively by at least one processor (130), the head-mounted display device (100) can shift the position of a first region of interest corresponding to a region of interest in the left-eye image by a shift value in a first direction along the X-axis, and shift the position of a second region of interest corresponding to a region of interest in the right-eye image by a shift value in a second direction opposite to the first direction along the X-axis.

[0230] In one embodiment of the present disclosure, based on the positional shift of the first region of interest and the second region of interest, the disparity between the first region of interest and the second region of interest may be increased by twice the shift value compared to the disparity before adjustment.

[0231] In one embodiment of the present disclosure, by executing one or more instructions individually or collectively by at least one processor (130), the head-mounted display device (100) may obtain a maximum value of disparity in a region of interest from a disparity map containing information regarding the disparity between a left-eye image and a right-eye image, and calculate a shift value based on the obtained maximum value of disparity. The shift value may be calculated as a value inversely proportional to the maximum value of disparity within the region of interest.

[0232] In one embodiment of the present disclosure, by executing one or more instructions individually or collectively by at least one processor (130), the head-mounted display device (100) can calculate a shift value based on a region of interest ratio value representing the ratio of the size of the region of interest to the size of the entire region of the stereo image. The shift value may be calculated as a value inversely proportional to the region of interest ratio value.

[0233] In one embodiment of the present disclosure, the head-mounted display device (100) can increase the parallax adjusted as the size of the region of interest is enlarged by a value obtained by dividing the ratio between the size of the region of interest before enlargement and the size of the stereo image before enlargement.

[0234] One aspect of the present disclosure provides a computer program product comprising a computer-readable storage medium. The storage medium may include instructions readable by a head-mounted display device (100) for the head-mounted display device (100) to perform the following operations: displaying a stereo image including a left eye image and a right eye image; receiving a user's zoom-in input to enlarge the size of the stereo image; determining the location and size of a region of interest within the entire area of ​​the stereo image based on the received zoom-in input; shifting the location of the region of interest in each of the left eye image and the right eye image by a shift value calculated to have an inverse relationship with the size of the region of interest to adjust the disparity between the region of interest in the left eye image and the right eye image; and increasing the size of the region of interest to the size of the entire area of ​​the stereo image, thereby increasing the adjusted disparity.

[0235] One aspect of the present disclosure provides a method for a head-mounted display device (100) to change the disparity of a stereo image. In one embodiment of the present disclosure, the method of operation of the head-mounted display device (100) may include a step (S1810) of receiving a user's zoom-out input to reduce the size of the stereo image while displaying a stereo image including a left-eye image and a right-eye image. The method of operation of the head-mounted display device (100) may include a step (S1820) of determining the location and size of a region of interest within the entire area of ​​the stereo image based on the received zoom-out input. The method of operation of the head-mounted display device (100) may include a step (S1830) of reducing the size of the region of interest based on the zoom-out input. The method of operation of the head-mounted display device (100) may include a step (S1840) of adjusting the disparity between the region of interest in each of the left-eye image and the right-eye image as the size of the region of interest is reduced. The size of the region of interest can be determined to be inversely proportional to the amount of movement of the received zoom-out input. The minimum value of the size of the region of interest can be determined based on the minimum value of the resolution of the preset stereo image.

[0236] One aspect of the present disclosure provides a computer program product comprising a computer-readable storage medium. The storage medium may include instructions readable by a head-mounted display device (100) for the head-mounted display device (100) to perform the following operations: displaying a stereo image including a left eye image and a right eye image; receiving a user's zoom-out input to reduce the size of the stereo image; determining the location and size of a region of interest within the entire area of ​​the stereo image based on the received zoom-out input; reducing the size of the region of interest based on the zoom-out input; and adjusting the disparity between the region of interest in each of the left eye image and the right eye image as the size of the region of interest is reduced.

[0237] A program executed by the head-mounted display device (100) described in the present disclosure may be implemented as a hardware component, a software component, and / or a combination of a hardware component and a software component. The program may be executed by any system capable of executing computer-readable instructions.

[0238] Software may include a computer program, code, instructions, or a combination of one or more of these, and may configure a processing unit to operate as desired or command the processing unit independently or collectively.

[0239] Software can be implemented as a computer program containing instructions stored on a computer-readable storage medium. Examples of computer-readable recording media include magnetic storage media (e.g., ROM (read-only memory), RAM (random-access memory), floppy disks, hard disks, etc.) and optical reading media (e.g., CD-ROMs, DVDs (Digital Versatile Discs)). Computer-readable recording media can be distributed across networked computer systems, allowing computer-readable code to be stored and executed in a distributed manner. The medium is readable by a computer, stored in memory, and can be executed by a processor.

[0240] Computer-readable storage media may be provided in the form of non-transitory storage media. Here, 'non-transitory' means only that the storage medium does not contain a signal and is tangible, and does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium. For example, a 'non-transitory storage medium' may include a buffer in which data is stored temporarily.

[0241] In addition, the program according to the embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product.

[0242] A computer program product may include a software program and a computer-readable storage medium on which the software program is stored. For example, the computer program product may be from the manufacturer of the head-mounted display device (100) or an electronics market (e.g., Samsung Galaxy Store). TMIt may include a product in the form of a software program that is distributed electronically through ). For electronic distribution, at least a portion of the software program may be stored on a storage medium or temporarily created. In this case, the storage medium may be a server of the manufacturer of the head-mounted display device (100), a server of an electronic market, or a storage medium of a relay server that temporarily stores the software program.

[0243] A computer program product may include a storage medium of a server or a storage medium of a head-mounted display device (100) in a system composed of a head-mounted display device (100) and / or a server. Alternatively, if there is a third device (e.g., a mobile device such as a 'smartphone' or 'tablet PC') that is communicationally connected to the head-mounted display device (100), the computer program product may include a storage medium of the third device. Alternatively, the computer program product may include a software program itself that is transmitted from the head-mounted display device (100) to the third device or from the third device to the head-mounted display device (100).

[0244] In this case, either the head-mounted display device (100) or one of the third devices may execute a computer program product to perform the method according to the disclosed embodiments. Alternatively, at least one of the head-mounted display device (100) and the third device may execute a computer program product to perform the method according to the disclosed embodiments in a distributed manner.

[0245] For example, a head-mounted display device (100) can execute a computer program product stored in memory (140, see FIG. 3) to control another electronic device (e.g., a mobile device) that is connected to the head-mounted display device (100) in communication to perform a method according to the disclosed embodiments.

[0246] As another example, a third device (e.g., a mobile device) may execute a computer program product to control an electronic device connected to the third device in communication to perform the method according to the disclosed embodiment.

[0247] When the third device executes a computer program product, the third device may download the computer program product from the head-mounted display device (100) and execute the downloaded computer program product. Alternatively, the third device may execute a computer program product provided in a pre-loaded state to perform the method according to the disclosed embodiments.

[0248] Although the embodiments have been described above with reference to limited examples and drawings, those skilled in the art can make various modifications and variations from the description above. For example, appropriate results can be achieved even if the described techniques are performed in a different order than described, and / or components such as the described computer system or module are combined or assembled in a form different from described, or replaced or substituted by other components or equivalents.

Claims

1. A method for a head-mounted display (HMD) device to change the disparity of a stereo image, A step (S210) of receiving a user's zoom-in input to enlarge the size of the stereo image while displaying the stereo image including the left eye image and the right eye image; A step (S220) of determining the location and size of a region of interest within the entire area of ​​the stereo image based on the received zoom-in input; A step (S230) of adjusting the disparity between the region of interest in the left eye image and the right eye image by shifting the position of the region of interest in at least one of the left eye image and the right eye image by a shift value calculated to have an inverse relationship with the size of the region of interest; and Step (S240) of increasing the size of the region of interest to the size of the entire region of the stereo image to increase the adjusted disparity; A method including 2. In Paragraph 1, The zoom-in input of the above user includes a positional movement of the head-mounted display device (100) caused by the movement of the user wearing the head-mounted display device (100), and The step (S220) of determining the location and size of the region of interest is, A step (S420) of recognizing a gaze point where the gaze directions of the two eyes converge by obtaining information regarding the gaze direction of the user's two eyes using an eye tracking sensor (116); A step (S430) of determining the position coordinates of the center point of the region of interest based on the position of the recognized gaze point; A step (S440) of measuring the distance traveled due to the user's movement by obtaining a change in position including the position, rotation, and translation of the head-mounted display device (100) using at least one of a GPS sensor and an IMU sensor (inertial measurement unit); and A step (S450) of determining the size of the region of interest centered on the position coordinates of the determined center point based on the measured travel distance; A method including 3. In Paragraph 1, The step (S210) of receiving the user's zoom-in input is, The method includes the step (S610) of receiving pinch zoom input from the user to enlarge the stereo image using a hand gesture, and The step (S220) of determining the location and size of the region of interest is, A method comprising the step (S620) of determining the size of the region of interest based on a value inversely proportional to the drag distance of the received pinch zoom input.

4. In any one of paragraphs 1 to 3, The step (S230) of adjusting the parallax of the stereo image above is, A method comprising the step (S1020) of shifting the position of a first region of interest corresponding to the region of interest in the left eye image by the shift value in a first direction along the X-axis, and shifting the position of a second region of interest corresponding to the region of interest in the right eye image by the shift value in a second direction opposite to the first direction along the X-axis.

5. In any one of paragraphs 1 through 4, The step (S230) of adjusting the parallax of the stereo image above is, A step (S1220) of obtaining the maximum value of the disparity of the region of interest from a disparity map containing information regarding the disparity between the left eye image and the right eye image; and A step of calculating the shift value based on the maximum value of the time difference obtained above (S1230); Includes, A method in which the above shift value is calculated as a value inversely proportional to the maximum value of the parallax within the region of interest.

6. In any one of paragraphs 1 through 4, The step (S230) of adjusting the parallax of the stereo image above is, The method includes the step (S1240, S1250) of calculating the shift value based on an interest region ratio value representing the ratio of the size of the interest region to the size of the entire area of ​​the stereo image. A method in which the above shift value is calculated as a value inversely proportional to the above interest region ratio value.

7. In any one of paragraphs 1 through 6, The step (S240) of increasing the above-mentioned adjusted time difference is, A method of increasing the adjusted disparity as the size of the region of interest is enlarged by a value obtained by dividing the ratio between the size of the region of interest before enlargement and the total size of the stereo image.

8. In a head-mounted display device (100) that changes the parallax of a stereo image, A sensor (110) comprising at least one of a position sensor (112) and an IMU (Inertial Measurement Unit) sensor (114); A left eye display (150L) for displaying a left eye image included in the stereo image and a right eye display (150R) for displaying a right eye image included in the stereo image; At least one processor (130) including processing circuitry; and Memory (140) for storing one or more instructions; Includes, By executing the above one or more instructions individually or collectively by the at least one processor (130), the head-mounted display device (100) is: Receiving a user's zoom-in input to enlarge the size of the above stereo image, and Based on the received zoom-in input, the location and size of the region of interest within the entire area of ​​the stereo image are determined, and The position of the region of interest in at least one of the left eye image and the right eye image is shifted by a shift value calculated to have an inverse relationship with the size of the region of interest, thereby adjusting the disparity between the region of interest in the left eye image and the right eye image. A head-mounted display device (100) that increases the adjusted parallax by expanding the size of the area of ​​interest to the size of the entire area of ​​the stereo image.

9. In Paragraph 8, The zoom-in input of the above user includes a positional movement of the head-mounted display device (100) caused by the movement of the user wearing the head-mounted display device (100), and By executing the above one or more instructions individually or collectively by the at least one processor (130), the head-mounted display device (100) is: By obtaining information regarding the direction of gaze of the user's two eyes using an eye tracking sensor (116), a gaze point where the direction of gaze of the two eyes converges is recognized, and Based on the location of the recognized gaze point, the location coordinates of the center point of the region of interest are determined, and By using the sensor (110) to obtain a change in position including the position, rotation, and translation of the head-mounted display device (100), the distance traveled due to the user's movement is measured. A head-mounted display device (100) that determines the size of the region of interest centered on the position coordinate value of the determined center point based on the measured travel distance.

10. In Paragraph 8, By executing the above one or more instructions individually or collectively by the at least one processor (130), the head-mounted display device (100) is: Receiving pinch zoom input from the user to enlarge the stereo image using a hand gesture, A head-mounted display device (100) that determines the size of the region of interest based on a value inversely proportional to the drag distance of the received pinch zoom input.

11. In any one of paragraphs 8 through 10, By executing the above one or more instructions individually or collectively by the at least one processor (130), the head-mounted display device (100) is: A head-mounted display device (100) that shifts the position of a first region of interest corresponding to the region of interest in the left-eye image by the shift value in a first direction along the X-axis, and shifts the position of a second region of interest corresponding to the region of interest in the right-eye image by the shift value in a second direction opposite to the first direction along the X-axis.

12. In Paragraph 11, A head-mounted display device (100) in which, based on the positional shift of the first region of interest and the second region of interest, the time difference between the first region of interest and the second region of interest is increased by twice the shift value compared to the time difference before adjustment.

13. In any one of paragraphs 8 through 12, By executing the above one or more instructions individually or collectively by the at least one processor (130), the head-mounted display device (100) is: The maximum value of the disparity of the region of interest is obtained from a disparity map containing information regarding the disparity between the left eye image and the right eye image, and Calculate the shift value based on the maximum value of the above-mentioned time difference, and A head-mounted display device (100) in which the above shift value is calculated as a value inversely proportional to the maximum value of the parallax within the region of interest.

14. In any one of paragraphs 8 through 12, By executing the above one or more instructions individually or collectively by the at least one processor (130), the head-mounted display device (100) is: The shift value is calculated based on the region of interest ratio value, which represents the ratio of the size of the region of interest to the size of the entire region of the stereo image, and A head-mounted display device (100) in which the above shift value is calculated as a value inversely proportional to the above interest region ratio value.

15. In any one of paragraphs 8 through 14, By executing the above one or more instructions individually or collectively by the at least one processor (130), the head-mounted display device (100) is: A head-mounted display device (100) that increases the adjusted disparity as the size of the area of ​​interest is enlarged by a value obtained by dividing the ratio between the size of the area of ​​interest before enlargement and the total size of the stereo image.