Head-mounted display device and operation method thereof

The head-mounted display device stabilizes stereoscopic images by tracking camera and user movements, addressing user discomfort and enhancing immersion in virtual reality environments.

WO2025225956A1PCT designated stage Publication Date: 2025-10-30SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/005140
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-26
Filing Date
2025-04-15
Publication Date
2025-10-30

AI Technical Summary

Technical Problem

Users experience discomfort and visual disturbances due to camera and head-mounted display device movements while viewing stereoscopic images in virtual reality environments.

Method used

A head-mounted display device that acquires stereo images, tracks camera and user movements, and adjusts image rendering to stabilize the displayed content based on these movements, using image stabilization techniques and rendering modes to minimize visual discomfort.

Benefits of technology

The solution effectively reduces visual discomfort and enhances user immersion by stabilizing stereoscopic images, providing a more comfortable and immersive virtual reality experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025005140_30102025_PF_FP_ABST
    Figure KR2025005140_30102025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed is an operation method of a head-mounted display device. The method may comprise the steps of: acquiring a stereo image; acquiring information on movement of a camera that captured the stereo image; acquiring information on movement of a user wearing the head-mounted display device; identifying movement of the stereo image with respect to the user, on the basis of the information on the movement of the camera and the information on the movement of the user; rendering image objects for outputting the stereo image on the basis of the movement of the stereo image; and displaying the rendered image objects.
Need to check novelty before this filing date? Find Prior Art

Description

Head-mounted display device and its operating method

[0001] A head-mounted display device and an operating method thereof are disclosed. Specifically, a head-mounted display device and an operating method thereof for stabilizing the display of stereoscopic images in a virtual environment are disclosed.

[0002] Video See-Through (VST) on head-mounted display (HMD) devices is a feature that allows users to observe the real environment through video in a virtual reality (VR) or augmented reality (AR) environment.

[0003] Users who watch videos through a head-mounted display device may experience discomfort while watching the video due to the movement of the camera that captures the video and the movement of the head-mounted display device itself.

[0004] According to one aspect of the present disclosure, a method of operating a head-mounted display device may be provided. In one embodiment, the method may include a step of acquiring a stereo image. In one embodiment, the method may include a step of acquiring information regarding the movement of a camera that captured the stereo image. In one embodiment, the method may include a step of acquiring information regarding the movement of a user wearing the head-mounted display device. In one embodiment, the method may include a step of identifying the movement of a stereo image relative to the user based on information regarding the movement of the camera and information regarding the movement of the user. In one embodiment, the method may include a step of rendering a video object that outputs a stereo image based on the movement of the stereo image. In one embodiment, the method may include a step of displaying the rendered video object.

[0005] According to one aspect of the present disclosure, a head-mounted display device is disclosed. The head-mounted display device may include a memory storing one or more display instructions and at least one processor executing one or more instructions stored in the memory. In one embodiment, the head-mounted display device may acquire a stereo image by executing one or more instructions. In one embodiment, the at least one processor may acquire information regarding the movement of a camera that captured the stereo image by executing one or more instructions. In one embodiment, the at least one processor may acquire information regarding the movement of a user wearing the head-mounted display device by executing one or more instructions. In one embodiment, the at least one processor may identify the movement of a stereo image relative to the user by executing one or more instructions based on information regarding the movement of the camera and information regarding the movement of the user. In one embodiment, the at least one processor may render a video object that outputs a stereo image based on the movement of the stereo image by executing one or more instructions. In one embodiment, the at least one processor may display the rendered video object by executing one or more instructions.

[0006] According to one aspect of the present disclosure, a computer-readable recording medium having recorded thereon a program for performing any one of the above-described and below-described methods for performing an operation of a head-mounted display device can be provided.

[0007] FIG. 1 is a drawing schematically illustrating the operation of a head mounted display device according to one embodiment of the present disclosure.

[0008] FIG. 2 is a flowchart for explaining the operation of a head mounted display device according to one embodiment of the present disclosure.

[0009] FIG. 3 is a diagram for explaining information related to stereo images and camera movement according to one embodiment of the present disclosure.

[0010] FIG. 4 is a diagram illustrating an image stabilization module and a rendering module according to one embodiment of the present disclosure.

[0011] FIG. 5 is a drawing for explaining a first rendering mode according to one embodiment of the present disclosure.

[0012] FIG. 6 is a drawing for explaining a second rendering mode according to one embodiment of the present disclosure.

[0013] FIG. 7 is a diagram for explaining an operation of selecting a rendering mode by a head mounted display device according to one embodiment of the present disclosure.

[0014] FIG. 8A is a flowchart illustrating a method for a head mounted display device according to one embodiment of the present disclosure to obtain motion data representing camera shake through filtering.

[0015] FIG. 8B is a flowchart illustrating a method for a head-mounted display device according to one embodiment of the present disclosure to obtain data indicating camera shake based on predicted movement.

[0016] FIG. 9 is a diagram illustrating a shaking data acquisition module according to one embodiment of the present disclosure.

[0017] FIG. 10 is a perspective view of a head mounted display device according to one embodiment of the present disclosure.

[0018] FIG. 11 is a detailed configuration diagram of a head mounted display device according to one embodiment of the present disclosure.

[0019] FIG. 12 is a detailed configuration diagram of an external electronic device according to one embodiment of the present disclosure.

[0020] The terms used in the embodiments of this specification have been selected from widely used, current terms, taking into account the functions of the present disclosure. However, these terms may vary depending on the intentions of those skilled in the art, precedents, the emergence of new technologies, etc. Furthermore, in certain cases, terms may be arbitrarily selected by the applicant, and in such cases, their meanings will be described in detail in the description of the relevant embodiments. Therefore, the terms used in this specification should not be defined simply as names of terms, but rather based on their meanings and the overall content of the present disclosure.

[0021] Unless the context clearly dictates otherwise, the singular forms "a," "an," and "the" are to be understood to include plural referents. Thus, for example, the description "a constituent surface" may also include reference to one or more of such surfaces.

[0022] Terms used herein, including technical or scientific terms, have the same meaning as commonly understood by one of ordinary skill in the art described herein.

[0023] Throughout this disclosure, when a part is said to "include" a component, this does not exclude other components, but rather implies the inclusion of other components, unless otherwise specifically stated. Furthermore, terms such as "part," "module," and the like used herein refer to a unit that processes at least one function or operation, which may be implemented in hardware or software, or a combination of hardware and software.

[0024] As used herein, the expression "configured to" can be used interchangeably with, for example, "suitable for," "having the capacity to," "designed to," "adapted to," "made to," or "capable of." The term "configured to" does not necessarily mean something is "specifically designed to" in hardware. Instead, in some contexts, the expression "a system configured to" can mean that the system, in conjunction with other devices or components, is "capable of." For example, the phrase "a processor configured to perform A, B, and C" can mean a dedicated processor for performing the operations (e.g., an embedded processor), or a generic-purpose processor (e.g., a CPU or application processor) that can perform the operations by executing one or more software programs stored in memory.

[0025] It should be understood that the blocks and combinations of flowcharts in each flowchart can be executed by one or more computer programs containing computer-executable instructions. The one or more computer programs may be stored entirely in a single memory, or may be stored in separate portions across multiple different memories.

[0026] All functions or operations described in this document may be performed by a single processor or a combination of processors. A single processor or a combination of processors is a circuitry that performs processing, and may include circuitry such as an Application Processor (AP), a Communication Processor (CP), a Graphical Processing Unit (GPU), a Neural Processing Unit (NPU), a Microprocessor Unit (MPU), a System on Chip (SoC), or an Integrated Chip (IC).

[0027] In the present disclosure, 'Augmented Reality (AR)' means displaying a virtual image together with a real environment (or real world), which is a space that physically exists in the real world, or displaying a virtual image together with a real object existing in the real environment.

[0028] In this disclosure, 'virtual reality (VR)' means showing an image of a virtual environment (or virtual world) created using computer graphics technology that is a space separate from the real environment.

[0029] In this disclosure, 'Mixed Reality (MR)' means providing an experience that transcends the virtual and real worlds by allowing objects existing in a real environment and objects in a virtual environment to interact with each other.

[0030] In the present disclosure, a 'head-mounted display device' may refer to an augmented reality device capable of expressing augmented reality, a virtual reality device capable of expressing virtual reality, or a mixed reality device capable of expressing mixed reality. In one embodiment, the head-mounted display device may include a shape of glasses worn on the user's face or a shape of a helmet worn on the user's head, but is not necessarily limited to the examples described above.

[0031] In the present disclosure, an 'artificial intelligence (AI) model' may refer to a set of functions or algorithms that are set to perform a desired characteristic (or purpose) by being learned using a plurality of learning data by a learning algorithm. Examples of learning algorithms include, but are not necessarily limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning. In one embodiment, the AI ​​model may be stored in the memory of a head-mounted display device. However, the AI ​​model is not limited thereto, and the head-mounted display device may transmit data input to the AI ​​model to the server and receive data output from the AI ​​model from the server.

[0032] In addition, the 'artificial intelligence model' in this specification may be composed of multiple neural network layers. Each of the multiple neural network layers has multiple weight values, and can perform neural network operations through operations between the operation results of the previous layer and the multiple weights. The multiple weights of the multiple neural network layers may be optimized based on the learning results of the artificial intelligence model. For example, the multiple weights may be updated so that the loss value or cost value obtained from the artificial intelligence model is reduced or minimized during the learning process. Examples of models including multiple neural network layers include, but are not limited to, a deep neural network (DNN), a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), a restricted boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), and deep Q-networks.

[0033] In the present disclosure, "rendering" may refer to an operation and function for visually representing a virtual object in a three-dimensional virtual space. In one embodiment, "rendering" may include an operation of determining the position and orientation of a virtual object in a virtual space (rendering space) and projecting the virtual object according to the user's position and viewpoint to create an image including the virtual object. In one embodiment, the image obtained through rendering may be displayed through a display.

[0034] In the present disclosure, data processing related to an image may mean data processing for each of a plurality of frames constituting an image.

[0035] Below, with reference to the attached drawings, embodiments of the present disclosure are described in detail so that those skilled in the art can easily practice the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein. In the drawings, portions irrelevant to the description have been omitted for clarity of explanation, and similar reference numerals have been used throughout the specification to designate similar parts.

[0036] The present disclosure will be described below with reference to the attached drawings.

[0037] FIG. 1 is a drawing schematically illustrating the operation of a head mounted display device according to one embodiment of the present disclosure.

[0038] Referring to FIG. 1, an external electronic device (2000) can capture a real environment (Real Environment) to obtain a stereo image (100). The Real Environment may refer to a physical space in the real world and may include various objects. For example, the Real Environment may include inanimate objects such as buildings and roads, as well as biological objects such as people and animals.

[0039] In one embodiment, the external electronic device (2000) may include a camera (or stereo camera) for capturing a stereo image (100). In one embodiment, the camera for capturing the stereo image (100) may include two cameras (or lenses and image sensors) for capturing a left image and a right image of the stereo image (100). For example, the external electronic device (2000) may include, but is not necessarily limited to, a smartphone or tablet PC including two cameras having different Fields of View (FoV).

[0040] In one embodiment, an external electronic device (2000) can obtain information (10) about the movement of a camera that captures a stereo image (100). In one embodiment, the external electronic device (2000) can include a motion sensor for measuring motion data related to the position, speed, direction, attitude, etc. of the camera that captures the stereo image (100). In one embodiment, the motion sensor can include an IMU (Inertial Measurement Unit) sensor. In one embodiment, the IMU sensor can be configured to measure the movement and direction of an object by combining an accelerometer, a gyroscope, a magnetometer, etc. In one embodiment, the IMU sensor can obtain 6 DoF (6 Degrees of Freedom) measurements including position coordinate values ​​(x-axis, y-axis, and z-axis coordinate values) and 3-axis angular velocity values ​​(roll, yaw, pitch) in a 3D space. However, the present invention is not necessarily limited to the above-described example, and the motion sensor can include a configuration that can obtain various data related to the movement of the electronic device.

[0041] In one embodiment, the external electronic device (2000) can obtain information (10) about the movement of the camera based on movement data measured from a motion sensor while the stereo image (100) is being captured. In one embodiment, the information (10) about the movement of the camera can include information about at least one of movement and rotation of the camera in three-dimensional space while capturing the stereo image (100). In one embodiment, the information (10) about the movement of the camera can be mapped to a plurality of frames of the stereo image (100) corresponding to the time at which the movement data is measured from the motion sensor and stored as metadata of the stereo image (100).

[0042] In one embodiment, an external electronic device (2000) may transmit information (10) about a stereo image (100) and movement of a camera to a head mounted display device (1000). The head mounted display device (1000) may store the received stereo image (100) and information (10) about movement of the camera in a memory. However, the present invention is not limited thereto, and the external electronic device (2000) may transmit information (10) about a stereo image (100) and movement of a camera to an external storage device (e.g., an external server), and the head mounted display device (1000) may receive information (10) about a stereo image (100) and movement of a stereo camera stored in the external storage device from the external storage device.

[0043] In one embodiment, the head mounted display device (1000) may include various types of devices that display a virtual environment and virtual objects. In one embodiment, the head mounted display device (1000) may model a virtual object (120) that outputs a stereo image (100) and a virtual environment (110) in which the video object (120) exists, and render the modeled virtual environment (110) and the video object (120). In one embodiment, the head mounted display device (1000) may display the rendered virtual environment (110) and the video object (120) through a display.

[0044] In one embodiment, the head mounted display device (1000) can obtain information (20) regarding the movement of a user (2) wearing the head mounted display device (1000). In one embodiment, the head mounted display device (1000) can include a movement sensor for measuring data regarding the position, speed, direction, posture, etc. of the head mounted display device (1000). In one embodiment, the movement sensor of the head mounted display device (1000) can correspond to the movement sensor of the external electronic device (2000) described above.

[0045] In one embodiment, the head mounted display device (1000) can obtain information (20) about the movement of the user (2) based on movement data measured from a movement sensor while the user (2) wears the head mounted display device (1000). In one embodiment, the information (20) about the movement of the user (2) can include information about the movement and rotation of the head (or body) of the user (2) wearing the head mounted display device (1000) in three-dimensional space.

[0046] In one embodiment, the head mounted display device (1000) can identify the movement (30) of a stereo image (100) for the user (2) (hereinafter, movement (30) of the stereo image (100)) based on information (10) about the movement of the camera and information (20) about the movement of the user (2). The movement (30) of the stereo image (100) can refer to the movement of the stereo image (100) felt by the user (2) when the user (2) views the stereo image (100) displayed at a fixed position in three-dimensional space through the head mounted display device (1000). In one embodiment, the movement (30) of the stereo image (100) can include information about the movement and rotation of the stereo image (100) in three-dimensional space. For example, if a camera that captured a stereo image (100) rotates 1 degree to the left at the time a specific frame is captured, and a user (2) viewing the stereo image (100) rotates 1 degree to the right at the time a video object (120) outputs a specific frame of the stereo image (100), the stereo image (100) may feel as if it has rotated 2 degrees to the left from the user's (2) perspective. Accordingly, the head mounted display device (1000) may identify the movement of the stereo image (100) with respect to the user (20) as a 2-degree rotation to the left. Meanwhile, the movement (30) of the stereo image (100) may be referred to by various expressions representing the same / similar concepts. For example, the movement (30) of the stereo image (100) may be replaced with expressions such as relative movement of the stereo image (100), movement of a scene, movement of a context, movement of a screen, etc., and is not necessarily limited to the examples described above.

[0047] The movement (30) of the stereo image (100) for the user (2) may be a hindrance to the user (2) viewing the stereo image (100). For example, the movement (30) of the stereo image (100) may include shaking and trembling of the stereo image (100) felt by the user, which is formed by unintended camera movement and unintended user movement. Accordingly, the user (2) may have difficulty immersing himself in viewing the stereo image (100) due to the movement (30) of the stereo image (100), and may feel visual discomfort and dizziness while viewing the stereo image (100).

[0048] In one embodiment, the head mounted display device (1000) can render a video object (120) that outputs a stereo image (100) based on the movement (30) of the stereo image (100). Here, the movement of the stereo image (100) can include at least one of the movement of a camera that captured the stereo image and the movement of a user (2) wearing the head mounted display device (1000). In one embodiment, the head mounted display device (1000) can reduce the movement of the stereo image (100) felt from the user's (2's) perspective by reconstructing one of the left and right images of the stereo image (100) based on the movement (30) of the stereo image (100) or adjusting the arrangement of the video object (120). Accordingly, the user (2) can view the stereo image (100) through the head mounted display device (1000) without any discomfort and can be more immersed in the content of the stereo image (100).

[0049] FIG. 2 is a flowchart for explaining the operation of a head mounted display device according to one embodiment of the present disclosure.

[0050] Referring to FIG. 2, the operation of the head mounted display device (1000) will be schematically described, and a detailed description of each operation will be described with reference to the drawings that follow. In addition, the operations of the head mounted display device (1000) described in the present disclosure can be understood as the operations of the processor (1700) of the head mounted display device (1000) illustrated in FIG. 9.

[0051] In step S210, the head mounted display device (1000) can acquire a stereo image. In one embodiment, the stereo image can be acquired by capturing it through a camera for capturing a stereo image. In one embodiment, the camera for capturing the stereo image can be included in an external electronic device (2000). In this case, the electronic device (1000) can receive a stereo image from the external electronic device (2000) and store the received stereo image in the memory of the electronic device (1000). In one embodiment, the camera for capturing the stereo image can be included in the head mounted display device (1000). In this case, the head mounted display device (1000) can capture a stereo image through the camera and store the captured stereo image in the memory of the electronic device (1000).

[0052] In one embodiment, a stereo image may include a left image corresponding to the user's left eye and a right image corresponding to the user's right eye. In one embodiment, the head mounted display device (1000) may include a display (or a region of the display) corresponding to the left image and a display (or a region of the display) corresponding to the right image. Here, the display corresponding to the left image may be referred to as a left-eye display, and the display corresponding to the right image may be referred to as a right-eye display. In one embodiment, a left image or an image object outputting a left image may be displayed through a display corresponding to the left image, and a right image or an image object outputting a right image may be displayed through a display corresponding to the right image. In one embodiment, the display corresponding to the left image and the display corresponding to the right image are positioned in front of the left and right eyes of a user wearing the head mounted display device (1000), so that the user can experience a three-dimensional effect of the stereo image.

[0053] In one embodiment, a stereo image may be images captured by two cameras having different FOVs. In other words, the cameras capturing the stereo image may include two cameras having different FOVs. In one embodiment, the head mounted display device (1000) may crop an area corresponding to an image captured by a camera having a relatively narrower FOV (e.g., a wide-angle camera) from an image captured by a camera having a relatively wider FOV (e.g., an ultra-wide camera) among the images captured by the two cameras having different FOVs. In one embodiment, the head mounted display device (1000) may acquire a stereo image from the image captured by the camera having the relatively narrower FOV and the cropped image. However, the present invention is not necessarily limited to the above-described example, and the FOVs of the two cameras capturing the stereo image may be the same, and the images captured by each camera may correspond to the left and right images of the stereo image. A detailed description of how to acquire stereo images using two cameras with different FOVs will be described again below with reference to Figure 3.

[0054] In step S220, the head mounted display device (1000) may obtain information about the movement of the camera that captured the stereo image. In one embodiment, the information about the movement of the camera may be received together with the stereo image from an external electronic device that captured the stereo image or from the memory of the head mounted display device (1000). In one embodiment, the information about the movement of the camera that captured the stereo image may include a plurality of movements corresponding to a plurality of frames of the stereo image. In one embodiment, the information about the movement of the camera may be included in the metadata of the stereo image by mapping and storing the movement of the camera corresponding to each frame of the stereo image. In one embodiment, the information about the movement of the camera that captured the stereo image may include information about the movement of the first camera that captured the left image of the stereo image and information about the movement of the second camera that captured the right image of the stereo image. The information about the movements of the first and second cameras that captured the stereo image will be described in detail below with reference to FIG. 3.

[0055] In one embodiment, the head mounted display device (1000) can obtain information about the movement of the camera based on data obtained through SLAM (Simultaneous Localization and Mapping) of the head mounted display device (1000). In one embodiment, a stereo image can be captured through a camera included in the head mounted display device (1000). In one embodiment, the head mounted display device (1000) can perform a SLAM algorithm based on at least one of an image obtained from the camera and movement data obtained from a movement sensor while capturing the stereo image through the camera, thereby performing location tracking (localization) and map creation (mapping) of the head mounted display device (1000). Here, the result of the location tracking can include information about at least one of the movement and rotation of the head mounted display device (1000). In one embodiment, the data obtained through SLAM can be used to identify in real time the direction or location that a user wearing the head mounted display device (1000) is currently viewing, or as data for interaction with a virtual object in a virtual space.

[0056] In one embodiment, the head mounted display device (1000) can acquire information about the movement of the head mounted display device (1000) obtained through SLAM while capturing a stereo image as information about the movement of the camera. For example, the head mounted display device (1000) can acquire information about at least one of the movement and rotation of the head mounted display device (1000) as a result of position tracking through a SLAM algorithm while capturing a stereo image. In this case, the head mounted display device (1000) can acquire the acquired information about the movement and rotation of the head mounted display device (1000) as information about the movement and rotation of the camera capturing the stereo image.

[0057] In one embodiment, the head mounted display device (1000) may acquire first motion data representing the movement of the camera at the time of capturing the stereo image. Here, the first motion data corresponds to the entire motion data (or raw data) of the camera, and may include data representing intentional motion of the camera and data representing unintentional motion of the camera. In addition, the first motion data representing the movement of the camera at the time of capturing the stereo image may correspond to data acquired through SLAM of the head mounted display device (1000) described above.

[0058] For example, while a user is shooting a stereo image (100) through a camera, intentional camera movements such as pan, tilt, and roll, which are intentional movements of the camera, may occur. Since such intentional camera movements are movements included in the context of the image captured by the camera itself, correction may not be performed according to the image stabilization described below. Accordingly, intentional camera movements may be referred to as movements that are not subject to correction. In addition, intentional camera movements may be referred to by various expressions representing the same / similar concepts. For example, intentional camera movements may be replaced with expressions such as intentional movement, low-frequency movement, and controlled movement, and are not limited to the examples described above.

[0059] In addition, unintentional camera movements (camera shake) may occur while the user is shooting a stereo image (100), such as shaking of the user's hand or arm, or vibration caused by external environmental factors (e.g., wind, etc.). Such unintentional camera movements correspond to noise unrelated to the movement included in the context of the image captured by the camera itself, and thus can be corrected according to the image stabilization described below. Accordingly, unintentional camera movements may be referred to as movements subject to correction. In addition, unintentional camera movements may be referred to by various expressions representing the same / similar concepts. For example, unintentional camera movements may be replaced with expressions such as unintended movement, high-frequency movement, noise-induced movement, etc., and are not limited to the examples described above.

[0060] In one embodiment, the camera movement may include camera shake. Here, the camera shake may correspond to the unintentional camera movement described above. However, the term "camera shake" may be replaced with various expressions representing the same or similar concepts. For example, the term "camera shake" may be replaced with expressions such as camera jitter, vibration, tremor, oscillation, and wobble, and is not limited to the examples described above.

[0061] In one embodiment, the head-mounted display device (1000) can obtain second motion data by filtering first motion data with a low-pass filter. In one embodiment, the head-mounted display device (1000) can obtain third motion data indicating camera shake based on a difference between the first motion data and the second motion data. In one embodiment, the head-mounted display device (1000) can adjust the cutoff frequency of the low-pass filter based on the amount of change in the movement of the camera.

[0062] In one embodiment, the head mounted display device (1000) may obtain second motion data representing camera motion at a subsequent time point predicted from the motion of the camera at a previous time point based on first motion data. In one embodiment, the head mounted display device (1000) may obtain third motion data representing camera shake based on a difference between the first motion data and the second motion data.

[0063] In one embodiment, the second motion data may include data indicating intentional camera movement, and the third motion data may include data indicating unintentional camera movement. In other words, the head mounted display device (1000) may obtain third motion data indicating camera shake (unintentional camera movement) based on the difference between the first motion data, which is raw data, and the second motion data, which indicates intentional camera movement. In one embodiment, the third motion data may correspond to information regarding camera movement for performing at least one of the first rendering mode and the second rendering mode, which is image stabilization, which will be described later.

[0064] In this way, the head-mounted display device (1000) according to one embodiment of the present disclosure can acquire third motion data representing camera shake as information regarding camera movement. Accordingly, image stabilization can be performed by considering only camera movements that interfere with viewing of stereo images from the user's perspective. The specific operation of the head-mounted display device (1000) acquiring information regarding camera movement, including camera shake, will be described again below with reference to FIGS. 8A and 8B .

[0065] In step S230, the head mounted display device (1000) can obtain information about the movement of a user wearing the head mounted display device (1000). In one embodiment, the information about the user's movement can be obtained while a video object outputting a stereo image is being rendered.

[0066] In one embodiment, information about the movement of a user wearing a head-mounted display device (1000) may include information about the movement of the user's left eye and information about the movement of the user's right eye. In one embodiment, the information about the movement of the user's left eye and the information about the movement of the user's right eye may include the movement of a point corresponding to the user's left eye and a point corresponding to the user's right eye in three-dimensional space. For example, the information about the user's movement may include information about at least one of a movement and a rotation of a point corresponding to the user's left eye and a point corresponding to the user's right eye in three-dimensional space, respectively.

[0067] In one embodiment, the head mounted display device (1000) can obtain information about movement of the user's left eye and information about movement of the user's right eye based on movement data measured from one motion sensor and relative positions of a point corresponding to the user's left eye and a point corresponding to the user's right eye with respect to the motion sensor in three-dimensional space.

[0068] For example, the head mounted display device (1000) may perform rotation transformation and translation transformation on data measured from the motion sensor based on data measured from the motion sensor and the relative positions of a point corresponding to the user's left eye and a point corresponding to the user's right eye with respect to the motion sensor in a three-dimensional space. In one embodiment, the head mounted display device (1000) may obtain information about the movement of the user's left eye and information about the movement of the point corresponding to the user's right eye based on the measured data on which rotation transformation has been performed. The above-described operation may be equally applied to a process of obtaining information about the movement of a camera capturing a left image of a stereo image and information about the movement of a camera capturing a right image based on data acquired from one motion sensor included in an external electronic device (2000).

[0069] In one embodiment, the head mounted display device (1000) may include a first motion sensor corresponding to the left eye of a user wearing the head mounted display device (1000) and a second motion sensor corresponding to the right eye of the user. In this case, the first motion sensor and the second motion sensor may be positioned near the left eye and right eye of the user, respectively. Accordingly, the head mounted display device (1000) may obtain information about the movement of the user's left eye based on the first motion sensor, and may obtain information about the movement of the user's right eye based on data measured from the second motion sensor. In one embodiment, the movement of the user's left eye and the movement of the user's right eye may be referred to as the movement of a point corresponding to the user's left eye and the movement of a point corresponding to the user's right eye, respectively.

[0070] In step S240, the head mounted display device (1000) can identify the movement of the stereo image for the user based on information about the movement of the camera and information about the movement of the user. In one embodiment, the head mounted display device (1000) can identify the movement corresponding to a specific frame based on information about the movement of the camera corresponding to a specific frame of the stereo image and information about the movement of the user acquired at the time when a specific frame output from the image object is rendered. In one embodiment, the head mounted display device (1000) can identify the difference between the movement of the camera and the movement of the user as the movement of the stereo image for the user.

[0071] In one embodiment, the motion of the stereo image relative to the user may include motion of the right image of the stereo image and motion of the left image.

[0072] In one embodiment, the head mounted display device (1000) can identify the movement of the left image for the user based on the difference between the movement of the user's left eye and the movement of the first camera that captured the left image, and can identify the movement of the right image for the user based on the difference between the movement of the user's right eye and the movement of the second camera that captured the right image. However, the present invention is not necessarily limited to the above-described example, and the head mounted display device (1000) can identify the movement of the left image and the movement of the right image based on the difference between the movement of the user's left eye and the movement of the user's right eye and the movement of the one camera when the information about the movement of the camera that captured the acquired stereo image includes only one movement (for example, the movement of an electronic device including a camera that captured the stereo image). In addition, when the acquired movement information of the user includes only one movement (for example, the movement of the user's head), the movement of the left image and the movement of the right image can be identified based on the difference between the movement of the one user and the movement of the first camera and the movement of the second camera.

[0073] In one embodiment, the head-mounted display device (1000) can identify the difference between the user's movement and the camera's movement as movement of a stereo image relative to the user if the magnitude of the user's movement is greater than a threshold value. In one embodiment, the head-mounted display device (1000) can identify the camera's movement as movement of a stereo image relative to the user if the magnitude of the user's movement is less than the threshold value.

[0074] In one embodiment, the head mounted display device (1000) can identify that the magnitude of the movement is greater than or equal to a threshold value if the magnitude of the translation vector during the movement is greater than or equal to a first threshold value, the magnitude of the rotation vector is greater than or equal to a second threshold value, or the magnitude of the translation vector and the magnitude of the rotation vector are greater than or equal to the first threshold value and the second threshold value, respectively. That is, the movement of the user wearing the head mounted display device (1000) may also include a movement that occurs to observe an object other than an image object in the virtual environment (e.g., an intentional movement of the user). In this case, by distinguishing whether the user's movement is an intentional movement based on whether the magnitude of the user's movement is greater than or equal to a threshold value, it is possible to determine whether the user's intentional movement is reflected in the first rendering mode or the second rendering mode described below.

[0075] At step S250, the head mounted display device (1000) can render a video object that outputs a stereo image based on the movement of the stereo image.

[0076] In one embodiment, the head mounted display device (1000) can model a video object capable of outputting (or displaying) an image in a virtual environment. In one embodiment, the video object can include various types of virtual objects capable of outputting an image, such as a monitor, a smartphone, a virtual window, or a wall existing in the virtual environment. In one embodiment, the video object can include a virtual object for adjusting the arrangement of a stereo image in a three-dimensional space and rendering a two-dimensional image of the adjusted arrangement stereo image. In one embodiment, the head mounted display device (1000) can render a scene of a virtual environment in which a video object exists based on the position (or eye position) of a user wearing the head mounted display device (1000) in a three-dimensional space.

[0077] In one embodiment, the head-mounted display device (1000) can render a video object corresponding to the user's left eye and a video object corresponding to the user's right eye, respectively. In this case, the video object corresponding to the user's left eye can output the left image of a stereo image, and the video object corresponding to the user's right eye can output the right image of the stereo image.

[0078] In one embodiment, the head mounted display device (1000) may perform at least one of a first rendering mode that adjusts the placement of image objects based on a difference in the magnitude of movement of the left image and the magnitude of movement of the right image, and a second rendering mode that reconstructs one of the left image and the right image.

[0079] In one embodiment, when performing the first rendering mode, the head mounted display device (1000) may adjust the arrangement of the image object based on the movement of the stereo image. In one embodiment, the head mounted display device (1000) may render the image object whose arrangement has been adjusted. In one embodiment, the head mounted display device (1000) may apply at least one of position movement and rotation of the image object so as to correspond to the movement of the stereo image in the space in which the image object is rendered.

[0080] In one embodiment, when performing the second rendering mode, the head mounted display device (1000) may obtain a disparity map of a stereo image. In one embodiment, the head mounted display device (1000) may select an image with a small movement size among the left and right images of the stereo image. In one embodiment, the head mounted display device (1000) may reconstruct an image that is not selected among the left and right images based on the disparity map and the selected image. In one embodiment, the head mounted display device (1000) may render a first image object that outputs the selected image and a second image object that outputs the reconstructed image.

[0081] In one embodiment, the head mounted display device (1000) can reconstruct the unselected images by warping the selected images according to the parallax map, or can reconstruct the unselected images by applying the parallax map and the selected images to an artificial intelligence model.

[0082] Specific details regarding the first rendering mode will be explained again through Fig. 5 below, and specific details regarding the second rendering mode will be explained again through Fig. 6 below.

[0083] At step S260, the head mounted display device (1000) can display the rendered image object.

[0084] In one embodiment, displaying a rendered image object may mean displaying an image representing a scene of a virtual environment that includes the image object. In one embodiment, the rendered image object may include an image object corresponding to the user's left eye and an image object corresponding to the user's right eye.

[0085] In one embodiment, displaying a rendered image object may mean displaying an image in which the image object is projected in two dimensions. In other words, the electronic device (1000) may display an image in which the image object is projected in two dimensions as a rendered image object. Here, the image object may be an image object whose arrangement is adjusted according to a first rendering mode to be described later, an image object that outputs an image reconstructed according to a second rendering mode to be described later, or an image object whose arrangement is adjusted according to the first rendering mode and outputs an image reconstructed according to the second rendering mode.

[0086] In one embodiment, the head mounted display device (1000) may perform perspective projection on an image object in a three-dimensional space to obtain an image in which the image object is projected two-dimensionally. In one embodiment, the head mounted display device (1000) may obtain an image of the image object by cropping an area of ​​an image in which the image object is projected two-dimensionally. In one embodiment, the head mounted display device (1000) may crop a preset crop area from an image in which the image object is projected two-dimensionally. In one embodiment, the crop area may have a preset size and shape based on the center of the image. In one embodiment, the size of the crop area may be determined based on at least one of a size and resolution of a display that displays a rendered stereoscopic image object.

[0087] In one embodiment, the head mounted display device (1000) may acquire an area corresponding to a crop area in an image in which an image object is projected in two dimensions as an image object image. In one embodiment, the image object image may include an image corresponding to a left image of a stereo image and an image corresponding to a right image. In other words, the electronic device (1000) may acquire an image object image corresponding to a left image based on an image object corresponding to a user's left eye, and may acquire an image corresponding to a right image based on an image object corresponding to the user's right eye.

[0088] In one embodiment, the crop area may extend beyond the two-dimensionally projected image of the video object. In other words, the crop area may include an area (hereinafter, a blank area) that does not include the two-dimensionally projected image of the video object. For example, if the video object, for which at least one of position translation and rotation of the video object has been performed by the first rendering mode, is projected into two dimensions, some areas of the crop area may extend beyond the two-dimensionally projected image of the video object, since the crop area has a preset position and size. Accordingly, the crop area may include an empty area that does not include the two-dimensionally projected image of the video object.

[0089] In one embodiment, the head mounted display device (1000) can interpolate a blank area of ​​a crop area. In one embodiment, the head mounted display device (1000) can interpolate an area not included in the crop area based on an opposite image of a stereo image. For example, a blank area may exist in the crop area for an image object of an image with a relatively narrower FOV among the left and right images of a stereo image, but a blank area may not exist in the crop area for an image object with a relatively wider FOV. In one embodiment, the head mounted display device (1000) can identify pixels corresponding to the blank area of ​​the crop area for an image object of an image with a relatively wider FOV in a two-dimensionally projected image of the image object of the image with a relatively narrower FOV. In one embodiment, the head mounted display device (1000) can interpolate the blank area of ​​the crop area by inpainting or synthesizing the identified pixels into the blank area of ​​the crop area. In one embodiment, the head mounted display device (1000) can obtain an image corresponding to a crop area in which a blank area is interpolated as a video object image.

[0090] However, it is not necessarily limited to the above-described example, and even if the FOVs of both images of a stereo image are the same, due to differences in the adjustment of the arrangement of the image objects of each image, a blank area may exist in the crop area of ​​one image object, but a blank area may not exist in the crop area of ​​the opposite image object. Even in this case, the head mounted display device (1000) can perform interpolation for the blank area of ​​the crop area in the same manner as described above, and obtain an image of the image object.

[0091] In one embodiment, the head mounted display device (1000) may select one of the left image and the right image of a stereo image and reconstruct the unselected image by performing a second rendering mode. In one embodiment, the head mounted display device (1000) may render a first image object that outputs the selected image and a second image object that outputs the reconstructed image. In one embodiment, the head mounted display device may display the first image object through a first display corresponding to the selected image and display the second image object through a second display corresponding to the reconstructed image. For example, if the left image is the selected image among the left and right images of a stereo image and the right image is the reconstructed image, the head mounted display device (1000) may display the image object that outputs the left image through a display corresponding to the left image and display the image object that outputs the reconstructed image through a display corresponding to the right image.

[0092] FIG. 3 is a diagram for explaining information related to stereo images and camera movement according to one embodiment of the present disclosure.

[0093] Referring to FIG. 3, the external electronic device (2000) may include a camera (2100) for capturing a stereo image (100). In one embodiment, the camera (2100) may include a first camera (2100-1) for capturing a left image of the stereo image (100) and a second camera (2100-2) for capturing a right image. In FIG. 3, the camera (2100) is described as being included in the external electronic device (2000), but is not necessarily limited thereto, and the camera (2100) may be a component included in the head mounted display device (1000). Accordingly, the operation of the external electronic device (2000) described below may be understood as the operation of the head mounted display device (1000).

[0094] In one embodiment, the first camera (2100-1) may be a camera including a lens that captures a relatively wider FOV than the second camera (2100-2). For example, the first image (101) captured by the first camera (2100-1) may have a wider FOV than the second image (103) captured by the second camera (2100-2). In other words, the left and right images of the stereo image may each have different FOVs.

[0095] In one embodiment, the external electronic device (2000) can obtain a third image (102) by cropping an area corresponding to the second image (103) from the first image (101). In one embodiment, the external electronic device (2000) can reduce distortion due to a difference in FOV between the first image (101) and the second image (103) based on parameter values ​​related to shooting of each of the first camera (2100-1) and the second camera (2100-2). The external electronic device (2000) can obtain a third image (102) by extracting an area corresponding to the second image (103) from the first image (101) in which distortion due to the difference in FOV is reduced. In this case, the third image (102) can correspond to a left image of the stereo image (100), and the second image (103) can correspond to a right image of the stereo image (100). However, it is not necessarily limited to the above-described example, and the first image (101) may correspond to the left image of the stereo image (100), and the second image (103) may correspond to the right image of the stereo image (100).

[0096] In one embodiment, the external electronic device (2000) may transmit a stereo image (100) including a second image (103) and a third image (102) to the head mounted display device (1000). However, the present invention is not limited to the above-described example, and the external electronic device (2000) may transmit the first image (101) and the second image (103) to the head mounted display device (1000), and the head mounted display device (1000) may obtain the stereo image (100) by obtaining the third image (102) based on the received first image (101) and second image (103).

[0097] In one embodiment, the external electronic device (2000) can obtain information (10) about the movement of the camera (2100) while capturing a stereo image (100) through the camera (2100). In one embodiment, the external electronic device (2000) can include a motion sensor and can obtain information about the movement of the camera (2100) based on data measured from the motion sensor while the stereo image (100) is captured. In one embodiment, the external electronic device (2000) can transmit information (10) about the movement of the camera (2100) to the head mounted display device (1000). In one embodiment, information (10) about the movement of the camera (2100) can be stored as metadata of the stereo image (100) and transmitted together with the stereo image (100). However, the present invention is not necessarily limited to the above-described example, and the camera (2100) and the motion sensor may be components included in the head mounted display device (1000), and the head mounted display device (1000) may obtain information (10) regarding the movement of the camera (2100) based on data measured from the motion sensor while the stereo image (100) is being captured.

[0098] In one embodiment, the information (10) regarding the movement of the camera (2100) may include information regarding the movement of the first camera (2100-1) and information regarding the movement of the second camera (2100-2). In one embodiment, the information regarding the movement of the first camera (2100-1) of the external electronic device (2000) and the information regarding the movement of the second camera (2100-2) may be different from each other. For example, when the external electronic device (2000) rotates and the center of rotation is close to the first camera (2100-1), the movement of the first camera (2100-1) that is closer to the center of rotation may be measured to be greater than the movement of the second camera (2100-1) that is farther from the center of rotation. In this way, if the movement of the first camera (2100-1) and the movement of the second camera (2100-2) are different, the movement of the left image and the movement of the right image of the stereo image described later may also be different from each other.

[0099] In one embodiment, the external electronic device (2000) may include a motion sensor for obtaining information regarding the movement of the camera (2100). In one embodiment, when there is one motion sensor, the external electronic device (2000) may obtain information regarding the movement of the first camera (2100-1) and information regarding the movement of the second camera (2100-2) based on data measured by the motion sensor and relative positions of the first camera (2100-1) and the second camera (2100-2) with respect to the motion sensor in three-dimensional space. However, it is not necessarily limited to the above-described example, and the external electronic device (2000) transmits data measured by the motion sensor and information on the relative positions of the first camera (2100-1) and the second camera (2100-2) with respect to the motion sensor to the head mounted display device (1000), and the head mounted display device (1000) may obtain information on the movement of the first camera (2100-1) and information on the movement of the second camera (2100-2) based on the data measured by the received motion sensor and information on the relative positions of the first camera (2100-1) and the second camera (2100-2).

[0100] In one embodiment, when the motion sensor includes a motion sensor for each of the first camera (2100-1) and the second camera (2100-2), the external electronic device (2000) can obtain information about the motion of the first camera (2100-1) based on data measured from the motion sensor of the first camera (2100-1), and can obtain information about the motion of the second camera (2100-2) based on data measured from the motion sensor of the second camera (2100-2).

[0101] In one embodiment, the information (10) regarding the movement of the camera (2100) may include data acquired from an IMU sensor of the camera. In one embodiment, the movement sensor included in the external electronic device (2000) may include an IMU sensor of the camera (2100). In this case, the data acquired from the IMU sensor of the camera (2100) may include data measured to perform OIS (Optical Image Stabilization) or EIS (Electronic Image Stabilization) captured by the camera. For example, the data acquired from the IMU sensor of the camera (2100) may include 6 DoF (6 Degrees of Freedom) measurements including position coordinate values ​​(x-axis, y-axis, and z-axis coordinate values) and 3-axis angular velocity values ​​(roll, yaw, and pitch) of the camera (2100). In one embodiment, the head mounted display device (1000) can identify the movement and rotation of the camera (2100) based on the 6 DoF measurements of the camera (2100).

[0102] In one embodiment, the information (10) regarding the movement of the camera (2100) may include metadata related to the EIS of the camera (2100). The metadata related to the EIS of the camera (2100) may include a cropping matrix indicating an area extracted in response to the movement of the camera (2100) from an image captured by the camera (2100) and a warping matrix for correcting distortion due to the movement of the camera (2100). In one embodiment, the head mounted display device (1000) may identify the movement and rotation of the camera (2100) based on the cropping matrix and the warping matrix received from the external electronic device (2000). For example, the head mounted display device (1000) may identify the movement of the camera (2100) based on a change in the center coordinates of an area extracted by the cropping matrix from a stereo image. As another example, the head mounted display device (1000) can identify the rotation of the camera (2100) based on the rotation about each axis of the three-dimensional space of the stereo image by the warping matrix.

[0103] FIG. 4 is a diagram illustrating an image stabilization module and a rendering module according to one embodiment of the present disclosure.

[0104] Referring to FIG. 4, the head mounted display device (1000) may include an image stabilization module (410) and a rendering module (420). The operation of the image stabilization module (410) and the rendering module (420) of the present disclosure may be understood as the operation of the head mounted display device (1000) or the processor of the head mounted display device (1000).

[0105] The image stabilization module (410) may be a component that outputs at least one of image stabilization information and a reconstructed image based on a stereo image (100), information about the movement of the camera (10), and information about the movement of the user (20).

[0106] In one embodiment, the image stabilization module (410) may include a high-speed image stabilization module (412). In one embodiment, when the head mounted display device (1000) performs the first rendering mode, the high-speed image stabilization module (412) may be a component that outputs image stabilization information based on information about the movement of the camera (10) and information about the movement of the user (20). The image stabilization information may include information related to adjusting the arrangement of an image object in three-dimensional space. For example, the image stabilization information may include a translation vector indicating a direction and a size in which an image object is moved in three-dimensional space, and a rotation vector indicating a direction and a size in which an image object is rotated.

[0107] In one embodiment, the high-speed image stabilization module (412) can identify the movement of a stereo image relative to the user based on information about the movement of the camera (10) and information about the movement of the user (20), and output image stabilization information having the same magnitude as the movement of the identified stereo image and indicating movement in the opposite direction. In one embodiment, when the movement of the stereo image includes the movement of the left image and the movement of the right image, the image stabilization module (412) can output image stabilization information for the left image and image stabilization information for the right image.

[0108] In one embodiment, the image stabilization module (410) may include a depth-adapted image stabilization module (414). In one embodiment, when the head-mounted display device (1000) performs the second rendering mode, the depth-adapted image stabilization module (414) may be a component that outputs a reconstructed image based on information about the movement of the camera (10) and information about the movement of the user (20). The reconstructed image may include a selected image, which is one of the left image and the right image of the stereo image, and an image that forms a parallax corresponding to a parallax map.

[0109] In one embodiment, the depth-adapted image stabilization module (414) can obtain a reconstructed image by warping the selected image according to a disparity map of a stereo image. In one embodiment, the depth-adapted image stabilization module (414) can map a plurality of pixels of the selected image to a plurality of pixels of the unselected image based on pixel values ​​of the disparity map. The depth-adapted image stabilization module (414) can interpolate the unmapped pixels among the plurality of pixels of the unselected image based on values ​​of surrounding pixels or can interpolate them through an inpainting model, which is an artificial intelligence model that performs image inpainting. The depth-adapted image stabilization module (414) can obtain a reconstructed image including the mapped pixels and the interpolated pixels.

[0110] In one embodiment, the depth-adapted image stabilization module (414) can reconstruct unselected images by applying the disparity map and the selected images to an artificial intelligence model. In one embodiment, the artificial intelligence model may be referred to as a stereo image reconstruction model, and may be a model trained to output images from different viewpoints based on one of the disparity map and the stereo images. The stereo image reconstruction model may include a convolutional neural network-based deep learning model, and may be trained based on a training data set including stereo images including left and right images and disparity maps corresponding to the stereo images.

[0111] The rendering module (420) may be a component that renders an image object. In one embodiment, the rendering module (420) may adjust the arrangement of an image object based on image stabilization information and render the image object with the adjusted arrangement. In one embodiment, adjusting the arrangement of an image object may mean applying at least one of position movement and rotation of the image object so as to correspond to the movement of a stereo image in the space where the image object is rendered. In one embodiment, the head mounted display device (1000) may adjust the arrangement of an image object for each of a plurality of frames of a stereo image output by the image object. In one embodiment, the rendering module (420) may adjust the arrangement of an image object that outputs a left image based on left image stabilization information, and may adjust the arrangement of an image object that outputs a right image based on right image stabilization information. In one embodiment, the rendering module (420) may render a first image object that outputs a selected image from among stereo images based on a reconstructed image, and a second image object that outputs a reconstructed image.

[0112] In one embodiment, the rendered image object may be displayed through the display (1100). In one embodiment, if the image object includes an image object corresponding to the left image of a stereo image and an image object corresponding to the right image, the image object corresponding to the left image may be displayed through the display corresponding to the left image, and the image object corresponding to the right image may be displayed through the display corresponding to the right image.

[0113] Hereinafter, the first rendering mode performed through the high-speed image stabilization module (412) and the second rendering mode performed through the depth application image stabilization module (414) will be described with reference to FIGS. 5 and 6.

[0114] FIG. 5 is a drawing for explaining a first rendering mode according to one embodiment of the present disclosure.

[0115] Referring to FIG. 5, the head mounted display device (1000) can identify the movement (30) of the stereo image based on the information (10) about the movement of the camera and the information (20) about the movement of the user. In one embodiment, the head mounted display device (1000) can identify the sum of the motion vector of the camera and the motion vector of the user for each of a plurality of frames of the stereo image as the motion vector of the stereo image. For example, the translation vector and rotation vector of the camera of a specific frame of the stereo image and , and the user's translation vector and rotation vector are and On the other hand, the motion vector of the stereo image is can be identified. Here, the motion vector of the camera for each frame corresponds to the movement of the camera acquired at the time when the corresponding frame is captured, and the motion vector of the user for each frame corresponds to the movement of the user acquired at the time when the corresponding frame is output through the video object.

[0116] In one embodiment, the head mounted display device (1000) can adjust the placement of the image object (120) based on the movement (30) of the stereo image. For example, the head mounted display device (1000) can adjust the rotation vector of the movement (30) of the stereo image in the virtual environment (110) in which the image object (120) is rendered. Based on the size, the image object (120-1) can be rotated in the opposite direction. In addition, the head mounted display device (1000) can be configured to calculate the movement vector (30) of the stereo image in the virtual environment (110), which is the space where the image object (120) is rendered. The image object (120-2) can be moved in the virtual environment (110) in the opposite direction and has the same size.

[0117] In one embodiment, the head mounted display device (1000) can render a video object (120) whose arrangement has been adjusted. For example, the head mounted display device (1000) can render a video object (120) whose arrangement has been adjusted in the opposite direction and whose size is the same as the motion (30) vector of a stereo image in a virtual environment (110). The rendered video object (120) can output a stereo image of a specific frame corresponding to the motion (30) vector of the stereo image. In one embodiment, the head mounted display device (1000) can render a video object whose arrangement has been adjusted and display the rendered video object by performing the above-described operation on all frames of the stereo image.

[0118] In this way, according to one embodiment of the present disclosure, the head mounted display device (1000) can adjust the arrangement of the image object (120) to correspond to the movement (30) of the stereo image in the virtual environment (110) through the first rendering mode, thereby preventing the user from feeling the movement of the stereo image. Unlike an image stabilization method that simply performs cropping and warping on the stereo image itself, this image stabilization method adjusts the arrangement of the image object (120) in the virtual environment (110) in which the image object (120) is rendered, so that the angle of view of the stereo image is not lost, and hardware resources required for image stabilization can be saved.

[0119] FIG. 6 is a drawing for explaining a second rendering mode according to one embodiment of the present disclosure.

[0120] Referring to FIG. 6, the head-mounted display device (1000) can obtain a parallax map (620) of a stereo image. The parallax map (620) may refer to an image that indicates how different positions of the same elements included in the left and right images are displayed in each image due to the difference in time between when the left image of the stereo image was captured and when the right image was captured. A plurality of pixels of the parallax map may have pixel values ​​proportional to parallax, and the parallax may be inversely proportional to the physical distance from the camera in a real environment.

[0121] In one embodiment, the head mounted display device (1000) can obtain a parallax map (620) of the stereo image based on the stereo image. In one embodiment, the head mounted display device (1000) can align the left image and the right image of the stereo image on the same plane so that the same y coordinate points to corresponding pixels. In one embodiment, the head mounted display device (1000) can obtain the parallax map (620) by applying a block matching or semi-global matching (SGM) algorithm to the aligned left image and right image.

[0122] In one embodiment, the head mounted display device (1000) can obtain a parallax map (620) based on data measured from a distance sensor. In one embodiment, the external electronic device (2000) or the head mounted display device (1000) can include a distance sensor. The distance sensor can be a component that measures data regarding a distance between an object in a real environment and the distance sensor. In one embodiment, the external electronic device (2000) or the head mounted display device (1000) can obtain a parallax map (620) corresponding to a stereo image based on data measured from the distance sensor while capturing a stereo image, and transmit the obtained parallax map (620) to the head mounted display device (1000).

[0123] In one embodiment, the head mounted display device (1000) may compare the magnitude of the movement of the left image and the magnitude of the movement of the right image during the movement of the stereo images. Here, the comparison of the magnitudes of the movements of the left and right images may mean comparing the magnitudes of the motion vectors during the movement, comparing the magnitudes of the rotation vectors, or comparing the sum of the magnitudes of the motion vectors and the rotation vectors. However, the present invention is not necessarily limited thereto, and the head mounted display device (1000) may apply preset weights to each of the magnitudes of the motion vector and the rotation vector, and compare the sum of the magnitudes of the motion vector to which the weights are applied and the magnitudes of the rotation vectors to which the weights are applied.

[0124] In one embodiment, the head mounted display device (1000) may select an image with a smaller motion magnitude among the left and right images based on the comparison result. In one embodiment, the head mounted display device (1000) may reconstruct an unselected image based on the parallax map (620) and the selected image. In one embodiment, the head mounted display device (1000) may render a first image object that outputs the selected image and a second image object that outputs the reconstructed image.

[0125] For example, the head mounted display device (1000) can select the left image if the magnitude of the movement (30-1) of the left image is smaller than the magnitude of the movement (30-2) of the right image. Then, when the left image (610) is selected, the head mounted display device (1000) can reconstruct the unselected right image based on the parallax map (620) and the left image (610). Then, the head mounted display device (1000) can render the image object (120-1) that outputs the left image (610) on the left rendering space (110-1), and render the image object (120-2) that outputs the reconstructed right image (630) on the right rendering space (110-2). Here, the left rendering space (110-1) may be a virtual environment displayed through a display corresponding to the user's left eye, and the right rendering space (110-2) may be a virtual environment displayed through a display corresponding to the user's right eye.

[0126] In this way, according to one embodiment of the present disclosure, the head mounted display device (1000) can reconstruct an image with a large movement size among stereo images and output it through an image object. That is, if an image with a large movement size is stabilized through the first rendering mode, the degree to which the arrangement of the image object is adjusted is large, which may cause great eye fatigue to the user. Accordingly, the head mounted display device (1000) can reduce the movement of the stereo image in a way that is less tiring for the user by performing image stabilization through the second rendering mode that reconstructs an image with a large movement size based on an image with a small movement size.

[0127] FIG. 7 is a diagram illustrating an operation of selecting a rendering mode by a head-mounted display device according to one embodiment of the present disclosure. Steps S240 and S260 of FIG. 7 correspond to steps S240 and S260 of FIG. 2 , and thus, a redundant description thereof will be omitted.

[0128] In step S710, the head-mounted display device (1000) may compare the difference in the magnitude of the movement of the left image and the magnitude of the movement of the right image. In one embodiment, the difference in the movements of the left image and the right image may mean the difference in the magnitude of the translation vector, the difference in the magnitude of the rotation vector, or the sum of the difference in the magnitude of the translation vector and the difference in the magnitude of the rotation vector.

[0129] In step S715, the head-mounted display device (1000) can determine whether the difference in the magnitude of the movement is greater than or equal to a threshold value. In one embodiment, the threshold value is a value for determining whether image stabilization is performed in the first rendering mode or the second rendering mode, and may be a preset value or a value determined based on user input.

[0130] In step S720, if the difference in the magnitude of the movement is greater than or equal to a threshold value (S715-Y), the head mounted display device (1000) can perform the first rendering mode and the second rendering mode (S720). In other words, if the difference in the magnitude of the movement of the left image and the right image is greater than or equal to a threshold value, the head mounted display device (1000) can perform image stabilization by adjusting the arrangement of the image object outputting the left image and the image object outputting the right image, and outputting the image with the larger magnitude of the movement among the left image and the right image as a reconstructed image.

[0131] At step S730, if the magnitude of the movement is less than the threshold value (S715-N), the head mounted display device (1000) can perform the first rendering mode. In other words, if the difference in the magnitude of the movement between the left image and the right image is less than the threshold value, the head mounted display device (1000) can perform image stabilization in which only the arrangement of the image object outputting the left image and the image object outputting the right image is adjusted.

[0132] However, it is not necessarily limited to the above-described example, and the head mounted display device (1000) can selectively perform the first rendering mode and the second rendering mode based on the difference in the magnitude of the movement of the left image and the magnitude of the movement of the right image. For example, the head mounted display device (1000) can perform the first rendering mode and the second rendering mode if the difference in the magnitude of the movement is greater than or equal to a threshold value (S715-Y), and not perform the first rendering mode and the second rendering mode if the difference in the magnitude of the movement is less than or equal to the threshold value (S715-N). As another example, the head mounted display device (1000) can perform the second rendering mode if the difference in the magnitude of the movement is greater than or equal to a threshold value (S715-Y), and can perform the first rendering mode if the difference in the magnitude of the movement is less than or equal to the threshold value (S715-N). In other words, by performing at least one of the first rendering mode and the second rendering mode depending on the hardware resources of the head mounted display device (1000) and the purpose of image stabilization, the movement of the stereo image perceived by the user can be reduced.

[0133] In this way, the head-mounted display device (1000) according to one embodiment of the present disclosure adjusts the arrangement of image objects based on the movement of a stereo image. However, when the difference in the magnitude of movement between the left and right images of the stereo image is large, the image object with the large magnitude of movement is reconstructed, thereby performing image stabilization. Accordingly, the user's visual discomfort that cannot be overcome by adjusting the arrangement of image objects alone due to the difference in the magnitude of movement can be reduced.

[0134] FIG. 8A is a flowchart illustrating a method for a head-mounted display device according to one embodiment of the present disclosure to obtain motion data representing camera shake through filtering. In one embodiment, step S800-1 of FIG. 8A may correspond to step S220 of FIG. 2 .

[0135] In step S810, the head-mounted display device (1000) may acquire first motion data representing the movement of the camera at the time of capturing the stereo image. In one embodiment, the first motion data corresponds to the entire motion data (or raw data) of the camera and may include intentional and unintentional camera movements.

[0136] In step S822, the head-mounted display device (1000) can obtain second motion data by filtering the first motion data with a low-pass filter. Here, the second motion data can include data indicating intentional camera motion. In other words, the head-mounted display device (1000) can obtain data indicating intentional camera motion based on filtering data regarding the overall motion of the camera with a low-pass filter.

[0137] In one embodiment, the low-pass filter may include an Infinite Impulse Response (IIR) filter. For example, the low-pass filter may include a second-order low-pass filter, such as a Butterworth filter, a Chebyshev filter, or an elliptic filter. However, the present invention is not necessarily limited to the examples described above, and the order or type of the low-pass filter may vary depending on the type of data being filtered (e.g., whether it is the movement of a camera or the movement of a user), the mathematical representation of the data (e.g., a quaternion, a rotation matrix, Euler angles, etc.), or the design of the electronic device performing the filtering.

[0138] In one embodiment, the head-mounted display device (1000) may adjust the cutoff frequency of the low-pass filter based on the amount of change in the movement of the camera. However, the adjustment of the cutoff frequency may be referred to by various expressions representing the same or similar concepts. For example, the adjustment of the cutoff frequency may be replaced with expressions such as "determine", "modify", and "calibrate" of the cutoff frequency, and is not limited to the examples described above.

[0139] In one embodiment, the head mounted display device (1000) may calculate a change in motion in a preset time interval based on the first motion data. Here, the change in motion may include an average change in the entire preset time interval or a maximum instantaneous change in the preset time interval. In one embodiment, the preset time interval may include the entire time interval in which the first motion data is measured, or may include each of a plurality of time intervals that divide the entire time interval into preset time intervals. In one embodiment, the head mounted display device (1000) may adjust a cutoff frequency of a low-pass filter applied to a corresponding time interval to increase as the calculated change amount increases. In one embodiment, the head mounted display device (1000) may adjust a cutoff frequency of a low-pass filter applied to a corresponding time interval to decrease as the calculated change amount decreases. In one embodiment, the head mounted display device (1000) may adjust a cutoff frequency of a low-pass filter applied to a corresponding time interval based on whether the calculated change amount exceeds a threshold value. For example, the head-mounted display device (1000) can adjust the cutoff frequency of the low-pass filter to a first frequency when the calculated change exceeds a threshold value, and can adjust the cutoff frequency of the low-pass filter to a second frequency lower than the first frequency when the calculated change is less than or equal to the threshold value.

[0140] In one embodiment, the lower the cutoff frequency of the low-pass filter applied to the first motion data, the more the second motion data can be obtained from which more high-frequency components representing the motion to be corrected are removed from the first motion data, but the phase delay phenomenon can be greater. Conversely, the higher the cutoff frequency of the low-pass filter applied to the first motion data, the less the high-frequency components can be removed from the first motion data, but the phase delay phenomenon can be smaller.

[0141] A head mounted display device (1000) according to one embodiment of the present disclosure can adaptively adjust the cutoff frequency of a low-pass filter applied to first motion data according to the magnitude of the change in motion. Accordingly, the head mounted display device (1000) can obtain second motion data that includes more data on intentional motion by applying a low-pass filter with a low cutoff frequency to a section in the first motion data in which the magnitude of the change in motion is small and the phase delay phenomenon is not greatly affected. Conversely, the head mounted display device (1000) can obtain second motion data in which the phase delay phenomenon occurs less by applying a low-pass filter with a high cutoff frequency to a section in the first motion data in which the magnitude of the change in motion is large and the phase delay phenomenon is greatly affected.

[0142] In step S830, the head mounted display device (1000) can obtain third motion data representing camera shake based on the difference between the first motion data and the second motion data. In one embodiment, the head mounted display device (1000) can calculate the difference between the first motion data and the second motion data and obtain the calculated difference as third motion data representing camera shake. Here, the camera shake may correspond to unintentional camera movement. In other words, the head mounted display device (1000) can obtain the difference between data on the overall movement of the camera and data on intentional camera movement as data on unintentional camera movement (data on camera shake).

[0143] In one embodiment, the head mounted display device (1000) may provide information about the movement of the camera, including third movement data representing camera shake obtained through step S830, to the image stabilization module (410) of FIG. 4.

[0144] In this way, the head mounted display device (1000) according to one embodiment of the present disclosure can obtain data on intentional camera movement (second movement data) from data on the overall movement of the camera (first movement data) by performing filtering through a low-pass filter. In addition, the head mounted display device (1000) can obtain information on the movement of the camera including data indicating camera shake (third movement data) by calculating the difference between the first movement data and the second movement data. Accordingly, when image stabilization including at least one of the first rendering mode and the second rendering mode is performed, image stabilization reflecting only unintentional camera movement can be performed.

[0145] FIG. 8B is a flowchart illustrating a method for a head-mounted display device according to an embodiment of the present disclosure to acquire data indicating camera shake based on predicted movement. In one embodiment, step S800-2 of FIG. 8B may correspond to step S220 of FIG. 2 . Furthermore, steps S810 and S830 of FIG. 8B correspond to steps S810 and S830 of FIG. 8A , and thus, a redundant description thereof will be omitted.

[0146] In step S824, the head mounted display device (1000) may obtain second motion data representing motion of a camera at a subsequent time point predicted from motion of a camera at a previous time point based on the first motion data. Here, the motions at each time point of the second motion data may be motions predicted based on the motion of the previous time point for each time point of the first motion data. In one embodiment, the head mounted display device (1000) may predict a next state based on a current state through a SLAM algorithm. In one embodiment, the head mounted display device (1000) may obtain data representing motion of a next state predicted based on the first motion data as second motion data.

[0147] For example, the head-mounted display device (1000) can use the state prediction equation of the Kalman filter below to perform the prediction step of the SLAM algorithm.

[0148]

[0149] Mathematical equation 1 is the state prediction equation of the Kalman filter, can correspond to the second motion data as a predicted state vector at time k, can correspond to the first motion data as the previous state vector at time k-1. In addition, is the state transition matrix, is the control input matrix, may mean a control input at time k. However, it is not necessarily limited to the above-described example, and the head-mounted display device (1000) may obtain second movement data representing movement at a subsequent time point predicted from movement at a previous time point based on first movement data, which is actually measured data, through data-based prediction using an artificial intelligence model or rule-based prediction using preset rules and algorithms.

[0150] In this way, an electronic device according to an embodiment of the present disclosure can obtain data predicted based on data on the overall movement of the camera (first movement data) as data on intentional camera movement (second movement data). In this case, the predicted data may include data that is predicted and obtained so as not to include data that is caused by factors that are difficult to predict from actually measured data. In addition, the predicted data may include data that does not cause a phase lag phenomenon. Accordingly, the head mounted display device (1000) can obtain data on intentional camera movement (second movement data) that does not cause a phase lag phenomenon while having a low possibility of including unintentional data, and can obtain information on camera movement that includes data indicating camera shake (third movement data) by calculating the difference between the first movement data and the second movement data. Accordingly, when image stabilization including the first rendering mode and the second rendering mode is performed, image stabilization that does not cause a phase lag phenomenon can be performed.

[0151] FIG. 9 is a diagram illustrating a shaking data acquisition module according to one embodiment of the present disclosure.

[0152] Referring to FIG. 9, the head mounted display device (1000) may include a motion data acquisition module (910) and a shake data acquisition module (920). The motion data acquisition module (910) and the shake data acquisition module (920) of the present disclosure may be understood as operations of the head mounted display device (1000) or a processor of the head mounted display device (1000).

[0153] In one embodiment, the motion data acquisition module (910) may be a component that outputs first motion data (911). In one embodiment, the motion data acquisition module (910) may output motion data representing the movement of a camera at the time of capturing a stereo image. In one embodiment, the first motion data (911) may be expressed in a 6 Dof data format, but is not necessarily limited thereto, and may include a quaternion including rotation information, a rotation matrix, and Euler angles, a transformation matrix including movement information, and a movement vector, etc. Hereinafter, for the convenience of describing the invention, the first motion data (911), the second motion data (923), and the shake data (925) will all be described as 'roll' data among the 6 Dof data.

[0154] In one embodiment, the motion data acquisition module (910) may output first motion data (911) representing the motion of the camera at the time of capturing the stereo image based on the metadata of the stereo image (100). In one embodiment, the metadata of the stereo image (100) may include motion data of the camera mapped and stored to each of a plurality of frames of the stereo image (100). In one embodiment, the motion data acquisition module (910) may acquire the first motion data (911) of a continuous time section by temporally reconstructing the motion of the camera of each of the plurality of frames included in the metadata of the stereo image.

[0155] In one embodiment, the movement data acquisition module (910) may output data acquired through SLAM of the head mounted display device (1000). In one embodiment, when the head mounted display device (1000) includes a camera that captures stereo images, data representing the movement of the head mounted display device (1000) acquired while capturing the stereo images may be output as first movement data (911).

[0156] In one embodiment, the shake data acquisition module (920) may be a component that outputs shake data (925) based on the first movement data (911). In one embodiment, the shake data acquisition module (920) may include a smoothing module (922) and a differential signal calculation module (924).

[0157] In one embodiment, the smoothing module (922) may be a component that outputs second motion data (923) based on the first motion data (911). In one embodiment, the smoothing module (922) may obtain second motion data (923) by filtering the first motion data (911) with a low-pass filter. In one embodiment, the smoothing module (922) may adjust a cutoff frequency of the low-pass filter based on the amount of change in the motion of the first motion data (911). In one embodiment, the smoothing module (922) may obtain second motion data (923) representing motion of a subsequent time point predicted from motion of a previous time point based on the first motion data (911). Since the operations and functions of the smoothing module (922) to obtain second movement data (923) based on the first movement data (911) correspond to the operations and functions of the head mounted display device described in FIGS. 8A and 8B, a redundant description will be omitted.

[0158] In one embodiment, the differential signal calculation module (924) may be a component that outputs shake data (925), which is the difference between the first motion data (911) and the second motion data (923). Here, the shake data (925) may correspond to motion data representing the shake of the camera acquired by the head mounted display device (1000) in step S830 of FIGS. 8A and 8B . In one embodiment, the differential signal calculation module (924) may calculate the difference between the motion of the first motion data (911) at each point in time and the motion of the second motion data (923) at each point in time, and may acquire the calculated difference as the shake data (925). For example, if the roll value of the first motion data at a point in time of 1 second is +0.2 and the roll value of the second motion data is -0.1, the roll value of the shake data (925) at a point in time of 1 second may be +0.3. Additionally, if the roll value of the first movement data at the 2-second point is -1 and the roll value of the second movement data is -1.2, the roll value of the shaking data (925) at the 2-second point may be -0.2.

[0159] Referring to FIG. 9, the first motion data (911) may include both low-frequency signal components representing changes in small movements over the entire time interval and high-frequency signal components representing large movements. In other words, the first motion data (911) may be raw data representing the entire movement of the camera, and thus may be data representing both intentional and unintentional camera movements.

[0160] On the other hand, the second motion data (923) may be data from which high frequency signal components indicating changes in small movements are removed compared to the first motion data (911). That is, the smoothing module (922) may use a low-pass filter or predicted data to exclude high frequency signal components indicating changes in small movements from the first motion data (911) and output the second motion data (923) including low frequency signal components indicating large movements.

[0161] In addition, since the shake data (925) is data indicating the difference between the first movement data (911) and the second movement data (923), it may include high-frequency signal components indicating changes in small movements. Since such shake data (925) is data indicating unnecessary movements unrelated to the user's intention, image stabilization based on the shake data (925) can be performed to reduce visual discomfort due to unintended camera movements.

[0162] FIG. 10 is a perspective view of a head mounted display device according to one embodiment of the present disclosure.

[0163] Referring to FIG. 10, a head-mounted display device (1000) may include a frame (1001), an optical system (1002), a display (1100), a memory (1200), a motion sensor (1400), and a processor (1700). However, the present invention is not limited to the above-described examples, and some components may be omitted or other components may be added.

[0164] In one embodiment, the frame (1001) includes other components of the head mounted display device (1000) and is configured to allow a user to mount the head mounted display device (1000), including, but not limited to, glasses temples, a nose bridge, and the like. In one embodiment, left-eye optical components and right-eye optical components may be arranged or attached to the left and right sides of the frame (1001), or the left-eye optical components and right-eye optical components may be integrally formed and mounted to the frame (1001). In another example, some of the optical components may be arranged or attached to only one of the left and right sides of the frame (1001).

[0165] In one embodiment, the optical system (1002) may be a component that transmits light of an image to a user's eye. In one embodiment, the optical system (1002) may include at least one lens having a refractive power (degree) to focus or redirect light of an image output from the display (1100). In one embodiment, light of an image output from the display (1100) may pass through the optical system (1002) and enter the user's eye.

[0166] The display (1100) is a component for displaying images and / or videos. The light of the image output from the display (1100) may be incident on the eyes of a user wearing the virtual reality head-mounted display device (1000). In one embodiment, the display (1100) may be configured as a physical device including at least one of a liquid crystal display, a thin film transistor-liquid crystal display, an organic light-emitting diode (OLED), a flexible display, a 3D display, and an electrophoretic display. In one embodiment, the display (1100) may include a display (1100-1) corresponding to the user's left eye (or left image) and a display (1100-2) corresponding to the user's right eye (or right image). The display (1100-1) corresponding to the left eye may display a left image of a stereo image, and the display (1100-2) corresponding to the right eye may display a right image of the stereo image. However, it is not necessarily limited thereto, and the head mounted display device (1000) may include a single display, and the left image may be displayed in one area of ​​the single display, and the right image may be displayed in another area.

[0167] In one embodiment, the head mounted display device (1000) may include electronic components such as a memory (1200), a motion sensor (1400), and a processor (1800), and the electronic components may be mounted on a PCB substrate, an FPCB substrate, or the like and positioned at one location of the frame (1001) or may be distributed and positioned at multiple locations. In one embodiment, the electronic components included in the head mounted display device (1000) may further include a communication interface, an input interface, an output interface, and the like, and specific details regarding the operation of the electronic components of the head mounted display device (1000) will be described again below with reference to FIG. 11.

[0168] FIG. 11 is a detailed configuration diagram of a head-mounted display device according to an embodiment of the present disclosure. Referring to FIG. 11, a head-mounted display device (1000) may include a display (1100), a memory (1200), a communication interface (1300), a motion sensor (1400), an input interface (1500), an output interface (1600), and a processor (1700). The display (1100), the memory (1200), the communication interface (1300), the motion sensor (1400), the input interface (1500), the output interface (1600), and the processor (1700) may each be electrically and / or physically connected to each other.

[0169] The components illustrated in FIG. 11 are merely according to one embodiment of the present disclosure, and the components included in the head mounted display device (1000) are not limited to those illustrated in FIG. 11. The head mounted display device (1000) according to one embodiment of the present disclosure may not include some of the components illustrated in FIG. 11, and may further include components not illustrated in FIG. 11. In addition, descriptions of overlapping content with those described in FIGS. 1 to 10 among the components illustrated in FIG. 11 will be omitted.

[0170] The memory (1200) may store instructions or program codes for performing functions or operations of the head-mounted display device (1000). In one embodiment, at least one instruction, algorithm, data structure, program code, and application program stored in the memory (1200) may be implemented in a programming or scripting language such as, for example, C, C++, Java, or an assembler.

[0171] In one embodiment, the memory (1200) may include at least one of a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., SD or XD memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a mask ROM, a flash ROM, etc.), a hard disk drive (HDD), or a solid state drive (SSD).

[0172] In one embodiment, the memory (1200) may include an image stabilization module (410) and a rendering module (420). The image stabilization module (410) may include a high-speed image stabilization module (412) and a depth-adapted image stabilization module (414). In one embodiment, the processor (1700) may perform at least one of a first rendering mode and a second rendering mode through the image stabilization module (410). In one embodiment, the processor (1700) may render an image object through the rendering module (420). Since the specific details of the image stabilization module (410) and the rendering module (420) have been described in FIG. 3, a redundant description will be omitted.

[0173] In one embodiment, the memory (1200) may include stereoscopic images and camera movement information. In one embodiment, the information regarding the stereoscopic images and camera movement may be received from an external electronic device that captured the stereoscopic images via a communication interface (1300) and stored in the memory (1200), or may be received from an external storage device (e.g., an external server) that stores stereoscopic images and camera movement information and stored in the memory (1200). In one embodiment, the memory (1200) may include information regarding the user's movement acquired based on data measured from a motion sensor (1400). In one embodiment, the memory (1200) may include an inpainting model and a stereo image reconstruction model used to acquire a reconstructed image. However, the present invention is not limited to the above-described examples, and the memory (1200) may further include various data necessary to perform the operations and functions of the head-mounted display device (1000) disclosed herein.

[0174] The communication interface (1300) is a component for the head mounted display device (1000) to communicate with an external electronic device. In one embodiment, the communication interface (1300) may perform data communication between the head mounted display device (1000) and the external electronic device using at least one of data communication methods including wired LAN, wireless LAN, Wi-Fi, Bluetooth, zigbee, Wi-Fi Direct (WFD), infrared Data Association (IrDA), Bluetooth Low Energy (BLE), Near Field Communication (NFC), Wireless Broadband Internet (Wibro), World Interoperability for Microwave Access (WiMAX), Shared Wireless Access Protocol (SWAP), Wireless Gigabit Alliance (WiGig), and RF communication.

[0175] In one embodiment, the communication interface (1300) can receive information about stereoscopic images and camera movements from an external electronic device. In one embodiment, the communication interface (1300) can receive a parallax map corresponding to the stereoscopic images. In one embodiment, the communication interface (1300) can transmit the stereoscopic images and movements to the external electronic device, and receive a rendered image object from the external electronic device. However, the present invention is not limited to the above-described examples, and various data necessary for performing the operations and functions of the head-mounted display device (1000) disclosed herein can be transmitted and received with the external electronic device through the communication interface (1300).

[0176] The motion sensor (1400) is a component for measuring the position, speed, direction, posture, etc. of the head mounted display device (1000). In one embodiment, the motion sensor (1400) may include an IMU (Inertial Measurement Unit) sensor. In one embodiment, the head mounted display device (1000) may obtain information about the movement of a user wearing the head mounted display device (1000) based on data measured from the motion sensor (1400). In one embodiment, the motion sensor (1400) may measure data related to at least one of movement and rotation of the user wearing the head mounted display device (1000) in a three-dimensional space. In one embodiment, data measured from the motion sensor (1400) may be provided to the processor (1700).

[0177] The input interface (1500) is a component for receiving various user inputs. In one embodiment, the input interface (1500) may include a touch panel, a physical button, a microphone, etc. In one embodiment, information input through the input interface (1500) may be provided to the processor (1700). In one embodiment, a user input for selecting one of a first rendering mode and a second rendering mode may be obtained through the input interface (1500). In this case, the head mounted display device (1000) may perform the selected rendering mode based on the user input. However, the present invention is not limited to the above-described example, and various data for performing the operations and functions of the head mounted display device (1000) disclosed herein may be input through the input interface (1500).

[0178] The output interface (1600) is a component for the head-mounted display device (1000) to provide various information to the user. In one embodiment, the output interface (1600) may include a speaker. In one embodiment, the output interface (1600) may output a voice corresponding to text displayed through the user interface based on a signal received from the processor (1700), or may output a voice indicating whether the rendering mode is activated. However, the present invention is not limited to the above-described example, and various voices for performing the operations and functions of the head-mounted display device (1000) disclosed herein may be output through the output interface (1600).

[0179] The processor (1700) can control the overall operations of the head mounted display device (1000). In one embodiment, the processor (1700) can include multiple processors. In one embodiment, at least one processor (1700) can perform the operations and functions of the head mounted display device (1000) disclosed herein by executing one or more instructions of a program stored in the memory (1200).

[0180] The processor (1700) may be configured as at least one of, for example, a Central Processing Unit, a microprocessor, a Graphic Processing Unit, Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), an Application Processor, a Neural Processing Unit, or an artificial intelligence processor designed with a hardware structure specialized for processing artificial intelligence models, but is not limited thereto.

[0181] When a method according to an embodiment of the present disclosure includes multiple operations, the multiple operations may be performed by a single processor or by multiple processors. For example, when a first operation, a second operation, and a third operation are performed by a method according to an embodiment, the first operation, the second operation, and the third operation may all be performed by a first processor, or the first and second operations may be performed by a first processor and the third operation may be performed by a second processor. However, the embodiments of the present disclosure are not limited thereto.

[0182] One or more processors according to the present disclosure may be implemented as a single-core processor or a multi-core processor. If a method according to an embodiment of the present disclosure includes multiple operations, the multiple operations may be performed by a single core or by multiple cores included in one or more processors.

[0183] In one embodiment, at least one processor (1700) can acquire a stereo image by executing one or more commands. In one embodiment, at least one processor (1700) can acquire information about the movement of a camera that captured a stereo image by executing one or more commands. In one embodiment, at least one processor (1700) can acquire information about the movement of a user wearing a head-mounted display device by executing one or more commands. In one embodiment, at least one processor (1700) can identify the movement of a stereo image for a user based on information about the movement of the camera and information about the movement of the user by executing one or more commands. In one embodiment, at least one processor (1700) can render an image object that outputs a stereo image based on the movement of the stereo image by executing one or more commands. In one embodiment, at least one processor (1700) can display the rendered image object by executing one or more commands.

[0184] In one embodiment, the motion of the stereo image may include the motion of the left image and the motion of the right image of the stereo image. In one embodiment, at least one processor (1700) may execute one or more instructions to perform at least one of a first rendering mode that adjusts the placement of image objects based on the difference in the magnitude of the motion of the left image and the magnitude of the motion of the right image, and a second rendering mode that reconstructs one of the left image and the right image.

[0185] In one embodiment, at least one processor (1700) may adjust the placement of image objects based on the movement of a stereo image by executing one or more commands. In one embodiment, at least one processor (1700) may perform a first rendering mode by rendering image objects whose placement has been adjusted by executing one or more commands.

[0186] In one embodiment, at least one processor (1700) may apply at least one of position translation and rotation of a video object to correspond to the movement of a stereo image in the space in which the video object is rendered by executing one or more instructions.

[0187] In one embodiment, at least one processor (1700) may obtain a disparity map of a stereo image by executing one or more commands. In one embodiment, at least one processor (1700) may select an image with a large motion magnitude among the left image and the right image by executing one or more commands. In one embodiment, at least one processor (1700) may reconstruct an image not selected among the left image and the right image based on the disparity map and the selected image by executing one or more commands. In one embodiment, at least one processor (1700) may render a first image object displaying the selected image and a second image object outputting the reconstructed image by executing one or more commands to perform a second rendering mode.

[0188] In one embodiment, at least one processor (1700) may obtain first motion data representing the movement of a camera at a time point when capturing a stereo image by executing one or more commands. In one embodiment, at least one processor (1700) may obtain second motion data by filtering the first motion data with a low-pass filter by executing one or more commands. In one embodiment, at least one processor (1700) may obtain third motion data representing camera shake based on a difference between the first motion data and the second motion data by executing one or more commands.

[0189] In one embodiment, at least one processor (1700) may adjust a cut-off frequency of a low-pass filter based on a change in movement of the camera by executing one or more instructions.

[0190] In one embodiment, at least one processor (1700) may obtain first motion data representing the motion of a camera at a time point when a stereo image is captured by executing one or more commands. In one embodiment, at least one processor (1700) may obtain second motion data representing the motion of a camera at a subsequent time point predicted from the motion of a camera at a previous time point based on the first motion data by executing one or more commands. In one embodiment, at least one processor (1700) may obtain third motion data representing camera shake based on a difference between the first motion data and the second motion data by executing one or more commands. In one embodiment, the display (1100) may include a first display and a second display. In one embodiment, at least one processor (1700) may display a first image object through a first display corresponding to a selected image, and display a second image object through a second display corresponding to a reconstructed image, by executing one or more commands.

[0191] In one embodiment, information about the movement of the camera may include data obtained from an Inertial Measurement Unit (IMU) sensor of the stereo camera or metadata related to Electronic Image Stabilization (EIS) of the camera.

[0192] Information about the user's movements may include information about the movement of the user's left eye and information about the movement of the user's right eye. In one embodiment, information about the camera's movements may include information about the movement of the left camera corresponding to the user's left eye and information about the movement of the right camera corresponding to the user's right eye.

[0193] In one embodiment, at least one processor (1700) can identify movement of a left image based on a difference between movement of a user's left eye and movement of a left camera, and can identify movement of a right image based on a difference between movement of a user's right eye and movement of a right camera by executing one or more instructions.

[0194] In one embodiment, at least one processor (1700) can reconstruct unselected images by warping the selected images according to the disparity map, or by applying the disparity map and the selected images to an artificial intelligence model, by executing one or more instructions.

[0195] However, it is not necessarily limited to the above-described examples, and at least one processor (1700) can perform various operations and functions of the head mounted display device (1000) disclosed in the present specification by executing one or more commands.

[0196] FIG. 12 is a detailed configuration diagram of an external electronic device according to one embodiment of the present disclosure.

[0197] Referring to FIG. 12, the external electronic device (2000) may include a camera (2100), a memory (2200), a motion sensor (2300), a communication interface (2400), and a processor (2500). The camera (2100), the memory (2200), the motion sensor (2300), the communication interface (2400), and the processor (2500) may each be electrically and / or physically connected to each other.

[0198] The components illustrated in FIG. 12 are merely in accordance with one embodiment of the present disclosure, and the components included in the external electronic device (2000) are not limited to those illustrated in FIG. 12. The external electronic device (2000) according to one embodiment of the present disclosure may not include some of the components illustrated in FIG. 12, and may further include components not illustrated in FIG. 12. In addition, descriptions of overlapping content with those described in FIGS. 1 to 11 among the components illustrated in FIG. 12 will be omitted.

[0199] The camera (2100) may be a component for capturing images of objects and backgrounds within a real environment by capturing a real environment. In one embodiment, the camera (2100) may capture stereoscopic images. In one embodiment, the camera (2100) for capturing stereoscopic images may include a camera for capturing a left image of the stereoscopic image and a camera for capturing a right image. In one embodiment, the camera (2100) may include a plurality of cameras (or lenses and image sensors) having different FOVs. In one embodiment, the camera (2100) may include a lens module, an image sensor, and an image processing module, and may capture images or videos obtained by the image sensor (e.g., CMOS or CCD). In one embodiment, images and videos captured by the camera (2100) may be transmitted to the processor (2500).

[0200] The memory (2200) may store instructions or program codes for performing functions or operations of the external electronic device (2000). In one embodiment, at least one instruction, algorithm, data structure, program code, and application program stored in the memory (2200) may be implemented in a programming or scripting language such as, for example, C, C++, Java, or an assembler.

[0201] In one embodiment, the memory (2200) may include a stereo image captured by the camera (2100). In one embodiment, the memory (2200) may include information regarding the movement of the camera acquired by the motion sensor (2300). In one embodiment, the memory (2200) may include a parallax map corresponding to the stereo image acquired by the distance sensor (not shown). However, the present invention is not limited to the above-described example, and the memory (2200) may further include various data necessary to perform the operation and function of the external electronic device (2000) disclosed herein.

[0202] The motion sensor (2300) is a component for measuring the position, speed, direction, posture, etc. of the external electronic device (2000). In one embodiment, the motion sensor (2300) may include an IMU (Inertial Measurement Unit) sensor. In one embodiment, the external electronic device (2000) may obtain information about the movement of the external electronic device (2000) or the camera (2100) based on data measured from the motion sensor (2300). In one embodiment, the motion sensor (2300) may measure data related to at least one of movement and rotation of the external electronic device (2000) or the camera (2100) in a three-dimensional space. In one embodiment, data measured from the motion sensor (2300) may be provided to the processor (2500).

[0203] The communication interface (2400) is a component for the external electronic device (2000) to communicate with the head mounted display device (1000). In one embodiment, the communication interface (2400) can transmit information about stereo images and camera movements to the head mounted display device (1000). In one embodiment, the communication interface (2400) can transmit a parallax map corresponding to the stereo images to the head mounted display device (1000). However, the present invention is not limited to the above-described example, and various data required to perform the operations and functions of the external electronic device (2000) disclosed herein can be transmitted and received with the head mounted display device (1000) through the communication interface (2400).

[0204] The processor (2500) can control the overall operation of the external electronic device (2000). In one embodiment, at least one processor (2500) can perform the operation and function of the head external electronic device (2000) disclosed herein by executing one or more instructions of a program stored in the memory (2200).

[0205] In one embodiment, at least one processor (2500) can obtain information about the movement of the camera (2100) based on data measured from the motion sensor (2300) by executing one or more commands of a program stored in the memory (2200). At least one processor (2500) can obtain a parallax map corresponding to a stereo image based on data measured from a distance sensor by executing one or more commands of a program stored in the memory (2200). In one embodiment, at least one processor (2500) can obtain a parallax map of a stereo image based on the stereo image by executing one or more commands of a program stored in the memory (2200). However, the present invention is not limited to the above-described example, and at least one processor (2500) can perform various operations and functions of the external electronic device (2000) disclosed herein by executing one or more commands of a program stored in the memory (2200).

[0206] Meanwhile, embodiments of the present disclosure may also be implemented in the form of a recording medium containing computer-executable instructions, such as program modules, executed by a computer. Computer-readable media may be any available media that can be accessed by a computer, and include both volatile and nonvolatile media, removable and non-removable media. Furthermore, computer-readable media may include computer storage media and communication media. Computer storage media include both volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Communication media may typically include computer-readable instructions, data structures, or other data in a modulated data signal, such as program modules.

[0207] Additionally, a computer-readable storage medium may be provided in the form of a non-transitory storage medium. Here, the term "non-transitory storage medium" simply means a tangible device that does not contain signals (e.g., electromagnetic waves). This term does not distinguish between cases where data is permanently stored in the storage medium and cases where data is temporarily stored. For example, a "non-transitory storage medium" may include a buffer in which data is temporarily stored.

[0208] The above description of the present disclosure is provided for illustrative purposes only, and those skilled in the art will readily appreciate that modifications to other specific forms can be made without altering the technical spirit or essential features of the present disclosure. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive. For example, components described as being single may be implemented in a distributed manner, and similarly, components described as being distributed may be implemented in a combined manner.

[0209] The scope of the present disclosure is indicated by the claims described below rather than the detailed description above, and all changes or modifications derived from the meaning and scope of the claims and their equivalent concepts should be interpreted as being included in the scope of the present disclosure.

Claims

1. In a method of operating a head-mounted display device, Step of acquiring a stereo image (S210); A step of obtaining information about the movement of the camera that captured the stereo image (S220); Step (S230) of acquiring information about the movement of a user wearing a head mounted display device; A step (S240) of identifying the movement of the stereo image for the user based on information about the movement of the camera and information about the movement of the user; A step (S250) of rendering an image object that outputs the stereo image based on the movement of the stereo image; and A method comprising a step (S260) of displaying the rendered image object.

2. In paragraph 1, The movement of the above stereo image is, Includes the motion of the left image and the motion of the right image of the above stereo image, The step of rendering the above video object is: A method comprising: performing at least one of a first rendering mode for adjusting the arrangement of the image object based on a difference between the magnitude of the movement of the left image and the magnitude of the movement of the right image, and a second rendering mode for reconstructing one of the left image and the right image.

3. In paragraph 2, The step of performing the above first rendering mode is: A step of adjusting the arrangement of the image object based on the movement of the stereo image; and A method comprising the step of rendering an image object whose arrangement has been adjusted.

4. In any one of paragraphs 2 to 3, The step of performing the above second rendering mode is: A step of obtaining a disparity map of the stereo image; A step of selecting an image having a small motion size among the left image and the right image; A step of reconstructing an unselected image among the left image and the right image based on the disparity map and the selected image; and A method comprising the step of rendering a first image object that outputs the selected image and a second image object that outputs the reconstructed image.

5. In any one of paragraphs 1 to 4, The step of obtaining information about the movement of the above camera is: A step of acquiring first movement data representing the movement of the camera at the time of shooting the stereo image; A step of filtering the first motion data with a low-pass filter to obtain second motion data; and A method comprising: obtaining third motion data representing camera shake based on a difference between the first motion data and the second motion data.

6. In paragraph 5, The step of filtering the above first movement data with the low-pass filter is: A method further comprising: a step of adjusting the cut-off frequency of the low-pass filter based on the amount of change in movement of the camera.

7. In any one of paragraphs 1 to 4, The step of obtaining information about the movement of the above camera is: A step of acquiring first movement data representing the movement of the camera at the time of shooting the stereo image; A step of obtaining second motion data representing the motion of the camera at a later time point predicted from the motion of the camera at a previous time point based on the first motion data; and A method comprising: obtaining third motion data representing camera shake based on a difference between the first motion data and the second motion data.

8. In the head mounted display device (1000), display (1100); A memory (1200) storing one or more instructions; and At least one processor (1700) for executing one or more instructions stored in the memory; The head mounted display device (1000) executes the one or more commands by the at least one processor (1700). Acquire stereo images, Obtain information about the movement of the camera that captured the above stereo image, Obtain information about the movements of a user wearing a head-mounted display device, Identifying the movement of the stereo image for the user based on information about the movement of the camera and information about the movement of the user, Rendering a video object that outputs the stereo image based on the movement of the stereo image, A head-mounted display device that displays the rendered image object.

9. In paragraph 8, The movement of the above stereo image is, Includes the motion of the left image and the motion of the right image of the above stereo image, The head mounted display device, wherein the at least one processor executes the one or more commands, A head mounted display device that performs at least one of a first rendering mode for adjusting the arrangement of the image object based on a difference in the magnitude of the movement of the left image and the magnitude of the movement of the right image, and a second rendering mode for reconstructing one of the left image and the right image.

10. In paragraph 9, The head mounted display device, wherein the at least one processor executes the one or more commands, Adjusting the placement of the image object based on the movement of the stereo image, A head mounted display device that performs the first rendering mode by rendering the image object whose arrangement has been adjusted.

11. In any one of paragraphs 9 to 10, The head mounted display device, wherein the at least one processor executes the one or more commands, Obtain a disparity map of the above stereo image, Select the image with the largest movement size among the above left and right images, Reconstructing an unselected image among the left image and the right image based on the parallax map and the selected image, A head-mounted display device that performs a second rendering mode by rendering a first image object displaying the selected image and a second image object outputting the reconstructed image.

12. In any one of paragraphs 8 to 11, The head mounted display device, wherein the at least one processor executes the one or more commands, Obtaining first motion data representing the movement of the camera at the time of shooting the stereo image, Filtering the above first motion data with a low-pass filter to obtain second motion data, A head mounted display device that obtains third motion data representing shake of the camera based on the difference between the first motion data and the second motion data.

13. In paragraph 12, The head mounted display device, wherein the at least one processor executes the one or more commands, A head-mounted display device that adjusts the cut-off frequency of the low-pass filter based on the amount of change in movement of the camera.

14. In any one of paragraphs 8 to 11, The head mounted display device, wherein the at least one processor executes the one or more commands, Obtaining first motion data representing the movement of the camera at the time of shooting the stereo image, Obtaining second motion data representing the motion of the camera at a later time point predicted from the motion of the camera at a previous time point based on the first motion data; A head mounted display device that obtains third motion data representing shake of the camera based on the difference between the first motion data and the second motion data.

15. A computer-readable recording medium recording a program for executing the method of any one of clauses 1 to 7 on a computer.

Citation Information

Patent Citations

  • Stereoscopic image display device and stereoscopic image display method

    JP2011141381A

  • Image processing device, method, and program

    JP2013020527A

  • Light control system using power line communication

    KR1020220001209A

  • Surgeon head-mounted display apparatuses

    US10013808B2

  • Information processing device, information processing method, and image display system

    WO2016017245A1