Head-mounted display device and control method thereof
The head-mounted display device addresses the issue of obstructed views by inpainting detected objects based on depth information, enhancing immersion in augmented and virtual reality experiences.
Patent Information
- Application Number
- PCT/KR2025/007181
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-31
- Filing Date
- 2025-05-27
- Publication Date
- 2025-12-04
AI Technical Summary
Existing head-mounted display devices struggle to provide immersive augmented and virtual reality experiences due to real-world objects obstructing the view, as they do not effectively inpaint or remove interfering elements based on spatial awareness.
A head-mounted display device that captures a real-world environment, detects objects using depth information, identifies target objects within a predefined distance range, and inpaints those objects to create an immersive experience by reconstructing the obstructed areas based on spatial information.
Enhances user immersion by removing obstructing objects from the view, providing a more engaging and unobstructed experience in augmented and virtual reality environments.
Smart Images

Figure KR2025007181_04122025_PF_FP_ABST
Abstract
Description
Head-mounted display device and control method thereof
[0001] A head-mounted display device and a control method thereof are disclosed. Specifically, a head-mounted display device and a control method thereof are disclosed that perform inpainting based on the distance between the head-mounted display device and an object in an image captured in a real-world environment.
[0002] Video See-Through (VST) on head-mounted display (HMD) devices is a feature that allows users to observe the real environment through video in a virtual reality (VR) or augmented reality (AR) environment.
[0003] By inpainting objects existing in the real environment displayed through VST on a head-mounted display (HMD) device, we can provide users with new experiences and immersion.
[0004] According to one aspect of the present disclosure, a method of operating a head-mounted display device may be provided. In one embodiment, the method may capture a real-world environment to obtain an original image. In one embodiment, the method may detect at least one object included in the original image. In one embodiment, the method may obtain depth information of the at least one object detected using the original image. In one embodiment, the method may identify a target object among the at least one object detected based on the depth information. In one embodiment, the method may inpaint an area corresponding to the identified target object in the original image. In one embodiment, the method may display the inpainted image.
[0005] According to one aspect of the present disclosure, a head-mounted display device is disclosed. The head-mounted display device may include a display, a stereo camera, a memory storing at least one instruction, and at least one processor executing the at least one instruction stored in the memory. In one embodiment, the head-mounted display device may capture a real-world environment through the stereo camera and acquire an original image by having the at least one processor execute at least one instruction, either alone or in cooperation with the at least one processor. In one embodiment, the head-mounted display device may acquire depth information of at least one object detected using the original image by having the at least one processor execute at least one instruction, either alone or in cooperation with the at least one processor. In one embodiment, the head-mounted display device may identify a target object among at least one object detected based on the depth information by having the at least one processor execute at least one instruction, either alone or in cooperation with the at least one processor. In one embodiment, the head-mounted display device may inpaint an area corresponding to a target object determined in the original image by having the at least one processor execute at least one instruction, either alone or in cooperation with the at least one processor. In one embodiment, the head mounted display device can display an inpainted image through the display by having at least one processor, either alone or cooperatively, execute at least one instruction.
[0006] According to one aspect of the present disclosure, a computer-readable recording medium having recorded thereon a program for performing any one of the above-described and below-described methods for performing an operation of a head-mounted display device can be provided.
[0007] The above and other aspects and / or features according to embodiments of the present disclosure will become more apparent through the following description with reference to the attached drawings.
[0008] FIG. 1 is a drawing schematically illustrating the operation of a head mounted display device according to one embodiment of the present disclosure.
[0009] FIG. 2 is a flowchart for explaining the operation of a head mounted display device according to one embodiment of the present disclosure.
[0010] FIG. 3 is a drawing for explaining an operation of a head mounted display device according to one embodiment of the present disclosure to detect an object and identify a distance of the detected object to the head mounted display device.
[0011] FIGS. 4A, 4B, and 4C are drawings for explaining an operation of a head mounted display device according to one embodiment of the present disclosure to determine an inpainting target.
[0012] FIG. 5 is a drawing for explaining a mask map according to one embodiment of the present disclosure.
[0013] FIGS. 6A and 6B are drawings for explaining a preset distance determined based on a virtual object according to one embodiment of the present disclosure.
[0014] FIG. 7 is a drawing for explaining a first FOV of an original image and a second FOV of an outer view image according to one embodiment of the present disclosure.
[0015] FIGS. 8A, 8B, and 8C are drawings illustrating a user interface according to an embodiment of the present disclosure.
[0016] FIG. 9 is a flowchart illustrating an operation of a head mounted display device according to one embodiment of the present disclosure to perform inpainting based on whether a situation requires inpainting.
[0017] FIG. 10 is a flowchart illustrating an operation of a head mounted display device performing noise canceling according to one embodiment of the present disclosure.
[0018] FIG. 11 is a perspective view of a head mounted display device according to one embodiment of the present disclosure.
[0019] FIG. 12 is a detailed configuration diagram of a head mounted display device according to one embodiment of the present disclosure.
[0020] FIG. 13 is a detailed configuration diagram of a server according to one embodiment of the present disclosure.
[0021] The terms used in the embodiments of this specification have been selected from widely used, current terms, taking into account the functions of the present disclosure. However, these terms may vary depending on the intentions of those skilled in the art, precedents, the emergence of new technologies, etc. Furthermore, in certain cases, the applicant may arbitrarily select terms, and in such cases, their meanings will be described in detail in the description of the relevant embodiments. Therefore, the terms used in this disclosure should not be defined simply as names of terms, but rather based on the meanings of the terms and the overall content of the present disclosure.
[0022] Unless the context clearly dictates otherwise, the singular forms "a," "an," and "the" are to be understood to include plural referents. Thus, for example, the description "a constituent surface" may also include reference to one or more of such surfaces.
[0023] Terms used herein, including technical or scientific terms, have the same meaning as commonly understood by one of ordinary skill in the art described herein.
[0024] Throughout this disclosure, when a part is said to "include" a component, this does not exclude other components, but rather implies the inclusion of other components, unless otherwise specifically stated. Furthermore, terms such as "part," "module," etc., used herein refer to a unit that processes at least one function or operation, which may be implemented in hardware or software, or a combination of hardware and software.
[0025] As used herein, the expression "configured to" can be used interchangeably with, for example, "suitable for," "having the capacity to," "designed to," "adapted to," "made to," or "capable of." The term "configured to" does not necessarily mean something is "specifically designed to" in terms of hardware. Instead, in some contexts, the expression "a system configured to" can mean that the system is "capable of" in conjunction with other devices or components. For example, the phrase "a processor configured to perform A, B, and C" can mean a dedicated processor for performing the operations (e.g., an embedded processor), or a general-purpose processor (e.g., a CPU or an application processor) that can perform the operations by executing one or more software programs stored in memory.
[0026] It should be understood that the blocks and combinations of flowcharts in each flowchart can be executed by one or more computer programs containing computer-executable instructions. The one or more computer programs may be stored entirely in a single memory, or may be stored in separate portions across multiple different memories.
[0027] All functions or operations described in the present disclosure may be processed by a single processor or a combination of processors. A single processor or a combination of processors may include circuitry that performs processing, such as an Application Processor (AP), a Communication Processor (CP), a Graphical Processing Unit (GPU), a Neural Processing Unit (NPU), a Microprocessor Unit (MPU), a System on Chip (SoC), or an Integrated Chip (IC).
[0028] In the present disclosure, 'Augmented Reality (AR)' means displaying a virtual image together with a real environment (or real world), which is a space that physically exists in the real world, or displaying a virtual image together with a real object existing in the real environment.
[0029] In this disclosure, 'virtual reality (VR)' means showing an image of a virtual environment (or virtual world) created using computer graphics technology that is a space separate from the real environment.
[0030] In this disclosure, 'Mixed Reality (MR)' means providing an experience that transcends the virtual and real worlds by allowing objects existing in a real environment and objects in a virtual environment to interact with each other.
[0031] In the present disclosure, a 'head-mounted display device' may refer to an augmented reality device capable of expressing augmented reality, a virtual reality device capable of expressing virtual reality, or a mixed reality device capable of expressing mixed reality. In one embodiment, the head-mounted display device may include a shape of glasses worn on the user's face or a shape of a helmet worn on the user's head, but is not necessarily limited to the examples described above.
[0032] In the present disclosure, 'inpainting' may mean changing or restoring pixels of a predetermined area designated as an inpainting target included in an image or video into pixels whose visual features are naturally connected to surrounding areas by applying an inpainting algorithm according to an embodiment described below.
[0033] In the present disclosure, an 'artificial intelligence (AI) model' may refer to a set of functions or algorithms that are set to perform a desired characteristic (or purpose) by being learned using a plurality of learning data by a learning algorithm. Examples of learning algorithms include, but are not necessarily limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning. In one embodiment, the AI model may be stored in the memory of a head-mounted display device. However, the AI model is not limited thereto, and the head-mounted display device may transmit data input to the AI model to the server and receive data output from the AI model from the server.
[0034] In the present disclosure, an 'artificial intelligence model' may be composed of a plurality of neural network layers. Each of the plurality of neural network layers has a plurality of weight values, and can perform neural network operations through operations between the operation results of the previous layer and the plurality of weights. The plurality of weights of the plurality of neural network layers may be optimized based on the learning results of the artificial intelligence model. For example, the plurality of weights may be updated so that the loss value or cost value obtained from the artificial intelligence model during the learning process is reduced or minimized. Examples of models including a plurality of neural network layers include, but are not limited to, a deep neural network (DNN), a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), a restricted boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), and deep Q-networks.
[0035] In the present disclosure, data processing related to an image may mean data processing for each of a plurality of frames constituting an image.
[0036] Below, with reference to the attached drawings, embodiments of the present disclosure are described in detail so that those skilled in the art can easily practice the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein. In the drawings, portions irrelevant to the description have been omitted for clarity of explanation, and similar reference numerals have been used throughout the present disclosure to indicate similar elements.
[0037] The present disclosure will be described below with reference to the attached drawings.
[0038] FIG. 1 is a drawing schematically illustrating the operation of a head mounted display device according to one embodiment of the present disclosure.
[0039] Referring to FIG. 1, a head-mounted display device (1000) can capture a real environment (Real Environment) to obtain an original image (110). The Real Environment may refer to a physical space in the real world where a user (1) exists, and may include various objects. For example, the Real Environment may include inanimate objects such as buildings and roads, as well as biological objects such as people and animals.
[0040] In one embodiment, the original image (110) may be an image acquired by shooting with a preset field of view (FOV) (100). In one embodiment, the original image (110) may include an object observable through the preset FOV (100) in a real environment. For example, the original image (110) may include, but is not limited to, a first person (10), a second person (20), a third person (30), a fourth person (40), and a pigeon (50) observed through the preset FOV (100) in a real environment.
[0041] In one embodiment, the head-mounted display device (1000) may include various types of devices that display the original image (110). For example, the head-mounted display device (1000) may include, but is not limited to, a mixed reality device that displays an image acquired in real time through a camera through a display, or a virtual reality device that displays a pre-stored image through a display.
[0042] In one embodiment, the head mounted display device (1000) can detect at least one object included in the original image (110). In one embodiment, the head mounted display device (1000) can obtain depth information of the at least one object detected using the original image. In one embodiment, the head mounted display device (1000) can determine (e.g., identify) a target object among the at least one object detected based on the depth information. Here, the target object may include an object to be subjected to inpainting, which will be described later, among the at least one object detected. In one embodiment, the head mounted display device (1000) can identify at least one object located within a preset distance range from the head mounted display device (1000) among the at least one object detected based on the depth information, and determine the target object among the at least one object identified.
[0043] For example, a user (1) may want to view a third person (30) and a fourth person (40) performing busking in a real environment through a head mounted display device (1000). In this case, the first person (10), the second person (20), and the pigeon (50) are obscuring the third person (30) and the fourth person (40), and therefore, they are elements that interfere with the viewing from the user's (1) perspective. Accordingly, the head mounted display device (1000) may detect the first person (10), the second person (20), the third person (30), the fourth person (40), and the pigeon (50) included in the original image (110), and, based on depth information of the detected objects, determine the first person (10), the second person (20), and the pigeon (50) located between the head mounted display device (1000) and the third person (30) and the fourth person (40) as target objects.
[0044] In one embodiment, the head mounted display device (1000) can inpaint an area corresponding to a target object in an original image (110). In one embodiment, the head mounted display device (1000) can display an inpainted image (120).
[0045] For example, the head mounted display device (1000) can obtain an inpainted image (120) by inpainting areas corresponding to a first person (10), a second person (20), and a pigeon (50) included in an original image (110), thereby reconstructing (or restoring) pixels representing areas of the first person (10), the second person (20), and the pigeon (50) into pixels representing areas of a real environment covered by the first person (10), the second person (20), and the pigeon (60). The inpainted image (120) can further include parts of a third person (30) and a fourth person (40) covered by the first person (10), the second person (20), and the pigeon (50) in the original image (110). Accordingly, the user (1) can immerse himself in the enjoyment of the busking performances of the third person (30) and the fourth person (40) through the inpainted video (120).
[0046] In this way, according to one embodiment of the present disclosure, by determining a target object based on depth information of at least one object included in an original image (110), inpainting can be performed in a real environment while taking into account the physical distance between the head mounted display device (1000) and the object. That is, since inpainting is performed while taking into account spatial information of the real environment in which the user (1) exists, the user (1) can experience immersion in a real environment independent of unnecessary elements.
[0047] FIG. 2 is a flowchart for explaining the operation of a head mounted display device according to one embodiment of the present disclosure.
[0048] Referring to FIG. 2, the operation of the head mounted display device (1000) will be briefly described, and a detailed description of each operation will be described with reference to the drawings that follow. The operations of the head mounted display device (1000) described in the present disclosure can be understood as the operations of the processor (1800) of the head mounted display device (1000) illustrated in FIG. 12 and the processor (2300) of the server (2000) illustrated in FIG. 13.
[0049] At step S210, the head mounted display device (1000) can capture a real environment to obtain an original image.
[0050] In one embodiment, the head mounted display device (1000) can acquire an original image based on an image (e.g., a stereo image) acquired through a stereo camera included in the head mounted display device (1000) or a pre-stored panoramic image.
[0051] In one embodiment, the head mounted display device (1000) may include a stereo camera. In one embodiment, the stereo camera may include a left camera and a right camera. The left camera and the right camera are positioned at a certain distance from the head mounted display device (1000) and capture a real environment in which a user wearing the head mounted display device (1000) is located from different angles, thereby obtaining a left-eye image and a right-eye image. In one embodiment, the head mounted display device may obtain the left-eye image and the right-eye image obtained through the left-eye camera and the right-eye camera as original images. In one embodiment, the obtained left-eye image and the right-eye image may be displayed on a display (or a first area of the display) of the head mounted display device corresponding to the user's left eye and on a display (or a second area of the display) of the head mounted display device corresponding to the user's right eye, respectively.
[0052] In one embodiment, the head mounted display device (1000) can acquire a pre-stored panoramic image. The pre-stored panoramic image may include an image that was captured by capturing a real-world environment before the head mounted display device (1000) is used. In one embodiment, the pre-stored panoramic image may include an image captured with a field of view (FOV) wider than the field of view (FOV) of the image displayed through the head mounted display device (1000). For example, the pre-stored panoramic image may include an image captured by a 360-degree camera capable of capturing the entire real-world environment simultaneously or a panoramic camera capable of capturing a field of view wider than the field of view of the original image. As another example, the pre-stored panoramic image may include a panoramic image generated based on images captured from various angles by changing the shooting angles of the stereo cameras of the head mounted display device (1000). In one embodiment, the head mounted display device (1000) can identify a point at which a user is looking or a point at which the user is looking in a three-dimensional space. In one embodiment, the head mounted display device (1000) can extract an area of a panoramic image corresponding to a point at which the user is gazing or looking, and obtain an image of the extracted area as an original image.
[0053] In step S220, the head-mounted display device (1000) can detect at least one object included in the original image. In one embodiment, the head-mounted display device (1000) can identify the class and location of each of at least one object included in the original image based on the acquired original image. Here, the class may include a category or label indicating the type of object to be identified in the image or video. In addition, the location of the object may include the location of an area corresponding to the object in the original image.
[0054] In one embodiment, the head-mounted display device (1000) can detect at least one object included in an original image by applying the original image to an object detection model. The object detection model may include an artificial intelligence model that takes an image as input and identifies the class and location of an object included in the image.
[0055] In one embodiment, the object detection model may output, as an object detection result, location information of a bounding box surrounding an object detected from an image based on an input image and class information of an object located within the bounding box. In one embodiment, the object detection model may be trained based on training images including objects of various classes and training metadata corresponding to the training images. In one embodiment, the training metadata may include location information of a bounding box surrounding an object included in the training image and class information of the object.
[0056] In one embodiment, the head-mounted display device (1000) can detect at least one object included in an original image by applying the original image to a segmentation model. The segmentation model may include an artificial intelligence model that takes an image as input and assigns a plurality of pixels included in the image to one of a plurality of preset classes.
[0057] In one embodiment, the segmentation model may include a semantic segmentation model and an instance segmentation model. The semantic segmentation model may output a segmentation map, in which a plurality of pixels of an input image are assigned unique values that distinguish them by a plurality of preset classes, as an object detection result. The instance segmentation model may output a segmentation map, in which a plurality of pixels of an input image are assigned unique values that distinguish them by a plurality of preset classes and by different objects of the same class, as an object detection result. In one embodiment, the head mounted display device (1000) may detect at least one object included in an original image by classifying pixels assigned the same value in the segmentation map.
[0058] In one embodiment, a segmentation model may be trained based on training images containing objects of various classes and training segmentation maps corresponding to the training images. In one embodiment, a training segmentation map input to a semantic segmentation model may have a plurality of pixels assigned unique values that distinguish between a plurality of predetermined classes. In one embodiment, a training segmentation map input to an instance segmentation model may have a plurality of pixels assigned unique values that distinguish between a plurality of classes and between different objects of the same class.
[0059] In one embodiment, the head-mounted display device (1000) can track at least one detected object. Here, tracking the object may mean continuously detecting a specific object in a plurality of frames and identifying changes in the position of the detected object. The head-mounted display device (1000) can track the object by assigning a unique ID to an object detected in each of a plurality of frames included in an original image and identifying changes in the position of an object assigned the same ID.
[0060] In one embodiment, the head-mounted display device (1000) can track at least one detected object by applying an object detection result obtained from an object detection model to an object tracking model. In one embodiment, the object tracking model can include a rule-based algorithm model or an artificial intelligence model that assigns a unique ID to each detected object based on the object detection result and identifies a change in the position of the object. In one embodiment, the object tracking model can include the aforementioned object detection model or a sub-model that can obtain the object detection result. In this case, the object tracking model can detect an object by inputting images of a plurality of frames and simultaneously track the detected object.
[0061] In one embodiment, the object tracking model may output position change information and identification information of the tracked object as a tracking result based on the object detection result or the images of a plurality of frames. In one embodiment, the object tracking model may be trained based on training images including objects of various classes and training metadata corresponding to the training images. In one embodiment, the training metadata input to the object tracking model may include position information of a bounding box surrounding an object included in a plurality of frames of the training images, the class of the object, and a unique ID assigned to each object.
[0062] In one embodiment, the head-mounted display device (1000) may acquire an outer view image representing a second FOV wider than a first FOV of the original image. Since the outer view image is an image representing a second FOV wider than the first FOV of the original image, objects not included in the original image may be included in the outer view image.
[0063] In one embodiment, the head mounted display device (1000) can obtain an original image captured with a first FOV through a stereo camera included in the head mounted display device (1000), and can obtain an outer view image captured with a second FOV through a sub camera included in the head mounted display device (1000). In one embodiment, the sub camera may include a plurality of cameras for capturing a user's hand or a real environment (e.g., the user's hand) outside the first FOV. In this case, the head mounted display device (1000) can obtain an outer view image based on images captured from the plurality of cameras. As another example, the sub camera may include a wide-angle camera capable of capturing with a second FOV that is wider than the first FOV.
[0064] In one embodiment, the head mounted display device (1000) can obtain an original image captured with a second FOV by obtaining an original image captured with a second FOV from a pre-stored panoramic image and extracting a portion of the panoramic image representing a first FOV from the pre-stored panoramic image.
[0065] In one embodiment, at least one object moving between the inside and the outside of the first FOV can be tracked based on the original image and the outer view image. In one embodiment, the head mounted display device (1000) can detect at least one object from the outer view image and track the detected at least one object. Since the operation of the head mounted display device (1000) detecting and tracking an object from the outer view image corresponds to the operation of detecting and tracking an object from the original image described above, a redundant description will be omitted.
[0066] In one embodiment, the head mounted display device (1000) can track at least one object moving between the inside and the outside of the first FOV based on a result of tracking at least one object acquired from an original image and a result of tracking at least one object acquired from an outer view image. In one embodiment, the head mounted display device (1000) can track at least one object moving between the inside and the outside of the first FOV, thereby including at least one of information about whether at least one object detected outside the first FOV is an object previously detected from the original image and information about a moving direction of the object.
[0067] For example, the head-mounted display device (1000) can assign the same unique ID to an object tracked from an original image and an object tracked from an outer view image. Accordingly, even if an object located inside the first FOV moves to an outside of the first FOV, the existing tracking can be maintained based on the tracking result of the object acquired from the outer view image. When an object is detected outside the first FOV, it can be identified whether the object is an object detected from the original image based on the tracking result of the object acquired from the original image.
[0068] In one embodiment, the head mounted display device (1000) may display a user interface indicating identification information of at least one object located outside the first FOV. In one embodiment, the identification information may include a tracking result or a detection result of at least one object moving between the inside and the outside of the first FOV. For example, if an object located inside the first FOV moves to the left of the first FOV and is detected outside the first FOV of the outer view image, the head mounted display device (1000) may display an indicator indicating that the object has moved to the left in the original image. As another example, if a new object that has not been detected in the original image is detected outside the first FOV of the outer view image, an indicator indicating that a new object has been detected may be displayed in the original image. However, the present invention is not necessarily limited thereto, and the identification information may also include information on the class and location of at least one object located outside the first FOV.
[0069] In step S230, the head mounted display device (1000) can obtain depth information of at least one object detected using the original image.
[0070] In one embodiment, the depth information may include information representing the depth of an object or background in a three-dimensional space as a two-dimensional image. For example, the depth information may include a depth map corresponding to multiple frames of the original image. However, the depth information is not necessarily limited thereto, and may also include the distance of each of at least one detected object to the head-mounted display device (1000).
[0071] In one embodiment, the head-mounted display device (1000) can obtain depth information of at least one detected object by applying an original image to a depth estimation model. In one embodiment, the depth estimation model may include an artificial intelligence model that inputs an image and outputs a depth map corresponding to the input image. In one embodiment, the depth estimation model may be trained based on training images and ground truth depth maps corresponding to the training images.
[0072] In one embodiment, the head mounted display device (1000) can obtain depth information of at least one object detected based on a stereo image. In one embodiment, the head mounted display device (1000) can obtain a stereo image of an original image and calculate a disparity between the obtained stereo images. In one embodiment, the head mounted display device (1000) can generate a depth map of the original image by calculating depth values of a plurality of pixels using the calculated disparity.
[0073] In one embodiment, the head-mounted display device (1000) can acquire depth information of at least one object based on data acquired based on a distance detection sensor. In one embodiment, the distance detection sensor can measure the distance between objects in a real environment where a distance source image is captured and the distance detection sensor. In one embodiment, the head-mounted display device (1000) can generate a depth map of the original image by mapping data acquired from the distance detection sensor to a plurality of pixels included in the original image.
[0074] In step S240, the head mounted display device (1000) can identify a target object among at least one object detected based on depth information.
[0075] In one embodiment, the head mounted display device (1000) can identify at least one object located within a preset distance range from the head mounted display device among at least one object detected based on depth information. In one embodiment, the head mounted display device (1000) can calculate an average of depth values of a plurality of pixels corresponding to at least one detected object included in a depth map, and identify at least one object whose distance corresponding to the calculated average of depth values is within the preset distance range. However, the present invention is not necessarily limited to the above-described example, and the head mounted display device (1000) can identify at least one object located within a preset distance range based on whether the distance corresponding to a maximum value or a minimum value of a plurality of pixels corresponding to at least one detected object included in the depth map is within the preset distance range.
[0076] In one embodiment, the head-mounted display device (1000) may determine a preset distance range based on a user input. For example, the head-mounted display device (1000) may obtain a user input that determines a first distance or less as the preset distance range, a second distance or more as the preset distance range, or a distance from the first distance to the second distance as the preset distance range, but is not necessarily limited thereto.
[0077] In one embodiment, the head mounted display device (1000) can display a virtual object. The virtual object is an object existing in a virtual environment of a three-dimensional space and can be positioned in the three-dimensional space. In one embodiment, the head mounted display device (1000) can generate a three-dimensional space including at least one object detected from the original image based on depth information of at least one object detected from the original image. In one embodiment, the head mounted display device (1000) can map the virtual object onto the generated three-dimensional space. In one embodiment, the head mounted display device (1000) can display the virtual object on the original image or the inpainted image based on the location where the virtual object is mapped in the three-dimensional space.
[0078] In one embodiment, the head-mounted display device (1000) can determine a preset distance range based on a virtual object. Details of how the head-mounted display device (1000) determines the preset distance range based on the virtual object will be described further below with reference to FIGS. 6A and 6B .
[0079] In one embodiment, the head mounted display device (1000) can determine a target object from among at least one object identified as being located within a preset distance range.
[0080] In one embodiment, the head mounted display device (1000) may determine an object classified into a preset class among at least one object identified as being located within a preset distance range as a target object. In one embodiment, the head mounted display device (1000) may determine an object selected as an inpainting target among at least one object identified as being located within a preset distance range as a target object. Specific details regarding the head mounted display device (1000) determining a target object among at least one object identified as being located within a preset distance range will be described again with reference to FIGS. 4A, 4B, and 4C below.
[0081] In step S250, the head mounted display device (1000) can inpaint an area corresponding to the determined target object in the original image. In one embodiment, the head mounted display device (1000) can obtain an inpainted image including an area corresponding to the target object in the original image, which is a reconstructed (or restored) area, by inpainting the target object based on the original image. Here, the reconstructed area may mean an area corresponding to the target object in the original image that is inpainted as an area that is estimated to be observed when the target object does not exist in a real environment.
[0082] In one embodiment, the head mounted display device (1000) can obtain a mask map indicating an area corresponding to a target object in an original image. In one embodiment, the head mounted display device (1000) can obtain an inpainted image by applying the original image and the mask map to an inpainting model for inpainting the target object. In one embodiment, the inpainting model can include an artificial intelligence model that inpaints an area indicated by the mask map in an image based on an image and a mask map corresponding to the image. In one embodiment, the inpainting model can output an image including an inpainted area in which an area indicated by the mask map is indicated based on the image and the mask map corresponding to the image. Specific details regarding the mask map indicating an area corresponding to the target object will be described again below with reference to FIG. 5.
[0083] In one embodiment, the inpainting model can perform inpainting based on spatial features contained in a single image so that the area indicated by the mask map matches the context within the image. The inpainting model can transform pixels in an area corresponding to a target object in the original image to have similar values (e.g., color, texture, etc.) to surrounding pixels so that the inpainted image does not include unnatural boundaries or patches. In one embodiment, the inpainting model can perform inpainting based on temporal features acquired from consecutive frames so that the inpainting matches motion that occurs between images of adjacent frames. The inpainting model can transform pixels in an area corresponding to a target object in the original image to have values that match the motion that occurs between adjacent frames so that the motion does not appear unnatural in the inpainted image. In one embodiment, the inpainting model can be trained based on a training image, a mask map corresponding to the training image, and a ground truth image from which the area indicated by the mask map is removed from the training image.
[0084] In one embodiment, the head-mounted display device (1000) identifies (or determines) whether a situation requires inpainting of a target object, and if identified as a situation requiring inpainting, performs inpainting of the determined target object. Details of how the head-mounted display device (1000) performs inpainting based on a situation requiring inpainting will be further described below with reference to FIG. 9.
[0085] In one embodiment, the head mounted display device (1000) may acquire a peripheral audio signal in a section where a target object determined from an original image is detected, and may acquire an audio signal in which an audio signal corresponding to the target object is reduced from the acquired peripheral audio signal. In one embodiment, specific details regarding the head mounted display device (1000) acquiring an audio signal in which an audio signal corresponding to the target object is reduced will be described again with reference to FIG. 10 below.
[0086] In step S260, the head mounted display device (1000) may display an inpainted image. In one embodiment, the head mounted display device (1000) may display an original image before the inpainting mode is activated, and may display an inpainted image instead of the original image after the inpainting mode is activated. In one embodiment, the head mounted display device (1000) may obtain a user input for controlling the activation of the inpainting mode, and may determine the activation of the inpainting mode based on the obtained user input. However, the present invention is not necessarily limited thereto, and the head mounted display device (1000) may display the original image or the inpainted image based on whether a situation requires inpainting for the aforementioned target object.
[0087] FIG. 3 is a drawing for explaining an operation of a head mounted display device according to one embodiment of the present disclosure to detect an object and identify a distance of the detected object to the head mounted display device.
[0088] In one embodiment, a head-mounted display device (1000) may acquire an original image (110). The original image (110) is an image acquired by photographing a real environment and may include various objects existing in the real environment. In one embodiment, the original image (110) may include various objects existing in the real environment. For example, the original image (110) may include a first person (10), a second person (20), a third person (30), a fourth person (40), and a pigeon (50) existing in the photographed real environment.
[0089] In one embodiment, the head mounted display device (1000) can detect at least one object included in the original image (110). For example, by detecting at least one object included in the original image (110), the head mounted display device (1000) can identify the class of the first person (10), the second person (20), the third person (30), and the fourth person (40) as being 'person' and the class of the pigeon (50) as being 'bird'. The head mounted display device (1000) can identify areas corresponding to a first person (10), a second person (20), a third person (30), a fourth person (40), and a pigeon (50) as a first bounding box (310), a second bounding box (320), a third bounding box (330), a fourth bounding box (340), and a fifth bounding box (350) by detecting at least one object included in the original image (110). However, the present invention is not limited thereto, and the areas corresponding to the first person (10), the second person (20), the third person (30), the fourth person (40), and the pigeon (50) may be identified on a pixel basis based on the boundary of each object.
[0090] In one embodiment, the head mounted display device (1000) can track at least one object detected in the original image (110). In one embodiment, the head mounted display device (1000) can track at least one detected object and identify a change in position of the at least one detected object by applying the original image (110) to an object tracking model. For example, the head mounted display device (1000) can assign unique IDs to a first person (10), a second person (20), a third person (30), a fourth person (40), and a pigeon (50) by tracking at least one object detected in the original image (110), and can identify a change in position of an object assigned the same ID between consecutive frames of the original image (110).
[0091] In one embodiment, the head mounted display device (1000) can obtain depth information (360) of at least one object included in the original image (110). In one embodiment, the head mounted display device (1000) can identify a distance of at least one object detected from the original image (110) to the head mounted display device (1000) based on the obtained depth information (360). For example, the depth information (360) can include a depth map of the original image (110). The head mounted display device (1000) can identify that the distances to the head mounted display devices of the first person (10), the second person (20), the third person (30), the fourth person (40), and the pigeon (50) are '6 m', '8 m', '16 m', '17 m', and '11 m', respectively, based on the depth values of the areas corresponding to the first person (10), the second person (20), the third person (30), the fourth person (40), and the pigeon (50) on the depth map.
[0092] FIGS. 4A, 4B, and 4C are drawings for explaining an operation of a head mounted display device according to one embodiment of the present disclosure to determine an inpainting target.
[0093] In one embodiment, the head mounted display device (1000) may obtain object recognition information (400). The object recognition information (400) may include a detection result and a tracking result of at least one object included in the original image (110). For example, the object recognition information (400) may include information about a location, a class, and a unique ID of at least one object detected from the original image (110). In addition, the object recognition information (400) may include a distance of at least one object detected from the original image (110) to the head mounted display device (1000). In one embodiment, the head mounted display device (1000) may sequentially obtain a detection result and a tracking result of at least one object included in each of a plurality of frames of the original image (110) and update the detection result and the tracking result with the object recognition information (400).
[0094] In one embodiment, the head mounted display device (1000) can determine a target object based on object recognition information (400).
[0095] Referring to FIG. 4A, the head mounted display device (1000) can determine a preset distance range for determining a target object. For example, the head mounted display device (1000) can determine the preset distance range as '7 m or less', '10 m or less', or '13 m or less'.
[0096] In one embodiment, the head mounted display device (1000) can identify an object located within a preset distance range from the head mounted display device (1000) among at least one object detected from the original image (110), and determine the identified object as a target object.
[0097] For example, if the preset distance range is '7 m or less', the head mounted display device (1000) can determine the first person (10) whose distance to the head mounted display device (1000) is '6 m' as the target object based on the object recognition information (400). Then, the head mounted display device (1000) can obtain an inpainted image (410-1) by inpainting the first person (10) determined as the target object in the original image (110). The inpainted image (410-1) can include an area reconstructed by inpainting an area corresponding to the first person (10) in the original image (110). That is, the inpainted image (410-1) can include a background such as a street or building that is obscured by the first person (10) in the real environment where the original image (110) was captured, and a part of the third person (30).
[0098] As another example, if the preset distance range is '10 m or less', the head mounted display device (1000) may determine a first person (10) whose distance is '6 m' and a second person (20) whose distance is '8 m' as target objects based on the object recognition information (400). Then, the head mounted display device (1000) may obtain an inpainted image (420-1) by inpainting the first person (10) and the second person (20) determined as target objects in the original image (110). The inpainted image (420-1) may include an area reconstructed by inpainting an area corresponding to the first person (10) and an area corresponding to the second person (20) in the original image (110). That is, the inpainted image (420-1) may include a background such as a street, a building, etc. covered by the first person (10) and the second person (20) in the real environment where the original image (110) was filmed, as well as a part of the third person (30) and a part of the fourth person (40).
[0099] As another example, if the preset distance range is '13 m or less', the head mounted display device (1000) may determine a first person (10) whose distance is '6 m', a second person (20) whose distance is '8 m', and a pigeon (50) whose distance is '11 m' as target objects based on the object recognition information (400). Then, the head mounted display device (1000) may obtain an inpainted image (430-1) by inpainting the first person (10), the second person (20), and the pigeon (50) determined as target objects in the original image (110). The inpainted image (430-1) may include an area reconstructed by inpainting an area corresponding to the first person (10), an area corresponding to the second person (20), and an area corresponding to the pigeon (50) in the original image (110). That is, the inpainted image (430-1) may include backgrounds such as streets, buildings, etc. covered by the first person (10), the second person (20), and the pigeon (50) in the real environment where the original image (110) was filmed, as well as parts of the musical instruments played by the third person (30), the fourth person (40), and the fourth person (40).
[0100] Referring to FIG. 4b, the head mounted display device (1000) can identify an object located within a preset distance range from the head mounted display device (1000) among at least one object detected from the original image (110), and determine an object classified into a preset class among the identified at least one object as a target object.
[0101] In one embodiment, the head mounted display device (1000) can determine a preset class based on a user input. In one embodiment, the head mounted display device (1000) can obtain a user input for selecting a preset class. For example, the head mounted display device (1000) can display a user interface indicating a plurality of classes that can be detected by an object detection model, and obtain a user input for selecting at least one of the displayed plurality of classes as an inpainting class. As another example, the head mounted display device (1000) can display a user interface indicating a plurality of classes of a plurality of objects detected in at least one of an original image and an outer view image. The head mounted display device (1000) can obtain a user input for selecting at least one of the plurality of classes displayed through the user interface, and determine the selected at least one class as a preset class. However, the present invention is not necessarily limited thereto, and the preset class may include a preset class regardless of the user input.
[0102] In the following FIG. 4b, when the preset distance range is '13 m or less' and the first person (10), the second person (20), and the pigeon (50) are identified as objects located within the preset distance range from the head mounted display device (1000), the head mounted display device (1000) determines the target object based on the preset class.
[0103] For example, if the preset class is 'person', the head mounted display device (1000) can determine the first person (10) and the second person (20) whose class is 'person' among the first person (10), the second person (20), and the pigeon (50) as target objects based on the object recognition information (400). Then, the head mounted display device (1000) can obtain an inpainted image (410-2) by inpainting the first person (10) and the second person (20) determined as target objects in the original image (110).
[0104] As another example, if the preset class is 'bird', the head mounted display device (1000) can determine a pigeon (50) whose class is 'bird' among the first person (10), the second person (20), and the pigeon (50) as a target object based on the object recognition information (400). Then, the head mounted display device (1000) can obtain an inpainted image (410-2) by inpainting the pigeon (50) determined as the target object in the original image (110).
[0105] As another example, if the preset class is 'bird', the head mounted display device (1000) can determine a pigeon (50) whose class is 'bird' among the first person (10), the second person (20), and the pigeon (50) as a target object based on the object recognition information (400). Then, the head mounted display device (1000) can obtain an inpainted image (420-2) by inpainting the pigeon (50) determined as the target object in the original image (110).
[0106] As another example, if the preset classes are 'person' and 'bird', the head mounted display device (1000) can determine the first person (10) and the second person (20) whose class is 'person' and the pigeon (50) whose class is 'bird' as target objects among the first person (10), the second person (20), and the pigeon (50) whose class is 'bird', based on the object recognition information (400). Then, the head mounted display device (1000) can obtain an inpainted image (430-2) by inpainting the first person (10), the second person (20), and the pigeon (50) determined as target objects in the original image (110).
[0107] Referring to FIG. 4C, the head mounted display device (1000) may identify an object located within a preset distance range from the head mounted display device (1000) among at least one object detected from the original image (110), and determine an object selected as an inpainting target among the at least one identified object as the target object. In one embodiment, the head mounted display device (1000) may identify an ID of an object selected as an inpainting target based on object recognition information (400), and determine an object assigned the identified ID among at least one object detected from the original image (110) as the target object.
[0108] In one embodiment, the head mounted display device (1000) can select an inpainting target based on a user input. The head mounted display device (1000) can obtain a user input for selecting at least one of at least one detected object as an inpainting target. For example, the head mounted display device (1000) can display a user interface representing a plurality of objects detected in at least one of an original image and an outer view image, and obtain a user input for selecting at least one of the displayed plurality of objects as an inpainting target. In one embodiment, the head mounted display device (1000) can select an object located outside a first FOV and detected only in the outer view image as an inpainting target. When the selected target object moves to within the first FOV, the head mounted display device (1000) can identify an object selected as an inpainting target among at least one object included in the original image based on a tracking result of the object acquired from the outer view image and a tracking result of the object acquired from the original image.
[0109] Hereinafter, in FIG. 4c, when the preset distance range is '13 m or less' and the first person (10), the second person (20), and the pigeon (50) are identified as objects located within the preset distance range from the head mounted display device (1000), the head mounted display device (1000) will be described as determining the target object based on the inpainting target.
[0110] For example, if the inpainting target is selected as a first person (10), the head mounted display device (1000) can determine the first person (10) with an ID of 1 among the first person (10), the second person (20), and the pigeon (50) as the target object based on the object recognition information (400). Then, the head mounted display device (1000) can obtain an inpainted image (410-3) by inpainting the first person (10) determined as the target object in the original image (110).
[0111] As another example, if the inpainting target is selected as a second person (20), the head mounted display device (1000) can determine the second person (20) with an ID of 2 among the first person (10), the second person (20), and the pigeon (50) as the target object based on the object recognition information (400). Then, the head mounted display device (1000) can obtain an inpainted image (420-3) by inpainting the second person (20) determined as the target object in the original image (110).
[0112] As another example, if the inpainting target is selected as a pigeon (50), the head mounted display device (1000) can determine the pigeon (50) with ID 5 among the first person (10), the second person (20), and the pigeon (50) as the target object based on the object recognition information (400). Then, the head mounted display device (1000) can obtain an inpainted image (430-3) by inpainting the pigeon (50) as the target object in the original image (110).
[0113] As another example, if the inpainting target is selected as a third person (30) or a fourth person (40), the head mounted display device (1000) may not determine the third person (30) or the fourth person (40) as a target object because the third person (30) or the fourth person (40) is not located at a distance of 13 m or less from the head mounted display device (1000). In one embodiment, if the object recognition information (400) is updated to indicate that the third person (30) or the fourth person (40) selected as the inpainting target is located at a distance of 13 m or less from the head mounted display device (1000), the head mounted display device (1000) may determine the third person (30) or the fourth person (40) selected as the inpainting target as a target object.
[0114] FIG. 5 is a drawing for explaining a mask map according to one embodiment of the present disclosure.
[0115] In one embodiment, the head mounted display device (1000) may obtain a mask map (510) representing an area corresponding to a target object in an original image (110). The mask map (510) representing an area corresponding to the target object may be expressed as binary data in which a pixel value of an area corresponding to the target object is 1 and a pixel value of a remaining area excluding the area corresponding to the target object is 0. However, the present invention is not limited thereto, and the values of the binary data may be opposite to each other, and may be expressed in a data format other than binary data.
[0116] For example, the head mounted display device (1000) can identify a first bounding box (310), a second bounding box (320), a third bounding box (330), a fourth bounding box (340), and a fifth bounding box (350) surrounding a first person (10), a second person (20), a third person (30), a fourth person (40), and a pigeon (50), respectively, by detecting at least one object included in the original image (110). In this case, if the first person (10), the second person (20), and the pigeon (50) are target objects, the head mounted display device (1000) can obtain a first mask map (510-1) representing the first bounding box (310), the second bounding box (320), and the fifth bounding box (330).
[0117] As another example, the head mounted display device (1000) can identify a plurality of pixels included within the boundaries of each of the first person (10), the second person (20), the third person (30), the fourth person (40), and the pigeon (50) by detecting at least one object included in the original image (110). In this case, if the first person (10), the second person (20), and the pigeon (50) are target objects, the head mounted display device (1000) can obtain a second mask map (510-2) representing the first person (10), the second person (20), and the pigeon (50).
[0118] In one embodiment, the head mounted display device (1000) can obtain an inpainted image (120) by applying an original image (110) and a mask map (510) indicating an area corresponding to a target object to an inpainting model (520) for inpainting the target object. For example, the head mounted display device (1000) can input the original image (110) and a first mask map (510-1) or a second mask map (510-2) to the inpainting model (520) and obtain an inpainted image (120) from the inpainting model (520). The inpainted image (120) can include an area in which an area indicated by the first mask map (510-1) or an area indicated by the second mask map (510-2) in the original image (110) is reconstructed.
[0119] In one embodiment, the inpainting model (520) may include an encoder (521) and a plurality of neural network layers. In one embodiment, the encoder may output a feature map of a plurality of frames based on an input image. The feature map of the plurality of frames may include various features, such as color, texture, shape, etc., of the plurality of frames. In one embodiment, the plurality of neural network layers may include various layers, such as a convolutional neural network layer, an attention layer that performs an attention mechanism, etc. The plurality of neural network layers may extract spatial features, temporal features, and context information of the plurality of frames based on the feature map output from the encoder, thereby outputting a feature map of the plurality of frames in which an area indicated by a mask map is reconstructed. In one embodiment, the decoder may output an inpainted image (120) based on the feature map of the plurality of frames in which an area indicated by a mask map is reconstructed.
[0120] FIGS. 6A and 6B are drawings for explaining a preset distance determined based on a virtual object according to one embodiment of the present disclosure.
[0121] In one embodiment, the head mounted display device (1000) can display virtual objects (620-1, 620-2). For example, the head mounted display device (1000) can obtain an original image including a first person (641) and a second person (642) existing in a real environment. The head mounted display device (1000) can generate a three-dimensional space including the first person (641) and the second person (642) based on depth information of the first person (641) and the second person (642). The head mounted display device (1000) can map the virtual objects (620-1, 620-2) on the generated three-dimensional space, and display the virtual objects (620-1, 620-2) on the original image or the inpainted image based on the locations to which the virtual objects (620-1, 620-2) are mapped.
[0122] In one embodiment, the head mounted display device (1000) can determine a preset distance range based on a virtual object (620-1, 620-2).
[0123] In one embodiment, the head mounted display device (1000) can determine a preset distance range based on the distance between the virtual object (620-1, 620-2) and the head mounted display device (1000). For example, the head mounted display device (1000) can calculate a distance value between the virtual object (620-1, 620-2) and the head mounted display device (1000) based on the coordinates of the head mounted display device (1000) and the coordinates of the virtual object (620-1, 620-2) in the generated three-dimensional space. Then, the distance value of the preset distance range can be determined based on the calculated distance value.
[0124] In one embodiment, the head mounted display device (1000) may determine a preset distance range based on the type of the virtual object (620-1, 620-2). For example, the types of the virtual objects (620-1, 620-2) may include a first type (or, long-distance display type) that is displayed at a far distance from the user, a second type (or, medium-distance display type) that is displayed at a certain distance from the user, and a third type (or, short-distance display type) that is displayed at a close distance from the user. If the virtual object is the first type, the head mounted display device (1000) may determine the distance from a position that is the first distance from the head mounted display device (1000) to the virtual object as the preset distance range. If the virtual object is the second type, the distance from the head mounted display device (1000) to the virtual object may be determined as the preset distance range. If the virtual object is the third type, the preset distance range may be determined to be a distance further than the position of the virtual object. However, the present invention is not necessarily limited to the examples described above, and the type of the virtual object (620-1, 620-2) may be determined based on the content provided by the virtual object (620-1, 620-2) or based on user input. In addition, the method for determining the preset distance range may vary depending on the type of the virtual object (620-1, 620-2).
[0125] Referring to FIG. 6A, the virtual object (620-1) may be a type that determines the distance from the head mounted display device (1000) to the virtual object (620-1) as a preset distance range (630-1). For example, the virtual object (620-1) may be a screen on which a movie is displayed. In this case, the head mounted display device (1000) may calculate the distance value between the head mounted display device (1000) and the virtual object (620-1) as '13 m or less', and determine the preset distance range (630-1) as '13 m or less'. In addition, the head mounted display device (1000) may determine a first person (641) located within a range of '13 m or less' from the head mounted display device (1000) as a target object, and may not determine a second person (642) located outside the '13 m or less' range as a target object. A first person (641) that can obscure a virtual object (620-1) may be determined as a target object because the first person (641) may be an obstacle to the user (610) viewing the virtual object (620-1). A second person (642) located behind a virtual object (620-2) may not be determined as a target object because the second person (642) does not obstruct the user (610) viewing the virtual object (620-1).
[0126] Referring to FIG. 6B, the virtual object (620-2) may be a type that determines a distance greater than the virtual object (620-1) as a preset distance range (630-2). For example, the virtual object (620-2) may be a screen on which a document being worked on is displayed. In this case, the head mounted display device (1000) may calculate the distance value between the head mounted display device (1000) and the virtual object (620-1) as '0.5 m' and determine the preset distance range (630-2) as 'more than 0.5 m'. In addition, the head mounted display device (1000) may determine the first person (641) and the second person (642) located within a range of 'more than 0.5 m' from the head mounted display device (1000) as target objects. Objects in the real environment (e.g., a cup, a table, a laptop, etc.) located between the user (610) and the virtual object (620-2) may not be determined as target objects because they do not interfere with the user (610) performing tasks related to the virtual object (620-2). The first person (641) and the second person (642) located behind the virtual object (620-2) may be determined as target objects because they may interfere with the user (610) performing tasks related to the virtual object (620-2).
[0127] FIG. 7 is a drawing for explaining a first FOV of an original image and a second FOV of an outer view image according to one embodiment of the present disclosure.
[0128] In one embodiment, the head mounted display device (1000) can obtain an original image (110) representing a first FOV (710). In one embodiment, the head mounted display device (1000) can include a stereo camera (1210). In one embodiment, the stereo camera (1210) can obtain an image including an object located within the first FOV (710) by photographing a real environment with the first FOV (710). In one embodiment, the head mounted display device (1000) can obtain the original image (110) based on the image obtained through the stereo camera (1210).
[0129] In one embodiment, the head mounted display device (1000) can obtain an outer view image (115) representing a second FOV (720). In one embodiment, the outer view image may be obtained by the head mounted display device (1000) including a sub-camera (1220). In one embodiment, the sub-camera (1220) can obtain an image including an object located inside the second FOV (720) by capturing a real environment with a second FOV (720) that is wider than the first FOV (710). In one embodiment, the head mounted display device (1000) can obtain the outer view image (115) based on an image obtained through the sub-camera (1220).
[0130] In one embodiment, the head mounted display device (1000) can track at least one object moving between the outside and inside of the first FOV (710) based on the original image (110) and the outer view image (115).
[0131] For example, the head mounted display device (1000) can obtain information that the person (730) is a new object that was not previously detected in the original image (110) and is moving its position inside the first FOV (710) based on the tracking result of the person (730) obtained from the original image (110) and the tracking result of the person (730) obtained from the outer view image (115). As another example, the head mounted display device (1000) can obtain information that the person (730) is an object that was previously detected in the original image (110) and is moving its position outside the first FOV (710) based on the tracking result of the person (730) obtained from the original image (110) and the tracking result of the person (730) obtained from the outer view image (115).
[0132] In one embodiment, the head mounted display device (1000) can display a user interface indicating identification information of at least one object located outside the first FOV (710). The user cannot recognize the presence of a person (730) located outside the first FOV (710) through the original image (110). Accordingly, the head mounted display device (1000) can provide the user with information about the presence of the person (730), whether the person is a new object not previously detected in the original image (110), the location, movement direction, and class of the person (730), etc. by displaying a user interface indicating identification information of the person (730).
[0133] FIGS. 8A, 8B, and 8C are drawings illustrating a user interface according to an embodiment of the present disclosure.
[0134] Referring to FIG. 8A, the head mounted display device (1000) can display a user interface for selecting an inpainting target. For example, the head mounted display device (1000) can detect a first person (10) from an original image (110) and display a first indicator (801) representing the detected first person (10) on the first person (10). The head mounted display device (1000) can display a second indicator (802) representing an object pointed by a user. The second indicator (802) can represent a virtual line generated in a direction pointed by a user input unit or a user's hand in a three-dimensional space corresponding to the original image (110) and an object that intersects the line. When an object detected by the second indicator (802) is selected, the head mounted display device (1000) can display a first overlay interface (810) for determining the selected object as an inpainting target. The head mounted display device (1000) can obtain a user input for determining a first person (10) as a target object by the user selecting a remove item (811) of the first overlay interface (810). The head mounted display device (1000) can obtain a user input for selecting a new detected object by the user selecting a cancel item (812) of the first overlay interface (810).
[0135] Referring to FIG. 8B, the head mounted display device (1000) can display a user interface indicating identification information of an object detected outside the first FOV (710). For example, the head mounted display device (1000) can detect an object (e.g., a person (730) of FIG. 7) outside the first FOV (710) from an outer view image. If the detected object is an object not detected in the original image (110), the head mounted display device (1000) can display a second overlay interface (820) indicating that a new object has been detected. The head mounted display device (1000) can obtain a user input for determining an object detected outside the first FOV (710) as a target object by the user selecting a remove item (821) in the second overlay interface (820). The head mounted display device (1000) can obtain a user input not to determine an object detected outside the first FOV (710) as a target object by the user selecting a cancel item (822).
[0136] Referring to FIG. 8C, the head-mounted display device (1000) may display a user interface for determining a preset distance range. For example, the head-mounted display device (1000) may identify whether the number of target objects is greater than or equal to a threshold value. Here, the threshold value may be determined based on the hardware resources of the head-mounted display device (1000). If the number of target objects is too large, it may be difficult to perform proper inpainting due to limitations in hardware resources, and thus the number of target objects may be limited. The threshold value may be determined based on the ratio of the area corresponding to the target object to the original image. If the ratio of the target object to the original image is too large, the completeness of the inpainting result may be reduced, and thus the number of target objects may be limited.
[0137] In one embodiment, the head mounted display device (1000) may display a third overlay interface (830) for determining a preset distance range when the number of target objects identified is greater than a threshold value. The head mounted display device (1000) may obtain a user input for determining the preset distance range by the user selecting a setting item (831). In addition, the head mounted display device (1000) may obtain a user input for deactivating the inpainting mode by the user selecting a release item (832).
[0138] FIG. 9 is a flowchart illustrating an operation of a head-mounted display device according to one embodiment of the present disclosure to perform inpainting based on whether a situation requires inpainting. Steps S240 and S250 of FIG. 9 correspond to steps S240 and S250 of FIG. 2 , and thus, any redundant description will be omitted.
[0139] In step S910, the head mounted display device (1000) can identify whether a situation requires inpainting of a target object based on movement information of a user wearing the head mounted display device.
[0140] In one embodiment, when the head mounted display device (1000) identifies a situation in which inpainting of a target object is required (S910-Y), the head mounted display device (1000) can obtain an inpainted image by inpainting the target object determined from the original image (S230).
[0141] In one embodiment, step S910 is illustrated in FIG. 9 as being performed between steps S240 and S250, but is not limited thereto, and may be performed before step S240 or after step S250.
[0142] In one embodiment, the head mounted display device (1000) can obtain user movement information through a movement sensor. In one embodiment, the head mounted display device (1000) can identify whether the user is moving based on the user movement information. In one embodiment, if the head mounted display device (1000) identifies that the user is moving, it can identify that the situation does not require inpainting for a target object. In one embodiment, if the head mounted display device (1000) identifies that the user is not moving, it can identify that the situation does require inpainting for a target object. That is, the head mounted display device (1000) can prevent the user from colliding with the target object by not displaying an inpainted image when the user is moving for the sake of the user's safety.
[0143] In one embodiment, the head-mounted display device (1000) can identify whether the user is moving based on the user's movement information. It can also identify whether inpainting of a target object is required. In one embodiment, if the head-mounted display device (1000) identifies the user as moving based on the user's movement information, it can determine that inpainting is not required. In one embodiment, if the head-mounted display device (1000) identifies the user as stationary based on the user's movement information, it can determine that inpainting is required.
[0144] In one embodiment, the head mounted display device (1000) can obtain gaze direction information of the user's eyes through an eye tracking sensor. In one embodiment, the head mounted display device (1000) can identify whether the user is focused on a virtual object based on the gaze direction information. For example, if it is determined that the user has fixed his / her gaze on a displayed virtual object for a threshold period of time based on the user's gaze direction information, it can be identified as a state of focusing on the displayed virtual object. In one embodiment, if the head mounted display device (1000) is identified as a state of focusing on a virtual object, it can identify that a situation is required for inpainting. In one embodiment, if the head mounted display device (1000) is identified as a state of not focusing on a virtual object, it can identify that a situation is not required for inpainting. That is, the head mounted display device (1000) can perform inpainting only when the user is focused on a virtual object, thereby preventing unnecessary hardware resource consumption while providing an environment in which the user can focus.
[0145] FIG. 10 is a flowchart illustrating an operation of a head-mounted display device performing noise cancellation according to one embodiment of the present disclosure. Steps S240 and S250 of FIG. 10 correspond to steps S210 and S250 of FIG. 2 , and thus, any redundant description will be omitted.
[0146] In step S1010, the head mounted display device (1000) may acquire an ambient audio signal of a section in which a target object is detected in an original image. In one embodiment, the head mounted display device (1000) may include a microphone that acquires an ambient audio signal at a time when the original image was captured. The head mounted display device (1000) may acquire an ambient audio signal of a section in which a target object is detected in the original image based on the signal acquired from the microphone.
[0147] In step S1020, the head mounted display device (1000) may obtain an audio signal in which an audio signal corresponding to a target object is reduced from the obtained ambient audio signal. In one embodiment, the head mounted display device (1000) may classify and separate the obtained ambient audio signal into audio signals for each of at least one detected object based on the obtained ambient audio signal. For example, the head mounted display device (1000) may classify and separate the obtained ambient audio signal into sounds corresponding to a car, sounds corresponding to a person, sounds corresponding to a bird, etc.
[0148] In one embodiment, the head-mounted display device (1000) may identify an audio signal corresponding to a detected target object among the audio signals of each of at least one detected object. For example, if the class of the target object is 'bird', the sound corresponding to 'bird' may be identified among the classified audio signals.
[0149] In one embodiment, the head mounted display device (1000) can reduce an audio signal corresponding to a target object identified in an ambient audio signal. In one embodiment, the head mounted display device (1000) can reduce an audio signal corresponding to the target object through a noise canceling algorithm. For example, the noise canceling algorithm may include, but is not necessarily limited to, filtering, masking, active noise canceling (ANC), and digital signal processing (DSP). In one embodiment, the head mounted display device (1000) can obtain an audio signal in which an audio signal corresponding to a target object identified in an ambient audio signal is reduced.
[0150] In one embodiment, the head-mounted display device (1000) may output a reduced audio signal corresponding to the acquired target object. For example, the head-mounted display device (1000) may include a speaker that outputs an audio signal, and may output a reduced audio signal corresponding to the acquired target object through the speaker at the time when the inpainted image is displayed.
[0151] In this way, the head mounted display device (1000) according to one embodiment of the present disclosure can allow a user to focus on the inpainted image from which the target object has been removed by removing an audio signal for the target object together with the inpainted image of the target object.
[0152] In one embodiment, although step S250 is illustrated as being performed after steps S1010 and S1020 in FIG. 10, the present invention is not limited thereto, and step S250 may be performed before steps S1010 and S1020, or step S250 may not be performed, and only steps S1010 and S1020 may be performed.
[0153] FIG. 11 is a perspective view of a head mounted display device according to one embodiment of the present disclosure.
[0154] Referring to FIG. 11, a head-mounted display device (1000) may include a frame (1001), an optical system (1002), a display (1100-1, 1100-2), a stereo camera (1210-1, 1210-2), a sub-camera (1220-1, 1220-2), a memory (1300), a distance detection sensor (1520), and a processor (1800). However, the present invention is not limited to the above-described examples, and some components may be omitted or other components may be added.
[0155] In one embodiment, the frame (1001) includes other components of the head mounted display device (1000) and is configured to allow a user to mount the head mounted display device (1000), including, but not limited to, glasses temples, a nose bridge, and the like. In one embodiment, left-eye optical components and right-eye optical components may be arranged or attached to the left and right sides of the frame (1001), or the left-eye optical components and right-eye optical components may be integrally formed and mounted to the frame (1001). In another example, some of the optical components may be arranged or attached to only one of the left and right sides of the frame (1001).
[0156] In one embodiment, the optical system (1002) may be a component that transmits light of an image to a user's eye. In one embodiment, the optical system (1002) may include at least one lens having a refractive power (degree) to focus or change the path of light of an image output from a display (1100-1, 1100-2). In one embodiment, light of an image output from a display (1100-1, 1100-2) may pass through the optical system (1002) and enter the user's eye.
[0157] The display (1100-1, 1100-2) is a component for displaying images and / or videos. Light of an image output from the display (1100-1, 1100-2) may be incident on the eyes of a user wearing the virtual reality head-mounted display device (1000). In one embodiment, the display (1100-1, 1100-2) may be configured as a physical device including at least one of a liquid crystal display, a thin film transistor-liquid crystal display, an organic light-emitting diode (OLED), a flexible display, a 3D display, and an electrophoretic display. In one embodiment, the displays (1100-1, 1100-2) may include a left-eye display (1100-1) and a right-eye display (1100-2), and the left-eye image of a stereo image may be displayed on the left-eye display (1100-1), and the right-eye image of the stereo image may be displayed on the right-eye display (1100-2). However, the present invention is not limited thereto, and the head-mounted display device (1000) may include a single display, and the left-eye image may be displayed on one area of the single display, and the right-eye image may be displayed on another area.
[0158] The stereo cameras (1210-1, 1210-2) are configured to capture images of objects and backgrounds within a real environment by photographing the real environment. In one embodiment, the stereo cameras (1210-1, 1210-2) may include a lens module, an image sensor, and an image processing module, and may capture images or moving images obtained by the image sensor (e.g., CMOS or CCD). In one embodiment, the stereo cameras (1210-1, 1210-2) may include a left-eye camera (1210-1) and a right-eye camera (1210-2). In one embodiment, the stereo cameras (1210-1, 1210-2) may capture stereo images or stereo videos based on images captured from the left-eye camera (1210-1) and the right-eye camera (1210-2). In one embodiment, the stereo cameras (1210-1, 1210-2) can capture the real environment with the first FOV (710) to obtain the original image. However, the present invention is not necessarily limited thereto, and the head-mounted display device (1000) may include a single camera or three or more multi-cameras other than the stereo cameras (1210-1, 1210-2), and may also obtain the original image from the single camera or the multi-cameras.
[0159] The sub-cameras (1220-1, 1220-2) are configured to capture images of the surrounding environment and the user's hand movements by photographing the real environment. In one embodiment, the sub-cameras (1220-1, 1220-2) may include a lens module, an image sensor, and an image processing module, and may capture images or videos obtained by an image sensor (e.g., CMOS or CCD). In one embodiment, the sub-cameras (1220-1, 1220-2) may capture an outer view image by photographing the real environment with a second FOV (720). However, the present invention is not limited thereto, and the head-mounted display device (1000) may include a single camera or three or more multi-cameras other than the sub-cameras (1220-1, 1220-2), and may also capture an outer view image from the single camera or the multi-cameras.
[0160] The distance detection sensor (1520) is configured to obtain data regarding the distance between an object in a real environment and the sensor. In one embodiment, the distance detection sensor (1520) may include an ultrasonic sensor that measures distance data based on the reflection time of sound, a laser sensor that measures distance data based on the reflection time or phase change of light, or the like. In one embodiment, the head-mounted display device (1000) may generate a depth map of an original image based on the data obtained from the distance detection sensor (1520), or generate a three-dimensional space including at least one object detected from the original image and an outer view image.
[0161] In one embodiment, the head mounted display device (1000) may include electronic components such as a memory (1300) and a processor (1800), and the electronic components may be mounted on a PCB substrate, an FPCB substrate, or the like and positioned at one location of the frame (1001) or may be distributed and positioned at multiple locations. In one embodiment, the electronic components included in the head mounted display device (1000) may further include a communication interface, an input interface, an output interface, and the like, and specific details regarding the operation of the electronic components of the head mounted display device (1000) will be described again below with reference to FIG. 13.
[0162] FIG. 12 is a detailed configuration diagram of a head mounted display device according to an embodiment of the present disclosure. Referring to FIG. 12, a head mounted display device (1000) may include a display (1100), a stereo camera (1210), a sub-camera (1220), a memory (1300), a communication interface (1400), a motion detection sensor (1510), a distance detection sensor (1520), an input interface (1600), an output interface (1700), and a processor (1800). The display (1100), the stereo camera (1210), the sub-camera (1220), the memory (1300), the communication interface (1400), the motion sensor (1510), the distance detection sensor (1520), the gaze tracking sensor (1530), the input interface (1600), the output interface (1700), and the processor (1800) may each be electrically and / or physically connected to each other.
[0163] The components illustrated in FIG. 12 are merely in accordance with one embodiment of the present disclosure, and the components included in the head mounted display device (1000) are not limited to those illustrated in FIG. 12. The head mounted display device (1000) according to one embodiment of the present disclosure may not include some of the components illustrated in FIG. 12, and may further include components not illustrated in FIG. 12. In addition, descriptions of overlapping content with the components described in FIG. 11 among the components illustrated in FIG. 12 will be omitted.
[0164] The memory (1300) may store instructions or program codes for performing functions or operations of the head mounted display device (1000). In one embodiment, at least one instruction, algorithm, data structure, program code, and application program stored in the memory (1300) may be implemented in a programming or scripting language such as, for example, C, C++, Java, or an assembler.
[0165] In one embodiment, the memory (1300) may include at least one of a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., SD or XD memory, etc.), a RAM (Random Access Memory), a SRAM (Static Random Access Memory), a ROM (Read-Only Memory), an EEPROM (Electrically Erasable Programmable Read-Only Memory), a PROM (Programmable Read-Only Memory), a Mask ROM, a Flash ROM, etc.), a hard disk drive (HDD), or a solid state drive (SSD).
[0166] In one embodiment, the memory (1300) may include a pre-stored panoramic image. The pre-stored panoramic image may be acquired from a stereo camera (1210) and stored in the memory (1300), or may be received from an external electronic device via a communication interface (1400) and stored in the memory (1300). In one embodiment, the memory (1300) may include an original image, an outer view image, and a depth map of the original image. In one embodiment, the memory (1300) may include detection results and tracking results of objects included in the original image and the outer view image. In one embodiment, the memory (1300) may include an object detection model, a segmentation model, an object tracking model, and an inpainting model. However, the present invention is not limited to the above-described examples, and the memory (1300) may further include various data necessary to perform the operations and functions of the head mounted display device (1000) described in the present disclosure.
[0167] The communication interface (1400) is a component for the head mounted display device (1000) to communicate with an external electronic device. In one embodiment, the communication interface (1400) may perform data communication between the head mounted display device (1000) and the external electronic device using at least one of data communication methods including wired LAN, wireless LAN, Wi-Fi, Bluetooth, zigbee, Wi-Fi Direct (WFD), infrared Data Association (IrDA), Bluetooth Low Energy (BLE), Near Field Communication (NFC), Wireless Broadband Internet (Wibro), World Interoperability for Microwave Access (WiMAX), Shared Wireless Access Protocol (SWAP), Wireless Gigabit Alliance (WiGig), and RF communication.
[0168] In one embodiment, the communication interface (1400) may transmit or receive at least one of an original image, an outer view image, and an inpainted image from an external electronic device to or from the external electronic device. However, the present disclosure is not necessarily limited to the above-described example, and various data necessary to perform the operations and functions of the head-mounted display device (1000) described in the present disclosure may be transmitted and received with the external electronic device via the communication interface (1400).
[0169] The motion sensor (1510) is a component for measuring the position, speed, direction, posture, etc. of the head mounted display device (1000). In one embodiment, the motion sensor may include an IMU (Inertial Measurement Unit) sensor. The IMU sensor may obtain 6 DoF (6 Degrees of Freedom) measurement values including position coordinate values (x-axis, y-axis, and z-axis coordinate values) and 3-axis angular velocity values (roll, yaw, pitch) of a user wearing the head mounted display device (1000). However, the present invention is not necessarily limited thereto, and the motion sensor (1510) may include various sensors for obtaining data necessary for identifying whether a user wearing the head mounted display device (1000) is in a moving state.
[0170] The gaze tracking sensor (1530) is a component for acquiring information on the gaze direction of a user's eyes. The gaze tracking sensor (1530) can detect the user's gaze direction by detecting an image of a person's pupil or pupils, or by detecting the direction or amount of light reflected from the cornea by near-infrared light. The gaze tracking sensor (1530) includes a left-eye gaze tracking sensor and a right-eye gaze tracking sensor, and can detect the gaze direction of the user's left eye and the gaze direction of the user's right eye, respectively. Detecting the user's gaze direction may mean acquiring information on the user's gaze direction.
[0171] The input interface (1600) is a component for receiving various user inputs. In one embodiment, the input interface (1600) may include a touch panel, a physical button, a microphone, etc. In one embodiment, information input through the input interface (1600) may be provided to the processor (1800). In one embodiment, a user input for determining a preset distance range may be obtained through the input interface (1600). In one embodiment, a user input for selecting a preset class may be obtained through the input interface (1600). In one embodiment, a user input for selecting an inpainting target may be obtained through the input interface (1600). In one embodiment, a user input for activating or deactivating the inpainting mode may be obtained through the input interface (1600). In one embodiment, the input interface (1600) may obtain an ambient audio signal at the time when the original image was captured. However, it is not necessarily limited to the above-described examples, and various data for performing the operation and function of the head mounted display device (1000) described in the present disclosure may be input through the input interface (1600).
[0172] The output interface (1700) is a component for the head-mounted display device (1000) to provide various information to the user. In one embodiment, the output interface (1700) may include a speaker. In one embodiment, the output interface (1700) may output a voice corresponding to text displayed through the user interface based on a signal received from the processor (1800), or may output a voice indicating whether the inpainting mode is activated. However, the present invention is not limited to the above-described example, and various voices for performing the operations and functions of the head-mounted display device (1000) described in the present disclosure may be output through the output interface (1700).
[0173] The processor (1800) can control the overall operations of the head-mounted display device (1000). In one embodiment, the processor (1800) can include multiple processors. In one embodiment, at least one processor (1800) can perform the operations and functions of the head-mounted display device (1000) described in the present disclosure by executing one or more instructions of a program stored in the memory (1300).
[0174] The processor (1800) may be configured as at least one of, for example, a Central Processing Unit, a microprocessor, a Graphic Processing Unit, Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), an Application Processor, a Neural Processing Unit, or an artificial intelligence processor designed with a hardware structure specialized for processing artificial intelligence models, but is not limited thereto.
[0175] When a method according to an embodiment of the present disclosure includes multiple operations, the multiple operations may be performed by a single processor or by multiple processors. For example, when a first operation, a second operation, and a third operation are performed by a method according to an embodiment, the first operation, the second operation, and the third operation may all be performed by a first processor, or the first and second operations may be performed by a first processor and the third operation may be performed by a second processor. However, the embodiments of the present disclosure are not limited thereto.
[0176] One or more processors according to the present disclosure may be implemented as a single-core processor or a multi-core processor. If a method according to an embodiment of the present disclosure includes multiple operations, the multiple operations may be performed by a single core or by multiple cores included in one or more processors.
[0177] In one embodiment, at least one processor (1800) may acquire an original image by photographing a real environment through a stereo camera by executing at least one command. In one embodiment, at least one processor (1800) may detect at least one object included in the original image by executing at least one command. In one embodiment, at least one processor (1800) may acquire depth information of at least one object detected using the original image by executing at least one command. In one embodiment, at least one processor (1800) may determine a target object among at least one object detected based on the depth information by executing at least one command. In one embodiment, at least one processor (1800) may inpaint an area corresponding to the determined target object in the original image. In one embodiment, at least one processor (1800) may display the inpainted image through the display (1100) by executing at least one command.
[0178] In one embodiment, at least one processor (1800) may identify at least one object located within a preset distance range from the head mounted display device (1000) among at least one object detected based on depth information by executing at least one command. In one embodiment, at least one processor (1800) may identify a target object among the at least one identified object by executing at least one command.
[0179] In one embodiment, at least one processor (1800) may display a virtual object through the display (1100) by executing at least one command. In one embodiment, at least one processor (1800) may identify a preset distance range based on the distance based on the displayed virtual object by executing at least one command.
[0180] In one embodiment, at least one processor (1800) can identify an object classified into a preset class among the identified at least one object as a target object by executing at least one instruction.
[0181] In one embodiment, at least one processor (1800) may obtain a user input for selecting at least one of the identified at least one object as an inpainting target by executing at least one command. In one embodiment, at least one processor (1800) may identify an object selected as an inpainting target from among the identified at least one object as a target object by executing at least one command.
[0182] In one embodiment, at least one processor (1800) may obtain a mask map indicating an area corresponding to a target object identified in an original image by executing at least one command. In one embodiment, at least one processor (1800) may obtain an inpainted image by applying the original image and the mask map to a learning model for inpainting the target object by executing at least one command.
[0183] In one embodiment, at least one processor (1800) may acquire an outer view image representing a second FOV wider than a first FOV of an original image through a sub-camera by executing at least one command. In one embodiment, at least one processor (1800) may track at least one object moving between the inside and the outside of the first FOV based on the original image and the outer view image by executing at least one command.
[0184] In one embodiment, at least one processor (1800) may execute at least one command to output a user interface indicating identification information of at least one object located outside the first FOV through the display.
[0185] In one embodiment, at least one processor (1800) may, by executing at least one command, identify whether a situation requires inpainting of an identified target object based on acquired motion information. In one embodiment, at least one processor (1800), by executing at least one command, may, if a situation is identified as requiring inpainting, inpaint an area corresponding to the identified target object.
[0186] FIG. 13 is a detailed configuration diagram of a server according to one embodiment of the present disclosure.
[0187] Referring to FIG. 13, the server (2000) may include a memory (2100), a communication interface (2200), and a processor (2300). Each may be electrically and / or physically connected to each other.
[0188] The components illustrated in FIG. 13 are merely in accordance with one embodiment of the present disclosure, and the components included in the server (2000) are not limited to those illustrated in FIG. 13. The server (2000) according to one embodiment of the present disclosure may not include some of the components illustrated in FIG. 13, and may further include components not illustrated in FIG. 13.
[0189] The memory (2100) may store instructions or program codes for performing functions or operations of the server (2000). In one embodiment, at least one instruction, algorithm, data structure, program code, and application program stored in the memory (2100) may be implemented in a programming or scripting language such as, for example, C, C++, Java, or assembler.
[0190] In one embodiment, the memory (2100) may include at least one of a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., SD or XD memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a mask ROM, a flash ROM, etc.), a hard disk drive (HDD), or a solid state drive (SSD).
[0191] In one embodiment, the memory (2100) may include a pre-stored panoramic image. The pre-stored panoramic image may be acquired from the head mounted display device (1000) via the communication interface (2200) and stored in the memory (2100), or may be received from an external electronic device via the communication interface (2200) and stored in the memory (2100). In one embodiment, the memory (2100) may include an original image, an outer view image, and a depth map of the original image. In one embodiment, the memory (2100) may include detection results and tracking results of objects included in the original image and the outer view image. In one embodiment, the memory (2100) may include an object detection model, a segmentation model, an object tracking model, and an inpainting model. However, the present invention is not limited to the above-described examples, and the memory (2100) may further include various data necessary to perform the operations and functions of the head mounted display device (1000) described in the present disclosure.
[0192] The communication interface (2200) is a component for the server (2000) to communicate with the head mounted display device (1000) or an external electronic device. In one embodiment, the communication interface (2200) may perform data communication between the server and the head mounted display device (1000) or an external electronic device using at least one of data communication methods including wired LAN, wireless LAN, Wi-Fi, Bluetooth, zigbee, Wi-Fi Direct (WFD), infrared Data Association (IrDA), Bluetooth Low Energy (BLE), Near Field Communication (NFC), Wireless Broadband Internet (Wibro), World Interoperability for Microwave Access (WiMAX), Shared Wireless Access Protocol (SWAP), Wireless Gigabit Alliance (WiGig), and RF communication.
[0193] In one embodiment, the communication interface (2200) may transmit at least one of the original image, the peripheral view image, and the inpainted image to the head mounted display device (1000) or an external electronic device, or receive the same from the head mounted display device (1000) or the external electronic device. However, the present invention is not limited to the above-described example, and various data necessary for performing the operation and function of the head mounted display device (1000) described in the present disclosure may be transmitted and received with the head mounted display device (1000) or the external electronic device through the communication interface (1400).
[0194] The processor (2300) can control the overall operations of the server (2000). In one embodiment, the processor (2300) can include multiple processors. In one embodiment, at least one processor (2300) can execute one or more instructions of a program stored in the memory (2100) to perform the operations and functions of the head mounted display device (1000) described in the present disclosure. In one embodiment, at least one processor (2300) can obtain an original image by executing at least one instruction. In one embodiment, at least one processor (2300) can detect at least one object included in the original image by executing at least one instruction. In one embodiment, at least one processor (2300) can obtain depth information of at least one object detected using the original image by executing at least one instruction. In one embodiment, at least one processor (2300) can identify a target object from among at least one object detected based on the depth information by executing at least one instruction. In one embodiment, at least one processor (2300) may inpaint an area corresponding to a target object identified in an original image by executing at least one command. Since the operation and function of the processor (2300) correspond to the operation and function of the processor (1800) of the head-mounted display device (1000) described in FIG. 12, a redundant description will be omitted.
[0195] Meanwhile, embodiments of the present disclosure may also be implemented in the form of a recording medium containing computer-executable instructions, such as program modules, executed by a computer. Computer-readable media may be any available media that can be accessed by a computer, and include both volatile and nonvolatile media, removable and non-removable media. Furthermore, computer-readable media may include computer storage media and communication media. Computer storage media include both volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Communication media may typically include computer-readable instructions, data structures, or other data in a modulated data signal, such as program modules.
[0196] Additionally, a computer-readable storage medium may be provided in the form of a non-transitory storage medium. Here, the term "non-transitory storage medium" simply means a tangible device that does not contain signals (e.g., electromagnetic waves). This term does not distinguish between cases where data is permanently stored in the storage medium and cases where data is temporarily stored. For example, a "non-transitory storage medium" may include a buffer in which data is temporarily stored.
[0197] The above description of the present disclosure is provided for illustrative purposes only, and those skilled in the art will readily appreciate that modifications to other specific forms can be made without altering the technical spirit or essential features of the present disclosure. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive. For example, components described as being single may be implemented in a distributed manner, and similarly, components described as being distributed may be implemented in a combined manner.
[0198] The scope of the present disclosure is indicated by the claims described below rather than the detailed description above, and all changes or modifications derived from the meaning and scope of the claims and their equivalent concepts should be interpreted as being included in the scope of the present disclosure.
Claims
1. In a method of operating a head-mounted display device, Step of obtaining an original image by photographing a real environment (S210); A step of detecting at least one object included in the original image (S220); A step (S230) of obtaining depth information of at least one detected object using the original image; A step (S240) of identifying a target object among at least one object detected based on the depth information; A step of inpainting an area corresponding to the identified target object in the original image (S250); and A method comprising a step (S260) of displaying the inpainted image.
2. In paragraph 1, The step of identifying the above target object is: A step of identifying at least one object located within a preset distance range from the head mounted display device among the at least one detected object based on the depth information; A method comprising: a step of identifying the target object among at least one object identified above; 3. In paragraph 2, The above method of operation is, a step of displaying a virtual object; and A method comprising: identifying the preset distance range based on the displayed virtual object; 4. In paragraph 2, The step of identifying the above target object is: A method comprising: a step of identifying an object classified into a preset class among at least one of the identified objects as the target object.
5. In any one of paragraphs 1 to 4, The above method of operation is, A step of obtaining an outer view image representing a second FOV wider than the first FOV of the original image; and A method further comprising: a step of tracking at least one detected object moving between the inside and the outside of the first FOV based on the original image and the outer view image.
6. In paragraph 5, The above method, A method further comprising: displaying a user interface indicating identification information of at least one detected object located outside the first FOV; 7. In any one of paragraphs 1 to 6, The above method, Further comprising a step of identifying whether a situation requires inpainting for the identified target object based on movement information of a user wearing the head mounted display device; The step of inpainting the area corresponding to the identified target object is: A method comprising: when a situation is identified in which inpainting is required, a step of inpainting an area corresponding to the identified target object; 8. In the head mounted display device (1000), display (1100); Stereo camera (1210); A memory (1300) storing at least one instruction; and At least one processor (1800) that executes at least one instruction stored in the memory (1300); The head mounted display device (1000) executes the at least one instruction by the at least one processor (1800) alone or in cooperation. The original image is obtained by photographing the real environment through the above stereo camera (1210), Detecting at least one object contained in the original image, Obtaining depth information of at least one detected object using the original image, Identifying a target object among at least one detected object based on the depth information, Inpainting the area corresponding to the identified target object in the original image, A head-mounted display device that displays the inpainted image through the display (1100).
9. In paragraph 8, The head mounted display device, wherein the at least one processor executes the at least one instruction, either alone or in cooperation with the at least one processor, Identifying at least one object located within a preset distance range from the head mounted display device among the at least one detected object based on the depth information; A head mounted display device that identifies the target object among at least one object identified above.
10. In paragraph 9, The head mounted display device, wherein the at least one processor executes the at least one instruction, either alone or in cooperation with the at least one processor, A step of displaying a virtual object through the above display; A head-mounted display device that identifies the preset distance range based on the displayed virtual object.
11. In paragraph 9, The head mounted display device, wherein the at least one processor executes the at least one instruction, either alone or in cooperation with the at least one processor, A head mounted display device that identifies an object classified into a preset class among at least one of the identified objects as the target object.
12. In any one of paragraphs 8 to 11, Including a sub camera; The head mounted display device, wherein the at least one processor executes the at least one instruction, either alone or in cooperation with the at least one processor, Obtain an outer view image representing a second FOV wider than the first FOV of the original image through the above sub camera, A head mounted display device that tracks at least one object moving between the inside and the outside of the first FOV based on the original image and the outer view image.
13. In paragraph 12, The head mounted display device, wherein the at least one processor executes the at least one instruction, either alone or in cooperation with the at least one processor, A head mounted display device that outputs a user interface indicating identification information of at least one object located outside the first FOV through the display.
14. In any one of paragraphs 8 to 13, Further comprising a motion sensor for obtaining user movement information; The head mounted display device, wherein the at least one processor executes the at least one instruction, either alone or in cooperation with the at least one processor, Based on the acquired movement information, identify whether inpainting is required for the identified target object, A head-mounted display device that, when identified as a situation requiring inpainting, inpaints an area corresponding to the identified target object.
15. A computer-readable recording medium recording a program for executing the method of any one of clauses 1 to 7 on a computer.
Citation Information
Patent Citations
Total field of view classification for head-mounted display
KR1020140034252A
Production of peptides using Bacillus and their manufacturing method
KR1020250031662A
Cover having the structure of opening and closing the drain hole formed in the cover for the sink drain
KR1020250140323A
A partition connector
KR102466925B1
Appratus and method for ESG management that facilitates response to internal and external ESG needs
KR102711439B1