Electronic device and control method therefor
The electronic device uses head or eye tracking to dynamically adjust 3D display elements based on user movement, overcoming the limitations of existing 3D technologies by providing a more immersive and realistic glasses-free 3D experience.
Patent Information
- Application Number
- PCT/KR2024/019881
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-18
- Filing Date
- 2024-12-05
- Publication Date
- 2025-07-17
AI Technical Summary
Existing 3D display technologies, such as those using binocular parallax, require auxiliary devices like glasses or suffer from limitations in providing a realistic three-dimensional experience without them.
An electronic device equipped with a camera, processor, and display that tracks user head or eye movements to dynamically adjust the display position and depth of objects within a 3D virtual space, enabling a glasses-free 3D experience through light field technology.
Provides a more immersive and realistic 3D experience by dynamically adjusting the display position and depth of objects based on user movement, enhancing the perceived three-dimensionality without the need for additional devices.
Smart Images

Figure KR2024019881_17072025_PF_FP_ABST
Abstract
Description
Electronic device and method of controlling the same
[0001] The present disclosure relates to an electronic device and a method for controlling the same, and more particularly, to an electronic device providing a 3D image and a method for controlling the same.
[0002] Advances in electronic technology have led to the development and proliferation of various types of electronic devices. In particular, display devices, used in a variety of settings, including homes, offices, and public spaces, have been continuously evolving in recent years.
[0003] Stereoscopy refers to three-dimensional technology. Recently, commercialized 3D displays primarily utilize binocular parallax. Binocular parallax offers the advantage of creating a three-dimensional effect on a single screen, such as a TV or theater screen. Methods utilizing binocular parallax can be categorized into stereoscopic (using glasses or other auxiliary devices) and autostereocopic (glassless) methods.
[0004] Recently, commercialization of glasses-free light field displays and glasses-free 3D displays utilizing eye-tracking is being continuously researched.
[0005] According to various embodiments of the present invention, an electronic device comprises: a display; at least one camera; a memory storing instructions; at least one processor including a processing circuit; and instructions, when individually or collectively executed by the at least one processor, cause the electronic device to display a 3D image including an object located in a 3D virtual space through the display, identify movement information corresponding to at least one of a head or an eye of a user in a 3D space where the user is located based on a captured image acquired through the camera, identify position movement information corresponding to the 3D virtual space based on the movement information, and control the display to change and display a display position and depth of the object within the 3D virtual space included in the 3D image based on the position movement information.
[0006] According to various embodiments, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify first movement distance information and first movement direction information corresponding to the position movement information, and control the display to change and display the display position and depth of the object within the three-dimensional virtual space based on the first movement distance information and the first movement direction information.
[0007] According to various embodiments, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify second movement distance information corresponding to a difference between the first position and the second position and second movement direction information from the first position to the second position when the position of at least one of the head or eyes of the user changes from a first position to a second position in the captured image acquired through the camera, and to identify the first movement distance information and the first movement direction information based on the second movement distance information and the second movement direction information.
[0008] According to various embodiments, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify the first movement distance information and the first movement direction information by scaling the second movement distance information and the second movement direction information to correspond to the three-dimensional virtual space.
[0009] According to various embodiments, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify the second movement distance information based on a three-axis coordinate value corresponding to the first location and a three-axis coordinate value corresponding to the second location, and to identify the second movement direction information based on a three-axis angular velocity value from the first location to the second location.
[0010] According to various embodiments, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain a vector value based on the three-axis coordinate value and the three-axis angular velocity value, and to identify the second movement distance information and the second movement direction information corresponding to the relative movement of at least one of the user's head or eyes based on the obtained vector value.
[0011] According to various embodiments, each of the three-axis coordinate values corresponding to the first position and the three-axis coordinate values corresponding to the second position may include an X coordinate, a Y coordinate, and a Z coordinate in XYZ space, and the three-axis angular velocity values may include a roll value, a pitch value, and a yaw value.
[0012] According to various embodiments, the 3D image includes a plurality of objects located in the 3D virtual space, and the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify a moving target object among the plurality of objects based on a current position of at least one of the user's head or eyes, or to identify a moving target object among the plurality of objects based on a user selection command.
[0013] According to various embodiments, the camera includes a plurality of cameras spaced apart by a preset distance, and the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify disarity information based on a first captured image and a second captured image acquired through the plurality of cameras, and to identify movement information corresponding to at least one of the head or eyes of the user in a three-dimensional space where the user is located based on the disarity information.
[0014] According to various embodiments, the display may be implemented as a Light Field Display (LFD).
[0015] A method for controlling an electronic device according to various embodiments may include: displaying a 3D image including an object located in a 3D virtual space; identifying movement information corresponding to at least one of a head or an eye of a user in a 3D space where the user is located based on a captured image acquired through a camera; identifying position movement information corresponding to the 3D virtual space based on the movement information; and changing and displaying a display position and depth of the object within the 3D virtual space included in the 3D image based on the position movement information.
[0016] In accordance with various embodiments, a non-transitory computer-readable medium storing computer instructions that, when individually or collectively executed by a processor including a processing circuit of an electronic device, cause the electronic device to perform an operation, the operation may include: displaying a 3D image including an object located in a 3D virtual space; identifying movement information corresponding to at least one of a head or an eye of a user in a 3D space where the user is located based on a captured image acquired through a camera; identifying position movement information corresponding to the 3D virtual space based on the movement information; and changing and displaying a display position and depth of the object within the 3D virtual space included in the 3D image based on the position movement information.
[0017] The above and other aspects, features and advantages of specific embodiments of the present disclosure will become more apparent from the following description taken in conjunction with the accompanying drawings.
[0018] FIG. 1 is a drawing for explaining the operation of an electronic device according to various embodiments.
[0019] FIG. 2A is a block diagram showing the configuration of an electronic device according to one embodiment.
[0020] FIG. 2b is a block diagram specifically showing the configuration of an electronic device according to various embodiments.
[0021] FIG. 3 is a drawing for explaining the structure and operation of a display according to various embodiments.
[0022] FIG. 4 is a flowchart for explaining a method of controlling an electronic device according to various embodiments.
[0023] FIG. 5 is a drawing for explaining examples of 3D images according to various embodiments.
[0024] FIG. 6 is a diagram for explaining a user location tracking method according to various embodiments.
[0025] FIG. 7 is a flowchart for explaining a method of controlling an electronic device according to various embodiments.
[0026] FIG. 8 is a diagram for explaining a method of mapping user movement information and movement information in a three-dimensional virtual space according to various embodiments.
[0027] FIG. 9 is a drawing for explaining a method for controlling a three-dimensional virtual space according to various embodiments.
[0028] FIG. 10 is a drawing for explaining a method for controlling a three-dimensional virtual space according to various embodiments.
[0029] The terms used in this disclosure will be briefly explained, and the disclosure will be specifically described with reference to the drawings.
[0030] The terms used in the embodiments of this disclosure have been selected from widely used, current terms, taking into account the functions of this disclosure. However, these terms may vary depending on the intentions or cases of those skilled in the art, the emergence of new technologies, etc. Furthermore, in certain cases, terms may be arbitrarily selected by the applicant, and in such cases, their meanings will be described in detail in the description of the relevant disclosure. Therefore, the terms used in this disclosure should not be defined simply as names of terms, but rather based on the meanings of the terms and the overall content of this disclosure.
[0031] In this disclosure, expressions such as “has,” “can have,” “includes,” or “may include” indicate the presence of a feature (e.g., a number, function, operation, or component such as a part), and do not exclude the presence of additional features.
[0032] In this disclosure, expressions such as “A or B,” “at least one of A and / or B,” or “one or more of A or / and B” can include all possible combinations of the listed items. For example, “A or B,” “at least one of A and B,” or “at least one of A or B” can all refer to cases where (1) only A is included, (2) only B is included, or (3) both A and B are included.
[0033] The expressions “first,” “second,” “first,” or “second,” etc., used in this disclosure can describe various components, regardless of order and / or importance, and are only used to distinguish one component from another, but do not limit the components.
[0034] When it is said that a component (e.g., a first component) is “operatively or communicatively coupled with / to” or “connected to” another component (e.g., a second component), it should be understood that the component may be directly coupled to the other component, or may be connected through another component (e.g., a third component).
[0035] The expression "configured to" as used in the present disclosure may be used interchangeably with, for example, "suitable for," "having the capacity to," "designed to," "adapted to," "made to," or "capable of." The term "configured to" may not necessarily mean only "specifically designed to" in terms of hardware.
[0036] In some contexts, the phrase "a device configured to" may mean that the device, in conjunction with other devices or components, is "capable of" performing A, B, and C. For example, the phrase "a processor configured (or set) to perform A, B, and C" may refer to a dedicated processor (e.g., an embedded processor) for performing those operations, or a general-purpose processor (e.g., a CPU or application processor) that can perform those operations by executing one or more software programs stored in a memory device.
[0037] Singular expressions include plural expressions unless the context clearly dictates otherwise. In this application, terms such as "comprise" or "consist of" are intended to indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the disclosure, but should be understood not to preclude the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0038] In the present disclosure, a "module" or "part" performs at least one function or operation and may be implemented in hardware or software, or a combination of hardware and software. Furthermore, multiple "modules" or multiple "parts" may be integrated into at least one module and implemented as at least one processor (not shown), excluding any "modules" or "parts" that need to be implemented as specific hardware.
[0039] Meanwhile, the various elements and areas in the drawings are schematically drawn. Therefore, the technical concepts of the present disclosure are not limited by the relative sizes or spacings drawn in the attached drawings.
[0040] Various embodiments of the present disclosure will be described in more detail with reference to the attached drawings below.
[0041] FIG. 1 is a drawing for explaining the operation of an electronic device according to various embodiments.
[0042] Although not shown in the drawing, the electronic device (100) (e.g., see FIG. 2A) may be implemented as various types of display devices such as a TV, a monitor, a kiosk, a tablet PC, an electronic picture frame, a mobile phone, a large format display (LFD), a digital signage, a digital information display (DID), a video wall, a projector display, etc. However, in some cases, it may be implemented as an image processing device (e.g., a set-top box, one connected box) that is connected to a display device and provides an image.
[0043] According to one embodiment, the electronic device (100) may be equipped with a light field display. A light field display is a display technology that provides a more realistic visual experience by expressing light field information, unlike conventional 2D or 3D displays.
[0044] While 2D or 3D displays typically provide limited information about light direction and depth, light field displays can provide visual experiences similar to those observed in the real world by expressing additional information about light direction and depth using light field information. For example, light field displays can be utilized to provide more realistic environments in virtual reality (VR) and / or augmented reality (AR) devices.
[0045] Fig. 1 is a diagram for explaining the operation of a light field display using a lenticular lens method. According to Fig. 1, each lenticular lens, for example, a micro lens array, is assigned a series of display pixels, and light from each pixel is directed in a specific direction by the lens to form a light field expressed in terms of light intensity and direction. When the display is viewed within the light field thus formed, the user can experience a three-dimensional effect.
[0046] According to one embodiment, the electronic device (100) can display the depth of an object (10, 20) in a virtual 3D space provided through the display (110) (see FIG. 2a) by moving it based on the user's positional movement in the actual 3D space where the user is located, for example, the positional movement of the eyes or head.
[0047] Below, various embodiments are described for identifying the user's positional movement in the actual 3D space where the user is located and mapping it to a virtual 3D space.
[0048] FIG. 2A is a block diagram showing the configuration of an electronic device according to one embodiment.
[0049] According to FIG. 2A, the electronic device (100) includes a display (110), a memory (120), a camera (130), and at least one processor (140) (e.g., including a processing circuit).
[0050] The display (110) may be implemented as a display including a self-luminous element or a display including a non-luminous element and a backlight. For example, it may be implemented as a display in various forms such as a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a light emitting diode (LED), a micro LED, a mini LED, a plasma display panel (PDP), a quantum dot (QD) display, a quantum dot light-emitting diode (QLED), etc., but is not limited thereto. The display (110) may also include a driving circuit, a backlight unit, etc., which may be implemented in a form such as an a-si TFT, an LTPS (low temperature poly silicon) TFT, an OTFT (organic TFT), etc. In one example, a touch sensor that detects a touch operation in the form of a touch film, a touch sheet, a touch pad, etc. may be disposed on the front of the display (110) so as to be implemented so as to detect various types of touch inputs. For example, the display (110) can detect various types of touch inputs, such as a touch input by a user's hand, a touch input by an input device such as a stylus pen, and a touch input by a specific electrostatic material. Here, the input device can be implemented as a pen-type input device that can be referred to by various terms such as an electronic pen, a stylus pen, an S-pen, etc. According to an example, the display (110) can be implemented as a flat display, a curved display, a flexible display that can be folded or / and rolled, etc.
[0051] The memory (120) can store data required for various embodiments. The memory (120) may be implemented in the form of memory embedded in the electronic device (100') or in the form of memory that can be detachably attached to the electronic device (100) depending on the purpose of data storage. For example, data for driving the electronic device (100) may be stored in a memory embedded in the electronic device (100'), and data for expanding the functions of the electronic device (100) may be stored in a memory that can be detachably attached to the electronic device (100). In the case of memory embedded in the electronic device (100), it may be implemented as at least one of volatile memory (e.g., dynamic RAM (DRAM), static RAM (SRAM), or synchronous dynamic RAM (SDRAM)), non-volatile memory (e.g., one time programmable ROM (OTPROM), programmable ROM (PROM), erasable and programmable ROM (EPROM), electrically erasable and programmable ROM (EEPROM), mask ROM, flash ROM, flash memory (e.g., NAND flash or NOR flash), hard drive, or solid state drive (SSD)). In addition, in the case of memory that can be attached or detached to the electronic device (100'), it may be implemented as at least one of memory cards (e.g., compact flash (CF), secure digital (SD), micro secure digital (Micro-SD), mini secure digital (Mini-SD), extreme digital (xD), multi-media card (MMC), etc.), external memory that can be connected to a USB port (e.g., USB memory), etc. It can be implemented.
[0052] As an example, the memory (120) may store a computer program including at least one instruction or instructions for controlling the electronic device (100).
[0053] In another example, the memory (120) may store images received from an external device (e.g., a source device), an external storage medium (e.g., USB), an external server (e.g., web hard), etc., for example, an input image. Alternatively, the memory (120) may store images acquired through a camera provided in the electronic device (100).
[0054] According to another example, the memory (120) can store various information required for image quality processing, for example, information, algorithms, image quality parameters, etc. for performing at least one of Noise Reduction, Detail Enhancement, Tone Mapping, Contrast Enhancement, Color Enhancement, or Frame Rate Conversion.
[0055] According to one embodiment, the memory (120) may be implemented as a single memory that stores data generated from various operations according to the present disclosure. However, according to one embodiment, the memory (120) may also be implemented to include multiple memories that each store different types of data or each store data generated at different stages.
[0056] In the above-described embodiment, it has been described that various data are stored in the external memory (120) of the processor (140), but at least some of the above-described data may be stored in the internal memory of the processor (140) depending on an implementation example of at least one of the electronic device (100) or the processor (140).
[0057] One or more cameras (130) can be turned on and perform shooting according to a preset event. For example, one or more cameras (130) can perform shooting according to an event in which the electronic device (100) (or the display (110)) is turned on. The camera (130) can convert a captured image into an electrical signal and generate image data based on the converted signal. For example, a subject can be converted into an electrical image signal through a semiconductor optical element (CCD; Charge Coupled Device), and the converted image signal can be amplified and converted into a digital signal and then signal processed. For example, the one or more cameras (130) can include at least one of a general (or basic) camera, an ultra-wide-angle camera, and a depth camera.
[0058] In one example, one or more cameras (130) may be positioned to capture the front of the display (110). For example, one or more cameras (130) may be positioned in the central area of the top bezel of the display (110).
[0059] In one example, one or more cameras (130) may be positioned in a direction and angle that can capture the front of the display (110). In one example, the cameras (130) may be positioned in a direction and angle that can be recognized as facing the front of the display (110) when the user's gaze is directed forward in the captured image.
[0060] In one example, one or more cameras (130) may include multiple cameras spaced apart at preset intervals (e.g., a specific interval) to capture different viewpoints. For example, the preset interval may be, but is not limited to, the same / similar distance as the distance between a person's two eyes.
[0061] At least one processor (140) includes various processing circuits and controls the overall operation of the electronic device (100). For example, at least one processor (140) may be connected to each component of the electronic device (100) and may control the overall operation of the electronic device (100). For example, at least one processor (140) may be electrically connected to the display (110) and the memory (120) and may control the overall operation of the electronic device (100). At least one processor (140) may be composed of one or more processors. For example, at least one processor (140) may include various processing circuits and / or multiple processors. For example, the term “processor” as used in the present disclosure, including the claims, may include various processing circuits, including at least one processor, wherein one or more of the at least one processors may be configured to individually and / or collectively perform various functions described in the present disclosure in a distributed manner. When the terms "processor," "at least one processor," and "one or more processors" are used herein to describe a processor configured to perform a number of functions, these terms encompass, for example and without limitation, situations where one processor performs some of the recited functions and other processor(s) perform other recited functions, as well as situations where a single processor can perform all of the recited functions. Furthermore, the at least one processor may comprise a combination of processors that perform various recited / disclosed functions (e.g., in a distributed manner). At least one processor may execute program instructions to achieve or perform various functions.
[0062] At least one processor (140) can perform operations of the electronic device (100) according to various embodiments by executing at least one instruction stored in the memory (120).
[0063] In one example, the artificial intelligence related functions according to the present disclosure may be operated through the processor and memory of an electronic device.
[0064] At least one processor (140) may be composed of one or more processors. In this case, the one or more processors may include at least one of a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and an NPU (Neural Processing Unit), but are not limited to the examples of the processors described above.
[0065] CPUs are general-purpose processors capable of performing not only general calculations but also artificial intelligence calculations. Their multi-layered cache structure allows for the efficient execution of complex programs. CPUs are advantageous for serial processing, enabling organic linking of previous and subsequent calculation results through sequential calculations. General-purpose processors are not limited to the examples described above, except where specifically identified as CPUs.
[0066] A GPU is a processor designed for large-scale computations, such as floating-point operations used in graphics processing. It integrates a large number of cores to perform large-scale computations in parallel. In particular, GPUs may be advantageous over CPUs in parallel processing methods, such as convolution operations. Furthermore, GPUs can be used as coprocessors to supplement the functions of CPUs. Processors for large-scale computations are not limited to the examples described above, except in cases where they are specifically referred to as GPUs.
[0067] An NPU is a processor specialized in artificial intelligence computation using artificial neural networks, and each layer of the artificial neural network can be implemented in hardware (e.g., silicon). Since NPUs are designed specifically according to the company's specifications, they have less freedom than CPUs or GPUs, but can efficiently process the AI computations requested by the company. Meanwhile, as a processor specialized in AI computation, an NPU can be implemented in various forms, such as a Tensor Processing Unit (TPU), an Intelligence Processing Unit (IPU), or a Vision Processing Unit (VPU). Except as specifically designated as an NPU, an AI processor is not limited to the examples described above.
[0068] Additionally, at least one processor (140) may be implemented as a SoC (System on Chip). In this case, the SoC may further include, in addition to at least one processor (140), a memory (120), and a network interface such as a bus for data communication between the processor (140) and the memory (120).
[0069] When a System on Chip (SoC) included in an electronic device (100) includes multiple processors, the electronic device (100) may perform operations related to artificial intelligence (e.g., operations related to learning or inference of a neural network model) by using some of the multiple processors. For example, the electronic device may perform operations related to artificial intelligence by using at least one of a GPU, NPU, VPU, TPU, or hardware accelerator specialized in artificial intelligence operations such as convolution operations or matrix multiplication operations among the multiple processors. However, these are merely various embodiments, and it is of course possible to process operations related to artificial intelligence by using a CPU or other general-purpose processor.
[0070] Additionally, the electronic device (100) can perform operations related to functions related to artificial intelligence by utilizing multiple cores (e.g., dual cores, quad cores, etc.) included in a single processor. In particular, the electronic device can perform artificial intelligence operations, such as convolution operations and matrix multiplication operations, in parallel by utilizing multiple cores included in the processor.
[0071] At least one processor (140) is controlled to process input data according to predefined operating rules or neural network models (or artificial intelligence models) stored in memory (120). The predefined operating rules or neural network models are characterized by being created through learning.
[0072] Here, "created through learning" means that a predefined set of behavioral rules or a neural network model with desired characteristics is created by applying a learning algorithm to a large number of learning data. This learning may be performed on the device itself, where the artificial intelligence according to the present disclosure is implemented, or through a separate server / system.
[0073] A neural network model may be composed of multiple neural network layers. At least one layer has at least one weight value and performs its operation through the operation result of the previous layer and at least one defined operation. Examples of neural networks include a convolutional neural network (CNN), a recurrent neural network (RNN), a deep neural network (DNN), a restricted boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), deep Q-networks, and a transformer. The neural networks in the present disclosure are not limited to the above-described examples unless otherwise specified.
[0074] A learning algorithm is a method for training a target device (e.g., a robot) using a plurality of learning data sets so that the target device can make decisions or predictions on its own. Examples of learning algorithms include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning. The learning algorithms in the present disclosure are not limited to the aforementioned examples unless otherwise specified. For convenience of explanation, at least one processor (140) will be referred to as a processor (140) hereinafter.
[0075] According to one embodiment, the electronic device (100) can receive various compressed images or images of various resolutions. For example, the electronic device (100) can receive images in a compressed form such as MPEG (Moving Picture Experts Group) (e.g., MP2, MP4, MP7, etc.), JPEG (joint photographic coding experts group), AVC (Advanced Video Coding), H.264, H.265, HEVC (High Efficiency Video Codec), etc. The electronic device (100) can receive any one of SD (Standard Definition), HD (High Definition), Full HD, and Ultra HD images.
[0076] For example, the processor (140) may perform image processing on an input image to obtain an output image. Here, the image processing may include, but is not limited to, at least one of image enhancement, image restoration, image transformation, image analysis, image understanding, image compression, image decoding, or scaling.
[0077] In the present disclosure, a "region" is a term referring to, for example, a portion of an image, and means at least one pixel block or a set of pixel blocks. In addition, a pixel block means, for example, a set of adjacent pixels that include at least one pixel.
[0078] For example, the input image may include a 3D image. For example, the input image may include a side-by-side image. A side-by-side image may be an image in which two images are placed side by side on a single screen. For example, each image may occupy half of the horizontal space of the screen, with one image located in the left area and the other in the right area. For example, the image located in the left area may be a left-eye image, and the image located in the right area may be a right-eye image.
[0079] For example, when a plurality of frames included in an input image are sequentially input, the processor (140) may store the plurality of frames in the memory (120) and read the frames stored in the memory (120) to perform various processing. A frame is a basic image unit in image content, and each frame is composed of pixels and may include resolution and color information. Hereinafter, content may be one frame included in the image content or a preset number of multiple frames, but for the convenience of explanation, they will be collectively referred to as content.
[0080] FIG. 2b is a block diagram specifically showing the configuration of an electronic device according to various embodiments.
[0081] According to FIG. 2B, the electronic device (100') may include a display (110), a memory (120), a camera (130), at least one processor (140) (e.g., including a processing circuit), a user interface (150) (e.g., including various circuits), a communication interface (160) (e.g., including a communication circuit), and a speaker (170). For the configurations illustrated in FIG. 2B that overlap with the configuration illustrated in FIG. 2A, a detailed description will not be repeated.
[0082] The user interface (150) includes various circuits and may be implemented as devices such as buttons, touch pads, mice, and keyboards, or as a touch screen capable of performing the above-described display functions and operation input functions.
[0083] The communication interface (160) includes various communication circuits, and can be implemented as various interfaces depending on the implementation example of the electronic device (100'). For example, the communication interface (140) can communicate with an external device, an external storage medium (e.g., a USB memory), an external server (e.g., a web hard drive), etc. through a communication method such as Bluetooth, AP-based Wi-Fi (Wireless LAN network), Zigbee, wired / wireless LAN (Local Area Network), WAN (Wide Area Network), Ethernet, IEEE 1394, HDMI (High-Definition Multimedia Interface), USB (Universal Serial Bus), MHL (Mobile High-Definition Link), AES / EBU (Audio Engineering Society / European Broadcasting Union), optical, coaxial, etc. In one example, the communication interface (160) can communicate with another electronic device, an external server, and / or a remote control device.
[0084] The speaker (170) may be configured to output various audio data as well as various notification sounds or voice messages. The processor (130) may control the speaker (170) to output feedback or various notifications in audio format according to various embodiments of the present disclosure.
[0085] In addition, the electronic device (100') may include sensors, microphones, etc. according to various implementation examples.
[0086] Sensors may include various types of sensors, such as touch sensors, proximity sensors, acceleration sensors (or gravity sensors), geomagnetic sensors, gyro sensors, pressure sensors, position sensors, distance sensors, light sensors, etc.
[0087] A microphone is a device configured to receive user voice or other sounds and convert them into audio data. However, according to another embodiment, the electronic device (100') may receive user voice input via an external device through a communication interface (160).
[0088] FIG. 3 is a drawing for explaining the structure and operation of a display (110) according to various embodiments.
[0089] According to FIG. 3, the display (110) may include a display panel (111), a display separator (112), and a backlight unit (113). However, depending on the implementation example of the display (110), the backlight unit (113) may not be included in the display (110).
[0090] The display panel (111) includes a plurality of pixels composed of a plurality of sub-pixels. Here, the sub-pixels may be composed of R (Red), G (Green), and B (Blue). For example, pixels composed of R, G, and B sub-pixels may be arranged in a plurality of row and column directions to form the display panel (141).
[0091] The display panel (111) displays a binocular viewpoint image (or a multi-viewpoint image). For example, the display panel (111) can display an image in which multiple images of a right-eye image and a left-eye image are sequentially and repeatedly arranged.
[0092] The viewing area separator (112) is arranged on the front of the display panel (111) to provide different viewpoints, i.e., multi-views, for each viewing area. In this case, the viewing area separator (112) may be implemented as a lenticular lens or a parallax barrier. For example, the viewing area separator (112) may be implemented as a lenticular lens including a plurality of lens regions. Accordingly, the lenticular lens may refract an image displayed on the display panel (111) through the plurality of lens regions. Each lens region is formed to have a size corresponding to at least one pixel, so that light passing through each pixel may be differently dispersed for each viewing area. As another example, the viewing area separator (112) may be implemented as a parallax barrier. The parallax barrier is implemented as a transparent slit array including a plurality of barrier regions. Accordingly, light can be blocked through slits between barrier areas to allow images from different viewpoints to be output for each viewing area.
[0093] For example, the field of view separator (112) may be implemented to operate at a constant angle to improve image quality, i.e., to avoid resolution reduction. In this case, the processor (130) may divide the right-eye image and the left-eye image based on the angle at which the field of view separator (112) is tilted, and combine them to generate a multi-view image. Accordingly, the user does not view the image displayed vertically or horizontally on the sub-pixels of the display panel (111), but rather views the image displayed so that the sub-pixels have a constant angle of inclination.
[0094] In one example, the field of view separator (112) may be implemented as an active type. For example, the field of view separator (112) may be implemented as an active lenticular lens or an active parallax barrier. In one example, the field of view separator (112) may be implemented as a lenticular lens array as illustrated in FIG. 3b.
[0095] According to an example, the field separation unit (112) may be implemented as a lenticular lens array as illustrated in FIG. 3B. When the display (110) operates in 3D mode, the processor (140) may control the lenticular lens (112) to operate in 3D mode by applying a preset voltage corresponding to the 3D mode to the active lenticular lens (112). In addition, when the display (110) operates in 2D mode, the processor (140) may control the lenticular lens (112) to operate in 2D mode by applying a preset voltage corresponding to the 2D mode to the active lenticular lens (112).
[0096] According to an example, the lenticular lens (112) includes a plurality of micro lenticular lenses, and a lens pattern may be formed within the gap between the display panel (111) and the lenticular lens (112).
[0097] For example, as illustrated in Fig. 3, a transparent frame made of PI in the shape of a micro lens is filled with liquid crystals, and the outside can be made of a replica made of a material having the same refractive index as the liquid crystal molecules when voltage is applied. ITO electrodes to which voltage is applied can be positioned above and below the micro lenses of this structure. In the 3D mode where no voltage A is applied, a difference in refractive index occurs between the liquid crystal molecules inside and the external replica, resulting in the effect of passing through the lenticular lens. On the other hand, in the 2D mode where voltage B is applied, the state of the liquid crystal changes to have the same refractive index as the external replica, and allows the input light to pass through as is.
[0098] The backlight unit (113) includes a backlight and provides light to the display panel (111). By the light provided from the backlight unit (113), the left-eye image and the right-eye image 1, 2 formed on the display panel (111) are projected onto the viewing area separator (112), and the viewing area separator (112) can disperse the light of each projected image 1, 2 and transmit it toward the viewer. For example, the viewing area separator (112) can create exit pupils at the viewer's position, i.e., the viewing distance. As illustrated in FIG. 3A, when the viewing area separator (112) is implemented as a lenticular lens array, the thickness and diameter of the lenticular lens, and when it is implemented as a parallax barrier, the spacing of the slits, etc., can be designed so that the exit pupils created by each row are separated by an average binocular center distance of less than 65 mm.
[0099] FIG. 4 is a flowchart for explaining a method of controlling an electronic device according to various embodiments.
[0100] According to FIG. 4, in operation 410, the electronic device (100) can display a 3D image including an object located in a 3D virtual space through the display (110).
[0101] For example, a 3D image may be a UI screen that includes at least one object located in a three-dimensional virtual space. For example, the UI screen may be a home UI screen that includes multiple applications. In this case, the object may include UI elements corresponding to the multiple applications. However, the UI screen is not necessarily limited to the home UI screen and may include UI screens such as a menu settings screen.
[0102] For example, the 3D image may be a VR content screen, such as a game content screen, an entertainment content screen, or a fitness content screen. In this case, the object may be an input control element, such as the user's head, the user's hand, a pointer, or a cursor.
[0103] For example, the 3D image may be a regular content screen, such as a movie content screen. In this case, the object may be a UI element, such as a content control menu.
[0104] In operation 420, the electronic device (100) can identify user movement information in a three-dimensional space where the user is located based on a captured image acquired through the camera (130). In one example, the user movement information may be movement information corresponding to at least one of the user's head or eyes. For convenience of explanation, the following description will assume that the user movement information is movement information of the user's eyes.
[0105] For example, the electronic device (100) can identify at least one of the user's head or eyes within the captured image obtained by the camera (130). For example, the electronic device (100) can identify an object area included in the captured image through at least one of object recognition, object detection, object tracking, and image segmentation. For example, the electronic device (100) can identify a user by using a technique such as semantic segmentation, which classifies and extracts objects included in an input image by type as needed, instance segmentation, which recognizes objects by classifying them by object even if they are of the same type, and a bounding box in the shape of a rectangle that includes the detected object when detecting an object included in an image.
[0106] For example, the electronic device (100) can identify a user's face from a captured image and identify the user's eyes from the user's face. For example, various conventional methods can be used as a face region detection method. Specifically, a direct recognition method and a statistical method can be used. The direct recognition method creates rules using physical features of the face image, such as the outline, skin color, and size of components or distances between them, and compares, inspects, and measures according to the rules. The statistical method can detect the face region according to a pre-learned algorithm. That is, it is a method of digitizing the unique features included in the input face and comparing and analyzing them with a large prepared database (shapes of faces and other objects). For example, the face region can be detected according to a pre-learned algorithm, and methods such as a Multi Layer Perceptron (MLP) and a Support Vector Machine (SVM) can be used. The user's eye region can be identified using a similar method.
[0107] In one example, the electronic device (100) can identify a user's eyes from a captured image using a learned artificial intelligence model. For example, the artificial intelligence model can be implemented as a neural network including multiple neural network layers. The artificial intelligence model can be implemented as a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), or a deep Q-network, but is not limited thereto.
[0108] For example, user movement information, for example, user eye movement information, may include movement distance information and movement direction information. For example, user movement information may include movement distance information and movement direction information in a three-dimensional space using a coordinate system of the X-axis, Y-axis, and Z-axis.
[0109] In operation 430, the electronic device (100) can identify position movement information corresponding to a three-dimensional virtual space based on movement information of the identified user.
[0110] For example, the electronic device (100) may identify positional movement information in a three-dimensional virtual space based on, for example, movement information of the user's eyes. For example, each of the movement information of the user's eyes and the positional movement information in the three-dimensional virtual space may be relative movement information having distance and direction information.
[0111] In operation 440, the electronic device (100) can control the display (110) to change and display the display position and depth of the object within the included three-dimensional virtual space of the 3D image based on the identified position movement information.
[0112] For example, depth is data indicating how far each pixel in an image is from the camera. Depth is typically provided in pixels, and the value of each pixel can represent the distance to the object at that location. For example, if the depth of a specific object is large, the binocular disparity between the left and right eyes is large, resulting in a relatively greater sense of three-dimensionality. On the other hand, if the depth is small, the binocular disparity between the left and right eyes is small, resulting in a relatively less sense of three-dimensionality.
[0113] For example, the electronic device (100) can control the display (110) to change and display the display position and depth of an object based on first movement distance information and first movement direction information corresponding to a three-dimensional virtual space.
[0114] According to an example, the electronic device (100) may identify first movement distance information and first movement direction information corresponding to a three-dimensional virtual space based on second movement distance information and second movement direction information corresponding to user movement information, for example, movement information of the user's eyes. For example, an operation may be performed to map movement information of the user's eyes in a three-dimensional space to movement information in a three-dimensional virtual space provided through the display (110). For example, the electronic device (100) may obtain first movement distance information and first movement direction information by scaling the second movement distance information and the second movement direction information to correspond to the three-dimensional virtual space. For example, the electronic device (100) may scale the second movement distance information to the first movement distance information based on a first scaling factor. For example, the electronic device (100) may scale the second movement direction information to the first movement direction information based on a second scaling factor. The scaling factor can be referred to by various terms such as scaling value, conversion factor, conversion value, adjustment factor, and adjustment value, but hereinafter it will be referred to as the scaling factor.
[0115] According to an example, the electronic device (100) may obtain a left-eye image and a right-eye image for providing a 3D image based on first movement distance information and first movement direction information corresponding to a three-dimensional space. For example, the electronic device (100) may obtain a left-eye image and a right-eye image in which the display position and depth of an object are changed based on the first movement distance information and the first movement direction information. For example, the electronic device (100) may obtain an output image by alternately arranging the left-eye image and the right-eye image in which the display position and depth of an object are changed by sub-sampling them by half in the horizontal direction. For example, the electronic device (100) may provide an output image (60) in which the left-eye image and right-eye images 1 and 2 are sequentially and repeatedly arranged, as illustrated in FIG. 3.
[0116] According to one embodiment, the electronic device (100) may generate an output image such that the viewpoint changes smoothly and continuously between consecutive frames through filtering. According to one example, the processor (140) may generate an output image such that the viewpoint changes smoothly between consecutive frames through an IIR (Infinite Impulse Response) filter or an FIR (Finite Impulse Response) filter. Accordingly, even if the user's position changes, a natural output image in which the depth difference between the binocular images is maintained can be obtained.
[0117] According to one embodiment, a 3D image provided by an electronic device (100) may include a plurality of objects located in a 3D virtual space.
[0118] For example, the electronic device (100) may identify a moving target object among a plurality of objects based on the current position of the user's eyes. For example, when the current position of the user's eyes is identified in the 3D space where the user is located, the electronic device (100) may map the current position of the user's eyes to a virtual 3D space provided by the display (110) and identify an object at a corresponding position as a moving target object.
[0119] In one example, the electronic device (100) may identify a target object to be moved among a plurality of objects based on a user selection command. For example, a target object to be moved may be identified based on the user's eye movements based on a user selection command such as at least one of a voice input, a remote control input, and a touch input.
[0120] Figure 4 illustrates an exemplary sequence of operations for convenience of explanation. However, if not required, the steps may be performed in parallel, and similar operations may not necessarily be limited to the sequence. Figure 5 is a diagram illustrating an example of a 3D image according to various embodiments.
[0121] As illustrated in FIG. 5, the 3D image may be a UI screen (510) including at least one object located in a three-dimensional virtual space. For example, the UI screen (510) may be a home UI screen including UI elements corresponding to multiple applications. For example, the UI screen (510) may be a home UI screen of a mode that provides 3D VR content, but is not limited thereto.
[0122] FIG. 6 is a diagram for explaining a user location tracking method according to various embodiments.
[0123] According to one embodiment, the electronic device (100) can track the position of the user's eyes based on a plurality of captured images acquired through a plurality of cameras. In one example, the electronic device (100) can track the position of the user's eyes based on a first captured image and a second captured image acquired through a plurality of cameras. In one example, one or more cameras (130) may include a plurality of cameras spaced apart at a preset interval and capturing different viewpoints. For example, the preset interval may be, but is not limited to, a distance equal to / similar to the distance between a person's two eyes.
[0124] Typically, a person receives two-dimensional images with left / right differences from both eyes, and can perceive three-dimensional distances through a process in which the human brain fuses the input images. According to one example, an electronic device (100) can track the position of a user's eyes in three-dimensional space using the same mechanism.
[0125] According to one example, the electronic device (100) can identify disarity information based on the first captured image and the second captured image, and track the position of the user's eyes in a three-dimensional space where the user is located based on the disarity information. For example, the electronic device (100) can track the position of the user's eyes using disarity, focal length, and baseline.
[0126] Parallax information refers to the difference in x-axis position for an object (e.g., the user's eyes) that is identically included in the left and right images, and can be, for example, the 'x' value in Fig. 6. Focal length can be the distance between the image plane (e.g., CCD, CMOS sensor) and the camera lens. Baseline can be the distance between the left camera and the right camera.
[0127] Figure 6 shows a state in which a point (x, y, z) in a three-dimensional space is captured on the image planes of the left and right cameras. The formula for three-dimensional distance information can be derived in the following order.
[0128] The (x, y, z) point in 3D space is imaged on the image plane through the centers of the lenses of the left and right cameras. The Y-axis is assumed to be the same for the left and right cameras. In this case, proportional equations such as the following mathematical equations 1 and 2 can be derived for the right-angled dashed line on the left and right of Fig. 6, respectively.
[0129]
[0130] Here, b can be the baseline and f can be the focal length. However, since the baseline and focal length are physical elements, they can be fixed constants.
[0131]
[0132] In this case, mathematical expression 3 can be derived from mathematical expressions 1 and 2.
[0133]
[0134] z represents the actual 3D distance, where d can be the parallax.
[0135] When the coordinates (x, y, z) in the three-dimensional space corresponding to the user's eyes are identified using the above-described method, the electronic device (100) can obtain movement information of the user's eyes based on the previous coordinates and current coordinates of the user's eyes.
[0136] Meanwhile, in FIG. 6, a method of obtaining user eye movement information using multiple cameras is described, but it is of course possible to obtain user eye movement information using a single camera, for example, a depth camera.
[0137] FIG. 7 is a flowchart for explaining a method of controlling an electronic device according to various embodiments.
[0138] According to FIG. 7, in operation 710, the electronic device (100) can display a 3D image including an object located in a 3D virtual space through the display (110).
[0139] In operation 720, the electronic device (100) can identify whether the position of at least one of the user's head or eyes changes from a first position to a second position in the captured image obtained through the camera (130). According to an example, the electronic device (100) can track the user's head or the user's eyes, but for the sake of convenience of explanation, it will be described below assuming that the user's eyes are tracked.
[0140] For example, the first and second locations may be arbitrary locations with different coordinate values in the three-dimensional space where the user is located. For example, the first and second locations may be acquired in the manner described in FIG. 6.
[0141] When the user's eyes are identified as having changed from the first position to the second position (S720:Y), in operation 730, the electronic device (100) can identify second movement distance information corresponding to the difference between the first position and the second position and second movement direction information from the first position to the second position.
[0142] According to an example, the electronic device (100) can identify second movement distance information based on a three-axis coordinate value corresponding to a first position and a three-axis coordinate value corresponding to a second position. For example, each of the three-axis coordinate value corresponding to the first position and the three-axis coordinate value corresponding to the second position may include an X coordinate, a Y coordinate, and a Z coordinate in an XYZ space.
[0143] According to an example, the electronic device (100) can identify second movement direction information based on a three-axis angular velocity value from a first position to a second position. For example, the three-axis angular velocity value can include a roll value, a pitch value, and a yaw value.
[0144] In operation 740, the electronic device (100) can scale the second movement distance information and the second movement direction information to correspond to a three-dimensional virtual space.
[0145] For example, the second movement distance information and the second movement direction information can be scaled to correspond to a three-dimensional virtual space to obtain the first movement distance information and the first movement direction information. For example, the electronic device (100) can scale the second movement distance information to the first movement distance information based on a first scaling factor. For example, the electronic device (100) can scale the second movement direction information to the first movement direction information based on a second scaling factor.
[0146] For example, the first scaling factor and the second scaling factor may be preset values based on the specifications of the electronic device (100). For example, the specifications of the electronic device (100) may include specifications of hardware elements provided in the electronic device (110), such as the performance and resolution of the display (110), the angle of view and arrangement position of the camera (130), etc.
[0147] In one example, at least one of the first scaling factor and the second scaling factor may be adjusted according to user settings.
[0148] For example, at least one of the first scaling factor and the second scaling factor may be adjusted based on the size of the space in which the electronic device (100) is located (e.g., the size of the space in front of the electronic device (100), the type of space (e.g., private space, public space), etc.
[0149] For example, at least one of the first scaling factor and the second scaling factor may be adjusted based on content characteristics such as the type of content, the three-dimensional information of the content, etc.
[0150] In operation 750, the electronic device (100) can change the display position of an object in a three-dimensional virtual space from a first display position to a second display position and change the depth of the object from the first depth to the second depth and display it based on the scaled first movement distance information and the first movement direction information.
[0151] According to an example, the electronic device (100) may obtain a second vector value based on the three-axis coordinate values and the three-axis angular velocity values in operation 740. For example, the second vector value may be a vector value in the form of a six-axis space matrix, but is not limited thereto. For example, the second vector value may be in the form of a rotation matrix R and a position vector d. The rotation matrix R represents a rotation transformation in a three-dimensional space and may transform a coordinate system around an origin according to the rotation. The position vector d may be a motion vector representing how much to move along each axis.
[0152] The second vector value obtained based on the three-axis coordinate value and the three-axis angular velocity value in operation 740 may be an example of second movement distance information and second movement direction information corresponding to the relative movement of at least one of the user's head or eyes.
[0153] For example, the electronic device (100) may obtain a first vector value by scaling a second vector value. The first vector value may be a vector value corresponding to the first movement distance information and the first movement direction information. For example, the electronic device (100) may obtain a second vector value corresponding to the movement of a three-dimensional virtual space of a 3D image provided through the electronic device (100) by scaling a second vector value corresponding to the movement of the user's eyes.
[0154] Figure 7 illustrates an exemplary sequence of operations for convenience of explanation. However, if not required, the order of steps may be performed in parallel, and similar operations may not necessarily be limited to this order. Figure 8 is a diagram illustrating a method for mapping user movement information and movement information in a three-dimensional virtual space according to various embodiments.
[0155] According to one embodiment, the electronic device (100) can control the display position and depth of an object included in a three-dimensional virtual space provided on the display (110) based on the user's movement information. The user's movement information may be information on the movement of the user's head or eyes identified in the three-dimensional space where the user is located. However, for convenience of explanation, it will be collectively referred to as "user's movement information" hereinafter.
[0156] For example, the user's movement information may be 6-axis spatial information. The 6-axis space may include not only X-coordinates, Y-coordinates, and Z-coordinates in a 3-dimensional space of the X-axis, Y-axis, and Z-axis, but also 3-axis angular velocity values. For example, the 3-axis angular velocity value may be a value that measures the rotational speed of the user's eyes with respect to the X-axis, Y-axis, and Z-axis, and may represent the speed at which the user's eyes rotate with respect to the X-axis, Y-axis, and Z-axis. For example, the 3-axis angular velocity value may include a roll value, a pitch value, and a yaw value. For example, the roll value may be a speed at which the user's eyes rotate with respect to the Z-axis, the pitch value may be a speed at which the user rotates with respect to the X-axis, and the yaw value may be a speed at which the user rotates with respect to the Y-axis.
[0157] According to an example, when the electronic device (100) identifies the user's movement information, the electronic device (100) can obtain control information for controlling the movement of an object included in a three-dimensional virtual space provided on the display (110) based on the identified movement information. For example, the electronic device (100) can convert a three-axis coordinate value and a three-axis angular velocity value corresponding to the user's movement into a first vector value. For example, when the user's eyes move from a first location to a second location in the three-dimensional space, the electronic device (100) can obtain a three-axis coordinate value and a three-axis angular velocity value corresponding to the movement of the user's eyes. In this case, since the three-axis angular velocity value is a value calculated based on the user's current location (e.g., the user's current eye location), the movement angle can be calculated based on the user or in the user's direction. In this case, since the object can be moved within the user's field of view, the three-dimensional effect of the object can be adjusted while maintaining X-talk performance.
[0158] For example, the electronic device (100) may obtain a first vector value for controlling movement of an object by applying a preset scaling factor to a second vector value. The preset scaling factor may be a constant value set during manufacturing, but is not limited thereto. For example, the scaling factor may be a value that can be set and / or changed by the user. For example, the scaling factor may be determined according to the type or property of a target object to be moved in a three-dimensional virtual space of a 3D image. For example, different scaling factors may be applied depending on whether the target element to be moved is a manual UI element or an input control element. For example, the manual UI element may include a UI element corresponding to an application, a UI element corresponding to a control menu, etc. For example, the input control element may be an input control element such as a user's head, a user's hand, a pointer, a cursor, etc.
[0159] FIG. 9 is a drawing for explaining a method for controlling a three-dimensional virtual space according to various embodiments.
[0160] According to one embodiment, the electronic device (100) can display a moving target object by moving it based on the user's movement information. For example, the electronic device (100) can display the object by changing the display position and depth of the object.
[0161] For example, a specific object, for example, an application UI element (511), may be displayed by changing the display position and depth of the object on the home UI screen (510) corresponding to the virtual three-dimensional space illustrated in FIG. 9. For example, the electronic device (100) may change the display position and depth of the object by mapping the relative movement information of the user's eyes to the virtual three-dimensional space. For example, when the depth is large, the left-right binocular disparity becomes large, so the three-dimensional effect is felt relatively large, and when the depth is small, the left-right binocular disparity becomes small, so the three-dimensional effect is felt relatively small. In this case, since the application UI element (511) can be moved within the user's field of vision, it is possible to adjust the three-dimensional effect of the object while maintaining the X-talk performance.
[0162] For example, the electronic device (100) may adjust the speed at which the display position and depth of a specific object (511) change based on the movement speed of the user's eyes. For example, the speed at which the display position and depth of a specific object (511) change may be adjusted in proportion to the movement speed of the user's eyes.
[0163] FIG. 10 is a drawing for explaining a method for controlling a three-dimensional virtual space according to various embodiments.
[0164] According to one embodiment, the electronic device (100) can display a moving target object by moving it based on the user's movement information. For example, the electronic device (100) can display the object by changing the display position and depth of the object.
[0165] For example, a display position and depth of a specific object, for example, an object of a pointer (1011), can be changed and displayed on a home UI screen (510) corresponding to a virtual three-dimensional space as illustrated in FIG. 10. For example, the electronic device (100) can change and display the display position and depth of an object by mapping relative movement information of the user's eyes to a virtual three-dimensional space.
[0166] For example, if a user wants to move a pointer (1011) to a position corresponding to an application UI element (511) within a three-dimensional virtual space corresponding to a home UI screen (510), the user can move his / her eyes based on the display position and depth of the application UI element (511). In this case, an object included in the three-dimensional virtual space within the user's field of vision can be dragged to a desired position and depth, thereby enabling adjustment of the three-dimensional effect of the object while maintaining X-talk performance.
[0167] According to one embodiment, the electronic device (100) can control the depth of a specific object to change smoothly by adjusting the viewpoint to change smoothly between consecutive frames using an IIR (Infinite Impulse Response) filter or an FIR (Finite Impulse Response) filter. For example, when the depth of a specific object is adjusted based on the movement of the user in a scene section including a plurality of frames, IIR can be applied to a plurality of frames included in each scene to obtain a binocular image in which the depth value of the specific object changes smoothly in the binocular image corresponding to each frame.
[0168] According to one embodiment, the electronic device (100) can acquire an image (e.g., a binocular image) in which the display position and depth of an object are adjusted using a learned artificial intelligence model. For example, the electronic device (100) can be trained to output a UI screen in which the display position and depth of an object are adjusted when information on the movement of a user's eyes or head and the current UI screen are input.
[0169] For example, when an artificial intelligence model is trained, it means that a basic artificial intelligence model (e.g., an artificial intelligence model including any random parameters) is trained using a learning algorithm using a plurality of training data, thereby creating a predefined operation rule or artificial intelligence model set to perform a desired characteristic (or purpose). Such learning may be performed through a separate server and / or system, but is not limited thereto, and may also be performed in the electronic device (100). Examples of learning algorithms include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.
[0170] For example, the learned artificial intelligence model may be an on-device model included in the electronic device (100), but is not limited thereto. For example, the learned artificial intelligence model may be implemented on a server.
[0171] Each operation according to the various embodiments described above may be performed by the processor (140), but if necessary, a module for each operation may be utilized. For example, each module may be implemented using at least one software, at least one hardware, and / or a combination thereof. Each module may be implemented to utilize a predefined algorithm, a predefined formula, and / or a learned artificial intelligence model to perform the operation. However, at least some modules may be distributed to an external device.
[0172] According to the various embodiments described above, it is possible to calculate the angle of movement in the direction of the user using eye tracking and / or head tracking. Accordingly, it is possible to move an object within the user's field of view, thereby adjusting the three-dimensionality of the object while maintaining crosstalk (X-talk) performance.
[0173] Meanwhile, the methods according to the various embodiments of the present disclosure described above can be implemented only with a software upgrade or a hardware upgrade for an existing electronic device and / or server.
[0174] Additionally, the various embodiments of the present disclosure described above can also be performed through an embedded server provided in an electronic device, or an external server of the electronic device.
[0175] Meanwhile, according to a temporary example of the present disclosure, the various embodiments described above can be implemented as software including instructions stored in a machine-readable storage medium that can be read by a machine (e.g., a computer). The device is a device that can call instructions stored from the storage medium and operate according to the called instructions, and may include an electronic device (e.g., electronic device (A)) according to the disclosed embodiments. When an instruction is executed by a processor, the processor can perform a function corresponding to the instruction directly or by using other components under the control of the processor. The instruction may include code generated or executed by a compiler or interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. 'Non-transitory' means that the storage medium does not contain signals and can be tangible, and does not distinguish whether data is stored semi-permanently or temporarily in the storage medium.
[0176] Furthermore, according to one embodiment of the present disclosure, the method according to the various embodiments described above may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or online through an application store (e.g., Play Store™). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0177] In addition, each of the components (e.g., modules or programs) according to the various embodiments described above may be composed of a single or multiple entities, and some of the corresponding sub-components described above may be omitted, or other sub-components may be further included in various embodiments. Alternatively or additionally, some components (e.g., modules or programs) may be integrated into a single entity, which may perform the same or similar functions as those performed by each of the corresponding components prior to integration. Operations performed by modules, programs or other components according to various embodiments may be executed sequentially, in parallel, iteratively or heuristically, or at least some operations may be executed in a different order, omitted, or other operations may be added.
[0178] Although various embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above, and various modifications may be made by those skilled in the art to which the present disclosure pertains without departing from the gist of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the present disclosure. In addition, it should be understood that all embodiments described herein can be used in combination with all other embodiments described herein.
Claims
1. In electronic devices, display; At least one camera; Memory that stores instructions; At least one processor comprising a processing circuit; and The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Displaying a 3D image including an object located in a 3D virtual space through the display, Based on the captured image acquired through the camera, movement information corresponding to at least one of the user's head or eyes in the three-dimensional space where the user is located is identified, Based on the above movement information, position movement information corresponding to the three-dimensional virtual space is identified, An electronic device that controls the display to change the display position and depth of the object within the three-dimensional virtual space included in the 3D image based on the position movement information.
2. In paragraph 1, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to identify first movement distance information and first movement direction information corresponding to the position movement information, An electronic device that controls the display to change the display position and depth of the object within the three-dimensional virtual space based on the first movement distance information and the first movement direction information.
3. In paragraph 2, The instructions, when individually or collectively executed by the at least one processor, cause the electronic device to, when the position of at least one of the head or eyes of the user changes from a first position to a second position in the captured image acquired through the camera, identify second movement distance information corresponding to a difference between the first position and the second position and second movement direction information from the first position to the second position, An electronic device that identifies the first movement distance information and the first movement direction information based on the second movement distance information and the second movement direction information.
4. In paragraph 3, An electronic device wherein the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to identify the first movement distance information and the first movement direction information by scaling the second movement distance information and the second movement direction information to correspond to the three-dimensional virtual space.
5. In paragraph 3, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to identify the second movement distance information based on a three-axis coordinate value corresponding to the first position and a three-axis coordinate value corresponding to the second position, An electronic device that identifies the second movement direction information based on a three-axis angular velocity value from the first location to the second location.
6. In paragraph 5, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to obtain a vector value based on the three-axis coordinate values and the three-axis angular velocity values, An electronic device that identifies the second movement distance information and the second movement direction information corresponding to the relative movement of at least one of the user's head or eyes based on the acquired vector value.
7. In paragraph 5, Each of the three-axis coordinate values corresponding to the first position and the three-axis coordinate values corresponding to the second position includes an X-coordinate, a Y-coordinate, and a Z-coordinate in the XYZ space, The above three-axis angular velocity values are, An electronic device comprising roll values, pitch values, and yaw values.
8. In paragraph 1, The above 3D image is, Contains a plurality of objects located in the above three-dimensional virtual space, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to identify a moving target object among the plurality of objects based on a current position of at least one of the user's head or eyes, or An electronic device that identifies a moving target object among the plurality of objects based on a user selection command.
9. In paragraph 1, The above camera comprises a plurality of cameras spaced apart by a specific distance, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Disarity information is identified based on the first and second captured images acquired through the above multiple cameras, An electronic device that identifies movement information corresponding to at least one of the user's head or eyes in a three-dimensional space where the user is located based on the parallax information.
10. In paragraph 1, The above display is, An electronic device including a Light Field Display (LFD).
11. In a method for controlling an electronic device, A step of displaying a 3D image including an object located in a 3D virtual space; A step of identifying movement information corresponding to at least one of the user's head or eyes in a three-dimensional space where the user is located based on a captured image acquired through a camera; A step of identifying position movement information corresponding to the three-dimensional virtual space based on the movement information; and A control method comprising: a step of changing and displaying the display position and depth of the object within the three-dimensional virtual space included in the 3D image based on the position movement information.
12. In paragraph 11, The steps for changing the display position and depth of the above object are as follows: A step of identifying first movement distance information and first movement direction information corresponding to the above location movement information; and A control method, comprising: a step of changing and displaying the display position and depth of the object within the three-dimensional virtual space based on the first movement distance information and the first movement direction information.
13. In paragraph 12, The step of identifying the position movement information corresponding to the above three-dimensional virtual space is: A step of identifying second movement distance information corresponding to a difference between the first position and the second position and second movement direction information from the first position to the second position when the position of at least one of the head or eyes of the user is changed from the first position to the second position in the captured image acquired through the camera; The steps for changing the display position and depth of the above object are as follows: A control method, comprising: a step of identifying the first movement distance information and the first movement direction information based on the second movement distance information and the second movement direction information.
14. In paragraph 13, The steps for changing the display position and depth of the above object are as follows: A control method further comprising: a step of identifying the first movement distance information and the first movement direction information by scaling the second movement distance information and the second movement direction information to correspond to the three-dimensional virtual space; 15. A non-transitory computer-readable medium storing computer instructions that, when executed individually or collectively by a processor including a processing circuit of an electronic device, cause the electronic device to perform an operation, The above actions are, A step of displaying a 3D image including an object located in a 3D virtual space; A step of identifying movement information corresponding to at least one of the user's head or eyes in a three-dimensional space where the user is located based on a captured image acquired through a camera; A step of identifying position movement information corresponding to the three-dimensional virtual space based on the movement information; and A non-transitory computer-readable medium, comprising: a step of changing and displaying the display position and depth of the object within the three-dimensional virtual space included in the 3D image based on the position movement information;
Citation Information
Patent Citations
Wearable display device and method of controlling layer
KR1020150037254A
Precursor metal-silicon containing thin film, deposition method of thin film using the same, and semiconductor device comprising the same
KR1020250017913A
Stereoscopic display apparatus, and display method thereof
KR102070800B1
KR20190130770A
KR20210100690A