Electronic device and control method therefor
The described electronic device uses AI-driven user detection and gaze analysis to dynamically control 3D displays, addressing the challenge of suboptimal mode switching by ensuring seamless adaptation to user presence and content type, thereby enhancing viewing comfort and accuracy.
Patent Information
- Application Number
- PCT/KR2024/019986
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-16
- Filing Date
- 2024-12-06
- Publication Date
- 2025-07-17
AI Technical Summary
Existing 3D display technologies struggle to dynamically adapt to user presence and gaze direction, often resulting in suboptimal viewing experiences due to unnecessary conversion between 2D and 3D modes.
An electronic device equipped with cameras and processors that utilize artificial intelligence models to detect user presence and gaze direction, dynamically controlling the display to operate in 3D mode when the user is present and looking at the screen, and switching to 2D mode otherwise, while also identifying stereoscopic content sections for accurate mode switching.
Enhances user convenience by automatically adapting the display mode to user interaction, ensuring optimal viewing experiences by minimizing discomfort from inappropriate 2D-3D conversions and improving content accuracy.
Smart Images

Figure KR2024019986_17072025_PF_FP_ABST
Abstract
Description
Electronic device and method of controlling the same
[0001] The present disclosure relates to an electronic device and a method for controlling the same, and more particularly, to an electronic device providing a 3D image and a method for controlling the same.
[0002] Advances in electronic technology have led to the development and proliferation of various types of electronic devices. In particular, display devices, used in a variety of settings, including homes, offices, and public spaces, have been continuously evolving in recent years.
[0003] Stereoscopy refers to three-dimensional technology. Recently, commercialized 3D displays primarily utilize binocular parallax. Binocular parallax offers the advantage of creating a three-dimensional effect on a single screen, such as a TV or theater screen. Methods utilizing binocular parallax can be categorized into stereoscopic (using glasses or other auxiliary devices) and autostereocopic (glassless) methods.
[0004] Recently, commercialization of glasses-free light field displays and glasses-free 3D displays utilizing eye-tracking is being continuously researched.
[0005] Aspects of embodiments of the present disclosure will be set forth in part in the description that follows, and in part will be apparent from the description or may be learned by practice of the embodiments presented.
[0006] An electronic device according to one or more embodiments includes a display configured to operate in one of a 3D mode and a 2D mode; one or more cameras configured to capture an image in front of the display; a memory configured to store one or more commands; and one or more processors. The one or more processors, by executing the one or more commands, identify whether a user is positioned in front of the display based on the captured image, and if it is determined that the user is positioned in front of the display, identify whether the user's gaze is directed toward the front of the display, and if it is determined that the user's gaze is directed toward the front of the display, control the display to operate in the 3D mode, and if it is determined that the user is not positioned in front of the display or the user's gaze is not directed toward the front of the display, control the display to operate in the 2D mode.
[0007] According to one or more embodiments, the one or more processors, by executing the one or more commands, when the user's gaze is identified as being directed toward the front of the display, identify a probability of stereoscopic content for each content section of the content to be displayed, control the display to operate in the 3D mode in a content section in which the probability of the stereoscopic content is greater than or equal to a threshold value, and control the display to operate in the 2D mode in a content section in which the probability of the stereoscopic content is less than or equal to the threshold value.
[0008] According to one or more embodiments, the one or more processors can input content into a learned artificial intelligence model for each content section by executing the one or more commands and identify a probability that the content is stereoscopic content based on information output from the learned artificial intelligence model.
[0009] According to one or more embodiments, the one or more processors can input the captured image into a learned artificial intelligence model by executing the one or more commands and identify whether the user's gaze is directed toward the front of the display based on information output from the learned artificial intelligence model.
[0010] According to one or more embodiments, the content may include side-by-side content. The one or more processors may, by executing the one or more commands, identify whether the similarity between the left and right regions of the side-by-side content for each content section is a threshold value, control the display to operate in the 3D mode in a content section where the similarity between the left and right regions of the side-by-side content is greater than or equal to the threshold value, and control the display to operate in the 2D mode in a content section where the similarity between the left and right regions of the side-by-side content is less than the threshold value.
[0011] According to one or more embodiments, the one or more processors can control the display to switch the 3D mode to the 2D mode when the user is not positioned in front of the display or the user's gaze is not directed toward the front of the display while the display is operating in the 3D mode by executing the one or more instructions.
[0012] According to one or more embodiments, the one or more processors, by executing the one or more commands, when the user's gaze is identified as being directed toward the front of the display while the display is operating in the 2D mode, identify whether the content is stereoscopic content for each content section, and control the display to operate in the 3D mode if the content is the stereoscopic content.
[0013] According to one or more embodiments, the one or more processors may, by executing the one or more commands, identify a probability of the content being stereoscopic content for each content section of the content, identify whether the user's gaze is directed toward the front of the display based on the captured image in a content section in which the probability of the content being stereoscopic content is greater than or equal to a threshold value, and control the display to operate in the 3D mode when the user's gaze is directed toward the front of the display and the probability of the content being stereoscopic content is greater than or equal to the threshold value, and control the display to operate in the 2D mode in a content section in which the probability of the content being stereoscopic content is less than or equal to the threshold value.
[0014] According to one or more embodiments, the one or more processors may identify that the user is positioned in front of the display when, by executing the one or more commands, a specific body of the user is included in the captured image or an area of the user's body greater than a preset ratio is identified as being included in the captured image.
[0015] According to one or more embodiments, the display may be implemented as a light field display including a lenticular lens array. The one or more processors may control the display to operate in the 3D mode or the 2D mode by adjusting a voltage applied to the lenticular lens array included in the light field display.
[0016] A method for controlling an electronic device including a display implemented to be operable in one of a 3D mode and a 2D mode according to one or more embodiments and one or more cameras for capturing an image in front of the display may include: identifying whether a user is positioned in front of the display based on captured images acquired through the one or more cameras; if it is determined that the user is positioned in front of the display, identifying whether the user's gaze is directed toward the front of the display based on the captured images; if it is determined that the user's gaze is directed toward the front of the display, controlling the display to operate in the 3D mode; and, if it is determined that the user is not positioned in front of the display or the user's gaze is not directed toward the front of the display, controlling the display to operate in the 2D mode.
[0017] A non-transitory computer-readable medium storing computer instructions that, when executed by a processor of an electronic device including a display configured to operate in one of a 3D mode and a 2D mode according to one or more embodiments and one or more cameras for capturing an image in front of the display, cause the electronic device to perform an operation, the operation may include: identifying whether a user is positioned in front of the display based on captured images acquired through the one or more cameras; if the user is identified as being positioned in front of the display, identifying whether the user's gaze is directed toward the front of the display based on the captured images; if the user's gaze is identified as being directed toward the front of the display, controlling the display to operate in the 3D mode; and if the user is identified as not being positioned in front of the display or the user's gaze is not directed toward the front of the display, controlling the display to operate in the 2D mode.
[0018] The above and other aspects, features and advantages of specific embodiments of the present disclosure will become more apparent from the following description taken in conjunction with the accompanying drawings.
[0019] FIG. 1 is a diagram illustrating the operation of an electronic device according to one or more embodiments.
[0020] FIG. 2A is a block diagram showing the configuration of an electronic device according to an embodiment of the present disclosure.
[0021] FIG. 2b is a block diagram specifically illustrating a configuration of an electronic device according to one or more embodiments of the present disclosure.
[0022] FIG. 3A is a drawing for explaining the structure and operation of a display according to one or more embodiments of the present disclosure.
[0023] FIG. 3b is a drawing for explaining the structure and operation of a display according to one or more embodiments of the present disclosure.
[0024] FIG. 4 is a flowchart illustrating a method for controlling an electronic device according to one or more embodiments of the present disclosure.
[0025] FIG. 5A is a diagram illustrating a user identification method according to one or more embodiments of the present disclosure.
[0026] FIG. 5b is a diagram illustrating a user identification method according to one or more embodiments of the present disclosure.
[0027] FIG. 6 is a diagram illustrating a method for obtaining user gaze information according to one or more embodiments of the present disclosure.
[0028] FIG. 7 is a diagram illustrating a method for obtaining user gaze information according to one or more embodiments of the present disclosure.
[0029] FIG. 8 is a flowchart illustrating a method for controlling an electronic device according to one or more embodiments of the present disclosure.
[0030] FIG. 9A is a diagram illustrating a method for identifying stereoscopic content according to one or more embodiments of the present disclosure.
[0031] FIG. 9b is a diagram illustrating a method for identifying stereoscopic content according to one or more embodiments of the present disclosure.
[0032] FIG. 10 is a flowchart illustrating a method for controlling an electronic device according to one or more embodiments of the present disclosure.
[0033] FIG. 11A is a diagram illustrating a method for identifying stereoscopic content according to one or more embodiments of the present disclosure.
[0034] FIG. 11B is a diagram illustrating a method for identifying stereoscopic content according to one or more embodiments of the present disclosure.
[0035] FIG. 12 is a flowchart illustrating a method of controlling an electronic device according to one or more embodiments.
[0036] The terms used in this specification will be briefly explained, and the present disclosure will be described in detail.
[0037] The terms used in the embodiments of this disclosure have been selected from widely used, current terms, taking into account the functions of this disclosure. However, these terms may vary depending on the intentions or cases of those skilled in the art, the emergence of new technologies, etc. Furthermore, in certain cases, terms may be arbitrarily selected by the applicant, and in such cases, their meanings will be described in detail in the description of the relevant disclosure. Therefore, the terms used in this disclosure should not be defined simply as names of terms, but rather based on the meanings of the terms and the overall content of this disclosure.
[0038] In this specification, expressions such as “has,” “can have,” “includes,” or “may include” indicate the presence of a feature (e.g., a number, function, operation, or component such as a part), and do not exclude the presence of additional features.
[0039] In this disclosure, expressions such as “A or B,” “at least one of A and / or B,” or “one or more of A or / and B” can include all possible combinations of the listed items. For example, “A or B,” “at least one of A and B,” or “at least one of A or B” can all refer to cases where (1) only A is included, (2) only B is included, or (3) both A and B are included.
[0040] As used herein, the expressions “first,” “second,” “first,” or “second,” etc., may describe various components, regardless of order and / or importance, and are only used to distinguish one component from another, but do not limit the components.
[0041] When it is said that a component (e.g., a first component) is “operatively or communicatively coupled with / to” or “connected to” another component (e.g., a second component), it should be understood that the component may be directly coupled to the other component, or may be connected through another component (e.g., a third component).
[0042] The expression "configured to" as used in the present disclosure may be used interchangeably with, for example, "suitable for," "having the capacity to," "designed to," "adapted to," "made to," or "capable of." The term "configured to" may not necessarily mean only "specifically designed to" in terms of hardware.
[0043] In some contexts, the phrase "a device configured to" may mean that the device, in conjunction with other devices or components, is "capable of" performing A, B, and C. For example, the phrase "a processor configured (or set) to perform A, B, and C" may refer to a dedicated processor (e.g., an embedded processor) for performing those operations, or a general-purpose processor (e.g., a CPU or application processor) that can perform those operations by executing one or more software programs stored in a memory device.
[0044] Singular expressions include plural expressions unless the context clearly dictates otherwise. In this application, terms such as "comprise" or "consist of" are intended to indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but should be understood not to preclude the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0045] In the embodiments, a "module" or "part" performs at least one function or operation and may be implemented as hardware or software, or as a combination of hardware and software. Furthermore, a plurality of "modules" or "parts" may be integrated into at least one module and implemented as at least one processor (not shown), excluding any "module" or "part" that needs to be implemented as specific hardware.
[0046] Meanwhile, the various elements and areas in the drawings are schematically drawn. Therefore, the technical concept of the present invention is not limited by the relative sizes or spacing depicted in the attached drawings.
[0047] An embodiment of the present disclosure will be described in more detail with reference to the attached drawings below.
[0048] FIG. 1 is a diagram illustrating the operation of an electronic device according to one or more embodiments.
[0049] The electronic device (100) may be implemented as various types of display devices such as a TV, monitor, kiosk, tablet PC, electronic picture frame, mobile phone, large format display (LFD), digital signage, digital information display (DID), video wall, projector display, etc. However, in some cases, it may be implemented as an image processing device (e.g., set-top box, one connected box) that is connected to the display device and provides an image.
[0050] According to one embodiment, the electronic device (100) may be equipped with a light field display. A light field display is a display technology that provides a more realistic visual experience by expressing light field information, unlike conventional 2D or 3D displays.
[0051] While 2D or 3D displays typically provide limited information about light direction and depth, light field displays can provide visual experiences similar to those observed in the real world by expressing additional information about light direction and depth using light field information. For example, light field displays can be utilized to provide more realistic environments in virtual reality (VR) and / or augmented reality (AR) devices.
[0052] Fig. 1 is a diagram for explaining the operation of a light field display using a lenticular lens method. According to Fig. 1, each lenticular lens, for example, a micro lens array, is assigned a series of display pixels, and light from each pixel is directed in a specific direction by the lens to form a light field expressed in terms of light intensity and direction. When the display is viewed within the light field thus formed, the user can experience a three-dimensional effect.
[0053] According to one embodiment, the electronic device (100) may operate in a 3D mode providing a 3D image or in a 2D mode providing a 2D image depending on the presence of a user in front, the context of the electronic device (100), and / or the context of the user.
[0054] Below, various embodiments for performing switching between 3D mode and 2D mode depending on the context of the electronic device (100) and / or the context of the user will be described.
[0055] FIG. 2A is a block diagram showing the configuration of an electronic device according to one embodiment.
[0056] According to FIG. 2a, the electronic device (100) includes a display (110), a memory (120), a camera (130), and one or more processors (140).
[0057] The display (110) may be implemented as a display including a self-luminous element or a display including a non-luminous element and a backlight. For example, it may be implemented as various types of displays such as an LCD (Liquid Crystal Display), an OLED (Organic Light Emitting Diodes) display, an LED (Light Emitting Diodes), a micro LED, a Mini LED, a PDP (Plasma Display Panel), a QD (Quantum dot) display, a QLED (Quantum dot light-emitting diodes), etc. The display (110) may also include a driving circuit, a backlight unit, etc., which may be implemented in a form such as an a-si TFT, an LTPS (low temperature poly silicon) TFT, an OTFT (organic TFT), etc. According to one example, a touch sensor that detects a touch operation in the form of a touch film, a touch sheet, a touch pad, etc. may be disposed on the front of the display (110) so as to be implemented so as to detect various types of touch inputs. For example, the display (110) can detect various types of touch inputs, such as a touch input by a user's hand, a touch input by an input device such as a stylus pen, and a touch input by a specific electrostatic material. Here, the input device can be implemented as a pen-type input device that can be referred to by various terms such as an electronic pen, a stylus pen, an S-pen, etc. According to an example, the display (110) can be implemented as a flat display, a curved display, a flexible display that can be folded or / and rolled, etc.
[0058] The memory (120) can store data required for various embodiments. The memory (120) may be implemented in the form of memory embedded in the electronic device (100') or in the form of memory that can be detachably attached to the electronic device (100) depending on the purpose of data storage. For example, data for driving the electronic device (100) may be stored in a memory embedded in the electronic device (100'), and data for expanding the functions of the electronic device (100) may be stored in a memory that can be detachably attached to the electronic device (100). Meanwhile, in the case of memory embedded in the electronic device (100), it may be implemented as at least one of volatile memory (e.g., dynamic RAM (DRAM), static RAM (SRAM), or synchronous dynamic RAM (SDRAM)), non-volatile memory (e.g., one time programmable ROM (OTPROM), programmable ROM (PROM), erasable and programmable ROM (EPROM), electrically erasable and programmable ROM (EEPROM), mask ROM, flash ROM, flash memory (e.g., NAND flash or NOR flash), hard drive, or solid state drive (SSD). In addition, in the case of memory that can be attached or detached to the electronic device (100'), it may be implemented as at least one of memory cards (e.g., compact flash (CF), secure digital (SD), micro secure digital (Micro-SD), mini secure digital (Mini-SD), extreme digital (xD), multi-media card (MMC), etc.), external memory that can be connected to a USB port (e.g., USB memory), etc. It can be implemented in the form of.
[0059] As an example, the memory (120) may store a computer program including at least one instruction or instructions for controlling the electronic device (100).
[0060] According to another example, the memory (120) may store an image received from an external device (e.g., a source device), an external storage medium (e.g., USB), an external server (e.g., web hard), etc., i.e., an input image. Alternatively, the memory (120) may store an image acquired through a camera provided in the electronic device (100).
[0061] According to another example, the memory (120) can store various information required for image quality processing, for example, information, algorithms, image quality parameters, etc. for performing at least one of Noise Reduction, Detail Enhancement, Tone Mapping, Contrast Enhancement, Color Enhancement, or Frame Rate Conversion.
[0062] According to one embodiment, the memory (120) may be implemented as a single memory that stores data generated from various operations according to the present disclosure. However, according to another embodiment, the memory (120) may be implemented to include multiple memories that each store different types of data or each store data generated at different stages.
[0063] In the above-described embodiment, it has been described that various data are stored in the external memory (120) of the processor (140), but at least some of the above-described data may be stored in the internal memory of the processor (140) depending on an implementation example of at least one of the electronic device (100) or the processor (140).
[0064] One or more cameras (130) can be turned on and perform shooting according to a preset event. For example, one or more cameras (130) can perform shooting according to an event in which the electronic device (100) (or the display (110)) is turned on. The camera (130) can convert a captured image into an electrical signal and generate image data based on the converted signal. For example, a subject can be converted into an electrical image signal through a semiconductor optical element (CCD; Charge Coupled Device), and the converted image signal can be amplified and converted into a digital signal and then signal processed. For example, the one or more cameras (130) can include at least one of a general (or basic) camera, an ultra-wide-angle camera, and a depth camera.
[0065] In one example, one or more cameras (130) may be positioned to capture the front of the display (110). For example, one or more cameras (130) may be positioned in the central area of the top bezel of the display (110).
[0066] In one example, one or more cameras (130) may be positioned in a direction and angle that can capture the front of the display (110). In one example, the cameras (130) may be positioned in a direction and angle that can be recognized as facing the front of the display (110) when the user's gaze is directed forward in the captured image.
[0067] For example, one or more cameras (130) may include multiple cameras spaced apart at preset intervals to capture different viewpoints. For example, the preset interval may be, but is not limited to, the same / similar distance as the distance between a person's two eyes.
[0068] One or more processors (140) control the overall operation of the electronic device (100). Specifically, one or more processors (140) may be connected to each component of the electronic device (100) to control the overall operation of the electronic device (100). For example, one or more processors (140) may be electrically connected to the display (110) and the memory (110) to control the overall operation of the electronic device (100). One or more processors (140) may be configured as one or more processors.
[0069] One or more processors (140) may perform operations of the electronic device (100) according to various embodiments by executing at least one instruction stored in the memory (110).
[0070] In one example, the artificial intelligence related functions according to the present disclosure may be operated through the processor and memory of an electronic device.
[0071] One or more processors (140) may be composed of one or more processors. In this case, the one or more processors may include at least one of a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and an NPU (Neural Processing Unit), but are not limited to the examples of the processors described above.
[0072] CPUs are general-purpose processors capable of performing not only general calculations but also artificial intelligence calculations. Their multi-layered cache structure allows for the efficient execution of complex programs. CPUs are advantageous for serial processing, enabling organic linking of previous and subsequent calculation results through sequential calculations. General-purpose processors are not limited to the examples described above, except where specifically identified as CPUs.
[0073] A GPU is a processor designed for large-scale computations, such as floating-point operations used in graphics processing. It integrates a large number of cores to perform large-scale computations in parallel. In particular, GPUs may be advantageous over CPUs in parallel processing methods, such as convolution operations. Furthermore, GPUs can be used as coprocessors to supplement the functions of CPUs. Processors for large-scale computations are not limited to the examples described above, except in cases where they are specifically referred to as GPUs.
[0074] An NPU is a processor specialized in artificial intelligence computation using artificial neural networks, and each layer of the artificial neural network can be implemented in hardware (e.g., silicon). Since NPUs are designed specifically according to the company's specifications, they have less freedom than CPUs or GPUs, but can efficiently process the AI computations requested by the company. Meanwhile, as a processor specialized in AI computation, an NPU can be implemented in various forms, such as a Tensor Processing Unit (TPU), an Intelligence Processing Unit (IPU), or a Vision Processing Unit (VPU). Except as specifically designated as an NPU, an AI processor is not limited to the examples described above.
[0075] Additionally, one or more processors (140) may be implemented as a System on Chip (SoC). In this case, the SoC may further include, in addition to one or more processors (140), a memory (110), and a network interface such as a bus for data communication between the processor (140) and the memory (110).
[0076] When a plurality of processors are included in a SoC (System on Chip) included in an electronic device (100), the electronic device (100) may perform operations related to artificial intelligence (e.g., operations related to learning or inference of a neural network model) by using some of the plurality of processors. For example, the electronic device may perform operations related to artificial intelligence by using at least one of a GPU, NPU, VPU, TPU, or hardware accelerator specialized in artificial intelligence operations such as convolution operations or matrix multiplication operations among the plurality of processors. However, this is only one or more embodiments, and it is of course possible to process operations related to artificial intelligence by using a CPU or a general-purpose processor.
[0077] Additionally, the electronic device (100) can perform operations related to functions related to artificial intelligence by utilizing multiple cores (e.g., dual cores, quad cores, etc.) included in a single processor. In particular, the electronic device can perform artificial intelligence operations, such as convolution operations and matrix multiplication operations, in parallel by utilizing multiple cores included in the processor.
[0078] One or more processors (140) are controlled to process input data according to predefined operating rules or neural network models (or artificial intelligence models) stored in memory (110). The predefined operating rules or neural network models are characterized by being created through learning.
[0079] Here, "created through learning" means that a predefined set of behavioral rules or a neural network model with desired characteristics is created by applying a learning algorithm to a large number of learning data. This learning may be performed on the device itself, where the artificial intelligence according to the present disclosure is implemented, or through a separate server / system.
[0080] A neural network model may be composed of multiple neural network layers. At least one layer has at least one weight value and performs its operation through the operation result of the previous layer and at least one defined operation. Examples of neural networks include a convolutional neural network (CNN), a recurrent neural network (RNN), a deep neural network (DNN), a restricted boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), deep Q-networks, and a transformer. The neural networks in the present disclosure are not limited to the above-described examples unless otherwise specified.
[0081] A learning algorithm is a method for training a target device (e.g., a robot) using a plurality of learning data sets so that the target device can make decisions or predictions on its own. Examples of learning algorithms include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning. The learning algorithms in the present disclosure are not limited to the aforementioned examples unless otherwise specified. For convenience of explanation, one or more processors (140) will be referred to as "processors (140)" below.
[0082] According to one embodiment, the electronic device (100) can receive various compressed images or images of various resolutions. For example, the electronic device (100) can receive images in a compressed form such as MPEG (Moving Picture Experts Group) (e.g., MP2, MP4, MP7, etc.), JPEG (joint photographic coding experts group), AVC (Advanced Video Coding), H.264, H.265, HEVC (High Efficiency Video Codec), etc. Alternatively, the electronic device (100) can receive any one of SD (Standard Definition), HD (High Definition), Full HD, and Ultra HD images.
[0083] For example, the processor (140) may process an input image to obtain an output image. Here, the image processing may include at least one of image enhancement, image restoration, image transformation, image analysis, image understanding, image compression, image decoding, or scaling.
[0084] In this specification, "region" is a term referring to a portion of an image, and means at least one pixel block or a set of pixel blocks. In addition, a pixel block means a set of adjacent pixels that include at least one pixel.
[0085] For example, the input image may include a 3D image. For example, the input image may include a side-by-side image. A side-by-side image may be an image in which two images are placed side by side on a single screen. For example, each image may occupy half of the horizontal space of the screen, with one image located in the left area and the other in the right area. For example, the image located in the left area may be a left-eye image, and the image located in the right area may be a right-eye image.
[0086] For example, when a plurality of frames included in an input image are sequentially input, the processor (140) may store the plurality of frames in the memory (120) and read the frames stored in the memory (120) to perform various processing. A frame is a basic image unit in image content, and each frame is composed of pixels and may include resolution and color information. Hereinafter, content may be one frame included in the image content or a preset number of multiple frames, but for the convenience of explanation, they will be collectively referred to as content.
[0087] FIG. 2b is a block diagram specifically illustrating a configuration of an electronic device according to one or more embodiments.
[0088] According to FIG. 2b, the electronic device (100') may include a display (110), a memory (120), a camera (130), one or more processors (140), a user interface (150), a communication interface (160), and a speaker (170). Among the configurations illustrated in FIG. 2b, a detailed description of configurations that overlap with those illustrated in FIG. 2a will be omitted.
[0089] The user interface (150) may be implemented as a device such as a button, a touch pad, a mouse, and a keyboard, or as a touch screen that can also perform the display function and operation input function described above.
[0090] It goes without saying that the communication interface (160) can be implemented as various interfaces depending on the implementation example of the electronic device (100'). For example, the communication interface (140) can communicate with an external device, an external storage medium (e.g., a USB memory), an external server (e.g., a web hard drive), etc. through a communication method such as Bluetooth, AP-based Wi-Fi (Wireless LAN network), Zigbee, wired / wireless LAN (Local Area Network), WAN (Wide Area Network), Ethernet, IEEE 1394, HDMI (High-Definition Multimedia Interface), USB (Universal Serial Bus), MHL (Mobile High-Definition Link), AES / EBU (Audio Engineering Society / European Broadcasting Union), optical, coaxial, etc. In one example, the communication interface (160) can communicate with another electronic device, an external server, and / or a remote control device.
[0091] The speaker (170) may be configured to output various audio data as well as various notification sounds or voice messages. The processor (130) may control the speaker (170) to output feedback or various notifications in audio format according to various embodiments of the present disclosure.
[0092] In addition, the electronic device (100') may include sensors, microphones, etc., depending on the implementation example.
[0093] Sensors may include various types of sensors, such as touch sensors, proximity sensors, acceleration sensors (or gravity sensors), geomagnetic sensors, gyro sensors, pressure sensors, position sensors, distance sensors, light sensors, etc.
[0094] A microphone is a device configured to receive user voice or other sounds and convert them into audio data. However, according to another embodiment, the electronic device (100') may receive user voice input via an external device through a communication interface (160).
[0095] FIGS. 3A and 3B are drawings for explaining the structure and operation of a display (110) according to one or more embodiments.
[0096] According to FIG. 3A, the display (110) may include a display panel (111), a field separation unit (112), and a backlight unit (113). However, depending on the implementation example of the display (110), the backlight unit (113) may not be included in the display (110).
[0097] The display panel (111) includes a plurality of pixels composed of a plurality of sub-pixels. Here, the sub-pixels may be composed of R (Red), G (Green), and B (Blue). For example, pixels composed of R, G, and B sub-pixels may be arranged in a plurality of row and column directions to form the display panel (141).
[0098] The display panel (111) displays a binocular viewpoint image (or a multi-viewpoint image). For example, the display panel (111) can display an image in which multiple images of a right-eye image and a left-eye image are sequentially and repeatedly arranged.
[0099] The viewing area separator (112) is arranged on the front of the display panel (111) to provide different viewpoints, i.e., multi-views, for each viewing area. In this case, the viewing area separator (112) may be implemented as a lenticular lens or a parallax barrier. For example, the viewing area separator (112) may be implemented as a lenticular lens including a plurality of lens regions. Accordingly, the lenticular lens may refract an image displayed on the display panel (111) through the plurality of lens regions. Each lens region is formed to have a size corresponding to at least one pixel, so that light passing through each pixel may be differently dispersed for each viewing area. As another example, the viewing area separator (112) may be implemented as a parallax barrier. The parallax barrier is implemented as a transparent slit array including a plurality of barrier regions. Accordingly, light can be blocked through slits between barrier areas to allow images from different viewpoints to be output for each viewing area.
[0100] For example, the field of view separator (112) may be implemented to operate at a constant angle to improve image quality, i.e., to avoid resolution reduction. In this case, the processor (130) may divide the right-eye image and the left-eye image based on the angle at which the field of view separator (112) is tilted, and combine them to generate a multi-view image. Accordingly, the user does not view the image displayed vertically or horizontally on the sub-pixels of the display panel (111), but rather views the image displayed so that the sub-pixels have a constant angle of inclination.
[0101] For example, the field of view separator (112) may be implemented as an active type so that the display (110) can operate in 3D mode or 2D mode. For example, the field of view separator (112) may be implemented as an active lenticular lens or an active parallax barrier.
[0102] As an example, the field separation unit (112) can be implemented as a lenticular lens array as shown in FIG. 3b.
[0103] As illustrated in FIG. 3B, when the display (110) operates in 3D mode, the processor (140) can control the lenticular lens (112) to operate in 3D mode by applying a preset voltage corresponding to the 3D mode to the active lenticular lens (112). In addition, when the display (110) operates in 2D mode, the processor (140) can control the lenticular lens (112) to operate in 2D mode by applying a preset voltage corresponding to the 2D mode to the active lenticular lens (112).
[0104] According to an example, the lenticular lens (112) includes a plurality of micro lenticular lenses, and a lens pattern may be formed within the gap between the display panel (111) and the lenticular lens (112).
[0105] For example, as illustrated in Fig. 3b, a transparent frame made of PI in the shape of a microlens is filled with liquid crystals, and the outside can be made of a replica made of a material having the same refractive index as the liquid crystal molecules when voltage is applied. ITO electrodes to which voltage is applied can be positioned above and below the microlenses of this structure. In the 3D mode where no voltage A is applied, a difference in refractive index occurs between the liquid crystal molecules inside and the external replica, resulting in the effect of passing through the lenticular lens. On the other hand, in the 2D mode where voltage B is applied, the state of the liquid crystal changes to have the same refractive index as the external replica, and allows the input light to pass through as is.
[0106] The backlight unit (113) provides light to the display panel (111). By the light provided from the backlight unit (113), the left-eye image and the right-eye images 1 and 2 formed on the display panel (111) are projected onto the viewing area separator (112), and the viewing area separator (112) can disperse the light of each projected image 1 and 2 and transmit it toward the viewer. For example, the viewing area separator (112) can create exit pupils at the viewer's position, i.e., the viewing distance. As illustrated in FIG. 3A, when the viewing area separator (112) is implemented as a lenticular lens array, the thickness and diameter of the lenticular lens, and when it is implemented as a parallax barrier, the spacing of the slits, etc. can be designed so that the exit pupils created by each row are separated by an average binocular center distance of less than 65 mm.
[0107] FIG. 4 is a flowchart illustrating a method for controlling an electronic device according to one or more embodiments.
[0108] According to FIG. 4, in operation 410, the electronic device (100) can identify a user included in the captured image.
[0109] For example, the electronic device (100) can identify a user based on a human figure in an image captured by the camera (130). For example, the electronic device (100) can identify an object area included in the image through at least one of object recognition, object detection, object tracking, and image segmentation. For example, the electronic device (100) can identify a user by using a technique such as semantic segmentation, which classifies and extracts objects included in an input image by type as needed, instance segmentation, which recognizes objects by classifying them by object even if they are of the same type, and a bounding box in the shape of a rectangle that includes the detected object when detecting an object included in an image.
[0110] When a user is identified in the captured image, in operation 420, the electronic device (100) can identify whether the user is positioned in front of the display (110) based on the captured image.
[0111] For example, if the electronic device (100) identifies that the user is included in the captured image, it can identify that the user is positioned in front of the display (110).
[0112] For example, the electronic device (100) may identify a user as being included in a captured image if a specific body part of the user is included in the captured image. For example, the specific body part of the user may be the user's head.
[0113] For example, the electronic device (100) may identify a user as being included in a captured image if a threshold percentage or greater of the user's body area is detected within the captured image. For example, if 90% or more of the user's body area is detected, the electronic device (100) may identify a user as being included in the captured image.
[0114] According to one embodiment, whether a user is positioned in front of the display (110) may be determined differently depending on the angle of view of the camera (130). This is because the angle of view of the camera (130) is the width of the field of view that the camera (130) captures, and thus the shooting range included in the captured image differs depending on the angle of view of the camera (130). The angle of view of the camera (130) may include a horizontal angle of view (HOA) and a vertical angle of view (VA). For example, the camera angle of view may vary depending on the type of camera lens. Depending on the type of lens, it may have various angles of view, such as a wide-angle lens, a standard lens, and a telephoto lens.
[0115] For example, the angle of view of the camera (130) can be adjusted according to user settings.
[0116] For example, the electronic device (100) may adjust the angle of view of the camera (130), for example, at least one of the horizontal angle of view and the vertical angle of view, based on the size of the space in which the electronic device (100) is located (for example, the size of the space in front of the electronic device (100)), the type of space (for example, private space, public space), etc.
[0117] For example, the electronic device (100) may adjust the angle of view of the camera (130), for example, at least one of the horizontal angle of view and the vertical angle of view, based on content characteristics such as the type of content, the three-dimensional information of the content, etc.
[0118] For example, the electronic device (100) may adjust the angle of view of the camera (130), for example, at least one of the horizontal angle of view and the vertical angle of view, based on user context such as the user's profile information, user content viewing tendencies, etc.
[0119] For example, the electronic device (100) may adjust the angle of view of the camera (130), for example, at least one of the horizontal angle of view and the vertical angle of view, based on the context of the electronic device (100), such as the specifications of the electronic device (100), the specifications of the display (110), the resolution of the display (110), etc.
[0120] When it is identified that the user is positioned in front of the display (110) (S420:Y), in operation 430, the electronic device (100) can identify the user's gaze based on the captured image.
[0121] For example, the electronic device (100) can detect the position of the user's face from a captured image and identify the user's eyes from the user's face to obtain the user's gaze information in real time. For example, various conventional methods can be used as a method for detecting the face area. Specifically, a direct recognition method and a statistical method can be used. The direct recognition method creates rules using physical features of the face image, such as the outline, skin color, and size of components or distances between them, and compares, inspects, and measures them according to the rules. The statistical method can detect the face area according to a pre-learned algorithm. That is, it is a method of converting unique features included in the input face into data and comparing and analyzing them with a large prepared database (shapes of faces and other objects). In particular, the face area can be detected according to a pre-learned algorithm, and methods such as a Multi Layer Perceptron (MLP) and a Support Vector Machine (SVM) can be used. The user's eye area can be identified using a similar method.
[0122] For example, the electronic device (100) may obtain user gaze information from a captured image using a learned artificial intelligence model. For example, the artificial intelligence model may be implemented as a neural network including multiple neural network layers. The artificial intelligence model may be implemented as a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), or a deep Q-network, but is not limited thereto.
[0123] For convenience of explanation, the method described herein involves identifying the user's face and eyes in a captured video and then acquiring gaze information. However, the captured video can also be input into a trained artificial intelligence model to acquire the user's gaze information. For example, the artificial intelligence model can be trained to identify the user's face and eyes in a captured video and then acquire gaze information. For example, the artificial intelligence model can output probability information indicating whether the user is facing the front of the display (110) or whether the user is facing the front.
[0124] For example, the learned artificial intelligence model may be an on-device model included in the electronic device (100), but is not limited thereto. For example, the learned artificial intelligence model may be implemented on a server.
[0125] When the electronic device (100) identifies that the user's gaze is directed toward the front of the display (110) (S440:Y), the electronic device (100) can control the display (110) to operate in 3D mode based on the characteristics of the content in operation 450.
[0126] For example, content characteristics may include various information related to the content, such as the type of content, 3D characteristics of the content, etc.
[0127] For example, the content type may include at least one of general content and advertising content. However, this is not limited to these types, and the content type may include content delivery formats, such as real-time streaming content or OTT content, or content genres, such as game content or movie content.
[0128] For example, the 3D nature of content may include the probability that the content is stereoscopic content. Stereoscopic content may be content that uses two images to create a three-dimensional effect.
[0129] However, content characteristics are not limited thereto and may include information such as at least one of scene change, motion size, or frame rate.
[0130] For example, the electronic device (100) may identify the characteristics of content for each content section. For example, the content section may be a preset frame section unit. The preset frame section unit may include one frame unit, multiple frame units, or scene units. The multiple frame units may be identified based on a preset number of frames. For example, the preset number may be a value set during the manufacturing of the electronic device (100) and / or a value that can be set / changed by the user. A scene is a unit representing a series of consecutive events or situations and may include frames corresponding to a series of events occurring at a specific location during a specific time.
[0131] For example, the electronic device (100) can identify content characteristics through image analysis for each content section or identify content characteristics based on metadata included in the content.
[0132] For example, even if the content is 3D, 2D content may be provided in certain sections. For example, 2D advertisement content may be inserted into the middle of 3D content. In this case, if the 3D mode is operated while the 2D advertisement content is provided, the user may see a screen with distorted pixels or a stereoscopic 2-View format, which may cause discomfort. Accordingly, the electronic device (100) can identify the content characteristics for each section.
[0133] For example, when the display (110) is implemented as a light field display, the electronic device (100) can control the lenticular lens (112) to operate in 3D mode by applying a preset voltage corresponding to the 3D mode to the active lenticular lens (112).
[0134] For example, the electronic device (100) may obtain an output image corresponding to a 3D mode based on a left-eye image and a right-eye image. For example, the electronic device (100) may obtain an output image based on a side-by-side image. For example, the electronic device (100) may obtain an output image by alternately arranging the left-eye image and the right-eye image by sub-sampling them by 1 / 2 in the horizontal direction. For example, the electronic device (100) may provide an output image (60) in which the left-eye image and the right-eye images 1 and 2 are sequentially and repeatedly arranged, as illustrated in FIG. 3.
[0135] In one example, the electronic device (100) can track the position of the user's head and / or eyes to obtain a changed output image. For example, when obtaining an output image that is mapped to the user's eye position according to the movement of the user's position, the output image can be generated so that the viewpoint changes smoothly and continuously between consecutive frames through filtering. In one example, the processor (140) can generate an output image so that the viewpoint changes smoothly between consecutive frames through an IIR (Infinite Impulse Response) filter or an FIR (Finite Impulse Response) filter. Accordingly, even if the user's position changes, it is possible to obtain a natural output image in which the depth difference between the binocular images is maintained.
[0136] The electronic device (100) can control the display (110) to operate in 2D mode in operation 460 when it is determined that the user is not positioned in front of the display (110) (S420:N) or the user's gaze is not directed toward the front of the display (110) (S440:N).
[0137] For example, when the display (110) is implemented as a light field display, the electronic device (100) can control the lenticular lens (112) to operate in 2D mode by applying a preset voltage corresponding to the 2D mode to the lenticular lens (112) in an active form.
[0138] Meanwhile, in Fig. 4, the order is mapped for all steps for convenience of explanation, but it is of course not necessarily limited to the order of steps that are not related to the order or can be performed in parallel.
[0139] FIGS. 5A and 5B are drawings illustrating a user identification method according to one or more embodiments.
[0140] According to one embodiment, the electronic device (100) can identify that the user (511) is positioned in front of the display (110) when the user (511) is included in the captured image (510) as illustrated in FIG. 5A.
[0141] For example, the electronic device (100) may identify that the user is not positioned in front of the display (110) if only a specific body part, for example, not the user's head, or an area less than a threshold ratio of the user's body area is identified, even if a portion of the user's body area is included in the captured image (520), as illustrated in FIG. 5b.
[0142] FIGS. 6 and 7 are drawings for explaining a method for obtaining user gaze information according to one or more embodiments.
[0143] According to an example, the electronic device (100) can obtain a captured image (610, 620, 630, 640) of the user through the camera (130).
[0144] According to an example, as illustrated in FIG. 6, the electronic device (100) can sequentially input the first to fourth captured images (610, 620, 630, 640) acquired in real time into a learned artificial intelligence model to acquire user gaze information corresponding to the first to fourth captured images (610, 620, 630, 640).
[0145] According to an example, the electronic device (100) can obtain the user's head direction information and / or gaze information as illustrated in FIG. 7 based on the output information of the artificial intelligence model.
[0146] For example, in the first captured image (610), the head direction can be identified as being forward and the gaze direction can also be identified as being forward. In this case, the user's gaze can be identified as being directed toward the front of the display (110).
[0147] For example, in the second captured image (620), the head direction may be identified as being forward, but the gaze direction may not be forward. In this case, the user's gaze may be identified as being directed toward the front of the display (110). In this case, the user's gaze may be identified as not being directed toward the front of the display (110).
[0148] For example, in the case of the third captured image (630), the head direction may not be in the frontal direction, but the gaze direction may also be identified as being in the frontal direction. In this case, the user's gaze may be identified as being directed toward the front of the display (110).
[0149] For example, in the fourth captured image (630), the head direction can be identified as not being in the frontal direction and the gaze direction can also be identified as not being in the frontal direction. In this case, the user's gaze can be identified as not being directed toward the front of the display (110).
[0150] For example, an AI model can be trained to analyze the size of the user's eyes, pupil size, etc. when a video is input, and output information about the user's gaze. The user's gaze information can be information about the direction the user's eyes are looking.
[0151] For example, the artificial intelligence model can be trained to output probability information about whether the user's gaze is directed forward, i.e., toward the front of the display (110), and / or whether the user's gaze is directed forward.
[0152] For example, when an artificial intelligence model is trained, it means that a basic artificial intelligence model (e.g., an artificial intelligence model including any random parameters) is trained using a plurality of training data by a learning algorithm, thereby creating a predefined operation rule or artificial intelligence model set to perform a desired characteristic (or purpose). Such learning may be performed through a separate server and / or system, but is not limited thereto, and may also be performed in the electronic device (100). Examples of learning algorithms include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.
[0153] For example, the learned artificial intelligence model may be an on-device model included in the electronic device (100), but is not limited thereto. For example, the learned artificial intelligence model may be implemented on a server.
[0154] FIG. 8 is a flowchart illustrating a method of controlling an electronic device according to one or more embodiments.
[0155] Among the operations illustrated in Fig. 8, detailed descriptions of operations that overlap with the steps illustrated in Fig. 4 will be omitted.
[0156] According to FIG. 8, in operation 810, the electronic device (100) can identify a user included in the captured image.
[0157] When a user is identified in the captured video, in operation 820, the electronic device (100) can identify whether the user is positioned in front of the display (110).
[0158] When it is identified that the user is positioned in front of the display (110) (S820:Y), in operation 830, the electronic device (100) can identify the user's gaze based on the captured image.
[0159] In operation 840, the electronic device (100) can identify whether the user's gaze is directed toward the front of the display (110).
[0160] When the electronic device (100) identifies that the user's gaze is directed toward the front of the display (110) (S840:Y), the electronic device (100) can identify the probability that the content is stereoscopic content for each content section based on the characteristics of the content in operation 850.
[0161] For example, the electronic device (100) can identify the probability of stereoscopic content for each content section based on content characteristics. The content characteristics may include the content type, the content's media format, etc. The content type may include the content's genre, whether it is advertising content, etc. The content's media format may include configuration information such as images, music, and text that constitute the content.
[0162] For example, the electronic device (100) can identify the probability of stereoscopic content for each content section based on the similarity between the left-eye image and the right-eye image. For example, the electronic device (100) can measure the similarity between the left-eye image and the right-eye image based on at least one of structural similarity, color-based similarity, and feature-based similarity between the left-eye image and the right-eye image. For example, structural similarity can be measured by comparing structural characteristics of the image, such as edges, sharpness, and texture of the image. For example, color-based similarity can be measured by comparing color information of the image, such as pixel values or color histograms of the image. For example, feature-based similarity can be measured by extracting characteristic points or patterns from the image and measuring the degree of correspondence between these features.
[0163] For example, the electronic device (100) can input content into a learned artificial intelligence model for each content section and obtain a probability that the content is stereoscopic content based on information output from the artificial intelligence model.
[0164] For example, an artificial intelligence model learned according to an example can be learned to, when multiple images are input, characterize the multiple images, identify similarity between the multiple images based on the similarity, and output information about the identified similarity and / or probability information about stereoscopic content.
[0165] Artificial intelligence models can be implemented using, but are not limited to, Deep Neural Networks (DNN), Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN), Restricted Boltzmann Machines (RBM), Deep Belief Networks (DBN), Bidirectional Recurrent Deep Neural Networks (BRDNN), or Deep Q-Networks.
[0166] In operation 860, the electronic device (100) can determine whether the probability of stereoscopic content in the corresponding content section is greater than or equal to a first threshold value. For example, the first threshold value may be a value set during manufacturing or a value set by the user. For example, the first threshold value may be a fixed value or a changeable value.
[0167] For example, the first threshold value may be set according to the characteristics of the display (110). The characteristics of the display (110) may include various characteristics such as screen size, resolution, color expressiveness, scanning method, and scanning frequency.
[0168] For example, the first threshold value may vary depending on the characteristics of the content. These characteristics may include the content type, the content's media format, etc. The content type may include the content's genre, whether it is advertising content, etc. The content's media format may include information about the content's composition, such as images, music, and text.
[0169] For example, the first threshold value may vary depending on the user's context. The user's context may include information such as the user's profile and viewing environment. For example, in a viewing environment where 3D content is easily recognized, the first threshold value may be relatively low. For example, for users who have difficulty recognizing 3D content (e.g., older users or users with poor eyesight), the first threshold value may be relatively high.
[0170] If the probability that the content section is stereoscopic content is greater than or equal to the first threshold value (S860:Y), in operation 870, the electronic device (100) can control the display (110) to operate in 3D mode in the content section.
[0171] For example, when the display (110) is implemented as a light field display, the electronic device (100) can control the lenticular lens (112) to operate in 3D mode by applying a preset voltage corresponding to the 3D mode to the active lenticular lens (112).
[0172] The electronic device (100) can control the display (110) to operate in 2D mode in operation 880 when it is determined that the user is not positioned in front of the display (110) (S820:N) or the user's gaze is not directed toward the front of the display (110) (S840:N).
[0173] For example, when the display (110) is implemented as a light field display, the electronic device (100) can control the lenticular lens (112) to operate in 2D mode by applying a preset voltage corresponding to the 2D mode to the lenticular lens (112) in an active form.
[0174] For example, even if the probability of being stereoscopic content is greater than or equal to a first threshold value, the electronic device (100) may be able to improve the accuracy of determining whether the content is stereoscopic content by comparing the left and right areas of the content with the entire area, since the content may be a single 2D content image.
[0175] For example, even if the probability of stereoscopic content in a content section is greater than or equal to a first threshold, the display (110) may be controlled to operate in 2D mode in some cases. For example, if the content section is less than the threshold time and the probability that the content in the preceding and following content sections are all stereoscopic content is less than the first threshold, the electronic device (100) may control the display (110) to provide only one area among the left area and the right area in the content section and to operate in 2D mode. Accordingly, the user can comfortably enjoy the content.
[0176] Meanwhile, in Fig. 8, the order is mapped for all steps for convenience of explanation, but it is of course not necessarily limited to the order of steps that are not related to the order or can be performed in parallel.
[0177] FIGS. 9A and 9B are drawings illustrating a method for identifying stereoscopic content according to one or more embodiments.
[0178] For example, as illustrated in FIGS. 9A and 9B , the electronic device (100) may input content images (910, 920) into a learned artificial intelligence model to obtain probability information about whether the image is stereoscopic content. The content image (910) may be a single frame unit, but is not limited thereto. For example, the content image (910) may be a unit having the same meaning as a plurality of frame units or a scene unit. A scene is a unit representing a series of consecutive events or situations and may include frames corresponding to a series of events occurring at a specific location during a specific time.
[0179] For example, an artificial intelligence model may be trained to analyze the similarity of the left and right images based on the central vertical line of the content image (910, 920) and output similarity information and / or probability information of stereoscopic content.
[0180] For example, as illustrated in FIG. 9A, if the first content image (910) is a side-by-side image for providing a 3D image, the similarity between the left and right images may be high. Accordingly, the artificial intelligence model may output information such as a similarity of 95% (or 0.95) and / or a probability of 95% (or 0.95) that it is stereoscopic content.
[0181] For example, as illustrated in FIG. 9b, if the second content image (920) is a 2D image, the similarity between the left and right images may be very low. Accordingly, the artificial intelligence model may output information such as a similarity of 15% (or 0.15) and / or a probability of 15% (or 0.15) that it is stereoscopic content.
[0182] For example, the learned artificial intelligence model may be an on-device model included in the electronic device (100), but is not limited thereto. For example, the learned artificial intelligence model may be implemented on a server.
[0183] FIG. 10 is a flowchart illustrating a method of controlling an electronic device according to one or more embodiments.
[0184] Among the operations illustrated in Fig. 10, detailed descriptions of operations that overlap with the steps illustrated in Fig. 4 and Fig. 8 will be omitted.
[0185] According to FIG. 10, in operation 1010, the electronic device (100) can identify a user included in the captured image.
[0186] When a user is identified in the captured video, in operation 1020, the electronic device (100) can identify whether the user is positioned in front of the display (110).
[0187] When it is identified that the user is positioned in front of the display (110) (S1020:Y), in operation 1030, the electronic device (100) can identify the user's gaze based on the captured image.
[0188] In operation 1040, the electronic device (100) can identify whether the identified user's gaze is directed toward the front of the display (110).
[0189] If the electronic device (100) identifies that the user's gaze is directed toward the front of the display (110) (S1040:Y), in operation S1050, the electronic device can identify whether the similarity between the left and right regions based on the central vertical line of the content for each content section is greater than or equal to a second threshold value. For example, the second threshold value may be a value set during manufacturing or a value set by the user. For example, the second threshold value may be a fixed value or a changeable value. Other characteristics of the second threshold value may be similar to those of the first threshold value described above.
[0190] In one example, the electronic device (100) may measure the similarity between the left area and the right area based on information corresponding to each of a plurality of pixels included in a frame of the content. For example, the content received by the electronic device (100) may include a digital value corresponding to each pixel, and the electronic device (100) may encode the digital value to obtain a grayscale value (e.g., a value from 0 to 255) corresponding to a specific grayscale range corresponding to each pixel. In one example, the electronic device (100) may measure the similarity between the left area and the right area based on the digital value before obtaining the grayscale value. However, the present invention is not limited thereto, and the similarity between the left area and the right area may also be measured based on a grayscale value corresponding to a specific grayscale range.
[0191] For example, the electronic device (100) may identify the difference values of corresponding pixel information of the left and right areas of the frame, and identify the similarity between the left and right areas of the frame based on the identified difference values. For example, the electronic device (100) may identify the similarity between the left and right areas of the frame based on at least one of a sum of the difference values for each pixel, an average value of the difference values for each pixel, and a maximum value or minimum value among the difference values for each pixel.
[0192] For example, the electronic device (100) may convert the average value of the difference values per pixel into a value for comparison with a second threshold value to determine whether the similarity between the left and right areas of the content is greater than or equal to the second threshold value. For example, the higher the similarity between the left and right areas, the smaller the average value of the difference values may be, and therefore, the value for comparison with the second threshold value may be converted to be inversely proportional to the average value.
[0193] If the similarity between the left and right areas of the content is identified as being greater than the second threshold value for each content section (S1050:Y), in operation S1060, the display (110) can be controlled to operate in 3D mode in the corresponding content section.
[0194] The electronic device (100) can control the display (110) to operate in 2D mode in operation 1070 when it is determined that the user is not positioned in front of the display (110) (S1020:N) or the user's gaze is not directed toward the front of the display (110) (S1040:N).
[0195] Meanwhile, in Fig. 10, the order is mapped for all steps for convenience of explanation, but it is of course not necessarily limited to the order of steps that are not related to the order or can be performed in parallel.
[0196] FIGS. 11A and 11B are drawings illustrating a method for identifying stereoscopic content according to one or more embodiments.
[0197] According to one embodiment, the electronic device (100) may measure the similarity between the left and right regions based on information corresponding to each of a plurality of pixels included in a frame of content. According to one example, the content may be a side-by-side video frame.
[0198] For example, the electronic device (100) may calculate the difference value of pixel information of each of the left image (1111) and the right image (1112) based on the vertical center line of the received content (1110) as illustrated in FIG. 11A. For example, if the received content (1110) is an image with a resolution of n*m (vertical resolution*horizontal resolution), the left image may have n*m / 2 pixel information, and the right image may also have n*2 / m pixel information.
[0199] According to an example, the electronic device (100) can calculate a difference value for pixel values at corresponding positions of the left image (1111) and the right image (1112). For example, the electronic device (100) can calculate a difference value c11 between the pixel value a11 at position (1, 1) of the left image (1111) and the pixel value b11 at position (1, 1) of the right image (1112), and can calculate a difference value c12 between the pixel value a12 at position (1, 2) of the left image (1111) and the pixel value b12 at position (1, 2) of the right image (1112). The electronic device (100) can calculate all pixel value differences at corresponding positions of the left image (1111) and the right image (1112) in the same manner.
[0200] For example, in the case of 3D content (1110) as illustrated in FIG. 11a, since the pixel values of the left image (1111) and the right image (1112) are similar, each of the difference values (c11, c12, ......) calculated at the corresponding pixel positions can be calculated to be less than the threshold value.
[0201] For example, in the case of 2D content (1120) as illustrated in FIG. 11b, since the difference in pixel values between the left image and the right image is large, each of the difference values (f11, f12, ......) calculated at the corresponding pixel positions may be calculated to be greater than the threshold value.
[0202] Accordingly, the electronic device (100) can identify the similarity of the left and right areas of the content based on the difference in pixel values of the left and right areas of the content.
[0203] FIG. 12 is a flowchart illustrating a method of controlling an electronic device according to one or more embodiments.
[0204] Among the operations illustrated in Fig. 12, detailed descriptions of operations that overlap with the steps illustrated in Figs. 4, 8, and 10 will be omitted.
[0205] According to FIG. 12, in operation 1210, the electronic device (100) can identify whether the probability that the received content is stereoscopic content is greater than or equal to a first threshold value.
[0206] For example, unlike the embodiment of FIG. 8, where the electronic device (100) identifies the probability that the content is stereoscopic content when the user's gaze is directed toward the front of the display (110), when the content is received, the electronic device (100) can first identify the probability that the content is stereoscopic content. According to an example, the method for identifying the probability that the content is stereoscopic content is the same as the method described in the various embodiments described above, and thus a detailed description thereof will be omitted.
[0207] If the probability that the input content is stereoscopic content is greater than or equal to the first threshold value (S1210:Y), in operation 1220, the electronic device (100) can identify the user included in the captured image.
[0208] In operation 1230, the electronic device (100) can identify whether the user is positioned in front of the display (110).
[0209] When it is identified that the user is positioned in front of the display (110) (S1230:Y), in operation 1240, the electronic device (100) can identify the user's gaze based on the captured image.
[0210] At operation 1250, the electronic device (100) can identify whether the user's gaze is directed toward the front of the display (110).
[0211] When the electronic device (100) identifies that the user's gaze is directed toward the front of the display (110) (S1250:Y), in operation S1260, the electronic device (100) can control the display (110) to operate in 3D mode in the corresponding content section.
[0212] The electronic device (100) can control the display (110) to operate in 2D mode in operation 1270 when it is determined that the user is not positioned in front of the display (110) (S1230:N) or the user's gaze is not directed toward the front of the display (110) (S1250:N).
[0213] Although the embodiment illustrated in FIG. 12 is described based on the embodiment illustrated in FIG. 8, it is obvious that in the embodiment illustrated in FIG. 10, the operation of identifying the similarity of the left and right areas of the content may be performed prior to identifying the user's location and gaze.
[0214] Meanwhile, in Fig. 12, the order is mapped for all steps for convenience of explanation, but it is of course not necessarily limited to the order of steps that are not related to the order or can be performed in parallel.
[0215] In the various embodiments described above, the electronic device (100) may perform downscaling before inputting a captured image or content image into the learned artificial intelligence model, if necessary. This is to reduce the computational load of the artificial intelligence model.
[0216] Each operation according to the various embodiments described above may be performed by the processor (140), but if necessary, a module for each operation may be utilized. For example, each module may be implemented using at least one software, at least one hardware, and / or a combination thereof. Each module may be implemented to utilize a predefined algorithm, a predefined formula, and / or a learned artificial intelligence model to perform the operation. However, at least some modules may be distributed to an external device.
[0217] According to the various embodiments described above, user convenience can be improved because natural switching between 2D mode and 3D mode is automatically performed according to the user's gaze and content situation.
[0218] Meanwhile, the methods according to the various embodiments of the present disclosure described above can be implemented only with a software upgrade or a hardware upgrade for an existing electronic device and / or server.
[0219] Additionally, the various embodiments of the present disclosure described above can also be performed through an embedded server provided in an electronic device, or an external server of the electronic device.
[0220] Meanwhile, according to a temporary example of the present disclosure, the various embodiments described above can be implemented as software including instructions stored in a machine-readable storage medium that can be read by a machine (e.g., a computer). The device is a device that can call instructions stored from the storage medium and operate according to the called instructions, and may include an electronic device (e.g., electronic device (A)) according to the disclosed embodiments. When an instruction is executed by a processor, the processor can perform a function corresponding to the instruction directly or by using other components under the control of the processor. The instruction may include code generated or executed by a compiler or interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' means that the storage medium does not contain a signal and is tangible, but does not distinguish between data being stored semi-permanently or temporarily in the storage medium.
[0221] Furthermore, according to one embodiment of the present disclosure, the method according to the various embodiments described above may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or online through an application store (e.g., Play Store™). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0222] In addition, each of the components (e.g., modules or programs) according to the various embodiments described above may be composed of a single or multiple entities, and some of the corresponding sub-components described above may be omitted, or other sub-components may be further included in various embodiments. Alternatively or additionally, some components (e.g., modules or programs) may be integrated into a single entity, which may perform the same or similar functions as those performed by each of the corresponding components prior to integration. Operations performed by modules, programs or other components according to various embodiments may be executed sequentially, in parallel, iteratively or heuristically, or at least some operations may be executed in a different order, omitted, or other operations may be added.
[0223] Although the preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above, and various modifications may be made by a person having ordinary skill in the art to which the present disclosure pertains without departing from the gist of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the present disclosure.
Claims
1. In electronic devices, A display implemented to operate in either 3D or 2D mode; One or more cameras for capturing images in front of the display; A memory storing one or more instructions; and comprising one or more processors; The one or more processors, by executing the one or more instructions, Identifying whether a user is positioned in front of the display based on the captured image, When the user is identified as being positioned in front of the display, it is identified whether the user's gaze is directed toward the front of the display, When the user's gaze is identified as being directed toward the front of the display, the display is controlled to operate in the 3D mode; An electronic device that controls the display to operate in the 2D mode when it is determined that the user is not positioned in front of the display or that the user's gaze is not directed toward the front of the display.
2. In paragraph 1, The one or more processors, by executing the one or more instructions, When the user's gaze is identified as being directed toward the front of the display, the probability that the content to be displayed is stereoscopic content is identified for each content section. Controlling the display to operate in the 3D mode in a content section where the probability of the stereoscopic content being greater than a threshold value; An electronic device that controls the display to operate in the 2D mode in a content section where the probability of the stereoscopic content being less than the threshold value.
3. In paragraph 2, The one or more processors, by executing the one or more instructions, An electronic device that inputs content into a learned artificial intelligence model for each content section and identifies the probability that the content is stereoscopic content based on information output from the learned artificial intelligence model.
4. In paragraph 2, The one or more processors, by executing the one or more instructions, An electronic device that inputs the above-described captured image into a learned artificial intelligence model and identifies whether the user's gaze is directed toward the front of the display based on information output from the learned artificial intelligence model.
5. In paragraph 1, The above content is, Contains side by side content, The one or more processors, by executing the one or more instructions, Identify whether the similarity between the left and right areas of the side-by-side content is a threshold value for each of the above content sections, Controlling the display to operate in the 3D mode in a content section where the similarity between the left and right areas of the side-by-side content is greater than the threshold value; An electronic device that controls the display to operate in the 2D mode in a content section where the similarity between the left and right areas of the side-by-side content is less than the threshold value.
6. In paragraph 1, The one or more processors, by executing the one or more instructions, While the above display is operating in the above 3D mode An electronic device that controls the display to switch the 3D mode to the 2D mode when it is determined that the user is not positioned in front of the display or that the user's gaze is not directed toward the front of the display.
7. In paragraph 1, The one or more processors, by executing the one or more instructions, When the user's gaze is identified as being directed toward the front of the display while the display is operating in the 2D mode, the content is identified as stereoscopic content by content section, An electronic device that controls the display to operate in the 3D mode when the content is the stereoscopic content.
8. In paragraph 1, The one or more processors, by executing the one or more instructions, Identify the probability that each section of the above content is stereoscopic content, In a content section where the probability of the stereoscopic content is greater than a threshold value, it is identified whether the user's gaze is directed toward the front of the display based on the captured image, If it is determined that the user's gaze is directed toward the front of the display and the probability of the stereoscopic content is greater than the threshold value, the display is controlled to operate in the 3D mode, An electronic device that controls the display to operate in the 2D mode in a content section where the probability of the stereoscopic content being less than the threshold value.
9. In paragraph 1, The one or more processors, by executing the one or more instructions, An electronic device that identifies the user as being positioned in front of the display when the captured image identifies that a specific body of the user is included in the captured image or that an area greater than a preset percentage of the user's body area is included in the captured image.
10. In paragraph 1, The above display is, It is implemented as a light field display including a lenticular lens array, The above processor, An electronic device that controls the display to operate in the 3D mode or the 2D mode by adjusting the voltage applied to a lenticular lens array included in the light field display. A method for controlling an electronic device including a display implemented to be capable of operating in one of 11.3D mode and 2D mode and one or more cameras for capturing an image in front of the display, A step of identifying whether a user is positioned in front of the display based on a captured image acquired through one or more cameras; When it is identified that a user is positioned in front of the display, a step of identifying whether the user's gaze is directed toward the front of the display based on the captured image; a step of controlling the display to operate in the 3D mode when it is identified that the user's gaze is directed toward the front of the display; and A control method, comprising: a step of controlling the display to operate in the 2D mode when it is determined that the user is not positioned in front of the display or that the user's gaze is not directed toward the front of the display.
12. In paragraph 11, The above control method is, If it is identified that the user's gaze is directed toward the front of the display, a step of identifying the probability that the content to be displayed is stereoscopic content for each content section is further included; The step of controlling the display to operate in the above 3D mode is: Controlling the display to operate in the 3D mode in a content section where the probability of the stereoscopic content being greater than a threshold value; The step of controlling the display to operate in the above 2D mode comprises: A control method for controlling the display to operate in the 2D mode in a content section where the probability of the stereoscopic content being less than the threshold value.
13. In paragraph 12, The step of identifying the probability that the above content is stereoscopic content is: A control method for inputting content into a learned artificial intelligence model for each content section and identifying the probability that the content is stereoscopic content based on information output from the learned artificial intelligence model.
14. In paragraph 12, The step of identifying whether the user's gaze is directed toward the front of the display is: A control method for inputting the above-described captured image into a learned artificial intelligence model and identifying whether the user's gaze is directed toward the front of the display based on information output from the learned artificial intelligence model. A non-transitory computer-readable medium storing computer instructions that, when executed by a processor of an electronic device, cause the electronic device to perform an operation, the computer instructions comprising a display implemented to be operable in one of 15.3D mode and 2D mode and one or more cameras for capturing an image in front of the display, The above actions are, A step of identifying whether a user is positioned in front of the display based on a captured image acquired through one or more cameras; When it is identified that a user is positioned in front of the display, a step of identifying whether the user's gaze is directed toward the front of the display based on the captured image; a step of controlling the display to operate in the 3D mode when it is identified that the user's gaze is directed toward the front of the display; and A non-transitory computer-readable medium comprising: a step of controlling the display to operate in the 2D mode when it is determined that the user is not positioned in front of the display or that the user's gaze is not directed toward the front of the display.
Citation Information
Patent Citations
Video recording device and method, and video reproducing device and method
JP2010068315A
Signal processing apparatus and signal processing method
JP2012034138A
Information processing device and information processing method
JP2013150038A
Drainage system with renewable energy generation device for flood prevention on slopes and agricultural land
KR1020240036940A
Display apparatus using sight direction to adjust display mode and operation method thereof
US20220103805A1