Electronic apparatus and control method thereof

WO2025187929A8PCT designated stage Publication Date: 2025-10-02SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/096980
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-05
Filing Date
2024-12-13
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Existing 2D image conversion to 3D technologies face challenges in creating natural-looking 3D images due to unnaturalness in hole areas of new viewpoints, where pixel values are artificially generated and filled, leading to side effects.

Method used

An electronic device and method that identifies object regions in an input image using a depth map, assesses hole filling complexity, and determines a new viewpoint to generate a natural-looking binocular image by optimizing hole filling.

Benefits of technology

The solution effectively generates natural-looking 3D images with minimal side effects by optimizing hole filling based on hole complexity, enhancing the 3D image generation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024096980_02102025_PF_FP_ABST
    Figure KR2024096980_02102025_PF_FP_ABST
Patent Text Reader

Abstract

This electronic apparatus comprises a memory for storing one or more instructions and at least one processor for executing the one or more instructions, wherein the instructions, when executed by the at least one processor, instruct the electronic device to: identify a plurality of object regions included in an input image on the basis of a depth map corresponding to the input image; identify a hole filling complexity for a hole region generated in a boundary region between the plurality of object regions; identify a viewpoint of a new viewpoint image on the basis of the hole filling complexity; and acquire the new viewpoint image on the basis of the identified viewpoint.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device and method of controlling the same

[0001] The present disclosure relates to an electronic device and a method for controlling the same, and more particularly, to an electronic device for generating a new viewpoint image and a method for controlling the same.

[0002] Advances in electronic technology have led to the development and proliferation of various types of electronic devices. In particular, display devices, used in a variety of settings, including homes, offices, and public spaces, have been continuously evolving in recent years.

[0003] Stereoscopy refers to three-dimensional technology. Recently, commercialized 3D displays primarily utilize binocular parallax. Binocular parallax offers the advantage of creating a three-dimensional effect on a single screen, such as a TV or theater screen. Methods utilizing binocular parallax can be categorized into stereoscopic (using glasses or other auxiliary devices) and autostereocopic (glassless) methods.

[0004] Research is currently underway to commercialize glasses-free light field displays and glasses-free 3D displays utilizing eye-tracking. Furthermore, research is ongoing beyond displays to convert existing 2D images into 3D, enabling consumers to experience a variety of 3D images.

[0005] According to one embodiment, an electronic device includes a memory storing one or more instructions; and at least one processor executing the one or more instructions, wherein the instructions, when executed by the at least one processor, cause the electronic device to identify a plurality of object regions included in an input image based on a depth map corresponding to the input image, identify a hole filling complexity for a hole region occurring in a boundary region between the plurality of object regions, identify a viewpoint of a new viewpoint image based on the hole filling complexity, and obtain the new viewpoint image based on the viewpoint.

[0006] A method for controlling an electronic device according to an embodiment may include: identifying a plurality of object regions included in an input image based on a depth map corresponding to the input image; identifying a hole filling complexity for a hole region occurring in a boundary region between the plurality of object regions; and identifying a viewpoint of a new viewpoint image based on the hole filling complexity and obtaining the new viewpoint image based on the identified viewpoint.

[0007] In one embodiment, a non-transitory computer-readable medium storing computer instructions that, when executed by a processor of an electronic device, cause the electronic device to perform an operation, the operation may include: identifying a plurality of object regions included in an input image based on a depth map corresponding to the input image; identifying a hole filling complexity for a hole region occurring in a boundary region between the plurality of object regions; and identifying a viewpoint of a new viewpoint image based on the hole filling complexity and obtaining the new viewpoint image based on the identified viewpoint.

[0008] The above and other aspects, features and advantages of specific embodiments of the present disclosure will become more apparent from the following description taken in conjunction with the accompanying drawings.

[0009] FIG. 1 is a diagram illustrating a new viewpoint image generation technique according to one or more embodiments.

[0010] FIG. 2A is a block diagram showing the configuration of an electronic device according to one embodiment.

[0011] FIG. 2b is a block diagram specifically illustrating a configuration of an electronic device according to one or more embodiments.

[0012] FIG. 3 is a drawing for explaining the structure and operation of a display according to one or more embodiments.

[0013] FIG. 4 is a flowchart illustrating a method for controlling an electronic device according to one or more embodiments.

[0014] FIG. 5A is a diagram illustrating a method for identifying hole filling complexity according to one or more embodiments.

[0015] FIG. 5b is a diagram illustrating a method for identifying hole filling complexity according to one or more embodiments.

[0016] FIG. 5c is a diagram illustrating a method for identifying hole filling complexity according to one or more embodiments.

[0017] FIG. 6 is a drawing for explaining a control method of an electronic device according to one or more embodiments.

[0018] FIG. 7 is a drawing for explaining a control method of an electronic device according to one or more embodiments.

[0019] FIG. 8 is a drawing for explaining a control method of an electronic device according to one or more embodiments.

[0020] FIG. 9 is a drawing for explaining a control method of an electronic device according to one or more embodiments.

[0021] FIG. 10 is a drawing for explaining in detail a method for providing a 3D image according to one or more embodiments.

[0022] FIG. 11a is a diagram illustrating a method for obtaining information using an artificial intelligence model according to one or more embodiments.

[0023] FIG. 11b is a diagram illustrating a method for obtaining information using an artificial intelligence model according to one or more embodiments.

[0024] FIG. 12A is a diagram illustrating a method for generating a new viewpoint image according to one or more embodiments.

[0025] FIG. 12b is a diagram illustrating a method for generating a new viewpoint image according to one or more embodiments.

[0026] FIG. 12c is a diagram illustrating a method for generating a new viewpoint image according to one or more embodiments.

[0027] FIG. 13a is a diagram illustrating a method for generating a new viewpoint image according to one or more embodiments.

[0028] FIG. 13b is a diagram illustrating a method for generating a new viewpoint image according to one or more embodiments.

[0029] FIG. 13c is a diagram illustrating a method for generating a new viewpoint image according to one or more embodiments.

[0030] FIG. 13d is a diagram illustrating a method for generating a new viewpoint image according to one or more embodiments.

[0031] FIG. 13e is a diagram illustrating a method for generating a new viewpoint image according to one or more embodiments.

[0032] FIG. 14a is a diagram illustrating a method for generating a new viewpoint image according to one or more embodiments.

[0033] FIG. 14b is a diagram illustrating a method for generating a new viewpoint image according to one or more embodiments.

[0034] FIG. 14c is a diagram illustrating a method for generating a new viewpoint image according to one or more embodiments.

[0035] FIG. 14d is a diagram illustrating a method for generating a new viewpoint image according to one or more embodiments.

[0036] FIG. 14e is a diagram illustrating a method for generating a new viewpoint image according to one or more embodiments.

[0037] FIG. 15a is a diagram illustrating a method for generating a new viewpoint image according to one or more embodiments.

[0038] FIG. 15b is a diagram illustrating a method for generating a new viewpoint image according to one or more embodiments.

[0039] FIG. 15c is a diagram illustrating a method for generating a new viewpoint image according to one or more embodiments.

[0040] FIG. 16A is a diagram illustrating a method for generating a new viewpoint image according to one or more embodiments.

[0041] FIG. 16b is a diagram illustrating a method for generating a new viewpoint image according to one or more embodiments.

[0042] The terms used in this disclosure will be briefly explained, and the disclosure will be described in detail.

[0043] The terms used in the embodiments of this disclosure have been selected from widely used, current terms, taking into account the functions of this disclosure. However, these terms may vary depending on the intentions or cases of those skilled in the art, the emergence of new technologies, etc. Furthermore, in certain cases, terms may be arbitrarily selected by the applicant, and in such cases, their meanings will be described in detail in the description of the relevant disclosure. Therefore, the terms used in this disclosure should not be defined simply as names of terms, but rather based on the meanings of the terms and the overall content of this disclosure.

[0044] In this specification, expressions such as “has,” “can have,” “includes,” or “may include” indicate the presence of a feature (e.g., a number, function, operation, or component such as a part), and do not exclude the presence of additional features.

[0045] In this disclosure, expressions such as “A or B,” “at least one of A and / or B,” or “one or more of A or / and B” can include all possible combinations of the listed items. For example, “A or B,” “at least one of A and B,” or “at least one of A or B” can all refer to cases where (1) only A is included, (2) only B is included, or (3) both A and B are included.

[0046] As used herein, the expressions “first,” “second,” “first,” or “second,” etc., may describe various components, regardless of order and / or importance, and are only used to distinguish one component from another, but do not limit the components.

[0047] When it is said that a component (e.g., a first component) is “operatively or communicatively coupled with / to” or “connected to” another component (e.g., a second component), it should be understood that the component may be directly coupled to the other component, or may be connected through another component (e.g., a third component).

[0048] The expression "configured to" as used in the present disclosure may be used interchangeably with, for example, "suitable for," "having the capacity to," "designed to," "adapted to," "made to," or "capable of." The term "configured to" may not necessarily mean only "specifically designed to" in terms of hardware.

[0049] In some contexts, the phrase "a device configured to" may mean that the device, in conjunction with other devices or components, is "capable of" performing A, B, and C. For example, the phrase "a processor configured (or set) to perform A, B, and C" may refer to a dedicated processor (e.g., an embedded processor) for performing those operations, or a general-purpose processor (e.g., a CPU or application processor) that can perform those operations by executing one or more software programs stored in a memory device.

[0050] Singular expressions include plural expressions unless the context clearly dictates otherwise. In this application, terms such as "comprise" or "consist of" are intended to indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but should be understood not to preclude the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0051] In the embodiments, a "module" or "part" performs at least one function or operation and may be implemented as hardware or software, or as a combination of hardware and software. Furthermore, a plurality of "modules" or "parts" may be integrated into at least one module and implemented as at least one processor (not shown), excluding any "module" or "part" that needs to be implemented as specific hardware.

[0052] Meanwhile, the various elements and areas in the drawings are schematically drawn. Therefore, the technical concept of the present invention is not limited by the relative sizes or spacing depicted in the attached drawings.

[0053] An embodiment of the present disclosure will be described in more detail with reference to the attached drawings below.

[0054] FIG. 1 is a diagram illustrating a new viewpoint image generation technique according to one or more embodiments.

[0055] Novel view synthesis technology is a technology that creates an image of a new viewpoint different from the original viewpoint from a two-dimensional image acquired through a monocular camera.

[0056] As an example, according to FIG. 1, in order to obtain a 3D image from a 2D image, a depth map (20) is obtained by estimating the depth from a 2D image, which is an input image (10), and a binocular image can be obtained (40) by controlling the depth for each object (30) according to the obtained depth map (20).

[0057] For example, to obtain a 3D image from a 2D image, View Synthesis (or Novel View Synthesis) can be performed. View Synthesis refers to a technology that generates an image from a new viewpoint by inferring it from an image captured at a specific viewpoint.

[0058] For example, according to View Synthesis, binocular images can be obtained by: i) generating a new right-eye image using a left-eye image as an input image, ii) generating a new left-eye image using a right-eye image as an input image, or iii) generating a new left-eye image and a new right-eye image based on the input image.

[0059] As described above, when generating a new viewpoint image, unnaturalness may occur in the hole area (or occluded area or point area) of the new viewpoint image as pixel values ​​are artificially generated and filled in. The hole area may be an area that is not exposed in the foreground area of ​​the current viewpoint, i.e., the input image, but is exposed when the viewpoint moves, i.e., in the new viewpoint image.

[0060] Below, we will describe various embodiments of analyzing an input image to generate a new viewpoint image advantageous for hole filling, thereby obtaining a natural binocular image with few side effects.

[0061] FIG. 2A is a block diagram showing the configuration of an electronic device according to one embodiment.

[0062] According to FIG. 2a, the electronic device (100) includes a memory (110) and at least one processor (120).

[0063] The electronic device (100) may be implemented as various types of display devices such as a TV, monitor, PC, kiosk, tablet PC, electronic picture frame, mobile phone, HMD (Head mounted Display), NED (Near Eye Display), LFD (Large format display), Digital Signage (digital signage), DID (Digital Information Display), video wall, projector display, etc., or as an image processing device (e.g., set-top box, one connected box) that provides images to the display device.

[0064] The memory (110) can store data required for various embodiments. The memory (110) may be implemented in the form of memory embedded in the electronic device (100') or in the form of memory that can be detachably attached to the electronic device (100) depending on the purpose of data storage. For example, data for driving the electronic device (100) may be stored in a memory embedded in the electronic device (100'), and data for expanding the functions of the electronic device (100) may be stored in a memory that can be detachably attached to the electronic device (100). Meanwhile, in the case of memory embedded in the electronic device (100), it may be implemented as at least one of volatile memory (e.g., dynamic RAM (DRAM), static RAM (SRAM), or synchronous dynamic RAM (SDRAM)), non-volatile memory (e.g., one time programmable ROM (OTPROM), programmable ROM (PROM), erasable and programmable ROM (EPROM), electrically erasable and programmable ROM (EEPROM), mask ROM, flash ROM, flash memory (e.g., NAND flash or NOR flash), hard drive, or solid state drive (SSD). In addition, in the case of memory that can be attached or detached to the electronic device (100'), it may be implemented as at least one of memory cards (e.g., compact flash (CF), secure digital (SD), micro secure digital (Micro-SD), mini secure digital (Mini-SD), extreme digital (xD), multi-media card (MMC), etc.), external memory that can be connected to a USB port (e.g., USB memory), etc. It can be implemented in the form of.

[0065] In one example, the memory (110) may store a computer program including at least one instruction or instructions for controlling the electronic device (100).

[0066] In another example, the memory (110) may store an image received from an external device (e.g., a source device), an external storage medium (e.g., USB), an external server (e.g., a web hard drive), or the like, i.e., an input image. Alternatively, the memory (110) may store an image acquired through a camera provided in the electronic device (100). Here, the image may be a 2D video, but is not limited thereto.

[0067] As another example, the memory (110) may store various information required for image quality processing, such as information, algorithms, image quality parameters, etc. for performing at least one of Noise Reduction, Detail Enhancement, Tone Mapping, Contrast Enhancement, Color Enhancement, or Frame Rate Conversion. In addition, the memory (110) may also store an intermediate image generated by image processing and an image generated based on depth information.

[0068] According to one embodiment, the memory (110) may be implemented as a single memory that stores data generated from various operations according to the present disclosure. However, according to another embodiment, the memory (110) may be implemented to include multiple memories that each store different types of data or each store data generated at different stages.

[0069] In the above-described embodiment, it has been described that various data are stored in the external memory (110) of the processor (120), but at least some of the above-described data may be stored in the internal memory of the processor (120) according to an implementation example of at least one of the electronic device (100) or the processor (120).

[0070] At least one processor (120) controls the overall operation of the electronic device (100). Specifically, at least one processor (120) is connected to each component of the electronic device (100) and can control the overall operation of the electronic device (100). For example, at least one processor (120) is operatively connected to the memory (110) and can control the overall operation of the electronic device (100). At least one processor (120) may be composed of one or more processors.

[0071] At least one processor (120) can perform operations of the electronic device (100) according to various embodiments by executing at least one instruction stored in the memory (110).

[0072] At least one processor (120) may include one or more of a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), an Accelerated Processing Unit (APU), a Many Integrated Core (MIC), a Digital Signal Processor (DSP), a Neural Processing Unit (NPU), a hardware accelerator, or a machine learning accelerator. The at least one processor (120) may control one or any combination of other components of the electronic device, and may perform operations related to communication or data processing. The at least one processor (120) may execute one or more programs or instructions stored in a memory. For example, the at least one processor may perform a method according to one or more embodiments of the present disclosure by executing at least one instruction stored in the memory.

[0073] When a method according to one or more embodiments of the present disclosure includes multiple operations, the multiple operations may be performed by one processor or by multiple processors. For example, when a first operation, a second operation, and a third operation are performed by a method according to one or more embodiments, the first operation, the second operation, and the third operation may all be performed by the first processor, or the first operation and the second operation may be performed by the first processor (e.g., a general-purpose processor) and the third operation may be performed by the second processor (e.g., an artificial intelligence-specific processor).

[0074] At least one processor (120) may be implemented as a single core processor including one core, or may be implemented as one or more multicore processors including multiple cores (e.g., homogeneous multicores or heterogeneous multicores). When at least one processor (120) is implemented as a multicore processor, each of the multiple cores included in the multicore processor may include an internal processor memory, such as a cache memory or an on-chip memory, and a common cache shared by the multiple cores may be included in the multicore processor. In addition, each of the multiple cores (or some of the multiple cores) included in the multicore processor may independently read and execute a program instruction for implementing a method according to one or more embodiments of the present disclosure, or all (or some) of the multiple cores may be linked to read and execute a program instruction for implementing a method according to one or more embodiments of the present disclosure.

[0075] When a method according to one or more embodiments of the present disclosure includes a plurality of operations, the plurality of operations may be performed by one core among the plurality of cores included in a multi-core processor, or may be performed by the plurality of cores. For example, when a first operation, a second operation, and a third operation are performed by a method according to one or more embodiments, the first operation, the second operation, and the third operation may all be performed by a first core included in the multi-core processor, or the first operation and the second operation may be performed by a first core included in the multi-core processor, and the third operation may be performed by a second core included in the multi-core processor.

[0076] In embodiments of the present disclosure, a processor may mean a system on a chip (SoC) in which at least one processor and other electronic components are integrated, a single-core processor, a multi-core processor, or a core included in a single-core processor or a multi-core processor, wherein the core may be implemented as a CPU, a GPU, an APU, a MIC, a DSP, an NPU, a hardware accelerator, or a machine learning accelerator, but embodiments of the present disclosure are not limited thereto. Hereinafter, for convenience of description, at least one processor (120) will be referred to as a processor (120).

[0077] According to one embodiment, the electronic device (100) can receive various compressed images or images of various resolutions. For example, the electronic device (100) can receive images in a compressed form such as MPEG (Moving Picture Experts Group) (e.g., MP2, MP4, MP7, etc.), JPEG (joint photographic coding experts group), AVC (Advanced Video Coding), H.264, H.265, HEVC (High Efficiency Video Codec), etc. Alternatively, the electronic device (100) can receive any one of SD (Standard Definition), HD (High Definition), Full HD, and Ultra HD images.

[0078] According to one embodiment, the processor (120) can obtain depth information from an input image. Here, the input image may include a still image, a plurality of consecutive still images (or frames), or a video. For example, the input image may be a 2D image.

[0079] Depth information can take the form of a depth map. A depth map represents three-dimensional distance information for objects within an image, and can be assigned to each pixel of the image. For example, an 8-bit depth (or depth) can have a grayscale value from 0 to 255. For example, when expressed based on black and white, black (low value) can represent an area farther from the viewer, and white (high value) can represent an area closer to the viewer. However, this is only an example, and the depth map can be expressed with various values ​​based on various criteria.

[0080] For example, the processor (120) may obtain a depth map from a 2D input image based on a depth estimation algorithm, a mathematical formula, a learned artificial intelligence model, etc. For example, the artificial intelligence model may be implemented as a neural network including a plurality of neural network layers. The artificial intelligence model may be implemented as a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), or a deep Q-network, but is not limited thereto.

[0081] According to an example, the processor (120) may process an input image and then acquire depth information based on the processed image. Here, the image processing may be digital image processing including at least one of image enhancement, image restoration, image transformation, image analysis, image understanding, image compression, image decoding, or scaling.

[0082] For example, various preprocessing may be performed before acquiring depth information for an input image. However, for convenience of explanation, the input image and the preprocessed image will not be distinguished and will be referred to as an input image.

[0083] In this specification, "region" is a term referring to a portion of an image, and means at least one pixel block or a set of pixel blocks. In addition, a pixel block means a set of adjacent pixels that include at least one pixel.

[0084] The processor (120) may store depth information corresponding to the input image, for example, a depth map, in the memory (110). For example, when a first frame and a second frame are sequentially input, the processor (120) may sequentially acquire a first depth map and a second depth map corresponding to the first frame and the second frame while applying preprocessing and / or postprocessing to the first frame and the second frame, and store the acquired first depth map and the second depth map in the memory (110). For example, the first frame and the second frame may be 2D monocular image frames. The first frame and the second frame are frames constituting a video, and if the second frame is a frame currently being processed, the first frame may be a previously processed frame. For example, if the video is a real-time streaming video, the first frame may be a frame streamed before the second frame. However, the present invention is not limited thereto, and the first frame may be a frame prior to a preset frame interval (for example, a 2-frame interval) based on the second frame. For convenience of explanation, the first and second frames are referred to as the previous frame and the current frame, respectively, below.

[0085] For example, the processor (120) can obtain depth maps of the first frame and the second frame based on various image processing methods, such as algorithms, formulas, artificial intelligence models, etc.

[0086] According to one embodiment, the processor (120) can identify hole filling complexity for a hole region that occurs in a boundary region between multiple object regions.

[0087] According to one embodiment, the processor (120) may identify a viewpoint of a new viewpoint image based on the hole-filling complexity and acquire a new viewpoint image based on the identified viewpoint. For example, the processor (120) may identify the hole-filling complexity based on the size of the hole and the complexity of the hole-filling context.

[0088] For example, the processor (120) may identify a hole filling complexity for a hole region based on at least one of the complexity of a boundary region (hereinafter, boundary complexity), the size of a hole occurring in the boundary region, or the complexity of a reference region for filling the hole.

[0089] For example, the processor (120) can acquire a new viewpoint image based on a viewpoint with a relatively low hole-filling complexity among the right-eye viewpoint or the left-eye viewpoint.

[0090] For example, the processor (120) can acquire a new viewpoint image based on the same viewpoint in units of preset frame intervals. For example, the preset frame interval units may include one frame unit or a scene unit.

[0091] For example, the processor (120) may analyze the hole-filling complexity on a frame-by-frame basis and obtain a new viewpoint image based on the analysis results. For example, the processor (120) may analyze the hole-filling complexity for a first frame and obtain a right-eye image as a new viewpoint image using the first frame as a left-eye image. The processor (120) may analyze the hole-filling complexity for a second frame and obtain a left-eye image as a new viewpoint image using the second frame as a right-eye image. The processor (120) may analyze the hole-filling complexity for a third frame and obtain a right-eye image as a new viewpoint image using the third frame as a left-eye image.

[0092] For example, the processor (120) may analyze hole-filling complexity on a scene-by-scene basis (e.g., multiple frames constituting a scene) and obtain a new viewpoint image based on the analysis results. For example, if the processor (120) analyzes the hole-filling complexity for the first frame included in the scene and identifies a new viewpoint image, the processor (120) may apply the same new viewpoint to all frames included in the scene to obtain a new viewpoint image. However, the first frame in the scene is merely an example, and the frame for determining a new viewpoint image in a scene need not necessarily be the first frame. For example, the new viewpoints to be applied in a scene may be identified in various ways, such as generating many new viewpoints among the new viewpoints (e.g., left-eye viewpoints and right-eye viewpoints) identified for each of multiple frames in the scene. It will be understood that generating a new viewpoint image based on the same viewpoint in a scene in this way may be suitable for non-real-time conversion, but is not limited thereto. As another example, the processor (120) may generate an effective new viewpoint image on a frame-by-frame basis in a scene. For example, when the processor (120) generates a new viewpoint image based on a viewpoint identified for each frame within a scene, the processor (120) may generate the new viewpoint image so that the viewpoint changes smoothly and continuously between consecutive frames through filtering. According to an example, the processor (120) may generate the new viewpoint image so that the viewpoint changes smoothly between consecutive frames through an IIR (Infinite Impulse Response) filter or an FIR (Finite Impulse Response) filter. For example, the depth map may be an 8-bit map, and among the grayscale values ​​from 0 to 255 included in the bitmap, 127 (or 128) may be used as a reference value, that is, 0 (or focal plane), and a form may be implemented in which a value smaller than 127 is represented as a - value and a larger value is represented as a + value.For example, when analyzing the hole-filling complexity of each frame, it is assumed that the average depth value corresponding to the right-eye image acquired based on the first frame is 5 and the average depth value corresponding to the left-eye image is -5, the average depth value corresponding to the right-eye image acquired based on the second frame is 4 and the average depth value corresponding to the left-eye image is -5, and the average depth value corresponding to the right-eye image acquired based on the third frame is 3 and the average depth value corresponding to the left-eye image is -4. In this case, the processor (120) can obtain the final binocular images corresponding to the first frame, the second frame, and the third frame based on depth values ​​such as (5, -5), (4, -6), (3, -7), etc., or (5, -5), (6, -4), (7, -3) using IIR. For example, the processor (120) can obtain a binocular image that smoothly changes while maintaining the depth difference between the binocular images corresponding to each frame by applying IIR to multiple frames included in each scene. In this way, identifying an effective viewpoint for each frame within a scene and generating a new viewpoint image may be suitable for real-time conversion, but is not limited thereto.

[0093] FIG. 2b is a block diagram specifically illustrating a configuration of an electronic device according to one or more embodiments.

[0094] According to FIG. 2b, the electronic device (100') may include a memory (110), at least one processor (120), a display (130), a camera (140), a user interface (150), a communication interface (160), and a speaker (170). Among the configurations illustrated in FIG. 2b, a detailed description of configurations that overlap with those illustrated in FIG. 2a will be omitted.

[0095] The display (130) may be implemented as a display including a self-luminous element or a display including a non-luminous element and a backlight. For example, it may be implemented as various types of displays such as an LCD (Liquid Crystal Display), an OLED (Organic Light Emitting Diodes) display, an LED (Light Emitting Diodes), a micro LED, a Mini LED, a PDP (Plasma Display Panel), a QD (Quantum dot) display, a QLED (Quantum dot light-emitting diodes), etc. The display (130) may also include a driving circuit, a backlight unit, etc., which may be implemented in a form such as an a-si TFT, an LTPS (low temperature poly silicon) TFT, an OTFT (organic TFT), etc. According to an example, a touch sensor that detects a touch operation in the form of a touch film, a touch sheet, a touch pad, etc. may be disposed on the front of the display (130) so as to be implemented so as to detect various types of touch inputs. For example, the display (130) can detect various types of touch inputs, such as a touch input by a user's hand, a touch input by an input device such as a stylus pen, and a touch input by a specific electrostatic material. Here, the input device can be implemented as a pen-type input device that can be referred to by various terms such as an electronic pen, a stylus pen, an S-pen, etc. According to an example, the display (130) can be implemented as a flat display, a curved display, a flexible display that can be folded or / and rolled, etc.

[0096] In one example, the display (130) may provide an output image composed of a binocular image including at least one new viewpoint image.

[0097] The camera (140) can be turned on and take pictures according to a preset event. The camera (140) can convert the captured image into an electrical signal and generate image data based on the converted signal. For example, the subject can be converted into an electrical image signal through a semiconductor optical element (CCD; Charge Coupled Device), and the converted image signal can be amplified and converted into a digital signal and then signal processed. For example, the camera (120) can include at least one of a general (or basic) camera and an ultra-wide-angle camera.

[0098] For example, a camera (140) may acquire a user-captured image and provide it to a processor (120). The processor (120) may detect the user's face position from the user-captured image and identify the user's eyes from the user's face to acquire the user's gaze information in real time. The processor (120) may rearrange the binocular images based on the user's real-time gaze information (or eye tracking information) to acquire an output image.

[0099] Various conventional methods can be used for facial region detection. Specifically, direct recognition methods and statistical methods can be used. Direct recognition methods create rules using physical features of the facial image, such as the contours, skin color, and the sizes and distances between components, and compare, inspect, and measure based on these rules. Statistical methods can detect the facial region based on a pre-trained algorithm. That is, the unique features contained in the input face are digitized and compared and analyzed against a large database (of faces and other objects). In particular, methods such as Multi-Layer Perceptron (MLP) and Support Vector Machine (SVM) can be used to detect the facial region based on a pre-trained algorithm. A similar method can be used to identify the user's eye region.

[0100] The user interface (150) may be implemented as a device such as a button, a touch pad, a mouse, and a keyboard, or as a touch screen that can also perform the display function and operation input function described above.

[0101] It goes without saying that the communication interface (160) can be implemented as various interfaces depending on the implementation example of the electronic device (100'). For example, the communication interface (140) can communicate with an external device, an external storage medium (e.g., a USB memory), an external server (e.g., a web hard drive), etc. through a communication method such as Bluetooth, AP-based Wi-Fi (Wireless LAN network), Zigbee, wired / wireless LAN (Local Area Network), WAN (Wide Area Network), Ethernet, IEEE 1394, HDMI (High-Definition Multimedia Interface), USB (Universal Serial Bus), MHL (Mobile High-Definition Link), AES / EBU (Audio Engineering Society / European Broadcasting Union), optical, coaxial, etc. In one example, the communication interface (160) can communicate with another electronic device, an external server, and / or a remote control device.

[0102] The speaker (170) may be configured to output various audio data as well as various notification sounds or voice messages. The processor (130) may control the speaker (170) to output feedback or various notifications in audio format according to various embodiments of the present disclosure.

[0103] In addition, the electronic device (100') may include sensors, microphones, etc., depending on the implementation example.

[0104] Sensors may include various types of sensors, such as touch sensors, proximity sensors, acceleration sensors (or gravity sensors), geomagnetic sensors, gyro sensors, pressure sensors, position sensors, distance sensors, light sensors, etc.

[0105] A microphone is a device configured to receive user voice or other sounds and convert them into audio data. However, according to another embodiment, the electronic device (100') may receive user voice input via an external device through a communication interface (160).

[0106] FIG. 3 is a drawing for explaining the structure and operation of a display (130) according to one or more embodiments.

[0107] According to FIG. 3, the display (130) may include a display panel (131), a display separator (132), and a backlight unit (133). However, depending on the implementation example of the display (130), the backlight unit (133) may not be included in the display (130).

[0108] The display panel (131) includes a plurality of pixels composed of a plurality of sub-pixels. Here, the sub-pixels may be composed of R (Red), G (Green), and B (Blue). For example, pixels composed of R, G, and B sub-pixels may be arranged in a plurality of row and column directions to form the display panel (141).

[0109] The display panel (131) displays a binocular viewpoint image (or a multi-viewpoint image). For example, the display panel (131) can display an image in which multiple images of a right-eye image and a left-eye image are sequentially and repeatedly arranged.

[0110] The viewing area separator (132) is arranged on the front of the display panel (131) to provide different viewpoints, i.e., multi-views, for each viewing area. In this case, the viewing area separator (132) may be implemented as a lenticular lens or a parallax barrier. For example, the viewing area separator (132) may be implemented as a lenticular lens including a plurality of lens regions. Accordingly, the lenticular lens may refract an image displayed on the display panel (131) through the plurality of lens regions. Each lens region is formed to have a size corresponding to at least one pixel, so that light passing through each pixel may be differently dispersed for each viewing area. As another example, the viewing area separator (132) may be implemented as a parallax barrier. The parallax barrier is implemented as a transparent slit array including a plurality of barrier regions. Accordingly, light can be blocked through slits between barrier areas to allow images from different viewpoints to be output for each viewing area.

[0111] For example, the field of view separation unit (132) may operate at a certain angle to improve image quality. In this case, the processor (130) may divide the right-eye image and the left-eye image based on the angle at which the field of view separation unit (132) is tilted, and combine them to generate a multi-view image. Accordingly, the user does not view the image displayed vertically or horizontally on the sub-pixels of the display panel (131), but rather views the image displayed with a certain angle on the sub-pixels.

[0112] As an example, the field separation unit (132) can be implemented as a lenticular lens array as illustrated in FIG. 3.

[0113] According to FIG. 3, the display panel (131) includes a plurality of pixels divided into a plurality of columns. Images of different viewpoints, for example, left-eye images and right-eye images, may be alternately arranged in each column. For example, as illustrated in FIG. 3, left-eye images and right-eye images 1 and 2 may be sequentially and repeatedly arranged.

[0114] The backlight unit (133) provides light to the display panel (131). By the light provided from the backlight unit (133), the left-eye image and the right-eye image 1, 2 formed on the display panel (131) are projected to the viewing area separator (132), and the viewing area separator (132) can disperse the light of each projected image 1, 2 and transmit it toward the viewer. For example, the viewing area separator (132) can create exit pupils at the viewer's position, i.e., the viewing distance. As illustrated in FIG. 3, when the viewing area separator (132) is implemented as a lenticular lens array, the thickness and diameter of the lenticular lens, and when it is implemented as a parallax barrier, the spacing of the slits, etc. can be designed so that the exit pupils created by each row are separated by an average binocular center distance of less than 65 mm. The separated lights can each form a viewing area.

[0115] FIG. 4 is a flowchart illustrating a method for controlling an electronic device according to one or more embodiments.

[0116] According to FIG. 4, in operation 410, the electronic device (100) can identify multiple object areas included in the input image based on a depth map corresponding to the input image.

[0117] For example, the electronic device (100) can use the difference in depth values ​​from neighboring pixels to determine an object and background or a boundary area between two objects in a depth map. For example, the electronic device (100) can define a pixel whose depth difference from a center pixel among four neighboring pixels on the top, bottom, left, and right is greater than a certain value as a depth boundary area. For example, as illustrated in FIG. 5A, in the depth map (510), an area (511) including a pixel whose depth difference from a center pixel among four neighboring pixels on the top, bottom, left, and right is greater than a certain value can be identified as a depth boundary area (520).

[0118] For example, the electronic device (100) can identify the boundary area of ​​an object within a depth map and then identify the size of the object through a morphological operation. For example, the electronic device (100) can identify an object area included in an input image through at least one of object recognition, object detection, object tracking, or image segmentation. For example, the electronic device (100) can identify an object area by using a technique such as semantic segmentation, which classifies and extracts objects included in an input image by type as needed, instance segmentation, which recognizes objects by classifying them by object even if they are of the same type, and a bounding box in the shape of a rectangle that includes the detected object when detecting an object included in an image.

[0119] In operation 420, the electronic device (100) can identify the hole-filling complexity for a hole region (or occlusion region) occurring in a boundary region between multiple object regions. The hole region may be an area that is not exposed in the foreground region in the current viewpoint, i.e., the input image, and is exposed when the viewpoint moves, i.e., in a new viewpoint image. For example, the electronic device (100) can set a hole region for each boundary between objects extracted from the depth map.

[0120] In one example, the electronic device (100) can identify a hole filling complexity for a hole region based on at least one of a boundary complexity, a size of a hole occurring in the boundary region, or a complexity of a reference region for filling the hole.

[0121] For example, boundary complexity may include dense information about objects. For example, when multiple objects are mixed around a boundary, the boundaries between objects become dense. This may increase the probability that the hole area is excessively set or unnecessary information is included in the context area (the area adjacent to the hole area), resulting in unstable results when filling the hole area. The context area may be an area adjacent to an occluded area. For example, the context area may be used as a reference area for filling the hole area. For example, the electronic device (100) may identify boundary complexity based on the number of neighboring boundaries included in the context area. For example, the electronic device (100) may initialize an occluded area within a range determined based on a target boundary and a corresponding context area, and define the number of boundaries existing within the initial context area as boundary complexity. For example, as illustrated in FIG. 5B, a context area may be identified (540) based on an image (530) representing a target boundary, and two boundaries within the context area may be identified based on an image (550) representing a boundary within the context area. Accordingly, the boundary complexity can be identified as 2. For example, as illustrated in FIG. 5c, a context area is identified (570) based on a target boundary (560), and five boundaries within the context area can be identified based on an image (580) representing the boundary within the context area. Accordingly, the boundary complexity can be identified as 5. However, the configuration for calculating the boundary complexity illustrated in FIGS. 5b and 5c is only an example and is not limited thereto, and the boundary complexity can be calculated in various ways.

[0122] According to one embodiment, the electronic device (100) may calculate the hole filling complexity by applying the same weight or different weights to the boundary complexity, the hole size, and the complexity of the reference region. For example, the electronic device (100) may apply the same preset weight or different weights to the boundary complexity, the hole size, and the complexity of the reference region. For example, the electronic device (100) may apply the same preset weight or different weights to the boundary complexity, the hole size, and the complexity of the reference region depending on the type of the image. For example, the electronic device (100) may apply the same preset weight or different weights to the boundary complexity, the hole size, and the complexity of the reference region based on the characteristics of each frame section included in the image. For example, each frame section may include one frame unit or one scene unit.

[0123] In operation 430, the electronic device (100) can identify a viewpoint of a new viewpoint image based on the hole filling complexity.

[0124] In one example, the electronic device (100) can identify a viewpoint of a new viewpoint image based on a viewpoint with relatively low hole-filling complexity, which is at least one of the right-eye viewpoint or the left-eye viewpoint.

[0125] In operation 440, the electronic device (100) can acquire a new viewpoint image based on the identified viewpoint.

[0126] In one example, if the electronic device (100) determines that the right-eye image has a relatively lower hole-filling complexity than the left-eye image, the electronic device (100) may generate the right-eye image as a new viewpoint image. In another example, if the electronic device (100) determines that the left-eye image has a relatively lower hole-filling complexity than the right-eye image, the electronic device (100) may generate the left-eye image as a new viewpoint image. In another example, if the electronic device (100) determines that generating both the right-eye image and the left-eye image based on the input image has a relatively lower hole-filling complexity, the electronic device (100) may generate both the right-eye image and the left-eye image.

[0127] For example, the electronic device (100) can analyze hole filling complexity on a frame-by-frame basis and obtain a new viewpoint image based on the analysis results.

[0128] For example, the electronic device (100) can analyze the hole-filling complexity on a scene-by-scene basis and acquire a new viewpoint image based on the analysis results. For example, the electronic device (100) can acquire a new viewpoint image based on the same viewpoint within the same scene, and when the scene changes, it can acquire a new viewpoint image based on a new viewpoint.

[0129] As described above, depending on the characteristics of the scene composition of the image, etc., the size and complexity of the hole area that occurs when generating a left-eye image or a right-eye image are different, so generating a viewpoint where the hole size is relatively small and hole filling is easy is advantageous for generating a natural image.

[0130] Meanwhile, in Fig. 4, the order is mapped for all steps for convenience of explanation, but it is of course not necessarily limited to the order of steps that are not related to the order or can be performed in parallel.

[0131] FIGS. 6 and 7 are drawings for explaining a method of controlling an electronic device according to one or more embodiments.

[0132] FIG. 6 is a flowchart for explaining a control method of an electronic device (100) according to one or more embodiments, and FIG. 7 is a block diagram for explaining a control method of an electronic device (100) according to one or more embodiments. According to FIG. 7, the electronic device (100) may include a depth estimation module (710), a warping module (720), a viewpoint determination module (730) (or may be a viewpoint determination module), and a hole filling module (740). For example, each module may be implemented with at least one software, at least one hardware, and / or a combination thereof.

[0133] For example, at least one of the depth estimation module (710), the warping module (720), the viewpoint determination module (730), and the hole filling module (740) may be implemented to utilize a predefined algorithm, a predefined formula, and / or a learned artificial intelligence model. The depth estimation module (710), the warping module (720), the viewpoint determination module (730), and the hole filling module (740) may be included within the electronic device (100), but may be distributed to at least one external device according to an example.

[0134] According to FIG. 6, in operation 610, the electronic device (100) may obtain a depth map corresponding to the input image. In one example, the processor (130) may obtain the depth map using the depth estimation module (710) and identify an object area included in the input image, for example, a foreground area and a background area.

[0135] For example, the depth estimation module (710) can obtain a depth map according to a block matching-based method. The block matching-based method designates a block to be used for search in the current image, and then searches for a block with the most similar value to the block used for search in the previous image, determines it as the same point, and generates depth information based on the difference in coordinate values ​​between the block used for search and the searched block. However, the present invention is not limited thereto, and a depth map can be obtained according to various methods capable of estimating depth information.

[0136] For example, the depth estimation module (710) can identify object regions included in the input image, such as foreground regions and background regions, based on the depth map and perform warping. For example, in a scene where two regions meet at a boundary, the part that is the object of perception can be called the foreground, and the rest can be called the background. However, the identification of the foreground region and the background region can also be performed through the warping module (720).

[0137] In operation 620, the electronic device (100) may perform warping on at least one of the right-eye image or the left-eye image. In one example, the processor (130) may perform warping using the warping module (720).

[0138] Warping refers to a technology for generating a new viewpoint image, and a new viewpoint image can be generated by performing hole filling on the hole area included in the warping image.

[0139] In operation 630, the electronic device (100) may identify at least one of the boundary complexity, the size of a hole generated in a boundary area, or the complexity of a reference area in at least one of the right-eye image or the left-eye image generated in the warping process. In one example, the processor (130) may use the viewpoint determination module (730) to identify at least one of the boundary complexity, the size of a hole generated in a boundary area, or the complexity of a reference area in at least one of the right-eye image or the left-eye image generated in the warping process.

[0140] For example, the electronic device (100) can identify the size of a hole by counting holes in a warping image generated during a warping process.

[0141] In one example, the electronic device (100) can identify the complexity of the reference area based on at least one of texture information or edge information corresponding to the reference area. Texture information refers to a unique pattern or shape of an area in an image that is considered to have the same texture. There are various types of edges in an image, and in one example, the edge information can include information about complex edges with various directions and straight edges with a clear direction. In general, by applying a first or second edge detection filter to an input image, edge information including edge strength and edge direction information (perpendicular to the gradient) can be obtained.

[0142] In operation 640, the electronic device (100) may identify an image with a relatively low hole-filling complexity among the right-eye image or the left-eye image as a new viewpoint image. For example, the processor (130) may use the viewpoint determination module (730) to identify an image with a relatively low hole-filling complexity among the right-eye image or the left-eye image as a new viewpoint image. For example, the electronic device (100) may determine that the larger the hole area, the higher the hole-filling complexity. For example, if the reference area has a complex texture or / and pattern, the electronic device (100) may determine that the hole-filling complexity is higher compared to a flat surface.

[0143] In operation 650, the electronic device (100) can obtain a new viewpoint image identified in operation 630. In one example, the processor (130) can generate a new viewpoint image by performing hole filling on a hole area included in the processed warping image using a hole filling module (740).

[0144] Meanwhile, in Fig. 6, the order is mapped for all steps for convenience of explanation, but it is of course not necessarily limited to the order of steps that are not related to the order or can be performed in parallel.

[0145] FIGS. 8 and 9 are drawings for explaining a method of controlling an electronic device according to one or more embodiments.

[0146] Among the operations illustrated in Fig. 8, detailed descriptions of operations that overlap with those illustrated in Fig. 6 will be omitted.

[0147] According to FIG. 8, in operation 810, the electronic device (100) can obtain a depth map corresponding to the input image.

[0148] In operation 820, the electronic device (100) can identify multiple object regions included in the input image based on the depth map. In one example, the processor (130) can obtain a depth map using the depth estimation module (710) and identify object regions included in the input image, for example, a foreground region and a background region.

[0149] In operation 830, the electronic device (100) may identify a boundary region and a reference region in the depth map based on a plurality of object regions identified in the depth map. In one example, the processor (130) may identify a boundary region and a reference region in the depth map based on a plurality of object regions identified in the depth map using a depth estimation module (710). However, the identification of the boundary region and the reference region in the depth map may also be performed through a viewpoint determination module (910).

[0150] For example, when a foreground region and a background region are identified in a scene where two regions meet at boundary regions, the depth estimation module (710) can identify a reference region included in the background region. The reference region may be a region for filling a hole region in a warped image.

[0151] In operation 840, the electronic device (100) may identify at least one of boundary complexity, the size of a hole corresponding to a boundary area, or the complexity of a reference area in the depth map. In one example, the processor (130) may use the viewpoint determination module (910) to identify at least one of boundary complexity, the size of a hole corresponding to a boundary area, or the complexity of a reference area in the depth map.

[0152] For example, the electronic device (100) can estimate the size of a hole corresponding to a boundary area based on step information for the boundary area in the depth map.

[0153] In one example, the electronic device (100) can identify the complexity of the reference area based on at least one of texture information or edge information corresponding to the reference area.

[0154] In operation 850, the electronic device (100) may identify a viewpoint with a relatively low hole-filling complexity among the right-eye viewpoint or the left-eye viewpoint based on at least one of the boundary complexity identified in the depth map, the size of the hole corresponding to the boundary area, or the complexity of the reference area. In one example, the processor (130) may identify a viewpoint with a relatively low hole-filling complexity among the right-eye viewpoint or the left-eye viewpoint using the viewpoint determination module (910).

[0155] In operation 860, the electronic device (100) may obtain a new viewpoint image based on the identified viewpoint. In one example, the processor (130) may obtain a warping image based on the identified viewpoint using the warping module (720), and may perform hole filling on the warping-processed warping image using the hole filling module (740) to obtain a new viewpoint image.

[0156] Meanwhile, in Fig. 8, the order is mapped for all steps for convenience of explanation, but it is of course not necessarily limited to the order of steps that are not related to the order or can be performed in parallel.

[0157] FIG. 10 is a drawing for explaining in detail a method for providing a 3D image according to one or more embodiments.

[0158] Among the configurations illustrated in Fig. 10, a detailed description of the configurations that overlap with those illustrated in Fig. 7 will be omitted. In addition, although Fig. 10 depicts a detailed configuration of the module illustrated in Fig. 7, it goes without saying that the detailed configuration of the module illustrated in Fig. 9 that overlaps with the module illustrated in Fig. 7 can also be implemented similarly.

[0159] According to FIG. 10, the depth estimation module (710) may include a downscaling module (711), a deep neural network module (712), and an upsampling and refinement module (713).

[0160] The downscaling module (711) can downscale (or downsample) the input image (10). For example, the downscaling module (711) can obtain a downscaled image by filtering the input image (10) by applying a filter to the input image (10). Here, filtering the input image (10) may mean weighting filter information on the input image (10). For example, as a downscaling method, at least one interpolation technique among bilinear interpolation, nearest neighbor interpolation, bicubic interpolation, deconvolution interpolation, subpixel convolution interpolation, polyphase interpolation, trilinear interpolation, and linear interpolation may be used.

[0161] The deep neural network module (712) can obtain a depth map (11) based on the downscaled image. In one example, the deep neural network module (712) can obtain a depth map (11) by inputting the downscaled image into the deep neural network.

[0162] The upsampling and refinement module (713) can upsample (or upscale) and refine the depth map (11) obtained from the deep neural network module (712). For example, the upsampling and refinement module (713) can upsample the depth map (11) using a method identical to / similar to the downscaling method.

[0163] For example, the upsampling and refinement module (713) may perform depth map refinement based on object density information, object thickness, etc. included in the depth map. For example, the upsampling and refinement module (713) may perform depth map refinement through median filter-based refinement or / and color information comparison-based refinement. Median filter-based refinement is a method of assigning the median value of the depth values ​​of pixels within a window as an output value. In this process, pixels in the depth boundary area may be excluded. The median filter-based refinement method can obtain a stable depth boundary and is effective in alleviating the depth value inversion phenomenon that occurs at the depth boundary. Color information comparison-based refinement is a method of assigning the depth value of the pixel with the most similar color value to the refinement target pixel among the pixels within the window as an output value. In this process, pixels in the depth boundary area may be excluded from the comparison. For example, the degree of color value similarity may be calculated as the Euclidean distance between the refinement target pixel and the comparison target pixel. The color information comparison-based refinement method enables precise depth map refinement even when the object is thin.

[0164] The warping module (720) can perform warping processing based on the depth map (12) obtained through the upsampling and refinement module (713).

[0165] The hole filling module (740) can perform hole filling on a hole area included in a warping image to obtain a binocular image (50) including at least one new viewpoint image.

[0166] The Light Field View Mapping module can obtain an output image (60) to be displayed on the display (130) based on the binocular image (50). In one example, the output image (60) may be a Side By Side image. The Side By Side image is an image in which the left-eye image and the right-eye image are each sub-sampled by 1 / 2 in the horizontal direction, and the left-eye image and the right-eye image (50) may be alternately arranged in the output image (60). For example, as illustrated in FIG. 3, an output image (60) in which the left-eye image and right-eye images 1 and 2 are sequentially and repeatedly arranged may be provided.

[0167] FIGS. 11A and 11B are drawings for explaining a method of obtaining information using an artificial intelligence model according to one or more embodiments.

[0168] According to one example, the electronic device (100) can input an input image (10) and a depth map (20) into a learned artificial intelligence model, as illustrated in FIG. 11A, to obtain information about an image at a viewpoint with a relatively low hole-filling complexity among a plurality of different viewpoint images. For example, the electronic device (100) can obtain information about an image at a viewpoint with a relatively low hole-filling complexity among the left-eye image and the right-eye image from the learned artificial intelligence model.

[0169] According to one example, the electronic device (100) can obtain information on the hole-filling complexity for each of a plurality of different viewpoint images by inputting an input image (10) and a depth map (20) into a learned artificial intelligence model, as illustrated in FIG. 11B. For example, the electronic device (100) can obtain information on the hole-filling complexity for each of the left-eye image and the right-eye image from the learned artificial intelligence model.

[0170] Here, the learning of the artificial intelligence model means that a basic artificial intelligence model (e.g., an artificial intelligence model including any random parameters) is learned using a plurality of training data by a learning algorithm, thereby creating a predefined operation rule or artificial intelligence model set to perform a desired characteristic (or purpose). Such learning may be performed through a separate server and / or system, but is not limited thereto, and may also be performed in the electronic device (100). Examples of the learning algorithm include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.

[0171] Here, the artificial intelligence model can be implemented as, for example, a Convolutional Neural Network (CNN), a Recurrent Neural Network (RNN), a Restricted Boltzmann Machine (RBM), a Deep Belief Network (DBN), a Bidirectional Recurrent Deep Neural Network (BRDNN), or a Deep Q-Network, but is not limited thereto.

[0172] FIGS. 12A to 12C are drawings for explaining a method for generating a new viewpoint image according to one or more embodiments.

[0173] According to one embodiment, the electronic device (100) can acquire a new viewpoint image based on the same viewpoint in units of preset frame sections. For example, the preset frame section unit may include a single frame unit or a scene unit. According to one example, the electronic device (100) can identify a viewpoint effective for hole filling (or a favorable viewpoint) in units of frames and generate a new viewpoint image based on the identified viewpoint. According to one example, the electronic device (100) can identify a viewpoint effective for hole filling in units of scenes and generate a new viewpoint image based on the identified viewpoint.

[0174] According to FIG. 12a, if the electronic device (100) determines that generating a right-eye image using the input image as the left-eye image is advantageous for hole filling, the electronic device can generate the right-eye image as a new viewpoint image based on the input image.

[0175] According to FIG. 12b, if the electronic device (100) determines that generating a left-eye image using the input image as the right-eye image is advantageous for hole filling, the electronic device can generate the left-eye image as a new viewpoint image based on the input image.

[0176] According to FIG. 12c, if it is determined that generating both a right-eye image and a left-eye image based on an input image is effective for hole filling, the electronic device (100) can generate the left-eye image and the right-eye image as new viewpoint images based on the input image.

[0177] FIGS. 13A to 13E are drawings for explaining a method for generating a new viewpoint image according to one or more embodiments.

[0178] According to one embodiment, when an input image (1310) as illustrated in FIG. 13a is acquired, the electronic device (100) can acquire a depth map (1320) as illustrated in FIG. 13b based on the input image (1310).

[0179] According to an example, the electronic device (100) may acquire a first warping image (1330) corresponding to a right-eye viewpoint as illustrated in FIG. 13C, and identify hole-filling complexity based on the location of the hole area (1331) included in the first warping image (1330) and the size of the hole area (1331). For example, the electronic device (100) may identify the size of the hole area (1331) based on the number of pixels corresponding to the hole area (1331) included in the first warping image (1330). For example, the electronic device (100) may identify a reference area based on the location of the hole area (1331) and identify the complexity of the reference area. For example, the electronic device (100) may identify hole-filling complexity for the hole area (1331) based on the size of the hole area (1331) and the complexity of the reference area.

[0180] According to an example, the electronic device (100) may acquire a second warping image (1340) corresponding to a left eye viewpoint as illustrated in FIG. 13d and identify hole filling complexity based on the location and size of the hole area included in the second warping image (1340).

[0181] According to FIGS. 13c and 13d, a hole area (1331) larger than a certain size is generated in the first warping image (1330) corresponding to the right-eye viewpoint, but it can be confirmed that almost no hole area is generated in the second warping image (1340) corresponding to the left-eye viewpoint (1341). In this case, the electronic device (100) can identify the left-eye viewpoint corresponding to the second warping image (1340) that is effective for hole filling as a new viewpoint, and can identify the left-eye image as a new viewpoint image. For convenience of explanation, it is assumed that the second warping image (1340) illustrated in FIG. 13d does not generate a hole area and is therefore identical to the left-eye image.

[0182] According to FIG. 13e, the electronic device (100) can acquire a binocular image based on an input image (1310) and a left-eye image (1340) which is a new viewpoint image. In FIG. 13e, for convenience of explanation, it is assumed that the second warping image (1340) illustrated in FIG. 13d does not have a hole and is therefore identical to the left-eye image.

[0183] However, in FIGS. 13a to 13e, for convenience of explanation, the hole-filling complexity is described as being obtained through warping processing, but it is of course also possible to obtain the hole-filling complexity through depth map (1320) analysis before warping processing.

[0184] FIGS. 14A to 14E are drawings for explaining a method for generating a new viewpoint image according to one or more embodiments.

[0185] According to one embodiment, when an input image (1410) as illustrated in FIG. 14a is acquired, the electronic device (100) can acquire a depth map (1420) as illustrated in FIG. 14b based on the input image (1410).

[0186] For example, the electronic device (100) may acquire a third warping image (1430) corresponding to a left-eye viewpoint as illustrated in FIG. 14c, and identify hole-filling complexity based on the location of the hole area (1431) included in the third warping image (1430) and the size of the hole area (1431). For example, the electronic device (100) may identify hole-filling complexity for the hole area (1431) based on the size of the hole area (1431) and the complexity of the reference area.

[0187] According to an example, the electronic device (100) can obtain a fourth warping image (1440) corresponding to the right-eye viewpoint as illustrated in FIG. 14d and identify the hole filling complexity based on the location of the hole area (1441) included in the second warping image (1440) and the size of the hole area (1441).

[0188] According to FIGS. 14c and 14d, a hole area (1431) larger than a certain size is generated in the third warping image (1430) corresponding to the left-eye viewpoint, but in the fourth warping image (1440) corresponding to the right-eye viewpoint, it can be confirmed that almost no holes are generated in the area (1441) corresponding to the third warping image (1430), but holes are generated in other areas (1442). In this case, since the size of the hole area (1442) generated in the fourth warping image (1440) is relatively smaller than the hole area (1431) generated in the third warping image (1430), the electronic device (100) can identify the right-eye viewpoint corresponding to the fourth warping image (1440) that is effective for hole filling as a new viewpoint, and can identify the right-eye image as a new viewpoint image.

[0189] According to FIG. 14e, the electronic device (100) can acquire a binocular image based on the input image (1410) and the right-eye image (1440-1), which is a new viewpoint image. For example, the electronic device (100) can acquire the right-eye image (1440-1) by filling the hole area (1442) in the fourth warping image (1440).

[0190] However, in FIGS. 14a to 14e, for convenience of explanation, the hole-filling complexity is described as being obtained through warping processing, but it is of course also possible to obtain the hole-filling complexity through depth map (1420) analysis before warping processing.

[0191] FIGS. 15A to 15C are drawings for explaining a method for generating a new viewpoint image according to one or more embodiments.

[0192] According to one embodiment, the electronic device (100) may generate a left-eye image as a new viewpoint image when generating a left-eye image using the input image as the right-eye viewpoint is effective for hole filling. Accordingly, the electronic device (100) may acquire a binocular image including the input image and the left-eye image.

[0193] For example, in the case of FIGS. 15a, 15b, and 15c, when generating a left-eye image based on the input images (1510, 1520, 1530), the size of the hole area (1511, 1521, 1531, 1532) is very small, the background is simple, and the boundary is simple, so that the left-eye image can be generated with almost no side effects. In addition, accordingly, the electronic device (100) can generate a left-eye image based on the input image as the right-eye viewpoint, thereby obtaining a binocular image.

[0194] FIGS. 16A and 16B are drawings illustrating a method for generating a new viewpoint image according to one or more embodiments.

[0195] According to one embodiment, the electronic device (100) may generate a right-eye image as a new viewpoint image when generating a right-eye image using the input image as a left-eye viewpoint is effective for hole filling. Accordingly, the electronic device (100) can acquire a binocular image including the input image and the right-eye image.

[0196] For example, in the case of FIGS. 16a and 16b, when generating a right-eye image based on the input images (1610, 1620), the size of the hole areas (1611, 1612) is very small, the background is simple, and the border is also simple, so that a right-eye image can be generated with almost no side effects. Accordingly, the electronic device (100) can generate a right-eye image based on the input image as the left-eye viewpoint, thereby obtaining a binocular image.

[0197] According to the various embodiments described above, a natural 3D image can be obtained by generating a new viewpoint image from an image at a point in time when the hole size is small and hole filling is easy, depending on the characteristics of the image as described above.

[0198] Meanwhile, the methods according to the various embodiments of the present disclosure described above can be implemented only with a software upgrade or a hardware upgrade for an existing electronic device and / or server.

[0199] Additionally, the various embodiments of the present disclosure described above can also be performed through an embedded server provided in an electronic device, or an external server of the electronic device.

[0200] Meanwhile, according to a temporary example of the present disclosure, the various embodiments described above can be implemented as software including instructions stored in a machine-readable storage medium that can be read by a machine (e.g., a computer). The device is a device that can call instructions stored from the storage medium and operate according to the called instructions, and may include an electronic device (e.g., electronic device (A)) according to the disclosed embodiments. When an instruction is executed by a processor, the processor can perform a function corresponding to the instruction directly or by using other components under the control of the processor. The instruction may include code generated or executed by a compiler or interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' means that the storage medium does not contain a signal and is tangible, but does not distinguish between data being stored semi-permanently or temporarily in the storage medium.

[0201] Furthermore, according to one embodiment of the present disclosure, the method according to the various embodiments described above may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or online through an application store (e.g., Play Store™). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0202] In addition, each of the components (e.g., modules or programs) according to the various embodiments described above may be composed of a single or multiple entities, and some of the corresponding sub-components described above may be omitted, or other sub-components may be further included in various embodiments. Alternatively or additionally, some components (e.g., modules or programs) may be integrated into a single entity, which may perform the same or similar functions as those performed by each of the corresponding components prior to integration. Operations performed by modules, programs or other components according to various embodiments may be executed sequentially, in parallel, iteratively or heuristically, or at least some operations may be executed in a different order, omitted, or other operations may be added.

[0203] Although the preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above, and various modifications may be made by a person having ordinary skill in the art to which the present disclosure pertains without departing from the gist of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the present disclosure.

Claims

1. In electronic devices, memory for storing one or more instructions; and At least one processor for executing one or more of the above instructions; The above instructions, when executed by the at least one processor, cause the electronic device to: Identifying multiple object regions included in the input image based on a depth map corresponding to the input image, Identify the hole filling complexity for the hole area that occurs in the boundary area between the above multiple object areas, Identifying a viewpoint of a new viewpoint image based on the above hole filling complexity, An electronic device that obtains the new viewpoint image based on the above viewpoint.

2. In paragraph 1, The instructions, when executed by the at least one processor, cause the electronic device to identify a hole filling complexity for the hole region based on at least one of a complexity of the boundary region, a size of a hole generated in the boundary region, or a complexity of a reference region for filling the hole. An electronic device that acquires a new viewpoint image based on a viewpoint with a relatively low hole filling complexity among the right-eye viewpoint or the left-eye viewpoint.

3. In paragraph 2, The instructions, when executed by the at least one processor, cause the electronic device to perform warping on at least one of the right-eye image or the left-eye image based on the depth map to generate a warped right-eye image and a warped left-eye image, Identifying at least one of the complexity of the boundary region, the size of a hole occurring in the boundary region, or the complexity of the reference region in at least one of the warped right-eye image or the warped left-eye image, An electronic device that identifies an image with a relatively low hole-filling complexity among the right-eye image or the left-eye image as the new viewpoint image.

4. In paragraph 2, The instructions, when executed by the at least one processor, cause the electronic device to identify the plurality of object regions in the depth map based on depth information, Identifying a boundary region and a reference region in the depth map based on a plurality of object regions identified in the depth map, Identifying at least one of the complexity of the boundary area, the size of a hole corresponding to the boundary area, or the complexity of the reference area in the depth map, An electronic device that identifies a point in time in which the hole filling complexity is relatively low among the right eye viewpoint or the left eye viewpoint based on at least one of the complexity of the boundary area identified in the depth map, the size of a hole corresponding to the boundary area, or the complexity of the reference area.

5. In paragraph 2, The instructions, when executed by the at least one processor, cause the electronic device to estimate the size of a hole corresponding to the boundary area based on step information for the boundary area in the depth map, or An electronic device that counts holes in an image generated by the above warping to identify the size of the holes.

6. In paragraph 2, An electronic device wherein the instructions, when executed by the at least one processor, cause the electronic device to identify the complexity of each reference area based on at least one of texture information or edge information corresponding to the reference area.

7. In paragraph 1, The above instructions, when executed by the at least one processor, cause the electronic device to input the input image and the depth map into a learned artificial intelligence model to obtain information about an image of a viewpoint having a relatively low hole-filling complexity among a plurality of different viewpoint images, or An electronic device that inputs the input image and the depth map into a learned artificial intelligence model to obtain information on hole-filling complexity for each of the plurality of different viewpoint images.

8. In paragraph 1, The above instructions, when executed by the at least one processor, cause the electronic device to obtain the new viewpoint image based on the same viewpoint in units of preset frame intervals, The above preset frame interval unit is, An electronic device containing either a single frame unit or a single scene unit.

9. In paragraph 8, The above instructions, when executed by the at least one processor, cause the electronic device to generate the new viewpoint image by filtering so that the viewpoint changes smoothly between consecutive frames when the electronic device acquires the new viewpoint image based on the viewpoint identified for each frame or each scene.

10. In paragraph 1, The above electronic device, including display; The instructions, when executed by the at least one processor, cause the electronic device to obtain an output image including a right-eye output image based on the input image and the new viewpoint image and a left-eye output image based on the input image and the new viewpoint image, An electronic device that provides the above output image through the above display.

11. In a method for controlling an electronic device, A step of identifying a plurality of object regions included in an input image based on a depth map corresponding to the input image; A step of identifying the hole filling complexity for a hole area occurring in a boundary area between the plurality of object areas; and A control method comprising: a step of identifying a viewpoint of a new viewpoint image based on the hole filling complexity and obtaining the new viewpoint image based on the identified viewpoint.

12. In paragraph 11, The step of identifying the above hole filling complexity is: Identifying a hole filling complexity for the hole area based on at least one of the complexity of the boundary area, the size of a hole occurring in the boundary area, or the complexity of a reference area for filling the hole, The step of acquiring the above new viewpoint image is: A control method for acquiring a new viewpoint image based on a viewpoint with a relatively low hole filling complexity among the right-eye viewpoint or the left-eye viewpoint.

13. In paragraph 12, The step of identifying the above hole filling complexity is: A step of performing warping on at least one of the right-eye image or the left-eye image based on the depth map to generate a warped right-eye image and a warped left-eye image; and A step of identifying at least one of the complexity of the boundary region, the size of a hole generated in the boundary region, or the complexity of the reference region in at least one of the right-eye image or the left-eye image generated in the warping process; The step of acquiring the above new viewpoint image is: A control method comprising: a step of identifying an image having a relatively low hole-filling complexity among the right-eye image or the left-eye image as the new viewpoint image.

14. In paragraph 12, The step of identifying the above hole filling complexity is: A step of identifying the plurality of object regions in the depth map based on depth information included in the depth map; A step of identifying a boundary region and a reference region in the depth map based on a plurality of object regions identified in the depth map; and A step of identifying at least one of the complexity of the boundary area, the size of a hole corresponding to the boundary area, or the complexity of the reference area in the depth map; The step of acquiring the above new viewpoint image is: A control method comprising: a step of identifying a point in time in which the hole filling complexity is relatively low among the right eye viewpoint or the left eye viewpoint based on at least one of the complexity of the boundary area identified in the depth map, the size of a hole corresponding to the boundary area, or the complexity of the reference area.

15. A non-transitory computer-readable medium storing computer instructions that, when executed by a processor of an electronic device, cause the electronic device to perform an operation, The above action is, A step of identifying a plurality of object regions included in an input image based on a depth map corresponding to the input image; A step of identifying the hole filling complexity for a hole area occurring in a boundary area between the plurality of object areas; and A non-transitory computer-readable medium comprising: a step of identifying a viewpoint of a new viewpoint image based on the hole-filling complexity and obtaining the new viewpoint image based on the identified viewpoint.