Electronic device and control method thereof

By identifying foreground and background regions, predicting side effects, and optimizing depth maps, the viewpoint movement path is controlled, solving the problems of inaccurate boundary recognition and viewpoint limitation in novel view generation, and achieving high-quality novel view image generation.

CN121844558APending Publication Date: 2026-04-10SAMSUNG ELECTRONICS CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-07-17
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing technologies suffer from problems such as inaccurate foreground and background boundary recognition, disappearance of thin objects, limited viewpoint movement paths, and limited viewpoint generation when generating novel view images, resulting in significant side effects and affecting image quality.

Method used

By identifying foreground and background regions based on depth maps, predicting side effect information, controlling the viewpoint movement path, and applying opening operations and depth map optimization techniques, the depth map is optimized to reduce side effects, generating high-quality and novel view images.

Benefits of technology

It effectively reduces side effects in generating novel views, improves image quality, provides more viewpoint options, and avoids misidentification of thin objects and inaccurate boundary issues.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121844558A_ABST
    Figure CN121844558A_ABST
Patent Text Reader

Abstract

An electronic device is disclosed. The electronic device includes a memory to store at least one instruction; and one or more processors individually and / or collectively including processing circuitry, identifying a foreground region and a background region included in the input image based on a depth map corresponding to the input image, and converting a viewpoint based on the foreground region to generate a new viewpoint image. The one or more processors: identify side-effect prediction information based on the depth map, the side-effect prediction information comprising at least one of: whether an object of a maximum specified thickness is included in the input image or how dense in the object set; and generating a new viewpoint image by controlling the viewpoint movement path based on the identified side effect prediction information.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The disclosure relates to an electronic device and a control method thereof, and for example, to an electronic device generating a novel view image and a control method thereof. BACKGROUND

[0002] Various types of electronic devices are being developed and provided as electronic technology advances. In particular, display devices used in various places such as homes, offices, and public places have been developed in recent years.

[0003] A novel view synthesis technique is a technique of generating an image of a viewpoint different from a viewpoint of a 2-dimensional (2D) image obtained by a monocular camera. Using this technique, a continuous video of a novel view image according to movement of a viewpoint trajectory of a single image can be generated. SUMMARY

[0004] According to an example embodiment, a mobile electronic device includes a memory storing at least one instruction and at least one processor including processing circuitry, individually and / or collectively configured to: identify a foreground region and a background region included in an input image based on a depth map corresponding to the input image, and generate a novel view image by converting a viewpoint based on the foreground region; identify side effect prediction information based on the depth map, the side effect prediction information including at least one of whether an object less than or equal to a designated thickness is included in the input image or an object density degree; and generate the novel view image by controlling a viewpoint movement path based on the identified side effect prediction information.

[0005] According to an example embodiment, the at least one processor, individually and / or collectively, can be configured to: generate the novel view image by controlling the viewpoint movement path to be reduced to less than a threshold range based on the object less than or equal to the designated thickness being identified as included in the input image based on the depth map or the object density degree being identified as being greater than or equal to a threshold value based on the depth map.

[0006] According to an example embodiment, the at least one processor, individually and / or collectively, can be configured to: proportionally control the viewpoint movement path to be reduced by an area ratio occupied by the object less than or equal to the designated thickness in the input image based on the object less than or equal to the designated thickness being identified as included in the input image based on the depth map.

[0007] According to an example embodiment, the at least one processor can be configured to: obtain a depth map to which an opening operation is applied by applying the opening operation to the depth map, and identify a region including the object less than or equal to the designated thickness based on difference information between the depth map and the depth map to which the opening operation is applied.

[0008] According to an example embodiment, the at least one processor can be configured to calculate a standard deviation of depth values of pixels within a window excluding a depth boundary region by applying the window to the depth map, and perform a depth map refinement based on the calculated standard deviation being greater than or equal to a threshold value.

[0009] According to an example embodiment, the at least one processor can be configured to generate a novel view image by reducing a viewpoint movement path to be less than a threshold range based on the degree of object density being identified as being greater than or equal to a threshold value based on the depth map.

[0010] According to an example embodiment, the at least one processor can be configured to adjust an occlusion region and a context region based on a boundary complexity indicating the degree of object density, and the occlusion region can be a region that is not exposed from a current viewpoint by a foreground region and is exposed at a movement of the viewpoint, and the context region can be a region adjacent to the occlusion region.

[0011] According to an example embodiment, the at least one processor can be configured to identify the boundary complexity based on a number of adjacent boundaries included within the context region, and reduce a width of at least one of the occlusion region and the context region based on the number of adjacent boundaries being greater than or equal to a threshold number.

[0012] According to an example embodiment, the at least one processor can be configured to refine the depth map by replacing depth values of a region corresponding to an object less than or equal to a specified thickness with depth values of surrounding pixels having color information most similar to a target pixel.

[0013] According to an example embodiment, the at least one processor can be configured to refine the depth map by replacing depth values of other regions other than the region corresponding to the object less than or equal to the specified thickness with a median value of depth values of pixels within the window.

[0014] According to an example embodiment, a method of controlling an electronic device includes identifying a foreground region and a background region included in an input image and generating a novel view image by converting a viewpoint based on the foreground region based on a depth map corresponding to the input image, and the generating of the novel view image includes identifying side effect prediction information based on the depth map, the side effect prediction information including at least one of whether an object less than or equal to a specified thickness is included in the input image or a degree of object density, and generating the novel view image by controlling a viewpoint movement path based on the identified side effect prediction information.

[0015] According to an example embodiment, a non-transitory computer-readable medium storing computer commands, when executed by at least one processor of an electronic device individually and / or collectively, causes the electronic device to perform operations including: identifying a foreground region and a background region included in an input image based on a depth map corresponding to the input image, and generating a novel view image by converting a viewpoint based on the foreground region, and the generating of the novel view image includes identifying side effect prediction information based on the depth map, the side effect prediction information including at least one of whether an object smaller than or equal to a specified thickness is included in the input image or an object density degree; and generating the novel view image by controlling a viewpoint movement path based on the identified side effect prediction information. BRIEF DESCRIPTION OF DRAWINGS

[0016] The above and other aspects, features, and advantages of certain embodiments of the disclosure will be more apparent from the following detailed description, taken in conjunction with the accompanying drawings, in which:

[0017] Figure 1 is a diagram illustrating an example novel view synthesis technique according to various embodiments;

[0018] Figure 2a is a block diagram illustrating an example configuration of an electronic device according to various embodiments;

[0019] Figure 2b is a block diagram illustrating an example configuration of an electronic device according to various embodiments;

[0020] Figure 3 is a flowchart illustrating an example method of controlling an electronic device according to various embodiments;

[0021] Figure 4 and Figure 5 include diagrams and flowcharts illustrating example methods of controlling an electronic device according to various embodiments;

[0022] Figure 6 is a diagram illustrating an example method of identifying a depth boundary region according to various embodiments;

[0023] Figure 7a and Figure 7b is a diagram illustrating an example method of optimizing a depth map according to various embodiments;

[0024] Figure 8 is a diagram illustrating an example of an example process of identifying a thin object in a depth map using an open operation according to various embodiments;

[0025] Figure 9a and Figure 9b is a diagram illustrating an example depth boundary region of a typical object and a depth boundary region of a thin object according to various embodiments;

[0026] Figure 10 is a diagram illustrating an example of an occlusion region and a context region according to various embodiments;

[0027] Figure 11 is a diagram illustrating a side effect generated by an over-provisioned occlusion region according to various embodiments;

[0028] Figure 12 is a diagram illustrating an example identification method of an occlusion region according to various embodiments;

[0029] Figure 13a and Figure 13b is a diagram illustrating an example of a boundary complexity according to various embodiments; and

[0030] Figure 14 is a diagram illustrating an example method of adjusting an occlusion region according to various embodiments. DETAILED DESCRIPTION

[0031] The terms used in the disclosure will be briefly described, and the disclosure will be described in more detail.

[0032] The terms used in the disclosure are selected common terms currently widely used in consideration of their functions in the present document. However, the terms can be changed according to the intention of those skilled in the art, legal or technical interpretation, emergence of new technology, etc. Also, in some cases, there can be arbitrarily selected terms, and in this case, the meaning of the terms will be more specifically disclosed in the relevant specification. Therefore, the terms used herein should not be simply construed as its name, but based on the meaning of the terms and the overall context of the disclosure.

[0033] In the disclosure, expressions such as "have", "may have", "include", and "may include" are used to indicate the presence of a corresponding feature (for example, elements such as numerical values, functions, operations, or components), without excluding the presence or possibility of additional features.

[0034] In the disclosure, expressions such as "A or B", "at least one of A and / or B", or "one or more of A and / or B" can include all possible combinations of items listed together. For example, "A or B", "at least one of A and B", or "at least one of A or B" can refer to all cases including (1) only A, (2) only B, or (3) both A and B.

[0035] Expressions such as "1st", "2nd", "first", or "second" used in the disclosure do not limit various elements in terms of order and / or importance, and can be used only to distinguish one element from another element, without limiting the relevant elements.

[0036] When an element (e.g., a first element) is referred to as being "(operatively or communicatively) coupled with / to" or "connected to" another element (e.g., a second element), it can be understood that the element is directly coupled with / to or connected to the other element or is indirectly coupled with / to or connected to the other element through another element (e.g., a third element).

[0037] The expression "configured to" (or "set to") used in the disclosure can be interchangeably used, for example, based on the environment, with "adapted to", "having a capability to", "designed to", "made to", or "capable of". The term "configured to" (or "set to") can not necessarily refer to, in terms of hardware, for example, "designed as".

[0038] In some cases, the expression "device configured to" can refer to, for example, the device can "perform something" together with another device or component. For example, the phrase "processor configured to (or set to) perform A, B, or C" can refer to a dedicated processor (e.g., an embedded processor) for performing the relevant operations, or a general-purpose processor (e.g., a central processing unit (CPU) or an application processor) capable of performing the relevant operations by executing one or more software programs stored in a memory device.

[0039] The singular expression includes the plural expression unless otherwise specified. It should be understood that terms such as "form" or "include" are used herein to indicate the presence of characteristics, numbers, steps, operations, elements, components, or combinations thereof, and do not exclude the presence or possibility of adding one or more other characteristics, numbers, steps, operations, elements, components, or combinations thereof.

[0040] The term "module" or "part" used in various embodiments herein performs at least one function or operation and can be implemented in hardware or software, or implemented in a combination of hardware and software. In addition, a plurality of "modules" or a plurality of "parts" other than the "module" or the "part" that needs to be implemented as a specific hardware can be integrated in at least one module and implemented as at least one processor (not shown).

[0041] The various elements and regions of the drawings have been schematically shown. Accordingly, the technical spirit of the disclosure is not limited by the relative sizes and distances shown in the drawings.

[0042] Various example embodiments of the disclosure will be described in greater detail below with reference to the accompanying drawings.

[0043] Figure 1 is a diagram illustrating an example novel view synthesis technique according to various embodiments.

[0044] Novel view synthesis can refer to the technique of generating images from a new viewpoint different from that of a 2D image obtained through a monocular camera. Using this technique, it is possible to generate a continuous video of novel view images moving according to the viewpoint trajectory of a single image. (https: / / shihmengli.github.io / 3D-Photo-Inpainting / content.sniklaus.com / kenburns / video.mp4).

[0045] In the prior art, the novel view image 10 generated based on depth information may include, for example: Figure 1 The characteristics of the background object 11, which is slightly moved 13, and the foreground object 12, which is moved more 14, are shown.

[0046] In the foreground-background separation technique described in the example, the foreground and background are separated based on a depth map, the viewpoint is changed based on the foreground, and an image viewed from a random viewpoint can be generated by filling the occlusion area generated according to the viewpoint change with the background. In this case, accurate separation of the foreground and background is required to obtain a natural rendering effect. However, when estimating the depth map using monocular or stereo images, continuous values ​​are estimated near the boundary between the foreground and background due to the characteristics of regression techniques. Therefore, accurately identifying the boundary when separating the foreground and background may be a key technique for generating good-quality images.

[0047] To accurately identify foreground and background boundaries, deep learning-based alpha matting techniques can be used. However, because these techniques operate on the premise that only one foreground object exists in the image, they can lead to several side effects as the object becomes more complex. Furthermore, techniques that identify boundaries by finding deep boundary regions where the foreground and background boundaries are blurred and applying median filters to these regions can suffer from the problem of thin foreground objects disappearing if the region corresponding to the foreground in the filter is small, due to the characteristics of the median filter.

[0048] As mentioned above, issues such as inaccurate depth estimation and filling of occluded areas can cause several side effects, and these problems tend to be greatly amplified as the movement path of the newly generated viewpoint becomes larger. For example, issues such as halls being incorrectly filled or thin objects not being properly separated and processed may occur.

[0049] To minimize and / or reduce potential side effects during image generation, novel view synthesis techniques in related art may use a fixed-range viewpoint movement path for all images, or control the viewpoint movement path to a region where the occlusion area is small. However, if a fixed viewpoint movement path is used, side effects occurring within the movement path may be unavoidable. Furthermore, because images from only a limited number of viewpoints are generated, a limited stereoscopic image can be provided to the user. Additionally, if the viewpoint is only moved to a region where the occlusion area is small, there is a limitation on the number of viewpoints that can be generated.

[0050] For example, there is a problem where object boundaries are bent or stretched because the object boundaries are not accurately defined at the estimated depth through the monocular image. In response to the above, depth map optimization is performed as a post-processing operation; however, because existing depth map optimization techniques cannot adequately handle thin objects, they may cause various side effects on the novel view generation results.

[0051] Therefore, various embodiments for avoiding or minimizing / reducing side effects by controlling the movement path of the viewpoint, which is newly generated by pre-predicting various side effects that may occur from the generated image, will be described in more detail below.

[0052] Figure 2a This is a block diagram illustrating an example configuration of an electronic device according to various embodiments.

[0053] refer to Figure 2a The electronic device 100 may include a memory 110 and at least one processor (e.g., including processing circuitry) 120.

[0054] The memory 110 may be implemented as a memory embedded in the electronic device 100, or as a memory that can be attached to or detached from the electronic device 100, depending on the data storage usage. For example, data for driving the electronic device 100 may be stored in the memory embedded in the electronic device 100, and data for extending the functionality of the electronic device 100 may be stored in the memory that can be attached to or detached from the electronic device 100. The memory embedded in the electronic device 100 may be implemented as at least one of volatile memory (e.g., dynamic random access memory (DRAM), static RAM (SRAM), or synchronous dynamic RAM (SDRAM)) or non-volatile memory (e.g., one-time programmable read-only memory (OTPROM), programmable ROM (PROM), erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), masked ROM, flash ROM, flash memory (e.g., NAND flash or NOR flash), hard disk drive (HDD), or solid-state drive (SSD)). Furthermore, the memory that can be attached to or detached from the electronic device 100 can be implemented in the form of, but is not limited to, memory cards (e.g., Compact Flash (CF), Secure Digital (SD), Micro-Secure Digital (micro-SD), Mini-Secure Digital (mini-SD), Extreme Digital (xD), Multimedia Card (MMC), etc.) and external memory (e.g., USB memory) that can be connected to a USB port (e.g., USB memory).

[0055] In the example, memory 110 may store at least one instruction for controlling electronic device 100 or a computer program including the instruction.

[0056] In another example, memory 110 may store images, i.e., input images received from external devices (e.g., source devices), external storage media (e.g., USB), external servers (e.g., WEBHARD), etc. Alternatively, memory 110 may store images acquired by a camera (not shown) disposed in electronic device 100. Here, the images may be, but are not limited to, 2D moving images.

[0057] In yet another example, memory 110 may store various information necessary for image quality processing, such as information for performing at least one of noise reduction, detail enhancement, tone mapping, contrast enhancement, color enhancement, or frame rate conversion, algorithms, image quality parameters, etc. Additionally, memory 110 may store intermediate images generated through image processing, as well as images generated based on depth information.

[0058] According to an embodiment, memory 110 may be implemented as a single memory that stores data generated according to various operations of this disclosure. However, according to an embodiment, memory 110 may be implemented as a plurality of memories that store different types of data or data generated from different levels.

[0059] In the embodiments, various data have been described as being stored in the external memory 110 of the processor 120; however, depending on at least one implementation example of the electronic device 100 or the processor 120, at least a portion of the aforementioned data may be stored in the memory inside the processor 120.

[0060] At least one processor 120 may include various processing circuitry and / or multiple processors. For example, as used herein (including the claims), the term "processor" may include various processing circuitry, including at least one processor, wherein one or more of the at least one processor may be configured individually and / or collectively in a distributed manner to perform the various functions described herein. As used herein, when "processor," "at least one processor," and "one or more processors" are described as being configured to perform multiple functions, these terms cover, for example, but not limited to, cases where one processor performs some of the stated functions while another(s) of processors performs other stated functions, and cases where a single processor may perform all of the stated functions. Additionally, at least one processor may include, for example, a combination of processors performing various stated / disclosed functions in a distributed manner. At least one processor may execute program instructions to implement or perform various functions and control the overall operation of electronic device 100. For example, at least one processor 120 may control the overall operation of electronic device 100 by being connected to each configuration of electronic device 100. For example, at least one processor 120 may control the overall operation of electronic device 100 by being electrically connected to display 130 and memory 110. At least one processor 120 may be composed of one or more processors.

[0061] At least one processor 120 can perform the operation of the electronic device 100 according to various embodiments by executing at least one instruction stored in the memory 110.

[0062] At least one processor 120 may include at least one of a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), a multi-core integrated circuit (MIC), a digital signal processor (DSP), a neural processing unit (NPU), a hardware accelerator, or a machine learning accelerator. At least one processor 120 may control one or any combination of other elements of an electronic device and perform operations associated with communication or data processing. At least one processor 120 may execute at least one program or instruction stored in memory. For example, at least one processor may perform methods according to various embodiments of the present disclosure by executing at least one instruction stored in memory.

[0063] When the methods according to various embodiments of this disclosure include multiple operations, the multiple operations may be executed by one processor or by multiple processors. For example, when the first operation, the second operation, and the third operation are performed by the methods according to the various embodiments, the first operation, the second operation, and the third operation may all be executed by the first processor, or the first operation and the second operation may be executed by the first processor (e.g., a general-purpose processor), and the third operation may be executed by the second processor (e.g., an artificial intelligence-specific processor).

[0064] At least one processor 120 may be implemented as a single-core processor including one core, or as at least one multi-core processor including multiple cores (e.g., homogeneous multi-core or heterogeneous multi-core). If at least one processor 120 is implemented as a multi-core processor, each of the multiple cores included in the multi-core processor may include internal processor memory, such as cache memory and on-chip memory, and a common cache shared by the multiple cores may be included in the multi-core processor. Furthermore, each (or a portion of the multiple cores) included in the multi-core processor may independently read and execute program instructions for implementing the methods according to various embodiments, or may read and execute program instructions for implementing the methods according to various embodiments of the present disclosure due to the interconnection of all (or a portion) of the multiple cores.

[0065] When the methods according to various embodiments of this disclosure include multiple operations, the multiple operations may be executed by one core of a plurality of cores, or by a plurality of cores included in a multi-core processor. For example, when the first operation, the second operation, and the third operation are performed by the methods according to the various embodiments, the first operation, the second operation, and the third operation may all be executed by the first core included in the multi-core processor, or the first operation and the second operation may be executed by the first core included in the multi-core processor, and the third operation may be executed by the second core included in the multi-core processor.

[0066] According to various embodiments of this disclosure, a processor may refer to a system-on-a-chip (SoC), a single-core processor, or a multi-core processor that integrates at least one processor and other electronic components, or a core included in a single-core or multi-core processor. The core described herein may be implemented as a CPU, GPU, APU, MIC, DSP, NPU, hardware accelerator, machine learning accelerator, etc., but is not limited thereto. However, for ease of description, at least one processor 120 may be designated as the processor 120 described below.

[0067] Processor 120 can obtain depth information from an input image. Here, the input image can include a still image, multiple consecutive still images (or frames), or a moving image. For example, the input image can be a 2D image. Here, the depth information can be in the form of a depth map. A depth map can, for example, refer to a table including depth information for each region of the image. Regions can be categorized into pixel units or defined as preset regions larger than pixel units. In the example, the depth map can be in the form of indicating 127 or 128 of a grayscale value between 0 and 255 as a reference value (i.e., 0 (or the focal plane)) and indicating values ​​less than 127 or 128 as negative (-) values ​​and values ​​greater than 127 or 128 as positive (+) values. The reference value for the focal plane can be randomly selected between 0 and 255. Here, a - value can refer to, for example, a depression, and a + value can refer to, for example, a projection. However, the above is merely an example, and depth maps can represent depth with various values ​​according to various criteria.

[0068] In the example, processor 120 can obtain depth information based on the image-processed image after image processing of the input image. Here, image processing can be digital image processing, which includes at least one of image enhancement, image restoration, image transformation, image analysis, image understanding, image compression, image decoding, or scaling.

[0069] In the example, various preprocessing steps can be performed before obtaining depth information about the input image; however, for the convenience of the following description, the input image and the preprocessed image will not be distinguished, and the input image will be designated as such.

[0070] Processor 120 can store depth information corresponding to the input image in memory 110. In this example, processor 120 can apply preprocessing and / or postprocessing to the first and second image frames based on the sequentially input first and second image frames to obtain first and second depth information corresponding to the first and second image frames, and then store them sequentially in memory 110. Here, the first and second image frames can be 2D monocular image frames.

[0071] In the example, the processor 120 can obtain depth information of the first image frame and the second image frame based on various image processing methods, such as (but not limited to) algorithms, equations, artificial intelligence models, etc.

[0072] Figure 2b This is a block diagram illustrating an example configuration of an electronic device according to various embodiments.

[0073] refer to Figure 2b The electronic device 100' may include a memory 110, at least one processor (e.g., including processing circuitry) 120, a display 130, a camera 140, a user interface (e.g., including circuitry) 150, a communication interface (e.g., including communication circuitry) 160, and a speaker 170. These details are not repeated here. Figure 2b The configuration shown is with Figure 2a A detailed description of the overlapping configurations shown.

[0074] Display 130 can be implemented as a display including self-emissive devices or a display including non-emissive devices and a backlight. For example, display 130 can be implemented as various forms of display, such as, but not limited to, liquid crystal displays (LCDs), organic light-emitting diode (OLED) displays, light-emitting diodes (LEDs), micro LEDs, mini LEDs, plasma display panels (PDPs), quantum dot (QD) displays, quantum dot light-emitting diodes (QLEDs), etc. Display 130 may include driving circuitry, backlight units, etc., which can be implemented in the form of a-Si TFTs, low-temperature polycrystalline silicon (LTPS) TFTs, organic TFTs (OTFTs), etc. In the example, a touch sensor having the form of a touch film, touch sheet, or touchpad for sensing touch operations can be arranged on the front surface of display 130 to sense various types of touch input. For example, display 130 can sense various types of touch input, such as, but not limited to, touch input from a user's hand, touch input from an input device such as a stylus, touch input from a specific electrostatic material, etc. The input device can be implemented as a pen-type input device, which can be referred to by various terms, such as, but not limited to, electronic pen, pointer pen, S-pen, etc. In the example, the display 130 can be implemented as a flat panel display, a curved display, a foldable and / or rollable flexible display, etc.

[0075] Camera 140 can be activated and perform shooting according to a preset event. Camera 140 can convert the captured image into an electrical signal and generate image data based on the converted signal. For example, the object can be converted into an electrical image signal by a semiconductor optical device (charge-coupled device (CCD)), and the converted image signal as described above can be a signal that has been processed after being amplified and converted into a digital signal. For example, camera 140 can include at least one of a conventional (or basic) camera and an ultra-wide-angle camera.

[0076] User interface 150 may include various circuits and may be implemented as a device such as a button, touchpad, mouse and keyboard, or as a touch screen that can perform the above-mentioned display functions and operation input functions together.

[0077] Communication interface 160 may include various communication circuits and can be implemented in various interfaces depending on the implementation of electronic device 100'. For example, communication interface 160 may communicate with external devices, external storage media (such as USB memory), external servers (such as WEBHARD), etc., via communication methods such as, but not limited to, Bluetooth, Wi-Fi (wireless LAN) based on access point (AP), Zigbee, wired / wireless LAN, wide area network (WAN), Ethernet, IEEE 1394, high-definition multimedia interface (HDMI), universal serial bus (USB), mobile high-definition link (MHL), AES / EBU, fiber optic, coaxial cable, etc. In this example, communication interface 160 may perform communication with another electronic device, external server, and / or remote control device, etc.

[0078] The speaker 170 may include a configuration that outputs not only various audio data, but also various notification sounds, voice messages, etc. According to various embodiments of this disclosure, the processor 120 may control the speaker 170 to output feedback or various notifications in audio form.

[0079] In addition, depending on the implementation, the electronic device 100' may include a sensor and a microphone.

[0080] Sensors can include various types of sensors, such as, but not limited to, touch sensors, proximity sensors, accelerometers (or gravity sensors), geomagnetic sensors, gyroscopes, pressure sensors, position sensors, distance sensors, illuminance sensors, etc.

[0081] A microphone can be configured to receive user voice or other sound input and convert it into audio data. However, in this embodiment, electronic device 100' can receive user voice input via communication interface 160 from an external device.

[0082] Figure 3 This is a flowchart illustrating example methods for controlling electronic devices according to various embodiments.

[0083] refer to Figure 3 The electronic device 100 can obtain a depth map corresponding to the input image (S310). The depth map (or depth information) can indicate the 3D distance information of objects present in the image, and can provide a depth map (or depth information) for each pixel of the image. In the example, the 8-bit depth can have grayscale values ​​between 0 and 255. For example, when represented based on black and white, black (low value) can indicate a position far from the viewer, while white (high value) can indicate a position close to the viewer. In the example, the depth map can be obtained from the 2D input image based on existing depth estimation algorithms, equations, trained artificial intelligence models, etc. For example, the artificial intelligence model can be implemented as a neural network including multiple neural network layers. The artificial intelligence model can be implemented as a deep neural network (DNN), convolutional neural network (CNN), recurrent neural network (RNN), restricted Boltzmann machine (RBM), deep belief network (DBN), bidirectional recurrent deep neural network (BRDNN), deep Q network, etc., but is not limited to the above implementation methods.

[0084] The electronic device 100 can identify foreground and background regions included in the input image based on the depth map (S320).

[0085] The electronic device 100 can identify side effect prediction information based on a depth map (S330), the side effect prediction information including at least one of the following: whether objects with a thickness less than or equal to a preset (e.g., specified) thickness are included in the input image or the degree of object density. For example, if thin objects are included in the input image when performing novel view composition, side effect prediction may occur because of the presence of objects when multiple objects are grouped together, etc. Figure 1 The various possibilities of side effects described in the text can be used to define the above as side effect prediction information, and side effect prediction information can be identified in the input image.

[0086] The electronic device 100 can switch viewpoints based on a foreground region by controlling the viewpoint movement path based on identified side effect prediction information, thereby generating a novel view image (S340). For example, the viewpoint movement path can be controlled based on the level of the side effect prediction information. If a side effect is predicted to be severe, the viewpoint moves within a relatively limited range to minimize and / or reduce the side effect.

[0087] According to the example, electronic device 100 can generate a novel view image by controlling the viewpoint movement path to be reduced to a range below a threshold, based on whether an object with a thickness less than or equal to a preset thickness is identified as included in the input image based on a depth map or whether the object density is identified as greater than or equal to a threshold based on a depth map.

[0088] In the example, the electronic device 100 can proportionally control the reduction of the viewpoint movement path based on the ratio of the area occupied by objects with a thickness less than or equal to a preset thickness in the input image, which are identified as included in the input image based on the depth map.

[0089] In the example, the electronic device 100 can obtain a depth map with the opening operation applied by applying the opening operation to the depth map, and identify regions including objects with a preset thickness based on the difference information between the depth map and the depth map with the opening operation applied.

[0090] In the example, the electronic device 100 can optimize the depth map by replacing the depth value of the region corresponding to an object with a preset thickness with the depth value of the surrounding pixels that have color information most similar to the target pixel.

[0091] In the example, electronic device 100 can generate a novel view image by controlling the viewpoint movement path to be reduced to a range below a threshold based on the object density level identified as greater than or equal to a depth map.

[0092] In the example, the electronic device 100 can optimize the depth map by replacing the depth values ​​of other regions outside the region corresponding to an object with a preset thickness with the median depth values ​​of the pixels in the window.

[0093] In the example, the electronic device 100 can calculate the standard deviation of the depth values ​​of pixels excluding depth boundary regions in the window by applying a preset window to the depth map, and perform depth map optimization when the calculated standard deviation is greater than or equal to a threshold.

[0094] In the example, electronic device 100 can adjust the occlusion region and the context region based on boundary complexity indicating the density of objects. For example, the occlusion region could be an area that is not exposed in the foreground region at the current viewpoint but is exposed when the viewpoint moves. For example, the context region could be an area adjacent to the occlusion region.

[0095] In the example, electronic device 100 can identify boundary complexity based on the number of adjacent boundaries included in the context region. When the number of adjacent boundaries is greater than or equal to a threshold number, electronic device 100 can reduce the width of at least one of the occlusion region and the context region.

[0096] Figure 4 and Figure 5 This includes diagrams and flowcharts illustrating example methods for controlling electronic devices according to various embodiments.

[0097] refer to Figure 4 The processor 120 can estimate the depth map based on the input image (410).

[0098] Then, the processor 120 can optimize the depth map (420) based on the side effect prediction information (460). In the example, such as Figure 5 As shown, processor 120 can identify whether a foreground object with a thickness less than or equal to a preset thickness is included in the input image based on the depth map (S510). Based on the identification that a foreground object with a thickness less than or equal to the preset thickness is included in the input image, processor 120 can optimize the depth map by replacing the depth values ​​of the region corresponding to the object with a thickness less than or equal to the preset thickness with the depth values ​​of surrounding pixels having color information most similar to the target pixel (S520). For example, processor 120 can perform depth map optimization based on color information comparison based on side effects predicted according to the presence of thin objects in the input image. Furthermore, processor 120 can optimize the depth map by replacing the depth values ​​of regions outside the region corresponding to the object with a thickness less than or equal to the preset thickness with the median depth values ​​of the pixels in the window (S530).

[0099] In the example, processor 120 can accurately separate foreground and background regions by optimizing the boundary regions of objects in the predicted depth map. To do this, the boundary regions of objects in the depth map can be identified, and then the size of the objects can be identified through morphological operations. For example, processor 120 can identify object regions included in the input image using at least one technique selected from object recognition, object detection, object tracking, or image segmentation. For example, processor 120 can use techniques such as, but not limited to, semantic segmentation that classifies and extracts objects included in the input image by type as needed, instance segmentation that identifies objects by classifying them by object (even if the objects are of the same type), and bounding boxes that include the quadrilateral shape of the detected objects when they are detected in the image.

[0100] In the example, processor 120 can optimize the depth map based on the thickness (or size) of the identified object region by selectively applying median filter-based optimization and color information comparison-based optimization. For example, the same process can be repeated several times (e.g., five times) to obtain more accurate optimization results.

[0101] For example, processor 120 can use the depth value difference between adjacent pixels to determine the boundary region between an object and the background, or between two objects, in a depth map. For instance, processor 120 can define a pixel having a depth difference from the center pixel that is greater than or equal to one of the four adjacent pixels (top, bottom, left, and right) as a depth boundary region. For example, as... Figure 6 As shown, a region 611 comprising pixels having a depth difference from the center pixel that is greater than or equal to one of the four adjacent pixels in the depth map 610 (top, bottom, left, and right) can be identified as a depth boundary region (620).

[0102] In the example, processor 120 may repeatedly perform depth map optimization several times (e.g., five times) for the identified boundary region. Depending on whether thin objects exist during optimization, processor 120 may adaptively use either a median filter-based optimization method or an optimization method based on comparison of generated information. Each optimization method may be performed within a window of a preset size, and if the depth boundary region is not included in the window, then optimization may not be performed.

[0103] Figure 7a and Figure 7b This is a diagram illustrating example methods for optimizing depth maps according to various embodiments.

[0104] Figure 7a The optimization process based on the median filter, as shown in the example, is illustrated.

[0105] Median filter-based optimization can be a method of assigning the median depth values ​​of the pixels present within the window as the output value. In this process, pixels in depth boundary regions can be excluded. Median filter-based optimization methods can reliably obtain depth boundaries and are effective in mitigating depth value inversion phenomena that occur at depth boundaries. However, for thin objects, since most of the object is contained within the depth boundary region, the application of a median filter can lead to the problem of thin objects disappearing, as they are output as background depth values.

[0106] Figure 7b The optimization process based on color information comparison is shown in the example.

[0107] The optimization method based on color information comparison assigns the depth value of the pixel within the window that has the most similar color value to the pixel to be optimized as the output value. During this process, pixels in depth boundary regions are excluded from the comparison.

[0108] In the example, the similarity of color values ​​can be calculated using the Euclidean distance between the pixel to be optimized and the pixel to be compared. Using an optimization method based on color information comparison, accurate depth map optimization is possible even when the object is thin. However, during optimization, unstable depth boundaries may be generated when no similar pixels exist within the window.

[0109] Therefore, processor 120 can adaptively select an optimization method based on whether thin objects exist within the input depth map. For example, if the depth boundary region is identified as a thin object, optimization based on color information comparison can be performed, while if the depth boundary region is not identified as a thin object, optimization based on a median filter can be performed. Processor 120 can use opening operations, as an example of morphological operations, to determine thin objects. Opening operations are operations that apply dilation operations after erosion operations, and opening, erosion, and dilation operations can be expressed as Equation 1 below.

[0110] [Equation 1]

[0111]

[0112] The size of the structuring element S used for the opening operation can be the same as or similar to the size of the window used for optimization. The processor 120 can determine a thin object based on the difference between the result obtained by applying the opening operation and the input depth map being greater than or equal to a threshold.

[0113] Figure 8 This is an illustration of an example process for identifying thin objects in a depth map using an opening operation according to various embodiments.

[0114] For example, processor 120 can apply opening operations to, for example... Figure 8 The depth map 810 shown is used to obtain the image 820 to which the opening operation has been applied. The processor 120 can identify thin object regions based on the image 830, which indicates the depth difference before and after the opening operation. Then, the processor 120 can perform color information comparison-based optimization in the regions containing thin objects in the depth map 840, and median filter-based optimization in the remaining regions.

[0115] Figure 9a and Figure 9b This is a diagram illustrating examples of the depth boundary regions of a typical object and a thin object according to various embodiments.

[0116] In the example, Figure 9a The depth boundary region of a typical object is shown, while Figure 9b The depth boundary region of the thin object is shown.

[0117] like Figure 9b As shown, if at least two depth boundaries are adjacent to the depth boundary region of a thin object, the entire region between the two depth boundaries can be set as the depth boundary region, and there is a possibility that the entire relevant region will be optimized as background depth values. Therefore, when optimizing the first round (e.g., the first and second rounds) from several (e.g., five) optimization processes, depth map optimization can be performed only if the standard deviation of the depth values ​​of pixels excluding the depth boundary region within the window is calculated and the standard deviation is greater than or equal to a threshold. For example, because the depth of both the foreground object and the background depth exist around typical depth boundary regions, the standard deviation of the depth values ​​within the kernel may be large, while if most of the foreground object is included in the depth boundary region, the standard deviation may be small. Therefore, if the standard deviation is small, the presence of thin objects can be assumed, and side effects that might arise in median filter-based optimization methods can be avoided.

[0118] Refer again Figure 4 The processor 120 can determine the occlusion region (430) from the depth map based on side effect prediction information (460). In the example, the processor 120 can predict possible side effects based on object density levels and identify the occlusion region based on the predicted side effects. For example, the processor 120 can adjust the size of the occlusion region and the context region based on the object density levels.

[0119] In the example, processor 120 can set occlusion regions and context regions based on object density levels, and adjust the size of each region according to the boundary complexity indicating the object density level.

[0120] An occlusion region can refer to, for example, an area that is occluded by the foreground from the current viewpoint but will appear as the viewpoint moves. For example, processor 120 can set an occlusion region for each boundary between objects extracted from the depth map.

[0121] The context region can be an area adjacent to the occluded area and can be used as a cue to fill the occluded area.

[0122] Figure 10 This is a diagram illustrating examples of occlusion areas and context areas according to various embodiments.

[0123] refer to Figure 10The sizes of the occlusion region 1011 and the context region 1012 can be set based on the boundary complexity of the first boundary in the depth map 1010. Furthermore, the sizes of the occlusion region 1021 and the context region 1022 can be set based on the boundary complexity of the second boundary in the depth map 1020. Additionally, the sizes of the occlusion region 1031 and the context region 1032 can be set based on the boundary complexity of the third boundary in the depth map 1030.

[0124] For example, if multiple objects are blended in a boundary region, the boundaries between the objects can be concentrated. Based on the above, the likelihood of generating unstable results when filling the occlusion region (440) used to generate a novel view image (450) may increase due to the occlusion region being over-set or unnecessary information being included in the context region.

[0125] Figure 11 This is a diagram illustrating example side effects generated by oversetting the occlusion area according to various embodiments.

[0126] For example, if the target boundary is close to an adjacent boundary, the occlusion region of the target boundary may encroach on the occlusion region of the adjacent boundary, potentially leading to undesirable results when generating images from different viewpoints. In such cases, the problem can be addressed by reducing the size of the occlusion region of the target boundary. Therefore, by adjusting the size of the occlusion region and the context region using boundary complexity that indicates the degree of object density, side effects can be minimized and / or reduced.

[0127] Figure 12 This is a diagram illustrating an example method for identifying occluded areas according to various embodiments.

[0128] refer to Figure 12 The electronic device 100 can determine the occlusion region and the context region based on the depth map (1210), and identify the boundaries within the context region (1220). The electronic device 100 can calculate the boundary complexity of the identified boundaries (1230), and adjust the occlusion region and the context region according to the boundary complexity (1240).

[0129] In the example, boundary complexity can be defined by the number of boundaries surrounding the target boundary. For instance, the occlusion region and its corresponding context region within a defined range can be initialized based on the target boundary, and the number of boundaries existing within the initial context region can be defined as the boundary complexity.

[0130] Figure 13a and Figure 13b This is a diagram illustrating examples of boundary complexity according to various embodiments.

[0131] In the example, electronic device 100 can identify boundary complexity based on the number of adjacent boundaries included in the context area.

[0132] For example, such as Figure 13a As shown, the context region (1320) can be identified based on image 1310 indicating the target boundary, and the two boundaries within the context region can be identified based on image 1330 indicating the boundaries within the context region. Therefore, the boundary complexity can be identified as 2.

[0133] For example, such as Figure 13b As shown, the context region can be identified based on the target boundary 1340, and the five boundaries within the context region 1350 can be identified based on the image 1360 indicating the boundaries within the context region. Therefore, the boundary complexity can be identified as 5.

[0134] In the example, if the boundary complexity is greater than or equal to a certain level (e.g., a certain number), the width of the occlusion region and the context region can be reduced so that the occlusion region is not over-allocated or unnecessary information is not included in the context region. For example, electronic device 100 can be based on, for example... Figure 14 The context region 1410 shown is used to identify the boundary 1420 within the context region. Figure 14 In the example shown, if the complexity of the boundary 1420 within the context region is as high as 5, at least one of the size and position of the identified occlusion region 1430 can be adjusted, and the adjusted occlusion region 1440 can be identified.

[0135] However, Figure 13a , Figure 13b and Figure 14 The number of boundary complexities and the adjustments to the position / size of the occluded area shown are merely examples, and the above is not limited to these.

[0136] According to various embodiments of this disclosure as described above, novel view images can be generated based on the estimated depth map by proportionally reducing the area ratio of thin objects to the entire image and / or the range of viewpoint movement, in accordance with boundary complexity. Alternatively, novel view images can be generated by applying different depth map optimization methods to the thin object region and the remaining region. Therefore, potential side effects based on thin objects and / or object density can be minimized and / or reduced.

[0137] Furthermore, the methods of the various embodiments of the present disclosure described above can be implemented using only software or hardware upgrades of display devices and electronic devices based on the relevant technologies.

[0138] Furthermore, the various embodiments disclosed above can be executed by an embedded server located in the electronic device, or by an external server of the electronic device.

[0139] Furthermore, according to embodiments of this disclosure, the various embodiments described above can be implemented using software including instructions stored in a machine-readable storage medium (e.g., a computer). A machine can invoke the instructions stored in the storage medium and, as a device operable according to the invoked instructions, can include an electronic device (e.g., electronic device (A)) according to the embodiments described above. Based on commands executed by the processor, the processor can directly or using other elements under the processor's control perform functions related to the commands. Commands can include code generated by a compiler or executed by an interpreter. The machine-readable storage medium can be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means, for example, that the storage medium is tangible and does not include signals, and the term does not distinguish whether data is stored semi-permanently or temporarily in the storage medium.

[0140] Furthermore, according to embodiments of this disclosure, methods comprising computer program products according to the various embodiments described above can be provided. The computer program products can be exchanged as goods between sellers and buyers. The computer program products can be distributed in the form of machine-readable storage media (e.g., optical disc read-only memory (CD-ROM)) or through app stores (e.g., Playstore). TM Online distribution. In the case of online distribution, at least a portion of the computer program product may be stored, at least temporarily, in a storage medium such as the memory of the manufacturer's server, the application store's server, or a relay server, or may be temporarily generated.

[0141] Furthermore, each element (e.g., a module or program) according to the various embodiments described above can be formed as a single entity or multiple entities, and some of the sub-elements described above can be omitted, or other sub-elements can be further included in the various embodiments. Alternatively or additionally, some elements (e.g., modules or programs) can be integrated into one entity to perform the same or similar functions performed by the individual elements prior to integration. According to the various embodiments, operations performed by a module, program, or other element can be performed sequentially, in parallel, repeatedly, or heuristically, or at least some operations can be performed in a different order, or different operations can be omitted or added.

[0142] While this disclosure has been shown and described with reference to various exemplary embodiments thereof, it should be understood that these exemplary embodiments are intended to be illustrative and not restrictive. Those skilled in the art will understand that various changes in form and detail may be made therein without departing from the true spirit and full scope of this disclosure, including the appended claims and their equivalents. It will also be understood that any one or more embodiments described herein may be used in conjunction with any other one or more embodiments described herein.

Claims

1.An electronic device comprising: a memory storing at least one instruction; and at least one processor including processing circuitry, individually and / or collectively configured to: identify a foreground region and a background region included in an input image based on a depth map corresponding to the input image, and generate a novel view image by converting a viewpoint based on the foreground region, wherein the at least one processor, individually and / or collectively configured to: identify side effect prediction information based on the depth map, the side effect prediction information including at least one of whether an object less than or equal to a specified thickness is included in the input image or an object density degree; and generate the novel view image by controlling a viewpoint movement path based on the identified side effect prediction information. 2.The electronic device of claim 1, wherein the at least one processor, individually and / or collectively configured to: generate the novel view image by controlling the viewpoint movement path to decrease to less than a threshold range based on the object less than or equal to the specified thickness being identified as included in the input image based on the depth map or the object density degree being identified as greater than or equal to a threshold based on the depth map. 3.The electronic device of claim 2, wherein the at least one processor, individually and / or collectively configured to: proportionally control the viewpoint movement path to decrease based on a ratio of an area occupied by the object less than or equal to the specified thickness in the input image based on the object less than or equal to the specified thickness being identified as included in the input image based on the depth map. 4.The electronic device of claim 2, wherein the at least one processor, individually and / or collectively configured to: obtain an eroded depth map by applying an erosion operation to the depth map, and identify a region including the object less than or equal to the specified thickness based on difference information between the depth map and the eroded depth map. 5.The electronic device of claim 2, wherein the at least one processor, individually and / or collectively configured to: calculate a standard deviation of depth values of pixels excluding a depth boundary region within a specified window by applying the specified window to the depth map, and perform depth map optimization based on the calculated standard deviation being greater than or equal to a threshold. 6.The electronic device of claim 2, wherein the at least one processor, individually and / or collectively configured to: generate the novel view image by controlling the viewpoint movement path to decrease to less than a threshold range based on the object density degree being identified as greater than or equal to a threshold based on the depth map. 7.The electronic device of claim 2, wherein the at least one processor, individually and / or collectively configured to: adjust an occlusion region and a context region based on a boundary complexity indicating the object density degree, wherein the occlusion region includes a region not exposed from a current viewpoint by the foreground region and exposed upon movement of the viewpoint, and wherein the context region is a region adjacent to the occlusion region. 8.The electronic device of claim 7, wherein the at least one processor, individually and / or collectively configured to: identify the boundary complexity based on a number of adjacent boundaries included within the context region, and reducing a width of at least one of the occlusion region and the context region based on a number of adjacent boundaries being greater than or equal to a threshold number. 9.The electronic device of claim 1, wherein, the at least one processor, individually and / or collectively, is configured to: optimize the depth map by replacing depth values of a region corresponding to an object less than or equal to a specified thickness with depth values of surrounding pixels having color information most similar to that of a target pixel. 10.The electronic device of claim 1, wherein, the at least one processor, individually and / or collectively, is configured to: optimize the depth map by replacing depth values of other regions outside the region corresponding to the object less than or equal to the specified thickness with a median value of depth values of pixels within a specified window. 11.A method of controlling an electronic device, the method comprising: identifying foreground regions and background regions included in an input image based on a depth map corresponding to the input image; and generating a novel view image by converting a viewpoint based on the foreground regions, wherein the generating of the novel view image comprises: identifying side effect prediction information based on the depth map, the side effect prediction information including at least one of whether an object less than or equal to a specified thickness is included in the input image or an object density degree; and generating the novel view image by controlling a viewpoint movement path based on the identified side effect prediction information. 12.The method of claim 11, wherein, the generating of the novel view image comprises: based on the object less than or equal to the specified thickness being identified as included in the input image or the object density degree being identified as greater than or equal to a threshold based on the depth map, generating the novel view image by reducing the viewpoint movement path to less than a threshold range. 13.The method of claim 12, wherein, the generating of the novel view image comprises: based on the object less than or equal to the specified thickness being identified as included in the input image based on the depth map, proportionally controlling the reduction of the viewpoint movement path by a ratio of an area occupied by the object less than or equal to the specified thickness in the input image. 14.The method of claim 12, wherein, the generating of the novel view image comprises: obtaining an eroded depth map by applying an erosion operation to the depth map; and identifying a region including the object less than or equal to the specified thickness based on difference information between the depth map and the eroded depth map. 15.A non-transitory computer-readable storage medium storing computer commands, which, when executed by at least one processor of an electronic device, individually and / or collectively, cause the electronic device to perform operations, the operations comprising: identifying foreground regions and background regions included in an input image based on a depth map corresponding to the input image; and generating a novel view image by converting a viewpoint based on the foreground regions, wherein the generating of the novel view image comprises: identifying side effect prediction information based on the depth map, the side effect prediction information including at least one of whether an object less than or equal to a specified thickness is included in the input image or an object density degree; and generating the novel view image by controlling a viewpoint movement path based on the identified side effect prediction information. A novel view image is generated by controlling a viewpoint movement path based on the identified side effect prediction information.