Image processing device and operation method thereof

The image processing device addresses the challenge of converting 2D images to 3D by dynamically applying a depth estimation model and non-linearly changing the depth map based on content type and scene analysis, resulting in improved accuracy and user experience.

WO2025127774A1PCT designated stage expired Publication Date: 2025-06-19SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/096095
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-13
Filing Date
2024-08-29
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Current technologies face challenges in accurately converting 2D images into 3D due to limitations in relative depth estimation and differences in display specifications, leading to depth estimation errors that affect user experience.

Method used

An image processing device that dynamically applies a depth estimation model in real time based on the content type of an input 2D image, and non-linearly changes the depth map by analyzing objects in each scene, to improve depth estimation accuracy.

Benefits of technology

This approach reduces depth estimation errors, enhances the accuracy of depth estimation, and improves user satisfaction by providing a more immersive 3D experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024096095_19062025_PF_FP_ABST
    Figure KR2024096095_19062025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are an image processing device for performing 3D conversion and an operating method thereof. The image processing device comprises: a memory that stores one or more instructions; and at least one processor that executes the one or more instructions stored in the memory. The processor may analyze a content type of an input image. The processor obtains a depth estimation model corresponding to the input image in real time by using on-device learning on the basis of the result of analyzing the content type of the input image. The processor obtains, on the basis of the depth estimation model according to on-device learning, a depth map for the input image reflecting estimated depth information. The processor performs 3D conversion from the input image on the basis of the depth map for the input image.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing device and its operating method

[0001] The present disclosure relates to an image processing device and an operating method thereof, and more particularly, to an image processing device for performing 3D transformation and an operating method thereof.

[0002] Recently, a number of 3D displays (such as Light Field Displays and Hologram Displays) have been released, offering features such as 3D games, 3D movies, and the ability to experience 3D models. However, aside from some content produced in 3D, there is a lack of content that can be viewed in 3D using 3D displays. Furthermore, users are increasingly preferring to view 2D videos in 3D for a greater sense of realism and immersion.

[0003] Accordingly, among some companies providing 3D displays, research is actively underway on technology to convert 2D (Two-Dimensional) images into 3D images when they are input.

[0004] The most crucial technology in converting 2D images to 3D is extracting depth information similar to that found in the real world from 2D images lacking depth information. The accuracy of estimating depth information from 2D images is steadily increasing thanks to deep learning methods utilizing artificial intelligence and the growing availability of diverse image-depth map databases (DBs) for learning.

[0005] However, without a depth sensor, absolute depth estimation of objects in an image is difficult, and only relative depth estimation between objects or between objects and the background is possible. Due to the limitations of this relative depth estimation and differences in display specifications, 3D images converted from 2D images suffer from depth estimation errors, which differ from the three-dimensionality perceived in the real world.

[0006] In addition, when a 2D image is input, which is a complex image with low correlation, such as a complex image containing various types of content such as graphic images, games, documents, comments, subtitles, and seminar materials, there is a problem that many depth estimation errors occur.

[0007] Users experience increased fatigue and difficulty viewing for long periods of time due to various depth estimation errors that occur when converting 2D images to 3D. To address this issue, some companies are providing a slider-style user interface (UI) that allows users to manually adjust the overall intensity of the three-dimensional effect in 3D-converted images.

[0008] However, in order to further improve the satisfaction and convenience of users using 3D displays, a technology is needed to further improve the depth estimation accuracy of input 2D images.

[0009] The present disclosure aims to provide a method for dynamically applying a depth estimation model in real time based on analysis of the content type of an input 2D image, and a method for non-linearly changing a depth map by analyzing objects included in each scene of an input 2D image, in order to solve the above-mentioned problems.

[0010] An image processing device according to one embodiment of the present disclosure includes a memory storing one or more instructions and at least one processor executing one or more instructions stored in the memory. The processor analyzes the content type of an input image by executing the one or more instructions. Based on the result of analyzing the content type of the input image, the processor obtains in real time a depth estimation model corresponding to the content type of the input image using on-device learning. Based on the depth estimation model according to on-device learning, the processor obtains a depth map for the input image reflecting the estimated depth information. The processor performs 3D transformation on the input image based on the depth map for the input image.

[0011] An operating method of an image processing device according to one embodiment of the present disclosure includes a step of analyzing a content type of an input image. The operating method includes a step of obtaining, in real time, a depth estimation model corresponding to the content type of the input image using on-device learning based on a result of analyzing the content type of the input image. The operating method includes a step of obtaining a depth map for the input image reflecting estimated depth information based on the depth estimation model according to on-device learning. The operating method includes a step of performing 3D transformation on the input image based on the depth map for the input image.

[0012] One embodiment of the present disclosure provides a computer-readable recording medium having recorded thereon a program for executing at least one of the embodiments of the disclosed method on a computer as a technical means for achieving the above-described technical task.

[0013] Other technical features will be readily apparent to those skilled in the art from the following drawings, descriptions and claims.

[0014] Figure 1 is a reference diagram for explaining the concept of an image processing device according to one embodiment.

[0015] FIG. 2A, FIG. 2B and FIG. 2C are drawings for explaining an operation of an image processing device performing 3D conversion from a 2D input image according to an example.

[0016] FIG. 3 is a schematic diagram illustrating an operation of an image processing device according to one embodiment performing 3D conversion from a 2D input image.

[0017] Figure 4 is an internal block diagram of an image processing device according to one embodiment.

[0018] FIG. 5 is a diagram for explaining a process in which an image processing device according to one embodiment analyzes a content type.

[0019] FIG. 6 is a diagram for explaining a process in which an image processing device according to one embodiment obtains a depth estimation model corresponding to a 2D input image from a cloud server.

[0020] FIG. 7 is a diagram for explaining a process in which an image processing device according to one embodiment obtains a depth estimation model using on-device learning.

[0021] FIG. 8 is a diagram for explaining a process in which an image processing device according to one embodiment obtains a depth estimation model using on-device learning.

[0022] FIG. 9 is a flowchart illustrating a method for an image processing device according to one embodiment to perform 3D transformation from a 2D input image.

[0023] Fig. 10 is an internal block diagram of an image processing device according to one embodiment.

[0024] FIG. 11 is a diagram for explaining a process in which an image processing device according to one embodiment analyzes a scene object of a 2D input image.

[0025] FIG. 12A is a diagram for explaining a process in which an image processing device dynamically changes a depth map according to one embodiment.

[0026] FIG. 12B is a diagram illustrating an example of a process in which an image processing device according to one embodiment dynamically changes a depth map.

[0027] FIG. 12C is a drawing for explaining an example of a process in which an image processing device dynamically changes a depth map according to one embodiment.

[0028] FIG. 13 is a drawing for explaining a process in which an image processing device according to one embodiment controls a three-dimensional effect.

[0029] FIG. 14 is a drawing for explaining a process in which an image processing device according to one embodiment controls a three-dimensional effect.

[0030] FIG. 15 is a flowchart illustrating a method for an image processing device according to one embodiment to perform 3D transformation from a 2D input image.

[0031] FIG. 16 is a drawing for explaining an example of the effect of an image processing device according to one embodiment.

[0032] Fig. 17 is a block diagram of an image processing device according to one embodiment.

[0033] The terms used in this disclosure are described as currently used general terms in consideration of the functions mentioned in this disclosure; however, these may mean various other terms depending on the intentions of engineers working in the field, precedents, the emergence of new technologies, etc. In addition, in certain cases, there are terms arbitrarily selected by the applicant, and in such cases, the meanings thereof will be described in detail in the description of the relevant embodiments. Therefore, the terms used in this disclosure should not be interpreted solely on the basis of the name of the term, but should be interpreted based on the meaning of the term and the overall contents of this disclosure.

[0034] Additionally, the terminology used in this disclosure is for the purpose of describing specific embodiments only and is not intended to limit the present disclosure.

[0035] Singular expressions may include plural expressions unless the context clearly indicates otherwise. Terms used herein, including technical or scientific terms, have the same meaning as commonly understood by a person of ordinary skill in the art described herein.

[0036] When a part of this specification is said to "include" a component, unless otherwise specifically stated, this does not exclude other components but rather implies the inclusion of other components. Furthermore, terms such as "part," "module," etc., used herein refer to a unit that processes at least one function or operation, which may be implemented in hardware or software, or a combination of hardware and software.

[0037] In the present disclosure, a processor may include various processing circuits and / or multiple processors. For example, the term “processor” as used herein, including in the claims, may include various processing circuits, including at least one processor. At least one processor, one or more processors, may be configured to perform various functions described herein, individually and / or collectively, in a distributed fashion. As used herein, “processor,” “at least one processor,” and “one or more processors” may be configured to perform various functions. However, these terms encompass, without limitation, situations where one processor performs some of the functions and other processor(s) perform other parts of the functions, and situations where a single processor may perform all of the functions. Furthermore, the at least one processor may include a combination of processors that perform various functions of the disclosed functions in a distributed manner. At least one processor may execute program instructions to achieve or perform various functions.

[0038] In addition, each component described below may additionally perform some or all of the functions performed by other components in addition to its own main function, and some of the main functions performed by each component may be performed exclusively by other components.

[0039] The expression "configured to" as used herein can be used interchangeably with, for example, "suitable for", "having the capacity to", "designed to", "adapted to", "made to", or "capable of", depending on the context.

[0040] The term "configured (or set up) to" may not necessarily mean "specifically designed to" hardware. Instead, in some contexts, the phrase "a system configured to" may mean that the system, in conjunction with other devices or components, is "capable of" doing something.

[0041] When a component is referred to herein as being "connected" or "connected" to another component, it should be understood that the component may be directly connected or directly connected to the other component, but may also be connected or connected via another component in between, unless otherwise specifically stated. When a part is referred to herein as being "connected" to another part, this includes not only cases where the parts are "directly connected," but also cases where the parts are "electrically connected" with another element in between.

[0042] As used herein, and particularly in the claims, the terms "above" and "above" and similar referents may refer to both the singular and the plural. Furthermore, unless the order of steps in a method according to the present disclosure is explicitly specified, the steps described may be performed in any appropriate order. The present disclosure is not limited by the order in which the steps are described.

[0043] The appearances of phrases such as “in some embodiments” or “in an embodiment” in various places throughout this specification are not necessarily all referring to the same embodiment.

[0044] Additionally, in the specification, the expression "at least one of a, b, or c" may refer to "a," "b," "c," "a and b," "a and c," "b and c," "all of a, b, and c," or variations thereof. The numbers used in the description of the specification (e.g., first, second, third, etc.) are merely identifiers to distinguish one element from another.

[0045] Some embodiments of the present disclosure may be represented by functional block configurations and various processing steps. Some or all of these functional blocks may be implemented by various hardware and / or software configurations that perform specific functions. For example, the functional blocks of the present disclosure may be implemented by one or more microprocessors or by circuit configurations for a given function. Furthermore, for example, the functional blocks of the present disclosure may be implemented using various programming or scripting languages. The functional blocks may be implemented as algorithms that execute on one or more processors. Furthermore, the present disclosure may employ conventional techniques for electronic configuration, signal processing, and / or data processing.

[0046] In order to clearly explain the present disclosure in the drawings, parts that are not related to the description have been omitted, and similar parts have been designated with similar drawing reference numerals throughout the specification. In addition, the drawing reference numerals used in each drawing are only for the purpose of explaining each drawing, and different drawing reference numerals used in different drawings do not indicate different elements. In addition, the connecting lines or connecting members between components illustrated in the drawings are only examples of functional connections and / or physical or circuit connections. In an actual device, connections between components may be indicated by various functional connections, physical connections, or circuit connections that may be replaced or added.

[0047] Below, with reference to the attached drawings, embodiments of the present disclosure are described in detail so that those skilled in the art can easily implement them. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein. In describing the embodiments, detailed descriptions of related known technologies will be omitted if they are determined to unnecessarily obscure the gist of the present disclosure.

[0048] In this disclosure, the term "user" refers to a person who utilizes an image processing device, and may include a consumer, evaluator, viewer, administrator, or installer. Furthermore, the term "manufacturer" in this disclosure may refer to a manufacturer that manufactures an image processing device and / or components included in the image processing device.

[0049] In the present disclosure, 'image' may mean a still image, a picture, a frame, a moving image composed of a plurality of consecutive still images, or a video.

[0050] In the present disclosure, a '2D (Two-Dimensional) image' is an image composed of a two-dimensional planar form in which each pixel corresponds to a row and a column, and may mean an image that does not include depth / height information.

[0051] In the present disclosure, a '3D (Three-Dimensional)' image is an image composed of a three-dimensional spatial form in which each pixel corresponds to row, column, and depth / height information, and may mean an image including depth / height information.

[0052] In this disclosure, a "scene" may refer to a series of consecutive image frames that constitute a video. A video may be composed of various scenes, and each scene may be connected to the next to form the overall flow of the video. Each scene that constitutes a video may be divided into an event occurring at a specific location and time, a specific theme, or a specific narrative unit.

[0053] In the present disclosure, a 'neural network' is a representative example of a computing system that simulates human brain nerves, and is not limited to an artificial neural network model using a specific algorithm. A neural network may also be referred to as an 'artificial neural network (ANN)' or a 'deep neural network (DNN)'. A neural network may include, for example, a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), a restricted boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), or deep Q-networks (Deep Q-Networks), a histogram of oriented gradients (HOG), a scale-invariant feature transform (SHIFT), a long short-term memory (LSTM), a support vector machine (SVM), a softmax, etc., but is not limited to the examples described above.

[0054] In the present disclosure, a "neural network model" may refer to a neural network generated / trained to perform operations to achieve a specific purpose. A neural network generated / trained to perform a specific function may be expressed as a "functional term" + "model" (e.g., a content type analysis model, a depth estimation model, an object analysis model, a depth dynamic change model, a three-dimensional effect control model, etc.).

[0055] In the present disclosure, when a neural network model is newly generated / trained / obtained, it can be expressed as 'updating the neural network model' or 'updating the parameters of the neural network model.' 'Updating' the neural network model can be referred to as 'renew,' 'adapt,' 'adjust,' 'modify,' or 'change.'

[0056] In this disclosure, "machine learning" may refer to an algorithm that enables a neural network model to learn from data, or an algorithm that enables a neural network model to receive input data and predict output data. "Deep learning" may refer to performing machine learning using a deep neural network model.

[0057] In the present disclosure, a 'parameter' or 'weight' is an element included in a matrix corresponding to a neural network model, and may mean a value applied to input data for inference using the neural network model. Each of the plurality of layers constituting the neural network model has a plurality of parameters or weights, and inference can be performed through operations between the operation results of the previous layer and the plurality of parameters or weights. The plurality of parameters or weights of the plurality of layers constituting the neural network model can be optimized through training of the neural network model. For example, the plurality of parameters or weights can be updated so that the loss value or cost value obtained from the neural network model is reduced or minimized during the training process.

[0058] FIG. 1 is a reference diagram for explaining the concept of an image processing device (100) according to one embodiment.

[0059] Referring to FIG. 1, the image processing device (100) may be an electronic device capable of receiving a 2D image (110) and converting it into a 3D image (120) and outputting it. In one embodiment, the image processing device (100) may be implemented as various types of electronic devices including a display.

[0060] The image processing device (100) may be fixed or mobile, and may be a 3D display (e.g., Light Field Display) that provides functions for experiencing 3D games, 3D movies, and 3D models, but is not limited thereto.

[0061] The image processing device (100) may include at least one of a digital TV capable of receiving digital broadcasting, a desktop, a smartphone, a tablet personal computer, a mobile phone, a video phone, an e-book reader, a laptop personal computer, a netbook computer, a digital camera, a PDA (Personal Digital Assistants), a PMP (Portable Multimedia Player), a camcorder, a navigation device, a wearable device, a smart watch, a home network system, a security system, and a medical device.

[0062] The image processing device (100) can be implemented as a flat display device, a curved display device having a screen with curvature, or a flexible display device whose curvature can be adjusted.

[0063] In one embodiment, the image processing device (100) may utilize artificial intelligence (AI) technology to perform an operation of converting a 2D image (110) into a 3D image (120). In one embodiment, the image processing device (100) may be an edge device in which artificial intelligence is connected to an electronic device that provides a 3D image (120) to a user.

[0064] In one embodiment, the image processing device (100) can acquire a 2D image (110) by inputting or receiving it.

[0065] In one embodiment, the image processing device (100) can analyze the content type of the acquired 2D image (110). In one embodiment, the image processing device (100) can analyze the content type of each scene constituting the acquired 2D image (110). This will be discussed in detail in FIG. 5.

[0066] In the present disclosure, the 'content type (class)' of the 2D input image (110) may include, but is not limited to, movies, dramas, FPS (First Person Shooter) games, RPG (Role Playing Game) games, RTS (Real Time Strategy) games, MMORPG (Massively Multiplayer Online Role Playing Game) games, documents, complex contents, presentation materials, etc. For example, complex contents may mean a single content that includes contents with various characteristics such as 2D animation, PPT, comment images, documents, etc., such as Internet lecture videos. The 'content type (class)' may be referred to as a 'content type (type)' or a 'content category (category)'.

[0067] In one embodiment, the image processing device (100) can generate / acquire in real time one or more depth estimation models corresponding to the 2D image (110) or each scene constituting the 2D image (110) based on the results of analyzing the content type of the acquired 2D image (110). This will be discussed in detail in FIGS. 6 to 8.

[0068] In the present disclosure, a 'depth estimation model' may mean an artificial neural network model learned / trained to predict depth information of each pixel constituting a 2D image (110). For example, the depth estimation model may mean an artificial neural network model learned / trained to predict depth information of each pixel constituting a 2D image (110) using a technology such as CNN, DNN, RNN, RBM, DBN, BRDNN, or deep Q-network, HOG, SHIFT, LSTM, SVM, SoftMax, etc., but is not limited thereto.

[0069] Depth information can refer to information indicating the location of a pixel relative to a certain reference plane in 3D space when performing 3D conversion from a 2D image, and can be expressed in units of meters or pixels.

[0070] In one embodiment, the image processing device (100) can acquire or generate in real time one or more depth estimation models corresponding to each scene constituting the 2D image (110) or the 2D image (110) using cloud-based AI technology or on-device-based AI technology.

[0071] In the case of cloud-based AI technology, the neural network model itself, the training of the neural network model, or the inference using the neural network model can be performed on a cloud server. In one embodiment, the image processing device (100) can obtain one or more depth estimation models corresponding to the 2D image (110) or each scene constituting the 2D image (110) in real time from the cloud server, based on the results of analyzing the content type of the 2D image (110).

[0072] In the case of on-device based AI technology, data can be processed in real time on the edge device itself, so that training of a neural network model and inference using the neural network model can be performed on the edge device. In one embodiment, the image processing device (100) can, based on the result of analyzing the content type of the 2D image (110), collect data on its own, that is, use on-device learning, and train a depth estimation model, thereby generating / acquiring one or more depth estimation models corresponding to the 2D image (110) or each scene constituting the 2D image (110) in real time.

[0073] In one embodiment, the on-device AI technology may be performed by at least one processor included in the image processing device (100). In one embodiment, the on-device AI technology may be referred to as on-device learning.

[0074] In one embodiment, the image processing device (100) may obtain a depth map (e.g., a first depth map) for the 2D image (110) based on one or more depth estimation models corresponding to the 2D image (110) or each scene constituting the 2D image (110), which are obtained in real time from a cloud server or obtained in real time on their own using on-device learning. The 'depth map' may mean a 2D image in which depth information of each pixel constituting the image is expressed as a value such as brightness or color of the pixel. The depth estimation model may receive the 2D input image (110) as input data and output the depth map for the 2D input image (110) as output data.

[0075] In the present disclosure, a depth map initially acquired for a 2D image (110) acquired by an image processing device (100), i.e., a depth map before non-linearly changing the depth map, may be referred to as a 'depth map' or a 'first depth map'.

[0076] In one embodiment, the image processing device (100) can obtain a modified depth map (e.g., a second depth map) by non-linearly modifying a depth map (e.g., a first depth map) based on the results of analyzing the sizes and distributions of objects included in each of the scenes constituting the 2D image (110). This will be described in detail with reference to FIGS. 11 to 13.

[0077] In the present disclosure, a depth map obtained by nonlinearly changing a first depth map initially acquired for a 2D image (110) input / received by an image processing device (100) may be referred to as a ‘modified depth map’ or a ‘second depth map’.

[0078] In one embodiment, the image processing device (100) can perform 3D transformation from a 2D image (110) based on a depth map (e.g., a first depth map) or a modified depth map (e.g., a second depth map).

[0079] In one embodiment, the image processing device (100) can obtain a 3D image (120) converted from a 2D image (110). In one embodiment, the image processing device (100) can control a display to output the 3D image (120).

[0080] In one embodiment, the image processing device (100) can control the display to generate and output a user interface (UI) that indicates the content type and the three-dimensionality of each scene corresponding to each scene constituting the 3D output image (120) converted from the 2D input image (110). The user can determine the content type and the three-dimensionality level of each scene of the 3D output image (120) being viewed through the UI provided by the image processing device (100).

[0081] According to one embodiment of the present disclosure, the image processing device (100) dynamically applies a depth estimation model in real time based on analysis of the content type of an input 2D image, analyzes objects included in each scene, and nonlinearly changes a depth map, thereby reducing depth estimation errors, further improving the accuracy of depth estimation, and enhancing the satisfaction of a user using a 3D display.

[0082] However, the effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned can be clearly understood by a person having ordinary skill in the technical field to which the present disclosure belongs from the description below.

[0083] FIG. 2A, FIG. 2B and FIG. 2C are drawings for explaining an operation in which an image processing device performs 3D transformation from a 2D input image according to an example.

[0084] FIG. 2A is a diagram illustrating an example of an operation in which an image processing device performs 3D transformation from a 2D input image according to an example.

[0085] Referring to FIG. 2A, the image processing device can receive a 2D image, perform 3D conversion according to blocks 210 to 240, and output a 3D image.

[0086] When a 2D image is input, the image processing device can estimate the depth of objects included in the image based on the input 2D image (210), generate a new viewpoint view based on a depth map which is a result of the depth estimation (220), perform hole-filling (230) to fill an empty area (e.g., an occlusion area) around an object in the image according to the generation of the new viewpoint view, and perform pixel mapping (240) to arrange pixels of the display to match the hole-filled new viewpoint view generated in real time.

[0087] FIG. 2B is a drawing specifically explaining a process (220) in which an image processing device creates a new viewpoint according to an example.

[0088] Referring to Figure 2B, the principle of generating a new viewpoint view is to generate a stereoscopic image from a 2D image based on a depth map. A stereoscopic image is an image that combines two images with two different viewpoints by utilizing the disparity in the positions of objects in the images seen by the left and right eyes, thereby creating an optical illusion that allows the viewer to perceive the depth of objects in the images.

[0089] For example, in the case of retreat (222, focal length a < focal length b), the image processing device, based on the depth map acquired at 210, adjusts the interocular angle between the left and right eyes to make an object appear to be further away from its actual location. The angle between the two eyes of reality (221) To make it smaller, a stereoscopic image 220A can be created by generating an image corresponding to the left eye and an image corresponding to the right eye and combining them.

[0090] For example, in the case of projection (223, focal length a > focal length c), the image processing device, based on the depth map acquired at 210, calculates the interocular angle between the left and right eyes to make an object appear to pop out from its actual location. The angle between the two eyes of reality (221) To make it larger, a stereoscopic image 220B can be created by generating and combining an image corresponding to the left eye and an image corresponding to the right eye.

[0091] FIG. 2C is a drawing for explaining an example of a UI (user interface) for adjusting the three-dimensional effect in a slide manner provided by an image processing device according to an example.

[0092] Referring to FIG. 2C, the user can adjust the intensity of the stereoscopic effect of the 3D converted image using the UI 250A or UI 250B provided by the image processing device. However, these types of UIs are only for manually adjusting the stereoscopic effect of the entire image while the user is watching the image, and do not automatically adjust the stereoscopic effect of the image to suit the type of content of the image or to suit the type of content of each scene when the scene of the image changes, as in one embodiment of the present disclosure.

[0093] The present disclosure provides a method for dynamically applying a depth estimation model using cloud-based AI or on-device-based AI technology based on the results of analyzing the content type of an input 2D image in a depth estimation process (210) to reduce depth estimation errors and improve the accuracy of depth estimation.

[0094] The present disclosure provides a method for non-linearly changing a depth map and controlling a three-dimensional effect based on the results of analyzing the content type of an input 2D image in a new viewpoint view generation process (220) to improve the satisfaction and convenience of a user using a 3D display.

[0095] However, the effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned can be clearly understood by a person having ordinary skill in the technical field to which the present disclosure belongs from the description below.

[0096] FIG. 3 is a schematic diagram for explaining an operation of an image processing device (100) according to one embodiment of the present invention to perform 3D conversion from a 2D input image (110).

[0097] Referring to FIG. 3, in block 310, an image processing device (100) according to one embodiment can analyze the content type of a 2D input image (110). This will be discussed in detail in FIG. 5.

[0098] The image processing device (100) can use the result of analyzing the content type acquired in block 310 to acquire a depth estimation model corresponding to the 2D input image (110) in real time in block 320. The image processing device (100) can use the result of analyzing the content type acquired in block 310 to non-linearly change the depth map in block 350. The image processing device (100) can use the result of analyzing the content type acquired in block 310 to control the three-dimensional effect of the 2D input image (110) in block 360.

[0099] In block 320, an image processing device (100) according to one embodiment can obtain a depth estimation model corresponding to a 2D input image (110) in real time based on a result of analyzing the content type of the 2D input image (110).

[0100] The image processing device (100) can obtain a depth estimation model corresponding to the 2D input image (110) in real time through a cloud server based on the result of analyzing the content type of the 2D input image (110). The image processing device (100) can obtain a depth estimation model corresponding to the 2D input image (110) in real time through on-device learning based on the result of analyzing the content type of the 2D input image (110). This will be described in detail with reference to FIGS. 6 to 8.

[0101] The image processing device (110) can use a depth estimation model corresponding to the 2D input image (110) acquired in block 320 to perform depth estimation in block 330 to acquire a depth map for the 2D input image (110).

[0102] Block 330 may operate similarly to block 210 of FIG. 2A. Any content overlapping with that described in FIG. 2A is referred to FIG. 2A and is omitted here.

[0103] In block 340, the image processing device (100) according to one embodiment can analyze the size and distribution of objects included in each scene constituting the 2D input image (110) based on the 2D input image (110) and the depth map for the 2D input image (110). This will be described in detail in FIG. 11.

[0104] The image processing device (100) can use the results of the size and distribution analysis of objects included in each scene obtained in block 340 to nonlinearly change the depth map in block 350. The image processing device (100) can use the results of the size and distribution analysis of objects included in each scene obtained in block 340 to control the three-dimensional effect of the 2D input image (110) in block 360.

[0105] In block 350, an image processing device (100) according to one embodiment can obtain a modified depth map (e.g., a second depth map) (20) by non-linearly changing a depth map (e.g., a first depth map) (10).

[0106] The image processing device (100) can obtain a modified depth map (e.g., second depth map) (20) by non-linearly changing the depth map (e.g., first depth map) (10) based on at least one of the results of the size and distribution analysis of objects included in each scene constituting the 2D input image (110) obtained in block 340, the results of the content type analysis of the 2D input image (110) obtained in block 310, or additional information. This will be described in detail with reference to FIGS. 12A to 13.

[0107] The image processing device (100) can use the modified depth map (e.g., second depth map) for the 2D input image (110) acquired in block 350 to generate a new viewpoint view in block 370.

[0108] In block 360, the image processing device (100) according to one embodiment can determine the relative positions of objects included in each scene constituting the 2D input image (110) from a virtual convergence plane corresponding to the screen, in order to control the three-dimensional effect of the 2D input image (110).

[0109] The image processing device (100) can obtain information on the relative positions of objects included in each scene constituting the 2D input image (110) from a virtual convergence plane corresponding to the screen based on at least one of the results of the size and distribution analysis of objects included in each scene constituting the 2D input image (110) obtained in block 340, the results of the content type analysis of the 2D input image (110) obtained in block 310, or additional information. This will be described in detail with reference to FIGS. 14 and 15.

[0110] The image processing device (100) can use information about the relative positions of objects included in each scene constituting the 2D input image (110) from the convergence plane obtained in block 360 to create a new viewpoint view in block 370.

[0111] Blocks 370 to 390 may operate similarly to blocks 220 to 240 of FIG. 2A. Any content overlapping with that described in FIG. 2A is referred to FIG. 2A and will not be described herein.

[0112] According to blocks 310 to 390 of FIG. 3, an image processing device (100) according to one embodiment can perform 3D conversion from a 2D input image (110) to obtain a 3D output image (120). The image processing device (100) can control a display to output the 3D output image (120).

[0113] The present disclosure provides a method for dynamically applying a depth estimation model using cloud-based AI or on-device-based AI technology based on the results of analyzing the content type of an input 2D image in blocks 310 to 330 to improve the accuracy of depth estimation by reducing depth estimation errors.

[0114] The present disclosure provides a method for controlling a three-dimensional effect of a 2D input image by non-linearly changing a depth map based on a result of analyzing the content type of an input 2D image in blocks 340 to 370 and determining the relative positions of objects included in each scene constituting the 2D input image from a virtual convergence plane corresponding to a screen, in order to improve the satisfaction and convenience of a user using a 3D display.

[0115] However, the effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned can be clearly understood by a person having ordinary skill in the technical field to which the present disclosure belongs from the description below.

[0116] Figure 4 is an internal block diagram of an image processing device (100) according to one embodiment.

[0117] Referring to FIG. 4, an image processing device (100) according to one embodiment may include a content type analysis unit (410), a depth estimation model acquisition unit (420), a depth estimation unit (430), and a 3D conversion performing unit (440).

[0118] The content type analysis unit (410), the depth estimation model acquisition unit (420), the depth estimation unit (430), and the 3D transformation execution unit (440) may be implemented with at least one processor. The content type analysis unit (410), the depth estimation model acquisition unit (420), the depth estimation unit (430), and the 3D transformation execution unit (440) may operate according to at least one instruction stored in a memory (e.g., 102 of FIG. 18).

[0119] Although FIG. 4 illustrates the content type analysis unit (410), the depth estimation model acquisition unit (420), the depth estimation unit (430), and the 3D conversion performance unit (440) individually, the content type analysis unit (410), the depth estimation model acquisition unit (420), the depth estimation unit (430), and the 3D conversion performance unit (440) may be implemented through a single processor. In this case, the content type analysis unit (410), the depth estimation model acquisition unit (420), the depth estimation unit (430), and the 3D conversion performance unit (440) may be implemented through a dedicated processor, or may be implemented through a combination of a general-purpose processor such as an application processor (AP), a central processing unit (CPU), or a graphics processing unit (GPU) and software. In addition, in the case of a dedicated processor, it may include a memory for implementing an embodiment of the present disclosure, or may include a memory processing unit for utilizing an external memory. Additionally, in the case of AI-specific processors such as NPU (Neural Processing Unit), they can be designed with a hardware structure specialized for processing specific AI models.

[0120] The content type analysis unit (410), the depth estimation model acquisition unit (420), the depth estimation unit (430), and the 3D conversion execution unit (440) may be configured with multiple processors. In this case, the content type analysis unit (410), the depth estimation model acquisition unit (420), the depth estimation unit (430), and the 3D conversion execution unit (440) may be implemented by a combination of dedicated processors, or may be implemented by a combination of multiple general-purpose processors such as an AP, a CPU, or a GPU, and software.

[0121] In one embodiment, the content type analysis unit (410) can analyze the content type of the 2D input image (110). In one embodiment, the content type analysis unit (410) can include suitable logic, circuits, interfaces, and / or codes that can be operated to analyze the content type of the 2D input image (110). The content type analysis unit (410) can analyze the content type of the 2D input image (110) using at least one of an artificial neural network (e.g., the content type analysis model (500) of FIG. 5) learned / trained to analyze the content type of the input image or a policy-based algorithm. In one embodiment, the content type analysis unit (410) can obtain, as a result of the content type analysis, information related to the probability that the 2D input image (110) corresponds to each of the predefined content types.

[0122] In one embodiment, when there is a scene change in the 2D input image (110), the content type analysis unit (410) may analyze the content type of each scene constituting the 2D input image (110). 'Scene change' may mean a change in the scene in the image, i.e., a change from a specific scene to another scene. The content type analysis result may include the result of analyzing the content type of each scene constituting the 2D input image (110).

[0123] In one embodiment, the content type analysis unit (410) can transmit the content type analysis result of the 2D input image (110) to the depth estimation model acquisition unit (420).

[0124] In one embodiment, the depth estimation model acquisition unit (420) can receive a content type analysis result of a 2D input image (110) from the content type analysis unit (410).

[0125] In one embodiment, the depth estimation model acquisition unit (420) may acquire a depth estimation model corresponding to the 2D input image (110) in real time based on the content type analysis result of the 2D input image (110). In one embodiment, the depth estimation model acquisition unit (420) may include appropriate logic, circuits, interfaces, and / or codes that may be operable to acquire a depth estimation model corresponding to the 2D input image (110) in real time.

[0126] In one embodiment, the depth estimation model acquisition unit (420) can acquire a depth estimation model corresponding to the 2D input image (110) in real time from a cloud server based on the content type analysis result of the 2D input image (110).

[0127] In one embodiment, the depth estimation model acquisition unit (420) can newly learn / generate / acquire one or more depth estimation models corresponding to the 2D input image (110) in real time using on-device learning based on the content type analysis result of the 2D input image (110).

[0128] In one embodiment, when there is a scene change of the 2D input image (110), the depth estimation model acquisition unit (420) can acquire / generate depth estimation models corresponding to each scene constituting the 2D input image (110) in real time based on the result of analyzing the content type of each scene constituting the 2D input image (110). The depth estimation models corresponding to each scene constituting the 2D input image (110) may be the same or different from each other. If even some of the parameter values ​​of the filters used in each layer constituting the depth estimation models are different, the depth estimation models can be referred to as being different from each other.

[0129] The depth estimation model acquired in real time from the cloud server can be dynamically applied in real time according to the result of the content type analysis of the 2D input image (110) or the result of the content type analysis of each scene by scene transition of the 2D input image (110), but is not updated by itself by the depth estimation model acquisition unit (420). The depth estimation model acquired using on-device learning can be dynamically applied in real time according to the result of the content type analysis of the 2D input image (110) or the result of the content type analysis of each scene by scene transition of the 2D input image (110), and can be updated by itself by the depth estimation model acquisition unit (420).

[0130] In one embodiment, the depth estimation model acquisition unit (420) can transmit one or more acquired / generated depth estimation models to the depth estimation unit (430).

[0131] In one embodiment, the depth estimation unit (430) can receive one or more depth estimation models from the depth estimation model acquisition unit (420).

[0132] In one embodiment, the depth estimation unit (430) may obtain a depth map (e.g., a first depth map) (10) for the 2D input image (110) by performing depth estimation on the 2D input image (110) using one or more received depth estimation models. In one embodiment, the depth estimation unit (430) may include suitable logic, circuitry, interfaces, and / or codes that may be operable to obtain a depth map (e.g., a first depth map) (10) for the 2D input image (110).

[0133] In one embodiment, when there is a scene change of the 2D input image (110), the depth estimation unit (430) applies depth estimation models corresponding to each of the scenes constituting the 2D input image (110) to each scene constituting the 2D input image (110), thereby performing depth estimation for each of the scenes constituting the 2D input image (110), thereby obtaining a depth map (e.g., first depth map) (10) for the 2D input image.

[0134] In one embodiment, the depth estimation unit (430) can transmit the acquired depth map (e.g., first depth map) (10) to the 3D transformation performing unit (440).

[0135] In one embodiment, the 3D transformation performing unit (440) can receive a depth map (e.g., first depth map) (10) for a 2D input image (110) from the depth estimation unit (430).

[0136] In one embodiment, the 3D transformation performing unit (440) may perform 3D transformation from a 2D input image (110) based on a received depth map (e.g., a first depth map) (10) to obtain a 3D output image (120). In one embodiment, the 3D transformation performing unit (440) may include appropriate logic, circuits, interfaces, and / or codes that may be operable to perform 3D transformation from a 2D input image (110) to obtain a 3D output image (120).

[0137] The 3D transformation performing unit (440) can obtain a 3D output image (120) from a 2D input image (110) by performing new viewpoint view generation, hole filling, and pixel mapping based on a depth map (e.g., first depth map) (10). The new viewpoint view generation, hole filling, and pixel mapping processes can operate similarly to blocks 220 to 240 of FIG. 2A. For the overlapping content described in FIG. 2A, refer to FIG. 2A, and the description thereof will be omitted here.

[0138] An image processing device (100) according to one embodiment of the present disclosure provides a method of dynamically applying a depth estimation model in real time based on a result of analyzing the content type of an input 2D image, thereby reducing depth estimation errors, improving the accuracy of depth estimation, and enhancing the satisfaction and convenience of a user using a 3D display.

[0139] However, the effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned can be clearly understood by a person having ordinary skill in the technical field to which the present disclosure belongs from the description below.

[0140] The specific operations of the image processing device (100) according to one embodiment of the present disclosure to perform 3D conversion from an input 2D image will be described in more detail through the drawings and descriptions thereof described below.

[0141] FIG. 5 is a drawing for explaining a process in which an image processing device (100) according to one embodiment analyzes a content type.

[0142] Referring to FIG. 5, an image processing device (100) according to one embodiment (e.g., a content type analysis unit (410) of the image processing device (100)) can analyze the content type of a 2D input image (110). The image processing device (100) can analyze the probability that the 2D input image (110) corresponds to each predefined content type, and obtain a content type analysis result (510) of the 2D input image (110).

[0143] The types of content defined above may include, but are not limited to, movies, dramas, FPS games, RPG games, RTS games, MMORPG games, documents, complex content, presentation materials, etc.

[0144] According to one embodiment, the content type analysis result (510) may include information related to the probability that the 2D input image (110) corresponds to each of the predefined content types. For example, the content type analysis result (510) may include information indicating a specific content type among the predefined content types to which the input 2D image (110) corresponds with the highest probability. For example, the content type analysis result (510) may include probability values ​​indicating the probability that the 2D input image (110) corresponds to each of the predefined content types.

[0145] According to one embodiment, the image processing device (100) can analyze the content type of a 2D input image (110) using a content type analysis model (500).

[0146] The content type analysis model (500) may refer to an artificial neural network model learned / trained to predict the content type of a 2D image (110). For example, the content type analysis model (500) may refer to an artificial neural network model learned / trained to predict the content type of a 2D image (110) using techniques such as CNN, DNN, RNN, RBM, DBN, BRDNN, or deep Q-network, HOG, SHIFT, LSTM, SVM, SoftMax, etc., but is not limited thereto.

[0147] The content type analysis model (500) can receive a 2D input image (110) as input data and output a content type analysis result (510) of the 2D input image (110) as output data. The image processing device (100) can obtain the learned content type analysis model (500) from a cloud server. The image processing device (100) can newly learn / generate / obtain the content type analysis model (500) using on-device learning.

[0148] According to one embodiment, the image processing device (100) may analyze the content type of a 2D input image (110) using a policy-based algorithm. The policy-based algorithm for analyzing the content type of a 2D image may be predefined / set by the manufacturer of the image processing device (100).

[0149] According to one embodiment, the image processing device (100) can analyze the content type of a 2D input image (110) by combining a content type analysis model (500) and a policy-based algorithm.

[0150] For example, if the 2D input image (110) is a drama, the image processing device (100) can obtain a content type analysis result including information indicating that the content type that the 2D input image (110) most likely corresponds to is a drama, or probability values ​​indicating the probability that the 2D input image (110) corresponds to a movie, drama, FPS game, RPG game, RTS game, MMORPG game, document, composite content, and presentation material, using a content type analysis model (500) or a policy-based algorithm.

[0151] According to one embodiment, when there is a scene transition of the 2D input image (110), the image processing device (100) can analyze the content type of the 2D input image (110) by analyzing the content type of each scene constituting the 2D input image (110).

[0152] For example, the image processing device (100) can detect a scene change of a 2D input image (110). When the image processing device (100) detects a scene change, the image processing device (100) can analyze the content type of each scene constituting the 2D input image (110).

[0153] In this case, the content type analysis result (510) may include information related to the probability that each of the scenes constituting the 2D input image (110) corresponds to each of the specific pre-defined content types. For example, the content type analysis result (510) may include information indicating the specific content types that each of the scenes constituting the 2D input image (110) corresponds to has the highest probability among the pre-defined content types. For example, the content type analysis result (510) may include probability values ​​that each of the scenes constituting the 2D input image (110) corresponds to each of the specific pre-defined content types.

[0154] For example, it can be assumed that a 2D input image (110) is a composite image composed of four scenes (S1, S2, S3, S4), and the content types of each scene are that scene S1 corresponds to a drama, scene S2 corresponds to a document, scene S3 corresponds to a drama, and scene S4 corresponds to an FPS game. The image processing device (100) can detect the scene transition S1->S2->S3->S4 of the 2D input image (110). The image processing device (100) can obtain a content type analysis result including information indicating that the content types that each scene S1, S2, S3, and S4 most likely correspond to are respectively a drama, a document, a drama, and an FPS game, or probability values ​​indicating the probability that each scene S1, S2, S3, and S4 corresponds to respectively a movie, a drama, an FPS game, an RPG game, an RTS game, an MMORPG game, a document, a composite content, a presentation material, etc., using a content type analysis model (500) or a policy-based algorithm. In Fig. 5, the 2D input image (110) is described as being composed of four scenes, but this is only an example, and the 2D input image (110) may be composed of numerous scenes or one scene.

[0155] The image processing device (100) can use the result of analyzing the type of content acquired (510) to acquire a depth estimation model corresponding to a 2D input image (110) in real time (see FIGS. 6 to 8), to nonlinearly change a depth map (see FIGS. 11 to 13), and to control a three-dimensional effect of a 2D input image (110) (see FIGS. 14 to 15).

[0156] An image processing device (100) according to one embodiment of the present disclosure provides a method of dynamically applying a depth estimation model and non-linearly changing a depth map based on a result of analyzing the content type of an input 2D image, thereby reducing a depth estimation error, improving the accuracy of depth estimation, and enhancing the satisfaction and convenience of a user using a 3D display.

[0157] However, the effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned can be clearly understood by a person having ordinary skill in the technical field to which the present disclosure belongs from the description below.

[0158] A 'depth estimation model' may refer to an artificial neural network model learned / trained to predict depth information of each pixel constituting a 2D image (110). For example, the depth estimation model may refer to an artificial neural network model learned / trained to predict depth information of each pixel constituting a 2D image (110) using a technology such as CNN, DNN, RNN, RBM, DBN, BRDNN, or deep Q-network, HOG, SHIFT, LSTM, SVM, SoftMax, etc., but is not limited thereto.

[0159] Hereinafter, with reference to FIG. 6, an operation of an image processing device (100) according to one embodiment to obtain a depth estimation model using cloud-based AI technology will be described, and with reference to FIGS. 7 and 8, an operation of an image processing device (100) according to one embodiment to obtain a depth estimation model using on-device-based AI technology will be described.

[0160] FIG. 6 is a diagram for explaining a process in which an image processing device (100) according to one embodiment obtains a depth estimation model corresponding to a 2D input image (110) from a cloud server (600).

[0161] In the case of cloud-based AI technology, the neural network model itself and the training of the neural network model can be performed on a cloud server. In one embodiment, the image processing device (100) can obtain one or more depth estimation models corresponding to the 2D image (110) or each scene constituting the 2D image (110) in real time from the cloud server (600) based on the results of analyzing the content type of the 2D image (110).

[0162] In one embodiment, depth estimation models for each content type can be pre-trained in a cloud server (600) using a learning database of 2D image-depth map pairs for each pre-defined content type. These can be expressed as depth estimation models trained offline. The depth estimation models trained offline can be installed in the image processing device (100) or server.

[0163] In one embodiment, the image processing device (100) may request the cloud server (600) to transmit depth estimation models for each content type learned offline, and the cloud server (600) may transmit only some of the requested depth estimation models to the image processing device (100) when there is a request from the image processing device (100). In one embodiment, the image processing device (100) may receive depth estimation models for each content type learned offline in advance from the cloud server (600) and store them within the image processing device (100), and may directly use some of the depth estimation models that are necessary without making a separate request to the cloud server (600).

[0164] For example, Model 1, Model 2, Model 3, and Model 4 learned offline may be stored in a cloud server (600) or an image processing device (100). Model 1 may be a model learned in advance using a learning database of movie or drama image-depth map pairs, Model 2 may be a model learned in advance using a learning database of FPS game image-depth map pairs, Model 3 may be a model learned in advance using a learning database of document-depth map pairs, and Model 4 may be a model learned in advance using a learning database of composite content-depth map pairs.

[0165] Referring to FIG. 6, at 610, an image processing device (100) (e.g., a depth estimation model acquisition unit (420) of the image processing device (100)) according to one embodiment can process a content type analysis result (510) received from a content type analysis unit (410).

[0166] For example, the image processing device (100) can process the content type analysis result using an IIR (Infinite Impulse Response) filter. The values ​​(probability values) representing the probability that each of the scenes constituting the 2D input image (110) included in the content type analysis result corresponds to each of the predefined content types may be a sequence of probability values ​​that change over time. The image processing device (100) can apply a specific weight to the content type analysis result, which is a sequence of probability values ​​that change over time, using a moving average method or an exponentially weighted moving average method, to calculate an accumulated weighted average for an arbitrary period of time (e.g., a time interval corresponding to a certain number of frames, a time interval at which a scene change occurs, etc.). At this time, the image processing device (100) can adjust the weight to determine how much the current average value is affected by new probability value data. The depth estimation model acquisition unit (420) can reduce the volatility of input probability value data by processing the content type analysis results using an IIR filter, thereby reducing the flicker phenomenon and improving stability.

[0167] In 620, the image processing device (100) according to one embodiment (e.g., the depth estimation model acquisition unit (420) of the image processing device (100)) can determine / predict the content type that the 2D input image (110) most likely corresponds to, based on the result of the processed content type analysis.

[0168] In 630, the image processing device (100) according to one embodiment can obtain in real time one depth estimation model corresponding to the type of content that the 2D input image (110) corresponds to with the highest probability from the cloud server (600).

[0169] An image processing device (100) according to one embodiment may receive and store depth estimation models for each content type that have been previously learned from the cloud server (600) in advance. The image processing device (100) may select / choose one depth estimation model corresponding to the content type that has the highest probability of being corresponded to the 2D input image (110) among the depth estimation models for each content type that have been previously stored, thereby obtaining a depth estimation model corresponding to the 2D input image (110) in real time.

[0170] For example, the image processing device (100) can select / choose / acquire a depth estimation model Mx corresponding to a 2D input image (110) according to [Mathematical Formula 1].

[0171] [Mathematical Formula 1]

[0172]

[0173] n may mean the total number of depth estimation models for each content type acquired by the image processing device (100) from the cloud server (600). The 2D input image (110) may mean the probability that each depth estimation model (M1, M2, M3, M4,,,,Mn) corresponds to a content type, and may correspond to data processed in 610 included in the content type analysis result (510).

[0174] For example, the image processing device (100) can determine / predict the content type that the 2D input image (110) most likely corresponds to as a drama based on the result of the processed content type analysis. The image processing device (100) can select / choose / acquire one depth estimation model M1, i.e., Model 1, corresponding to the content type that the 2D input image (110) most likely corresponds to as a drama, from the cloud server (600).

[0175] An image processing device (100) according to one embodiment can obtain, in real time, one or more corresponding depth estimation models for each of the scenes constituting a 2D input image (110) from a cloud server (600).

[0176] An image processing device (100) according to one embodiment selects / chooses one or more depth estimation models corresponding to content types that have the highest probability of corresponding to each of the scenes constituting the 2D input image (110) from among depth estimation models for each type of content stored in advance, thereby obtaining one or more corresponding depth estimation models in real time for each of the scenes constituting the 2D input image (110).

[0177] Accordingly, rather than applying a single depth estimation model to a 2D input image, the image processing device (100) can dynamically apply depth estimation models to each scene according to the content type of each scene constituting the input 2D image. In particular, the image processing device (100) can reduce errors in depth estimation and improve the accuracy of depth estimation when a 2D image including various scenes with different content types is input.

[0178] For example, when a scene change (640) of a 2D input image (110) is detected, at 610, the image processing device (100) can process the content type analysis result by initializing the accumulated content type analysis result data. The image processing device (100) can initialize the accumulated content type analysis result data whenever there is a scene change (640) of the 2D input image (110). At 620, the image processing device (100) can determine / predict the content types that each of the scenes constituting the 2D input image (110) corresponds to with the highest probability, based on the processed content type analysis result.

[0179] At 630, the image processing device (100) can obtain, in real time, one or more corresponding depth estimation models for each of the scenes constituting the 2D input image (110) from the cloud server (600).

[0180] The image processing device (100) selects / chooses one or more depth estimation models corresponding to the content type that each scene has the highest probability of corresponding to, among the depth estimation models for each content type that are stored in advance, for each scene that constitutes a 2D input image (110), thereby obtaining depth estimation models corresponding to each scene that constitutes a 2D input image (110) in real time.

[0181] If the type of content that each scene constituting the 2D input image (110) is most likely to correspond to is the same, the depth estimation models corresponding to each scene may be the same, and if the type of content that each scene is most likely to correspond to is different, the depth estimation models corresponding to each scene may be different.

[0182] For example, it can be assumed that a 2D input image (110) is a composite image composed of four scenes (S1, S2, S3, S4), and the content types of each scene are as follows: scene S1 corresponds to a drama, scene S2 corresponds to a document, scene S3 corresponds to a drama, and scene S4 corresponds to an FPS game. The image processing device (100) can detect the scene transition S1->S2->S3->S4 of the 2D input image (110). The image processing device (100) can select / select / obtain Model 1 as a depth estimation model corresponding to 'drama' which has the highest probability of corresponding to scene S1 among Model 1, Model 2, Model 3, and Model 4 acquired from the cloud server (600), select / select / obtain Model 4 as a depth estimation model corresponding to 'document' which has the highest probability of corresponding to scene S2, select / select / obtain Model 1 as a depth estimation model corresponding to 'composite content' which has the highest probability of corresponding to scene S3, and select / select / obtain Model 2 as a depth estimation model corresponding to 'FPS game' which has the highest probability of corresponding to scene S4. The image processing device (100) can select / select / acquire in real time one or more depth estimation models Model 1, Model 2, Model 1, and Model 4 corresponding to each of the scenes S1, S2, S3, and S4 constituting the 2D input image (110).

[0183] The image processing device (100) can use one depth estimation model corresponding to the acquired 2D input image (110) or one or more depth estimation models corresponding to each of the scenes constituting the 2D input image (100) to obtain a depth map (e.g., a first depth map) (10) for the 2D input image (110).

[0184] For example, the image processing device (100) can obtain a depth map (e.g., first depth map) (10) for the 2D input image (110) using Model 1 corresponding to the acquired 2D input image (110).

[0185] For example, a depth map (e.g., a first depth map) (10) for a 2D input image (110) can be obtained by using Model 1 corresponding to the acquired scene S1 to obtain a depth map for S1, using Model 4 corresponding to S2 to obtain a depth map for S2, using Model 1 corresponding to S3 to obtain a depth map for S3, and using Model 2 corresponding to S4 to obtain a depth map for S4.

[0186] In Fig. 6, only four depth estimation models for each content type are shown, but this is only an example, and the image processing device (100) can obtain fewer or more depth estimation models for each content type than four from the cloud server (600). In addition, in Fig. 6, the 2D input image (110) is described as being composed of four scenes, but this is only an example, and the 2D input image (110) may be composed of numerous scenes or a single scene.

[0187] An image processing device (100) according to one embodiment of the present disclosure provides a method of dynamically applying a depth estimation model in real time using cloud-based AI technology according to the content type of an input 2D image or the content type of each scene of an input 2D image, thereby reducing depth estimation errors, improving the accuracy of depth estimation, and enhancing the satisfaction and convenience of a user using a 3D display.

[0188] However, the effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned can be clearly understood by a person having ordinary skill in the technical field to which the present disclosure belongs from the description below.

[0189] FIG. 7 is a diagram for explaining a process in which an image processing device (100) according to one embodiment obtains a depth estimation model using on-device learning.

[0190] In the case of on-device based AI technology, data can be processed in real time on the edge device itself, so that training of a neural network model and inference using the neural network model can be performed on the edge device. In one embodiment, the image processing device (100) can collect data and train a depth estimation model on its own, that is, using on-device learning without going through the cloud server (600), based on the result of analyzing the content type of the 2D image (110), thereby generating / acquiring one or more depth estimation models corresponding to the 2D image (110) or each scene constituting the 2D image (110) in real time.

[0191] The image processing device (100) can learn / create a new depth estimation model using on-device learning by utilizing a general-purpose processor such as a central processing unit (CPU), a graphic processing unit (GPU), or a dedicated processor such as a neural processing unit (NPU) mounted within the image processing device (100).

[0192] Referring to FIG. 7, at 710, an image processing device (100) according to an embodiment (e.g., a depth estimation model acquisition unit (420) of the image processing device (100)) may process a content type analysis result (510) received from a content type analysis unit (410). For example, the image processing device (100) may process the content type analysis result using an IIR (Infinite Impulse Response) filter. For content overlapping with that described at 610 of FIG. 6, refer to FIG. 6, and a description thereof will be omitted here.

[0193] In 720, the image processing device (100) according to one embodiment can obtain a depth estimation model Model Z corresponding to a 2D input image (110) in real time by updating the parameters of the depth estimation model (701) installed in the image processing device (100) or server and training it anew based on the result (510) of analyzing the type of processed content.

[0194] A depth estimation model (701) for each content type can be pre-learned offline using a learning database of 2D image-depth map pairs for each pre-defined content type, and can be installed in an image processing device (100) or a server.

[0195] In one embodiment, the image processing device (100) may request the cloud server to transmit depth estimation models for each content type learned offline in order to obtain a newly learned depth estimation model, and the cloud server may transmit only some of the requested depth estimation models to the image processing device (100) when there is a request from the image processing device (100). In one embodiment, the image processing device (100) may receive depth estimation models for each content type learned offline in advance from the cloud server and store them in the image processing device (100) in order to obtain a newly learned depth estimation model, and may directly use some of the depth estimation models that are necessary without making a separate request to the cloud server.

[0196] For example, Model 1, Model 2, Model 3, and Model 4 learned offline can be stored in a cloud server or image processing device (100). Model 1 can be a model learned in advance using a learning database of movie or drama image-depth map pairs, Model 2 can be a model learned in advance using a learning database of FPS game image-depth map pairs, Model 3 can be a model learned in advance using a learning database of document-depth map pairs, and Model 4 can be a model learned in advance using a learning database of composite content-depth map pairs.

[0197] The image processing device (100) can determine / predict the content type with the highest probability of corresponding to the 2D input image (110) based on the processed content type analysis result (510).

[0198] An image processing device (100) can learn / generate a depth estimation model Model Z corresponding to a 2D input image (110) in real time by updating parameters of a depth estimation model Model X corresponding to a content type with the highest probability of corresponding to a 2D input image (110) among depth estimation models (701) using learning data (700) composed of a pair of a 2D input image (110) generated by the image processing device (100) and a depth map (e.g., a first depth map) (10) for the 2D input image (110).

[0199] The image processing device (100) can obtain a depth map (e.g., first depth map) (10) for the 2D input image (110) included in the learning data 700 by inputting it into the depth estimation model Model X through a forward propagation process, thereby outputting from the depth estimation model (701).

[0200] The image processing device (100) can compare the depth map (e.g., first depth map) (10) for the 2D input image (110) and the 2D input image (110) included in the learning data through a backward propagation process to calculate the loss, which is the difference between the two images, and adjust the parameters of Model X so that the loss is minimized. For example, the image processing device (100) can differentiate the loss with respect to the parameters of Model X to calculate the gradient of the loss. The image processing device (100) can update the parameters of Model X in a direction in which the gradient of the calculated loss decreases, and repeat until the loss is minimized. For example, the image processing device (100) can optimize the parameters of the depth estimation model (701) within the image processing device (100) using a gradient descent algorithm. The image processing device (100) can repeatedly update the parameters of Model X through several iterations to minimize loss.

[0201] For example, the image processing device (100) can determine / predict the content type that the 2D input image (110) most likely corresponds to as a drama based on the result of the processed content type analysis. The image processing device (100) uses learning data (700) composed of a pair of a 2D input image (110) and a depth map (e.g., a first depth map) (10) for the 2D input image (110), and, through a forward propagation process and a backward propagation process, updates the parameters of the depth estimation model Model 1 corresponding to the content type 'drama' that is most likely to correspond to the 2D input image (110) among the depth estimation models (701) within the image processing device (100), thereby obtaining the depth estimation model Model Z corresponding to the 2D input image (110) in real time.

[0202] An image processing device (100) according to one embodiment can obtain, in real time, one or more corresponding depth estimation models for each of the scenes constituting a 2D input image (110) using on-device learning.

[0203] An image processing device (100) according to one embodiment updates parameters of one or more depth estimation models corresponding to content types that have the highest probability of corresponding to each of the scenes constituting the 2D input image (110), among depth estimation models for each content type, thereby obtaining one or more corresponding depth estimation models in real time for each of the scenes constituting the 2D input image (110).

[0204] Accordingly, rather than applying a single depth estimation model to a 2D input image, the image processing device (100) can dynamically apply depth estimation models to each scene according to the content type of each scene constituting the input 2D image. In particular, the image processing device (100) can reduce errors in depth estimation and improve the accuracy of depth estimation when a 2D image including various scenes with different content types is input.

[0205] For example, when a scene change (730) of a 2D input image (110) is detected, at 710, the image processing device (100) can process the content type analysis result by initializing the accumulated content type analysis result data. The image processing device (100) can initialize the accumulated content type analysis result data whenever there is a scene change (730) of the 2D input image (110). At 720, the image processing device (100) can determine / predict the content types that each of the scenes constituting the 2D input image (110) has the highest probability of corresponding to, based on the processed content type analysis result. The image processing device (100) can obtain, in real time, one or more depth estimation models corresponding to the content types that each of the scenes constituting the 2D input image (110) has the highest probability of corresponding to, by updating parameters of one or more depth estimation models corresponding to the content types that each of the scenes constituting the 2D input image (110) has the highest probability of corresponding to.

[0206] For example, it can be assumed that a 2D input image (110) is a composite image composed of four scenes (S1, S2, S3, S4), and the content types of each scene are as follows: scene S1 corresponds to a drama, scene S2 corresponds to a document, scene S3 corresponds to a drama, and scene S4 corresponds to an FPS game. The image processing device (100) can detect the scene transition S1->S2->S3->S4 of the 2D input image (110). The image processing device (100) updates the parameters of the depth estimation model Model1 corresponding to 'drama', which scene S1 most likely corresponds to, to Model We can generate / obtain and update the parameters of the depth estimation model Model 4 corresponding to the 'document' that has the highest probability of being scene S2. We can generate / obtain and update the parameters of the depth estimation model Model 3 corresponding to the 'composite content' that scene S3 most likely corresponds to. We can generate / obtain and update the parameters of the depth estimation model Model 2 corresponding to 'FPS game', which is most likely to be scene S4. The image processing device (100) can generate / acquire one or more depth estimation models corresponding to each of the scenes S1, S2, S3, and S4 constituting the 2D input image (110). , Model , Model , Model can be created / obtained in real time.

[0207] The image processing device (100) can use one depth estimation model corresponding to the acquired 2D input image (110) or one or more depth estimation models corresponding to each of the scenes constituting the 2D input image (100) to obtain a depth map (e.g., a first depth map) (10) for the 2D input image (110).

[0208] For example, the image processing device (100) can obtain a depth map (e.g., a first depth map) (10) for the 2D input image (110) using Model Z corresponding to the obtained 2D input image (110).

[0209] For example, Model corresponding to the acquired scene S1 Obtain a depth map for S1 using , and a Model corresponding to S2 Obtain a depth map for S2 using , and a Model corresponding to S3 Obtain a depth map for S3 using , and a Model corresponding to S4 By obtaining a depth map for S4 using , a depth map (e.g., first depth map) (10) for a 2D input image (110) can be obtained.

[0210] In Fig. 7, only four depth estimation models for each content type are shown, but this is only an example, and the image processing device (100) may store fewer or more depth estimation models for each content type than four. In addition, in Fig. 7, the 2D input image (110) is described as being composed of four scenes, but this is only an example, and the 2D input image (110) may be composed of numerous scenes or a single scene.

[0211] An image processing device (100) according to one embodiment of the present disclosure provides a method of dynamically applying a depth estimation model in real time using on-device AI technology according to the type of content of an input 2D image or the type of content by scene of an input 2D image, thereby reducing depth estimation errors, improving the accuracy of depth estimation, and enhancing the satisfaction and convenience of a user using a 3D display.

[0212] However, the effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned can be clearly understood by a person having ordinary skill in the technical field to which the present disclosure belongs from the description below.

[0213] FIG. 8 is a diagram for explaining a process in which an image processing device (100) according to one embodiment obtains a depth estimation model using on-device learning.

[0214] Referring to FIG. 8, at 810, an image processing device (100) according to one embodiment (e.g., a depth estimation model acquisition unit (420) of the image processing device (100)) may process a content type analysis result (510) received from a content type analysis unit (410). For example, the image processing device (100) may process the content type analysis result using an IIR (Infinite Impulse Response) filter. For content overlapping with that described in 710 of FIG. 7, refer to FIG. 7, and a description thereof will be omitted here.

[0215] At 820, the image processing device (100), according to one embodiment, can generate / obtain a depth estimation model Model Mz corresponding to a 2D input image (100) in real time by interpolating a depth estimation model (801) mounted on the image processing device (100) or a server based on a result (510) of analyzing the type of processed content.

[0216] A depth estimation model (801) for each content type can be pre-trained offline using a learning database of 2D image-depth map pairs for each pre-defined content type, and can be installed in an image processing device (100) or a cloud server.

[0217] In one embodiment, the image processing device (100) may request the cloud server to transmit depth estimation models for each content type learned offline in order to obtain an interpolated depth estimation model, and the cloud server may transmit only some of the requested depth estimation models to the image processing device (100) when there is a request from the image processing device (100). In one embodiment, the image processing device (100) may receive depth estimation models for each content type learned offline in advance from the cloud server in order to obtain an interpolated depth estimation model and store them in the image processing device (100), and may directly use some of the depth estimation models that are necessary without making a separate request to the cloud server.

[0218] For example, Model 1, Model 2, Model 3, and Model 4 learned offline can be stored in a cloud server or image processing device (100). Model 1 can be a model learned in advance using a learning database of movie or drama image-depth map pairs, Model 2 can be a model learned in advance using a learning database of FPS game image-depth map pairs, Model 3 can be a model learned in advance using a learning database of document-depth map pairs, and Model 4 can be a model learned in advance using a learning database of composite content-depth map pairs.

[0219] Interpolation between neural network models refers to the process of creating a model with intermediate performance between two or more models using linear interpolation, polynomial interpolation, or various other interpolation methods. Interpolation between neural network models is possible when the network structures and learning methods are the same between the models. The network structures of the neural network models are the same, which may mean that the number of layers constituting the neural network, the number of neurons in each layer, the activation functions, etc. are the same. The learning methods of the neural network models are the same, which may mean that the learning algorithm, hyperparameters, and preprocessing methods of the learning data are the same. In the following, it is assumed that the depth estimation models for each content type (801, e.g., Model 1, Model 2, Model 3, Model 4) have the same network structures and learning methods, and thus are in a relationship where interpolation is possible.

[0220] The image processing device (100) adjusts weight values ​​for interpolation between models based on the probability that the 2D input image (110) corresponds to each content type based on the processed content type analysis result (510), thereby interpolating parameters of the depth estimation models (801), thereby generating / obtaining a depth estimation model Model Mz corresponding to the 2D input image (100) in real time.

[0221] The image processing device (100) can adjust weight values ​​for interpolation between models based on the probability that the 2D input image (110) corresponds to each content type using a policy-based algorithm. The policy-based algorithm for defining / adjusting weight values ​​for interpolation between models can be predefined / set by the manufacturer of the image processing device (100).

[0222] For example, the image processing device (100) can generate / obtain a depth estimation model Mz corresponding to a 2D input image (110) by interpolating two models according to [Mathematical Formula 2].

[0223] [Equation 2]

[0224]

[0225] and may mean a weight value for interpolation. can mean the parameters of Model A, may mean a parameter of Model B. Model A and Model B may mean one of the depth estimation models (801) for each content type. may refer to parameters of a depth estimation model Mz corresponding to a generated / obtained 2D input image (110).

[0226] For example, the image processing device (100) can determine / predict that the 2D input image (110) has a 20% probability of being a 'drama' and an 80% probability of being an 'FPS game' based on the processed content type analysis result (510). The image processing device (100) uses a policy-based algorithm to determine a weight value for interpolation based on the 20% probability that the 2D input image (110) corresponds to a 'drama' and the 80% probability that the 2D input image (110) corresponds to an 'FPS game'. and The image processing device (100) can obtain a depth estimation model Model Mz corresponding to a 2D input image (110) in real time by interpolating Model 1 and Model 2 according to [Mathematical Formula 2] by substituting Model 1 corresponding to a 'drama' and Model 2 corresponding to an 'FPS game' into Model A and Model B, respectively.

[0227] An image processing device (100) according to one embodiment can obtain, in real time, one or more corresponding depth estimation models for each of the scenes constituting a 2D input image (110) using on-device learning.

[0228] An image processing device (100) according to one embodiment adjusts weight values ​​for interpolation between models according to the probability that each scene constituting a 2D input image (110) corresponds to each content type among depth estimation models for each content type, thereby interpolating parameters of depth estimation models (801), thereby obtaining one or more corresponding depth estimation models in real time for each of the scenes constituting the 2D input image (110).

[0229] Accordingly, rather than applying a single depth estimation model to a 2D input image, the image processing device (100) can dynamically apply depth estimation models to each scene according to the content type of each scene constituting the input 2D image. In particular, the image processing device (100) can reduce errors in depth estimation and improve the accuracy of depth estimation when a 2D image including various scenes with different content types is input.

[0230] For example, when a scene change (830) of a 2D input image (110) is detected, at 810, the image processing device (100) can process the content type analysis result by initializing the accumulated content type analysis result data. The image processing device (100) can initialize the accumulated content type analysis result data whenever there is a scene change (830) of the 2D input image (110). At 820, the image processing device (100) can determine / predict the probability that each scene constituting the 2D input image (110) corresponds to each content type based on the processed content type analysis result (510). The image processing device (100) adjusts weight values ​​for interpolation between models according to the probability that each scene constituting the 2D input image (110) corresponds to each content type, thereby interpolating parameters of depth estimation models (801), thereby obtaining one or more corresponding depth estimation models in real time for each of the scenes constituting the 2D input image (110).

[0231] For example, it can be assumed that a 2D input image (110) is a composite image composed of four scenes (S1, S2, S3, S4), and the content types of each scene are that scene S1 corresponds to a drama, scene S2 corresponds to a document, scene S3 corresponds to a drama, and scene S4 corresponds to an FPS game. The image processing device (100) can detect the scene transition S1->S2->S3->S4 of the 2D input image (110). The image processing device (100) adjusts the weight values ​​for interpolation of Model 1, Model 2, Model 3, and Model 4 based on the probability that scene S1 corresponds to a 'drama', the probability that scene S1 corresponds to an 'FPS game', the probability that scene S1 corresponds to a 'composite content', and the probability that scene S1 corresponds to a 'document', thereby interpolating Model 1, Model 2, Model 3, and Model 4. In a similar manner, the image processing device (100) can generate / obtain one or more depth estimation models Model corresponding to each of the scenes S1, S2, S3, and S4 constituting the 2D input image (110). , Model , Model , Model can be created / obtained in real time.

[0232] The image processing device (100) can use one depth estimation model corresponding to the acquired 2D input image (110) or one or more depth estimation models corresponding to each of the scenes constituting the 2D input image (100) to obtain a depth map (e.g., a first depth map) (10) for the 2D input image (110).

[0233] For example, the image processing device (100) can obtain a depth map (e.g., a first depth map) (10) for the 2D input image (110) by using Model Mz corresponding to the acquired 2D input image (110).

[0234] For example, Model corresponding to the acquired scene S1 Obtain a depth map for S1 using , and a Model corresponding to S2 Obtain a depth map for S2 using , and a Model corresponding to S3 Obtain a depth map for S3 using , and a Model corresponding to S4 By obtaining a depth map for S4 using , a depth map (e.g., first depth map) (10) for a 2D input image (110) can be obtained.

[0235] In FIG. 8, the image processing device (100) is illustrated as interpolating two depth estimation models, but this is only an example, and the image processing device (100) can interpolate two or more depth estimation models in a similar manner.

[0236] In Fig. 8, only four depth estimation models for each content type are shown, but this is only an example, and the image processing device (100) may store fewer or more depth estimation models for each content type than four. In addition, in Fig. 8, the 2D input image (110) is described as being composed of four scenes, but this is only an example, and the 2D input image (110) may be composed of numerous scenes or a single scene.

[0237] An image processing device (100) according to one embodiment of the present disclosure provides a method of dynamically applying a depth estimation model in real time using on-device AI technology according to the type of content of an input 2D image or the type of content by scene of an input 2D image, thereby reducing depth estimation errors, improving the accuracy of depth estimation, and enhancing the satisfaction and convenience of a user using a 3D display.

[0238] However, the effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned can be clearly understood by a person having ordinary skill in the technical field to which the present disclosure belongs from the description below.

[0239] FIG. 9 is a flowchart for explaining a method (900) in which an image processing device (100) according to one embodiment performs 3D conversion from a 2D input image (110).

[0240] Referring to FIG. 9, in step 910, an image processing device (100) according to one embodiment may analyze the content type of an input image (110). For example, the input image (110) may be a 2D image. For example, the content type may include, but is not limited to, movies, dramas, FPS games, RPG games, RTS games, MMORPG games, documents, complex content, presentation materials, etc.

[0241] An image processing device (100) according to one embodiment can analyze the content type of an input image (110) using a content type analysis model (500). The content type analysis model (500) may refer to an artificial neural network model learned / trained to predict the content type of an input image (110).

[0242] An image processing device (100) according to one embodiment can analyze the content type of an input image (110) using a policy-based algorithm. The policy-based algorithm for analyzing the content type of an input image can be predefined / set by the manufacturer of the image processing device (100).

[0243] According to one embodiment, the image processing device (100) can analyze the content type of an input image (110) by combining a content type analysis model (500) and a policy-based algorithm.

[0244] According to one embodiment, the content type analysis result (510) may include information related to the probability that the input image (110) corresponds to each of the predefined content types. For example, the content type analysis result (510) may include information indicating a specific content type among the predefined content types to which the input image (110) corresponds with the highest probability. For example, the content type analysis result (510) may include probability values ​​indicating the probability that the input image (110) corresponds to each of the predefined content types.

[0245] In step 920, the image processing device (100) according to one embodiment can obtain a depth estimation model corresponding to the content type of the input image (110) in real time based on the result (510) of analyzing the content type of the input image (110).

[0246] A depth estimation model may refer to an artificial neural network model learned / trained to predict depth information of each pixel constituting a 2D image (110).

[0247] An image processing device (100) according to one embodiment can obtain a depth estimation model corresponding to the content type of an input image (110) in real time through a cloud server (600) based on the result (510) of analyzing the content type of an input image (110).

[0248] An image processing device (100) according to one embodiment can obtain a depth estimation model corresponding to the content type of an input image (110) in real time using on-device learning based on the result (510) of analyzing the content type of an input image (110).

[0249] An image processing device (100) according to one embodiment can obtain a depth estimation model corresponding to an input image (110) in real time by updating and learning (720) the parameters of a depth estimation model installed in the image processing device (100) or a server based on the result (510) of analyzing the content type of an input image (110).

[0250] An image processing device (100) according to one embodiment can obtain a depth estimation model corresponding to an input image (110) in real time by interpolating (820) depth estimation models installed in the image processing device (100) or a server based on a result (510) of analyzing the content type of an input image (110).

[0251] An image processing device (100) according to one embodiment can obtain depth estimation models corresponding to each scene constituting an input image (110) in real time based on a result (510) of analyzing the content type of an input image (110).

[0252] In step 930, the image processing device (100) according to one embodiment can obtain a depth map (e.g., a first depth map) (10) for the input image (110) based on a depth estimation model according to on-device learning. The depth map (10) reflects depth information estimated by the depth estimation model. The image processing device (100) can input the input image (110) as input data to the depth estimation model obtained in step 920, and output the depth map for the input image (110) as output data. The depth map may mean a 2D image in which depth information of each pixel constituting the image is expressed as a value such as brightness or color of the pixel.

[0253] An image processing device (100) according to one embodiment can obtain a depth map (e.g., a first depth map) (10) for an input image (110) in real time by applying depth estimation models corresponding to each scene constituting the input image (110) for each scene constituting the input image (110).

[0254] The depth map initially acquired by the image processing device (100) for the input image (110), i.e., the depth map before non-linearly changing the depth map, may be referred to as a 'depth map' or a 'first depth map'.

[0255] In step 940, an image processing device (100) according to one embodiment can perform 3D transformation from an input image (110) based on a depth map (e.g., a first depth map) (10) for the input image (110).

[0256] An image processing device (100) according to one embodiment can perform 3D transformation from an input image (110) by performing a new viewpoint creation process (370), a hole filling process (380), and a pixel mapping process (390) based on a depth map (e.g., a first depth map) (10) for an input image (110).

[0257] In one embodiment, the image processing device (100) can obtain a 3D converted output image (120) from an input image (110). The image processing device (100) can control a display to output the output image (120).

[0258] In one embodiment, the image processing device (100) can control the display to generate and output a UI that indicates the content type and the three-dimensionality of each scene corresponding to each scene that constitutes the 3D-converted output image (120) from the input image (110). The user can determine the content type and the three-dimensionality level of each scene of the 3D output image (120) being viewed through the UI provided by the image processing device (100).

[0259] According to one embodiment of the present disclosure, the image processing device (100) can further improve the accuracy of depth estimation by reducing depth estimation errors by dynamically applying a depth estimation model in real time based on analysis of the content type of an input 2D image, and can improve the satisfaction of a user using a 3D display.

[0260] However, the effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned can be clearly understood by a person having ordinary skill in the technical field to which the present disclosure belongs from the description below.

[0261] Fig. 10 is an internal block diagram of an image processing device (100) according to one embodiment.

[0262] Referring to FIG. 10, an image processing device (100) according to one embodiment may include a content type analysis unit (1010), a depth estimation model acquisition unit (1020), a depth estimation unit (1030), a scene object analysis unit (1040), a depth map dynamic change unit (1050), a three-dimensional effect control unit (1060), and a 3D conversion performing unit (1070).

[0263] The content type analysis unit (1010), the depth estimation model acquisition unit (1020), the depth estimation unit (1030), the scene object analysis unit (1040), the depth map dynamic change unit (1050), the three-dimensional effect control unit (1060), and the 3D transformation execution unit (1070) may be implemented with at least one processor. The content type analysis unit (1010), the depth estimation model acquisition unit (1020), the depth estimation unit (1030), the scene object analysis unit (1040), the depth map dynamic change unit (1050), the three-dimensional effect control unit (1060), and the 3D transformation execution unit (1070) may operate according to at least one instruction stored in a memory (e.g., 102 of FIG. 18).

[0264] FIG. 10 illustrates the content type analysis unit (1010), the depth estimation model acquisition unit (1020), the depth estimation unit (1030), the scene object analysis unit (1040), the depth map dynamic change unit (1050), the three-dimensional effect control unit (1060), and the 3D conversion execution unit (1070) individually, but the content type analysis unit (1010), the depth estimation model acquisition unit (1020), the depth estimation unit (1030), the scene object analysis unit (1040), the depth map dynamic change unit (1050), the three-dimensional effect control unit (1060), and the 3D conversion execution unit (1070) may be implemented through one processor. In this case, the content type analysis unit (1010), depth estimation model acquisition unit (1020), depth estimation unit (1030), scene object analysis unit (1040), depth map dynamic change unit (1050), three-dimensional effect control unit (1060), and 3D transformation performing unit (1070) may be implemented as a dedicated processor, or may be implemented through a combination of a general-purpose processor such as an application processor (AP), a central processing unit (CPU), or a graphic processing unit (GPU) and software. In addition, in the case of a dedicated processor, it may include a memory for implementing an embodiment of the present disclosure, or a memory processing unit for utilizing an external memory. In addition, in the case of an artificial intelligence-dedicated processor such as an NPU (Neural Processing Unit), it may be designed as a hardware structure specialized for processing a specific artificial intelligence model.

[0265] The content type analysis unit (1010), the depth estimation model acquisition unit (1020), the depth estimation unit (1030), the scene object analysis unit (1040), the depth map dynamic change unit (1050), the three-dimensional effect control unit (1060), and the 3D conversion execution unit (1070) may be configured with a plurality of processors. In this case, the content type analysis unit (1010), the depth estimation model acquisition unit (1020), the depth estimation unit (1030), the scene object analysis unit (1040), the depth map dynamic change unit (1050), the three-dimensional effect control unit (1060), and the 3D conversion execution unit (1070) may be implemented by a combination of dedicated processors, or may be implemented by a combination of a plurality of general-purpose processors such as an AP, a CPU, or a GPU, and software.

[0266] In one embodiment, the content type analysis unit (1010), the depth estimation model acquisition unit (1020), and the depth estimation unit (1030) may operate similarly to the content type analysis unit (410), the depth estimation model acquisition unit (420), and the depth estimation unit (430) of FIG. 4, respectively. Any overlapping content described in FIG. 4 is referred to FIG. 4, and description thereof will be omitted here.

[0267] In one embodiment, the scene object analysis unit (1040) can receive a depth map (e.g., a first depth map) (10) for a 2D input image (110) from the depth estimation unit (1030).

[0268] In one embodiment, the scene object analysis unit (1040) may analyze the size and distribution of objects included in each scene for each scene constituting the 2D input image (110). In one embodiment, the scene object analysis unit (1040) may analyze the size and distribution of objects included in each scene for each scene constituting the 2D input image (110) based on the 2D input image (110) and a depth map (e.g., a first depth map) (10) for the 2D input image (110).

[0269] In one embodiment, the scene object analysis unit (1040) may include suitable logic, circuitry, interfaces and / or code operable to analyze the size and distribution of objects included in each scene constituting the 2D input image (110).

[0270] An 'object' may refer to an independently identifiable subject or object within each scene constituting a 2D input image (110), and may refer to an element having visually distinct characteristics. For example, it may refer to a specific object, animal, person, etc. included in each scene constituting a 2D input image (110).

[0271] In one embodiment, the scene object analysis unit (1040) may analyze the size and distribution of objects included in each scene constituting the 2D input image (110) by using at least one of an artificial neural network (e.g., the scene object analysis model (1100) of FIG. 11) or a policy-based algorithm that is learned / trained to analyze the size and distribution of objects included in each scene constituting the 2D input image (110).

[0272] In one embodiment, the scene object analysis unit (1040) can obtain information related to size characteristics of each object included in each scene and distribution characteristics of each object included in each scene as a result of size and distribution analysis of objects included in each scene.

[0273] In one embodiment, the scene object analysis unit (1040) can transmit the results of analyzing the size and distribution of objects included in each scene constituting the 2D input image (110) to the depth map dynamic change unit (1050) and the three-dimensional effect control unit (1060).

[0274] In one embodiment, the depth map dynamic change unit (1050) may receive, from the scene object analysis unit (1040), the results of analyzing the sizes and distributions of objects included in each scene constituting the 2D input image (110). The depth map dynamic change unit (1050) may receive, from the content type analysis unit (1010), the results of analyzing the content type of the 2D input image (110) (510). The depth map dynamic change unit (1050) may receive, from the depth estimation unit (1030), a depth map (e.g., a first depth map) (10) for the 2D input image (110).

[0275] In one embodiment, the depth map dynamic change unit (1050) may obtain additional information including at least one of metadata information about the 2D input image (110), information about the viewing environment, or user setting information related to the 2D input image (110).

[0276] In one embodiment, the depth map dynamic change unit (1050) can obtain a modified depth map (e.g., second depth map) (20) by non-linearly changing the depth map (e.g., first depth map) (10) based on at least one of the results of analyzing the size and distribution of objects included in each scene, the results of analyzing the content type of the 2D input image (110) (510), or the acquired additional information.

[0277] In one embodiment, the depth map dynamic modification unit (1050) may include suitable logic, circuitry, interfaces and / or code operable to obtain a modified depth map (e.g., a second depth map) (20) by non-linearly modifying a depth map (e.g., a first depth map) (10).

[0278] In one embodiment, the depth map dynamic change unit (1050) can obtain a modified depth map (e.g., a second depth map) (20) by non-linearly changing the depth map (e.g., a first depth map) (10) using at least one of an artificial neural network (e.g., a depth map dynamic change model (1200) of FIG. 12A) learned / trained to non-linearly change the input depth map constituting the 2D input image (110), a policy-based algorithm, or a LUT (Lookup Table) method.

[0279] In one embodiment, the depth map dynamic change unit (1050) can transmit the modified depth map (e.g., second depth map) (20) to the 3D transformation performing unit (1070). The depth map dynamic change unit (1050) can transmit the modified depth map (e.g., second depth map) (20) to the three-dimensional effect control unit (1060).

[0280] In one embodiment, the stereoscopic effect control unit (1060) may receive the results of analyzing the sizes and distributions of objects included in each scene constituting the 2D input image (110) from the scene object analysis unit (1040). The stereoscopic effect control unit (1060) may receive the results of analyzing the content type of the 2D input image (110) (510) from the content type analysis unit (1010). The stereoscopic effect control unit (1060) may receive a depth map (e.g., a first depth map) (10) for the 2D input image (110) from the depth estimation unit (1030). The stereoscopic effect control unit (1060) may receive a modified depth map (e.g., a second depth map) (20) for the 2D input image (110) from the depth map dynamic change unit (1050).

[0281] In one embodiment, the stereoscopic effect control unit (1060) may obtain additional information including at least one of metadata information regarding the 2D input image (110), information regarding the viewing environment, or user setting information related to the 2D input image (110).

[0282] In one embodiment, the stereoscopic effect control unit (1060) can determine the relative positions of objects included in each scene constituting the 2D input image (110) from a virtual convergence plane corresponding to the screen based on at least one of the results of analyzing the size and distribution of objects included in each scene, the results of analyzing the content type of the 2D input image (110) (510), or the acquired additional information.

[0283] In one embodiment, the stereoscopic effect control unit (1060) may include suitable logic, circuitry, interfaces and / or code operable to determine the relative positions of objects included in each scene constituting the 2D input image (110) from a virtual convergence plane corresponding to the screen.

[0284] In one embodiment, the stereoscopic effect control unit (1060) may determine the relative positions of objects included in each scene constituting the 2D input image (110) from a virtual convergence plane corresponding to the screen by using at least one of an artificial neural network (e.g., the stereoscopic effect control model (1300) of FIG. 14) learned / trained to determine the relative positions of objects included in each scene constituting the 2D input image (110) from a virtual convergence plane corresponding to the screen, a policy-based algorithm, or a LUT (Lookup Table) method.

[0285] In one embodiment, the stereoscopic effect control unit (1060) can transmit information about the relative positions of objects included in each scene with respect to the convergence plane to the 3D transformation performing unit (1070).

[0286] In one embodiment, the 3D transformation performing unit (1070) may receive a modified depth map (e.g., second depth map) (20) for the 2D input image (110) from the depth map dynamic change unit (1050). In one embodiment, the 3D transformation performing unit (1070) may receive information about the relative positions of objects included in each scene with respect to the convergence plane from the stereoscopic effect control unit (1060).

[0287] In one embodiment, the 3D transformation performing unit (1070) may perform 3D transformation from a 2D input image (110) based on a modified depth map (e.g., a second depth map) (20) to obtain a 3D output image (120). In one embodiment, the 3D transformation performing unit (1070) may operate similarly to the 3D transformation performing unit (440) of FIG. 4. Any content overlapping with that described in FIG. 4 will be referred to FIG. 4, and its description will be omitted here.

[0288] In one embodiment, the 3D transformation performing unit (1070) may perform 3D transformation from the 2D input image (110) based on a modified depth map (e.g., second depth map) (20) for the 2D input image (110) and information about the relative positions of objects included in each scene with respect to the convergence plane. The 3D transformation performing unit (1070) may perform 3D transformation from the 2D input image (110) by performing a new viewpoint view generation process based on information about the relative positions of objects included in each scene with respect to the convergence plane.

[0289] According to one embodiment of the present disclosure, the image processing device (100) nonlinearly changes a depth map and controls a three-dimensional effect based on analysis of the content type of an input 2D image and analysis of objects included in each scene, thereby reducing depth estimation errors, further improving the accuracy of depth estimation, and enhancing the satisfaction of a user using a 3D display.

[0290] However, the effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned can be clearly understood by a person having ordinary skill in the technical field to which the present disclosure belongs from the description below.

[0291] FIG. 11 is a drawing for explaining a process in which an image processing device (100) according to one embodiment analyzes a scene object of a 2D input image (110).

[0292] Referring to FIG. 11, an image processing device (100) according to one embodiment (e.g., a scene object analysis unit (1040) of the image processing device (100)) can analyze the size and distribution of objects included in each scene for each scene constituting a 2D input image (110).

[0293] The image processing device (100) can analyze the size and distribution of objects included in each scene, for each scene constituting the 2D input image (110), based on the 2D input image (110) and the depth map (e.g., the first depth map) (10) for the 2D input image (110). The image processing device (100) can analyze the size and distribution of objects included in each scene, for each scene constituting the 2D input image (110), and obtain the size and distribution analysis result (1110) of the objects included in each scene. The 'size and distribution analysis result (1110) of objects included in each scene' can be referred to as 'information related to the size characteristics of each object included in each scene and the distribution characteristics of each object included in each scene.'

[0294] The image processing device (100) can measure the size of objects by calculating the area of ​​pixels occupied by each object included in each scene constituting the 2D input image (110), and analyze the size of objects included in each scene constituting the 2D input image (110) by comparing the relative sizes between objects included in one scene.

[0295] The image processing device (100) can analyze how each object included in each scene constituting the 2D input image (110) is distributed in space based on a depth map (e.g., first depth map) (10) for the 2D input image (110).

[0296] For example, in 1120, it can be assumed that the 2D input image (110) is composed of scenes A to F. The image processing device (100) can analyze, based on the 2D input image (110) and the depth map (e.g., the first depth map) (10), whether the sizes of the objects included in scenes A to F are large or small and whether the distribution of the objects is near or far. However, 1120 is only an example, and the image processing device (100) can obtain the size and distribution analysis results (1110) of the objects included in each scene as numerical expressions that can represent the sizes and distributions of the objects included in each scene constituting the 2D input image (110).

[0297] According to one embodiment, the image processing device (100) can analyze the size and distribution of objects included in each scene for each scene constituting the 2D input image (110) using a scene object analysis model (1100).

[0298] The scene object analysis model (1100) may refer to an artificial neural network model learned / trained to analyze the size and distribution of objects included in each scene constituting the 2D input image (110). For example, the scene object analysis model (1100) may refer to an artificial neural network model learned / trained to analyze the size and distribution of objects included in each scene constituting the 2D input image (110) using a technology such as CNN, DNN, RNN, RBM, DBN, BRDNN, or deep Q-network, HOG, SHIFT, LSTM, SVM, SoftMax, etc., but is not limited thereto.

[0299] The scene object analysis model (1100) can receive a 2D input image (110) and a depth map (e.g., a first depth map) (10) for the 2D input image as input data, and output the size and distribution analysis results (1110) of objects included in each scene as output data. The image processing device (100) can obtain the learned scene object analysis model (1100) from a cloud server. The image processing device (100) can learn / generate / obtain the scene object analysis model (1100) using on-device learning.

[0300] According to one embodiment, the image processing device (100) may analyze the size and distribution of objects included in each scene constituting the 2D input image (110) using a policy-based algorithm. The policy-based algorithm for analyzing the size and distribution of objects included in each scene constituting the 2D input image (110) may be predefined / set by the manufacturer of the image processing device (100).

[0301] According to one embodiment, the image processing device (100) can analyze the size and distribution of objects included in each scene constituting the 2D input image (110) by combining a scene object analysis model (1100) and a policy-based algorithm.

[0302] According to one embodiment of the present disclosure, the image processing device (100) nonlinearly changes a depth map and controls a three-dimensional effect based on analysis of the size and distribution of objects included in each scene constituting an input 2D image, thereby reducing depth estimation errors, further improving the accuracy of depth estimation, and enhancing the satisfaction of a user using a 3D display.

[0303] However, the effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned can be clearly understood by a person having ordinary skill in the technical field to which the present disclosure belongs from the description below.

[0304] FIGS. 12A to 12C are drawings for explaining a process in which an image processing device (100) according to one embodiment dynamically changes a depth map.

[0305] Referring to FIG. 12A, an image processing device (100) according to one embodiment (e.g., a depth map dynamic change unit (1050) of the image processing device (100)) can obtain a modified depth map (e.g., a second depth map) (20) by non-linearly changing a depth map (e.g., a first depth map) (10).

[0306] An image processing device (100) can obtain a modified depth map (e.g., a second depth map) (20) by nonlinearly changing the depth map (e.g., a first depth map) (10) for a 2D input image (110), a result of analyzing the size and distribution of objects included in each scene constituting the 2D input image (110) (1110), a result of analyzing the type of content of the 2D input image (110) (510), or additional information (1201).

[0307] In one embodiment, the image processing device (100) can obtain additional information (1201). The additional information (1201) can include at least one of metadata information associated with the 2D input image (110), information regarding a viewing environment, or user setting information.

[0308] Metadata information associated with a 2D input image (110) may refer to information indicating various details related to the 2D input image (110). For example, metadata information associated with a 2D input image (110) may include information about the title, description, content type and subject of the image, information about the date and time the image was created, information about the location where the image was shot, information about the producer, editor, copyright holder and usage rights of the image, information about the resolution, bitrate and codec of the image, information about the color gamut and file format of the image, the total number of frames or scenes constituting the image, and information about characters appearing in the image.

[0309] Information about the viewing environment may refer to information about the surrounding environment of a user viewing a 3D output image (120). For example, information about the viewing environment may include the luminance of the surrounding environment of the user measured by a sensor in the image processing device (100), or the screen brightness of the display on which the 3D image (120) is displayed.

[0310] The user setting information may refer to information reflecting the user's preference regarding the stereoscopic effect of the 2D input image (110), i.e., the degree of stereoscopic effect with which the user wishes to view the 2D input image (110). For example, the user setting information related to the 2D input image (110) may include information regarding the intensity of the stereoscopic effect manually adjusted by the user through UI 250A and UI 250B of FIG. 2C.

[0311] In one embodiment, the image processing device (100) may obtain additional information (1201) from another device existing outside the image processing device (100), or may obtain additional information (1201) through a process inside the image processing device (100).

[0312] In one embodiment, the image processing device (100) can obtain a modified depth map (e.g., a second depth map) (20) by non-linearly changing the depth map (e.g., a first depth map) (10) according to 1210A to 1230A.

[0313] In 1210A, an image processing device (100) according to one embodiment may generate an offset for each object included in each scene based on the result (1110) of analyzing the size and distribution of objects included in each scene constituting a 2D input image (110). The offset may mean a deviation in depth information of pixels constituting each object included in a depth map (e.g., a first depth map) (10), and may include a positive deviation or a negative deviation.

[0314] In 1220A, an image processing device (100) according to one embodiment can non-linearly change a depth map (eg, first depth map) (10) by adjusting depth information included in a depth map (eg, first depth map) (10) for a 2D input image (110) by an offset generated for each object included in each scene.

[0315] In 1230A, the image processing device (100) according to one embodiment can control the depth map output range based on at least one of the result (510) of analyzing the content type of the 2D input image (110) or the acquired additional information (1201). The output range of the depth map can mean a range from the minimum to the maximum of the depth information of pixels included in the depth map. In one embodiment, the image processing device (100) can control the depth map output range by giving priority to the information included in the additional information (1201) over the result (510) of analyzing the content type of the 2D input image (110).

[0316] For example, referring to FIG. 12B, the input depth map may be a depth map (e.g., a first depth map) (10), and the output depth map may be a modified depth map (e.g., a second depth map) (20). The image processing device (100) may obtain the output depth map by non-linearly modifying the input depth map through steps 1210B to 1230B.

[0317] For example, in 1210B, the image processing device (100) can generate a positive offset for some objects included in a specific scene constituting the 2D input image (110) and generate a negative offset for some other objects, based on the result (1110) of analyzing the size and distribution of objects included in each scene constituting the 2D input image (110).

[0318] For example, in 1220B, the image processing device (100) can adjust the depth information of the pixels constituting the corresponding objects included in the input depth map by the offset generated in 1210B. Adjusting the depth information by a positive offset may mean adjusting the object to appear more protruding forward, and adjusting the depth information by a negative offset may mean adjusting the object to appear more receded.

[0319] For example, in 1230B, the image processing device (100) can obtain, as additional information (1201), information about the viewing environment indicating that the user's surroundings are a dark indoor environment with a luminance close to 0 lux. In this case, the image processing device (100) can control the overall output range of the depth map reflecting the depth information adjusted in 1220B to be 0.5 times larger in order to reduce user fatigue.

[0320] For example, referring to FIG. 12C, at 1221, the image processing device (100) can nonlinearly change the input depth map by applying a positive deviation to the depth information of the pixels constituting the human object and applying a negative deviation to the depth information of the pixels constituting the background object. At 1222, the image processing device (100) can determine that the content type is text based on the content type analysis result (510) or the additional information (1201) and control the overall output range of the depth map to 0. At 1223, the image processing device (100) can nonlinearly change the input depth map by applying a positive deviation to the depth information of the pixels constituting the building object and applying a negative deviation to the depth information of the pixels constituting the background object.

[0321] Referring again to FIG. 12A, according to one embodiment, the image processing device (100) can obtain a modified depth map (e.g., a second depth map) (20) by non-linearly modifying a depth map (e.g., a first depth map) (10) using a depth map dynamic modification model (1200).

[0322] The depth map dynamic change model (1200) may refer to an artificial neural network model learned / trained to nonlinearly change a depth map (e.g., a first depth map) (10). For example, the depth map dynamic change model (1200) may refer to an artificial neural network model learned / trained to nonlinearly change a depth map (e.g., a first depth map) (10) using a technology such as CNN, DNN, RNN, RBM, DBN, BRDNN, or deep Q-network, HOG, SHIFT, LSTM, SVM, SoftMax, etc., but is not limited thereto.

[0323] The depth map dynamic change model (1200) can receive a depth map (e.g., a first depth map) (10) for a 2D input image as input data, and output a modified depth map (e.g., a second depth map) (20) for the 2D input image as output data. The image processing device (100) can obtain the learned depth map dynamic change model (1200) from a cloud server. The image processing device (100) can learn / generate / obtain the depth map dynamic change model (1200) using on-device learning.

[0324] According to one embodiment, the image processing device (100) can obtain a modified depth map (e.g., a second depth map) (20) by nonlinearly modifying a depth map (e.g., a first depth map) (10) using a policy-based algorithm. The policy-based algorithm for nonlinearly modifying a depth map (e.g., a first depth map) (10) can be predefined / set by the manufacturer of the image processing device (100).

[0325] According to one embodiment, the image processing device (100) can obtain a modified depth map (e.g., a second depth map) (20) by non-linearly modifying the depth map (e.g., a first depth map) (10) using a LUT (Lookup Table) method. A specific function or mapping table for the LUT method may be pre-defined / set by the manufacturer of the image processing device (100). For example, the image processing device (100) can use a mapping table that collectively adjusts depth information of pixels corresponding to a specific range in the depth map (e.g., a first depth map) (10) by a pre-defined offset. For example, in the case of a scene having a specific object distribution, the image processing device (100) can use a mapping table that collectively adjusts depth information of pixels constituting objects of the scene by a pre-defined offset.

[0326] According to one embodiment, the image processing device (100) can obtain a modified depth map (e.g., a second depth map) (20) by non-linearly modifying a depth map (e.g., a first depth map) (10) by combining two or more of a depth map dynamic change model (1200), a policy-based algorithm, and a LUT method.

[0327] According to one embodiment of the present disclosure, the image processing device (100) can further improve the accuracy of depth estimation by reducing depth estimation errors by nonlinearly changing the depth map, and improve the satisfaction of a user using a 3D display.

[0328] However, the effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned can be clearly understood by a person having ordinary skill in the technical field to which the present disclosure belongs from the description below.

[0329] FIG. 13 and FIG. 14 are drawings for explaining a process in which an image processing device (100) according to one embodiment controls a three-dimensional effect.

[0330] The operation of adjusting the 'depth', which is one of the main factors affecting the three-dimensional effect, is described in FIGS. 12A to 12C. FIGS. 13 to 14 describe the operation of adjusting the 'degree of projection and receding relative to the screen', which is one of the other main factors affecting the three-dimensional effect. Pop-up (negative disparity) may mean that an object is protruding forward from the screen, and receding (behind screen, positive disparity) may mean that an object is behind the screen. The reference position (in-focus, zero disparity) may mean that an object is on a virtual convergence plane corresponding to the screen.

[0331] Referring to FIG. 13, an image processing device (100) according to one embodiment (e.g., a three-dimensional effect control unit (1060) of the image processing device (100)) can determine the relative positions of objects included in each scene constituting a 2D input image (110) from a virtual convergence plane corresponding to a screen.

[0332] The image processing device (100) can determine the position of a virtual convergence plane corresponding to the screen based on at least one of the results of analyzing the size and distribution of objects included in each scene (1110), the results of analyzing the content type of the 2D input image (110) (510), or the acquired additional information (1201) and at least one of the depth map (10) or the modified depth map (20) for the 2D input image (110). As the position of the virtual convergence plane is determined, the image processing device (100) can determine the relative positions of objects included in each scene constituting the 2D input image (110) from the convergence plane. The relative positions of objects included in each scene from the convergence plane can be determined as one of a forward position, a reference position, or a backward position.

[0333] For example, the image processing device (100) may determine the relative position of the object from the convergence plane as a forward position if the object is large, the object distribution is near-field, and the near-field and far-field are clearly distinguished based on the analysis result (1110) of the size and distribution of the objects included in each scene. For example, the image processing device (100) may determine the relative position of the object from the convergence plane as a backward position if the object is small, the object distribution is far-field, based on the analysis result (1110) of the size and distribution of the objects included in each scene. For example, the image processing device (100) may determine the relative position of the object from the convergence plane as a reference position if the content type of the 2D input image (110) is a document based on the analysis result (510) of the content type of the 2D input image (110) or the acquired additional information (1201).

[0334] For example, referring to FIG. 14, 1410 to 1460 illustrate examples of determining relative positions of objects included in each scene constituting the 2D input image (110) from a convergence plane based on at least one of the results of analyzing the size and distribution of objects included in each scene (1110), the results of analyzing the content type of the 2D input image (110) (510), or the acquired additional information (1201) and at least one of the depth map (10) or the modified depth map (20) for the 2D input image (110).

[0335] Referring again to FIG. 13, according to one embodiment, the image processing device (100) can use a three-dimensional effect control model (1300) to determine the relative positions of objects included in each scene constituting the 2D input image (110) from a virtual convergence plane corresponding to the screen.

[0336] The three-dimensional effect control model (1300) may refer to an artificial neural network model learned / trained to determine the relative positions of objects included in each scene constituting the 2D input image (110) from a virtual convergence plane corresponding to the screen. For example, the three-dimensional effect control model (1300) may refer to an artificial neural network model learned / trained to determine the relative positions of objects included in each scene constituting the 2D input image (110) from a virtual convergence plane corresponding to the screen using a technology such as CNN, DNN, RNN, RBM, DBN, BRDNN, or deep Q-network, HOG, SHIFT, LSTM, SVM, SoftMax, etc., but is not limited thereto.

[0337] The three-dimensional effect control model (1300) can receive as input data at least one of the analysis result (1110) of the size and distribution of objects included in each scene, the analysis result (510) of the content type of the 2D input image (110), or the acquired additional information (1201) and at least one of the depth map (10) or the modified depth map (20) for the 2D input image (110), and output as output data information on the relative positions of objects included in each scene with respect to the convergence plane. The image processing device (100) can acquire the learned three-dimensional effect control model (1300) from a cloud server. The image processing device (100) can learn / generate / acquire the three-dimensional effect control model (1300) using on-device learning.

[0338] According to one embodiment, the image processing device (100) may determine the relative positions of objects included in each scene constituting the 2D input image (110) from a virtual convergence plane corresponding to the screen using a policy-based algorithm. The policy-based algorithm for determining the relative positions of objects included in each scene constituting the 2D input image (110) from a virtual convergence plane corresponding to the screen may be predefined / set by the manufacturer of the image processing device (100).

[0339] According to one embodiment, the image processing device (100) can determine the relative positions of objects included in each scene constituting the 2D input image (110) from a virtual convergence plane corresponding to the screen using a LUT (Lookup Table) method. A specific function or mapping table for the LUT method can be predefined / set by the manufacturer of the image processing device (100). For example, in the case of a scene having a specific type of content, the image processing device (100) can use a mapping table that determines the relative positions of objects of the scene from the virtual convergence plane as predefined positions.

[0340] According to one embodiment, the image processing device (100) can determine the relative positions of objects included in each scene constituting the 2D input image (110) from a virtual convergence plane corresponding to the screen by combining two or more of a stereoscopic effect control model (1300), a policy-based algorithm, and a LUT method.

[0341] According to one embodiment of the present disclosure, the image processing device (100) can further improve the accuracy of depth estimation by reducing depth estimation errors and enhance the satisfaction of a user using a 3D display by determining whether objects included in each scene constituting a 2D input image are to be advanced or retreated relative to the screen.

[0342] However, the effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned can be clearly understood by a person having ordinary skill in the technical field to which the present disclosure belongs from the description below.

[0343] FIG. 15 is a flowchart for explaining a method (1500) for performing 3D conversion from a 2D input image (110) by an image processing device (100) according to one embodiment.

[0344] Referring to FIG. 15, in step 1510, an image processing device (100) according to one embodiment may analyze the content type of an input image (110). Step 1510 may operate similarly to step 910 of FIG. 9. Any content overlapping with that described in FIG. 9 will be referred to FIG. 9, and its description will be omitted here.

[0345] In step 1520, the image processing device (100) according to one embodiment can obtain a depth estimation model corresponding to the input image (110) in real time based on the result (510) of analyzing the content type of the input image (110).

[0346] The image processing device (100) can acquire a depth estimation model corresponding to the input image (110) in real time from a cloud server based on the result (510) of analyzing the content type of the input image (110). The image processing device (100) can learn / generate / acquire one or more depth estimation models corresponding to the input image (110) in real time using on-device learning based on the result of analyzing the content type of the input image (110).

[0347] When the image processing device (100) uses on-device learning, step 1520 may operate similarly to step 920 of FIG. 9. Any content overlapping with that described in FIG. 9 is referred to FIG. 9, and description thereof is omitted here.

[0348] In step 1530, the image processing device (100) according to one embodiment may obtain a depth map (10) for the input image (110) based on the depth estimation model. Step 1530 may operate similarly to step 930 of FIG. 9. Any content overlapping with that described in FIG. 9 may be referred to FIG. 9, and a description thereof will be omitted here.

[0349] In step 1540, an image processing device (100) according to one embodiment can analyze the size and distribution of objects included in each scene constituting the input image (110) based on the input image (110) and the depth map (10) for the input image.

[0350] An image processing device (100) according to one embodiment can measure the size of objects by calculating the area of ​​pixels occupied by each object included in each scene constituting an input image (110), and analyze the size of objects included in each scene constituting a 2D input image (110) by comparing the relative sizes between objects included in one scene.

[0351] An image processing device (100) according to one embodiment can analyze how each object included in each scene constituting the input image (110) is distributed in space based on a depth map (10) for the input image (110).

[0352] According to one embodiment, the image processing device (100) can analyze the size and distribution of objects included in each scene constituting the input image (110) using a scene object analysis model (1100).

[0353] The scene object analysis model (1100) may refer to an artificial neural network model learned / trained to analyze the size and distribution of objects included in each scene constituting the input image (110).

[0354] An image processing device (100) according to one embodiment can analyze the size and distribution of objects included in each scene constituting an input image (110) using a policy-based algorithm. The policy-based algorithm for analyzing the size and distribution of objects included in each scene constituting an input image (110) can be predefined / set by the manufacturer of the image processing device (100).

[0355] According to one embodiment, the image processing device (100) can analyze the size and distribution of objects included in each scene constituting the input image (110) by combining a scene object analysis model (1100) and a policy-based algorithm.

[0356] In step 1550, an image processing device (100) according to one embodiment can obtain a modified depth map (20) by non-linearly changing a depth map (10) based on the size and distribution analysis results (1110) of objects included in each scene.

[0357] An image processing device (100) according to one embodiment can obtain additional information (1201). In one embodiment, the additional information (1201) can include at least one of metadata information associated with an input image (110), information regarding a viewing environment, or user setting information.

[0358] An image processing device (100) according to one embodiment can obtain a modified depth map (20) by non-linearly changing the depth map (10) based on at least one of a depth map (10) for an input image (110), a result of analyzing the size and distribution of objects included in each scene constituting the input image (110) (1110), a result of analyzing the type of content of the input image (110) (510), or additional information (1201).

[0359] An image processing device (100) according to one embodiment can generate an offset for each object included in each scene based on the result (1110) of analyzing the size and distribution of objects included in each scene constituting an input image (110).

[0360] An image processing device (100) according to one embodiment can nonlinearly change a depth map (10) by adjusting depth information included in a depth map (10) for an input image (110) by an offset generated for each object included in each scene.

[0361] An image processing device (100) according to one embodiment may control a depth map output range based on at least one of a result (510) of analyzing the content type of an input image (110) or acquired additional information (1201). The output range of the depth map may mean a range from a minimum to a maximum of depth information of pixels included in the depth map. In one embodiment, the image processing device (100) may control the depth map output range by giving priority to information included in the additional information (1201) over the result (510) of analyzing the content type of the input image (110).

[0362] According to one embodiment, the image processing device (100) can obtain a modified depth map (20) by nonlinearly modifying the depth map (10) using a depth map dynamic modification model (1200).

[0363] The depth map dynamic change model (1200) may refer to an artificial neural network model learned / trained to non-linearly change the depth map (10).

[0364] According to one embodiment, the image processing device (100) can obtain a modified depth map (20) by nonlinearly modifying the depth map (10) using a policy-based algorithm. The policy-based algorithm for nonlinearly modifying the depth map (10) can be predefined / set by the manufacturer of the image processing device (100).

[0365] According to one embodiment, the image processing device (100) can obtain a modified depth map (20) by nonlinearly modifying the depth map (10) using the LUT method. A specific function or mapping table for the LUT method can be predefined / set by the manufacturer of the image processing device (100).

[0366] According to one embodiment, the image processing device (100) can obtain a modified depth map (20) by nonlinearly modifying the depth map (10) by combining two or more of a depth map dynamic modification model (1200), a policy-based algorithm, and a LUT method.

[0367] In step 1560, the image processing device (100) according to one embodiment can determine the relative positions of objects included in each scene from a virtual convergence plane corresponding to the screen based on at least one of the result of analyzing the type of content of the input image (110) (510) or the result of analyzing the size and distribution of objects included in each scene (1110) and at least one of the depth map (10) or the modified depth map (20) for the input image.

[0368] An image processing device (100) according to one embodiment can determine the relative positions of objects included in each scene constituting an input image (110) from a virtual convergence plane corresponding to a screen using a three-dimensional effect control model (1300).

[0369] The three-dimensional effect control model (1300) may refer to an artificial neural network model learned / trained to determine the relative positions of objects included in each scene constituting a 2D input image (110) from a virtual convergence plane corresponding to the screen.

[0370] An image processing device (100) according to one embodiment can determine the relative positions of objects included in each scene constituting an input image (110) from a virtual convergence plane corresponding to a screen using a policy-based algorithm. A policy-based algorithm for determining the relative positions of objects included in each scene constituting an input image (110) from a virtual convergence plane corresponding to a screen can be predefined / set by a manufacturer of the image processing device (100).

[0371] An image processing device (100) according to one embodiment can determine the relative positions of objects included in each scene constituting an input image (110) from a virtual convergence plane corresponding to a screen using a LUT (Lookup Table) method.

[0372] An image processing device (100) according to one embodiment can determine the relative positions of objects included in each scene constituting an input image (110) from a virtual convergence plane corresponding to a screen by combining two or more of a three-dimensional effect control model (1300), a policy-based algorithm, and a LUT method.

[0373] In step 1570, an image processing device (100) according to one embodiment can perform 3D transformation from an input image based on a modified depth map (20) and relative positions (1310) of objects included in each scene with respect to a virtual convergence plane.

[0374] An image processing device (100) according to one embodiment can perform 3D transformation from an input image (110) by performing a new viewpoint view generation process (370), a hole filling process (380), and a pixel mapping process (390) based on a modified depth map (e.g., a first depth map) (10) for an input image (110) and relative positions (1310) of objects included in each scene with respect to a virtual convergence plane.

[0375] According to one embodiment of the present disclosure, the image processing device (100) nonlinearly changes a depth map and controls a three-dimensional effect based on analysis of the size and distribution of objects included in each scene constituting an input 2D image, thereby reducing depth estimation errors, further improving the accuracy of depth estimation, and enhancing the satisfaction of a user using a 3D display.

[0376] However, the effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned can be clearly understood by a person having ordinary skill in the technical field to which the present disclosure belongs from the description below.

[0377] FIG. 16 is a drawing for explaining an example of the effect of an image processing device (100) according to one embodiment.

[0378] Referring to FIG. 16, when a 2D input image (110) is a composite image composed of scenes having various types of content, an image processing device (100) according to an embodiment of the present disclosure can provide a 2D to 3D conversion method that allows a user to watch for a long time by providing a strong sense of three-dimensionality in an FPS game, providing a moderate sense of three-dimensionality in an RPG game, appropriately adjusting the sense of three-dimensionality according to a scene in drama and movie content, and reducing the sense of three-dimensionality to a 2D level in a document, taking into account the sense of three-dimensionality for each scene and the user's fatigue.

[0379] Specifically, the method provided by the image processing device (100) according to one embodiment of the present disclosure is different from the existing technology that continuously provides a fixed three-dimensional effect for an input image. The image processing device (100) according to one embodiment of the present disclosure dynamically controls the three-dimensional effect according to the type of content to provide a dynamic sense of depth according to the type of content, non-linearly adjusts the depth map according to the size and distribution of objects in each scene, and determines the relative position of the screen, thereby reducing depth estimation errors, improving depth estimation accuracy, and reducing user fatigue even when watching for a long time.

[0380] Fig. 17 is a block diagram of an image processing device (100) according to one embodiment.

[0381] Referring to FIG. 17, an image processing device (100) according to one embodiment may include at least one processor (101) and memory (102).

[0382] The memory (102) may store one or more instructions for performing the 3D transformation function disclosed in the present disclosure. The memory (102) may store at least one program executed by the processor (101). The memory (102) may store at least one neural network and / or a predefined policy-based algorithm or neural network model. In addition, the memory (102) may store data input to or output from the image processing device (100).

[0383] The memory (102) may include at least one type of storage medium among a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., SD or XD memory, etc.), a RAM (Random Access Memory), a SRAM (Static Random Access Memory), a ROM (Read-Only Memory), an EEPROM (Electrically Erasable Programmable Read-Only Memory), a PROM (Programmable Read-Only Memory), a magnetic memory, a magnetic disk, and an optical disk.

[0384] At least one processor (101) can control the overall operation of the image processing device (100). At least one processor (101) can control the image processing device (100) to function by executing one or more instructions stored in the memory (102).

[0385] For example, at least one processor (101) can perform the function of the image processing device (100) described in FIGS. 1 to 16 by executing one or more instructions stored in the memory 120.

[0386] At least one processor (101) may be composed of one or more processors. In this case, one or more processors may be a general-purpose processor such as a CPU, AP, or DSP (Digital Signal Processor), a graphics-only processor such as a GPU or VPU (Vision Processing Unit), or an artificial intelligence-only processor such as an NPU. For example, if one or more processors are artificial intelligence-only processors, the artificial intelligence-only processor may be designed with a hardware structure specialized for processing a specific artificial intelligence model.

[0387] According to one embodiment, at least one processor (101) can analyze the content type of an input image by executing one or more instructions stored in a memory (102). Based on the result of analyzing the content type of the input image by executing one or more instructions stored in the memory (102), the at least one processor (101) can obtain a depth estimation model corresponding to the input image in real time using on-device learning. The at least one processor (101) can obtain a depth map for the input image reflecting the estimated depth information based on the depth estimation model according to on-device learning by executing one or more instructions stored in the memory (102). The at least one processor (101) can perform 3D transformation on the input image based on the depth map for the input image.

[0388] According to one embodiment, the result of analyzing the content type of the input image may include information related to probabilities that the input image corresponds to each of the predefined content types.

[0389] According to one embodiment, the depth estimation model can be acquired in real time by training the parameters of the depth estimation model by updating them based on the results of analyzing the content type of the input image in the image processing device.

[0390] According to one embodiment, a depth estimation model can be obtained in real time by interpolating depth estimation models based on a result of analyzing the content type of an input image in an image processing device.

[0391] According to one embodiment, at least one processor (101) can obtain depth estimation models corresponding to each of the scenes constituting the input image in real time based on a result of analyzing the content type of the input image by executing one or more instructions stored in the memory (102). At least one processor (101) can obtain a depth map for the input image in real time by applying depth estimation models corresponding to each of the scenes constituting the input image to each of the scenes constituting the input image by executing one or more instructions stored in the memory (102).

[0392] According to one embodiment, at least one processor (101) may analyze the sizes and distributions of objects included in each scene constituting the input image based on an input image and a depth map for the input image by executing one or more instructions stored in the memory (102). At least one processor (101) may obtain a modified depth map by non-linearly modifying the depth map based on the results of the analysis of the sizes and distributions of objects included in each scene by executing one or more instructions stored in the memory (102). At least one processor (101) may perform 3D transformation from the input image based on the modified depth map by executing one or more instructions stored in the memory (102).

[0393] According to one embodiment, at least one processor (101) may obtain additional information including at least one of metadata information about an input image, information about a viewing environment, or user setting information by executing one or more instructions stored in a memory (102). At least one processor (101) may control an output range of the depth map or the modified depth map based on at least one of the additional information obtained by executing one or more instructions stored in the memory (102) or a result of analyzing the content type of the input image.

[0394] According to one embodiment, at least one processor (101) may generate an offset for each object included in each scene based on the analysis results of the sizes and distributions of objects included in each scene by executing one or more instructions stored in the memory (102). At least one processor (101) may adjust depth information included in a depth map for an input image by the offset generated for each object included in each scene by executing one or more instructions stored in the memory (102), thereby nonlinearly changing the depth map to obtain a modified depth map.

[0395] According to one embodiment, at least one processor (101) may determine relative positions of objects included in each scene from a virtual convergence plane corresponding to a screen based on at least one of the results of analyzing the type of content of an input image by executing one or more instructions stored in a memory (102), the results of analyzing the sizes and distributions of objects included in each scene, or the acquired additional information. At least one processor (101) may perform 3D transformation from an input image based on information about the relative positions of objects included in each scene with respect to the virtual convergence plane by executing one or more instructions stored in a memory (102).

[0396] According to one embodiment, at least one processor (101) can control a display to generate and output a UI (user interface) that represents the type of content corresponding to each scene and the three-dimensionality of each scene that constitutes an output image converted in 3D from an input image by executing one or more instructions stored in a memory (102).

[0397] A specific example for explaining an embodiment according to the present disclosure is only one combination of each criterion, method, detailed method, and operation, and through a combination of at least two or more techniques among the various techniques described, an image processing device dynamically controls a three-dimensional effect according to a type of content, thereby providing a dynamic depth effect according to a type of content, non-linearly adjusting a depth map according to the size and distribution of objects in each scene, and controlling the three-dimensional effect, thereby reducing a depth estimation error, improving depth estimation accuracy, and reducing user fatigue even when watching for a long time.

[0398] Additionally, at this time, the method may be performed in a manner determined by one or a combination of at least two of the aforementioned techniques. For example, it may be possible to perform a portion of the operation of one embodiment in combination with a portion of the operation of another embodiment.

[0399] A device-readable storage medium may be provided in the form of a non-transitory storage medium. Here, the term "non-transitory storage medium" simply means a tangible device that does not contain signals (e.g., electromagnetic waves). This term does not distinguish between cases where data is permanently stored in the storage medium and cases where data is temporarily stored. For example, a "non-transitory storage medium" may include a buffer in which data is temporarily stored.

[0400] According to one embodiment, the method according to various embodiments disclosed in the present document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., a downloadable app) may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.

Claims

1. In an image processing device, Memory that stores one or more instructions; and At least one processor configured to execute one or more instructions stored in the memory; The at least one processor executes the one or more instructions, Analyze the content type of the input video, Based on the results of analyzing the content type of the input image, a depth estimation model corresponding to the input image is obtained in real time using on-device learning. Based on the depth estimation model according to the above on-device learning, a depth map for the input image reflecting the estimated depth information is obtained, An image processing device that performs 3D transformation from an input image based on a depth map for the input image.

2. An image processing device according to claim 1, wherein the result of analyzing the content type of the input image includes information related to probabilities that the input image corresponds to each of the predefined content types.

3. In the first or second paragraph, the depth estimation model, An image processing device, which obtains in real time by updating and training the parameters of the depth estimation model based on the results of analyzing the content type of the input image in the image processing device.

4. In any one of clauses 1 to 3, the depth estimation model, An image processing device, which obtains in real time by interpolating depth estimation models based on the results of analyzing the content type of the input image in the image processing device.

5. In any one of paragraphs 1 to 4, the at least one processor executes the one or more instructions, Based on the results of analyzing the content type of the input image, depth estimation models corresponding to each scene constituting the input image are acquired in real time, An image processing device that obtains a depth map for the input image in real time by applying depth estimation models corresponding to each scene constituting the input image to each scene constituting the input image.

6. In any one of paragraphs 1 to 5, the at least one processor executes the one or more instructions, Based on the input image and the depth map for the input image, for each scene constituting the input image, the size and distribution of objects included in each scene are analyzed, Based on the results of the size and distribution analysis of the objects included in each of the above scenes, a modified depth map is obtained by nonlinearly changing the depth map, An image processing device that performs 3D transformation from the input image based on the modified depth map.

7. In any one of claims 1 to 6, the at least one processor executes the one or more instructions, Obtaining additional information including at least one of metadata information about the input image, information about the viewing environment, or user setting information, An image processing device that controls the output range of the depth map or the modified depth map based on at least one of the acquired additional information or the result of analyzing the content type of the input image.

8. In the 6th or 7th paragraph, the at least one processor executes the one or more instructions, Based on the results of the size and distribution analysis of the objects included in each of the above scenes, an offset is generated for each object included in each of the above scenes. An image processing device that nonlinearly changes the depth map by adjusting depth information included in a depth map for the input image by an offset generated for each object included in each of the above scenes, thereby obtaining the modified depth map.

9. In any one of paragraphs 1 to 8, the at least one processor executes the one or more instructions, As a result of analyzing the content type of the input image, the relative positions of the objects included in each scene are determined from a virtual convergence plane corresponding to the screen based on at least one of the results of analyzing the size and distribution of the objects included in each scene or the obtained additional information. An image processing device that performs 3D transformation from the input image based on information about the relative positions of objects included in each scene with respect to the virtual convergence plane.

10. In any one of claims 1 to 9, the at least one processor executes the one or more instructions, An image processing device that controls a display to generate and output a UI (user interface) that represents the type of content corresponding to each scene and the three-dimensionality of each scene that constitutes an output image converted into 3D from the input image.

11. In the operating method of the image processing device, A step of analyzing the content type of the input image; A step of obtaining a depth estimation model corresponding to the input image in real time using on-device learning based on the result of analyzing the content type of the input image; A step of obtaining a depth map for the input image reflecting the estimated depth information based on the depth estimation model according to the on-device learning; and A method comprising the step of performing 3D transformation from the input image based on a depth map for the input image.

12. In the 11th paragraph, the depth estimation model, A method for obtaining in real time by updating and training the parameters of the depth estimation model based on the results of analyzing the content type of the input image in the image processing device.

13. In clause 11 or 12, the depth estimation model, A method for obtaining in real time, by interpolating depth estimation models based on the results of analyzing the content type of the input image in the image processing device.

14. In any one of clauses 11 to 13, the method, A step of obtaining depth estimation models corresponding to each scene constituting the input image in real time based on the result of analyzing the content type of the input image; and A method further comprising a step of obtaining a depth map for the input image in real time by applying depth estimation models corresponding to each scene constituting the input image, for each scene constituting the input image.

15. A computer-readable recording medium having recorded thereon at least one program for implementing the method described in any one of claims 11 to 14.

Citation Information

Patent Citations

  • Responding to remote media classification queries using classifier models and context parameters

    KR1020180120146A

  • A paper filter-based color sensor that can detect viruses

    KR1020230017681A

  • Horizontal Form Pillow Packaging Machine For Packing Individually Packaged Article Into Pillow Packaging Bag

    KR1020240065764A

  • Method and system for verifying performance of vehicle active noise control

    KR1020250035224A

  • High-speed real-time scene reconstruction from input image data

    WO2023111909A1