Image providing apparatus and image providing method thereof, and display apparatus and display method thereof
By introducing AI-driven DNN and OOTF into the display device, the problem of the inability to personalize tone mapping in existing technologies has been solved, enabling high-quality display of HDR images.
Patent Information
- Application Number
- CN202080074016.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-11-15
- Filing Date
- 2020-11-12
- Publication Date
- 2026-01-16
- Estimated Expiration
- 2040-11-12
AI Technical Summary
Existing display devices cannot personalize their tone mapping curves according to the ambient light signal when processing high dynamic range (HDR) images, which limits the improvement of image quality.
A deep neural network (DNN) based on artificial intelligence (AI) and an optical transfer function (OOTF) are used to adjust tone mapping. By acquiring the encoded data of the image and AI meta-information, the processing of light signals is optimized to achieve personalized tone mapping.
It improves the quality of displayed images and can make personalized adjustments based on different ambient light signals, enhancing the brightness and detail of the image.
Smart Images

Figure CN114641793B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The disclosure relates to the field of image processing, and more particularly, to high dynamic range (HDR) for improving the quality of an image to be displayed. BACKGROUND
[0002] The human recognizable luminance range is approximately 10 -6 nits to 10 8 nits, but the luminance range that a human encounters in real life is much greater than the recognizable range. In order to maximize the realism of a video, various researches for providing a high dynamic range (HDR) greater than a dynamic range supported by a high definition television (HDTV) and related technical standardization have been made.
[0003] When assuming that the minimum representable luminance is 0.0 and the maximum representable luminance is 1.0, the existing display device uses an 8-bit fixed point value to express the luminance level of each channel. In HDR, when expressing the luminance level, a greater or smaller luminance value can be precisely expressed by using 16-bit, 32-bit, or 64-bit floating point data. In an HDR image, a bright object looks bright, a dark object looks dark, and the details of both the bright object and the dark object are viewable.
[0004] The luminance range of a light signal having a linear luminance value can be greater than the luminance range that a display device can achieve, and thus a tone mapping curve for tone mapping of the light signal is used. In the related art, because the tone mapping curve is uniformly applied to the light signal regardless of the environment of the light signal, the quality improvement of the image to be displayed is limited. SUMMARY
[0005] TECHNICAL PROBLEM
[0006] Provided is an image providing apparatus and an image providing method thereof and a display apparatus and a display method thereof that improve the quality of an image to be displayed through artificial intelligence (AI)-based tone mapping.
[0007] TECHNICAL SOLUTION
[0008] According to an aspect of the disclosure, a display apparatus includes a memory storing one or more instructions, and a processor configured to execute the stored one or more instructions to obtain encoding data of a first digital image and artificial intelligence (AI) meta information indicating a specification of a deep neural network (DNN), obtain a second digital image corresponding to the first digital image by decoding the encoding data, obtain an optical signal converted from the second digital image according to a predetermined electro-optical transfer function (EOTF), and obtain a display signal by processing the optical signal using an optical-optical transfer function (OOTF) and a high dynamic range (HDR) DNN set according to the AI meta information.
[0009] Advantageous Effects
[0010] Provided is an image providing apparatus and an image providing method thereof and a display apparatus and a display method thereof that improve the quality of an image to be displayed through artificial intelligence (AI)-based tone mapping. BRIEF DESCRIPTION OF DRAWINGS
[0011] The above and other aspects, features, and advantages of certain embodiments of the disclosure will be more apparent from the following description taken in conjunction with the following drawings, in which:
[0012] Figure 1 An image providing method and a display method according to an embodiment are illustrated;
[0013] Figure 2 is a block diagram of an image providing apparatus according to an embodiment;
[0014] Figure 3 is a graph illustrating an optical-optical transfer function (OOTF);
[0015] Figure 4 A method of determining a specification of a deep neural network (DNN) by an image providing apparatus according to an embodiment is illustrated;
[0016] Figure 5 A method of determining a specification of a DNN by an image providing apparatus according to another embodiment is illustrated;
[0017] Figure 6 A method of determining a specification of a DNN by an image providing apparatus according to another embodiment is illustrated;
[0018] Figure 7 A method of determining a specification of a DNN by an image providing apparatus according to another embodiment is illustrated;
[0019] Figure 8 is a table indicating various specifications of a DNN determined by an image providing apparatus;
[0020] Figure 9 a frame constituting a first digital image is shown;
[0021] Figure 10 AI display data according to an embodiment is shown;
[0022] Figure 11 structure of AI meta information included in Figure 10 AI display data shown;
[0023] Figure 12 AI display data according to another embodiment is shown;
[0024] Figure 13 structure of AI meta information included in Figure 12 AI display data shown;
[0025] Figure 14 is a flowchart of an image providing method according to an embodiment;
[0026] Figure 15 is a flowchart of an image providing method according to another embodiment;
[0027] Figure 16 is a block diagram of a display apparatus according to an embodiment;
[0028] Figure 17 tone mapping operation performed by a display apparatus according to an embodiment is shown;
[0029] Figure 18 tone mapping operation performed by a display apparatus according to another embodiment is shown;
[0030] Figure 19 tone mapping operation performed by a display apparatus according to another embodiment is shown;
[0031] Figure 20 tone mapping operation performed by a display apparatus according to another embodiment is shown;
[0032] Figure 21 is a flowchart of a display method according to an embodiment; and
[0033] Figure 22 is a flowchart of a display method according to another embodiment. DETAILED DESCRIPTION
[0034] Optimal mode
[0035] According to an aspect of the disclosure, a display apparatus includes a memory storing one or more instructions, and a processor configured to execute the stored one or more instructions to obtain encoding data of a first digital image and artificial intelligence (AI) meta information indicating a specification of a deep neural network (DNN), obtain a second digital image corresponding to the first digital image by decoding the encoding data, obtain an optical signal converted from the second digital image according to a predetermined electro-optical transfer function (EOTF), and obtain a display signal by processing the optical signal using an optical-optical transfer function (OOTF) and a high dynamic range (HDR) DNN set according to the AI meta information.
[0036] The HDR DNN can include a plurality of layers, and the AI meta information can include at least one of a number of layers, a type of layer, a number of filter kernels used in at least one layer, a size of a filter kernel used in at least one layer, a weight or a bias value of a filter kernel used in at least one layer.
[0037] The processor can be further configured to execute the one or more instructions to convert the optical signal according to the OOTF, input the optical signal to the HDR DNN, and obtain the display signal by adding a signal converted from the optical signal according to the OOTF to an output signal of the HDR DNN.
[0038] The processor can be further configured to execute the one or more instructions to convert the optical signal according to the OOTF, input a first intermediate image converted from the optical signal according to an optical-electrical transfer function (OETF) to the HDR DNN, convert a second intermediate image output from the HDR DNN according to the EOTF, and obtain the display signal by adding a signal converted from the optical signal according to the OOTF to a signal converted from the second intermediate image according to the EOTF.
[0039] The processor can be further configured to execute the one or more instructions to obtain the display signal by processing the optical signal according to one of the OOTF and the HDR DNN and processing a result of the processing of the one of the OOTF and the HDR DNN according to the other of the OOTF and the HDR DNN.
[0040] The processor can be further configured to execute the one or more instructions to obtain OOTF meta information to be used for setting the OOTF, and input the obtained OOTF meta information to the HDR DNN.
[0041] The second digital image can include a plurality of frames, and the processor can be further configured to execute the one or more instructions to obtain first AI meta information for frames in a first group of the plurality of frames and second AI meta information for frames in a second group of the plurality of frames, and to independently set an HDR DNN for the frames in the first group according to the first AI meta information and to independently set an HDR DNN for the frames in the second group according to the second AI meta information.
[0042] The first AI meta information can include first identification information of frames to which the first AI meta information is applied, and the second AI meta information can include second identification information of frames to which the second AI meta information is applied.
[0043] According to another aspect of the disclosure, an image providing apparatus includes a memory storing one or more instructions, and a processor configured to execute the stored one or more instructions to encode a first digital image, and transmit, to a display apparatus, encoded data of the first digital image and artificial intelligence (AI) meta information indicating a specification of a deep neural network (DNN) determined based on difference information between a result of processing an optical signal corresponding to the first digital image by using an optical-optical transfer function (OOTF) and the DNN and a labeled signal.
[0044] The labeled signal can be predetermined based on a signal converted from the optical signal according to the OOTF.
[0045] The first digital image can include a plurality of frames, and the processor can be further configured to execute the one or more instructions to independently determine a first specification of a DNN for frames in a first group of the plurality of frames and a second specification of the DNN for frames in a second group of the plurality of frames.
[0046] The processor can be further configured to execute the one or more instructions to divide the plurality of frames into the frames in the first group and the frames in the second group based on a histogram similarity or a variance of pixel values of each of the plurality of frames.
[0047] The processor can be further configured to execute the one or more instructions to determine a representative frame in the first group and a representative frame in the second group of the plurality of frames, determine a first specification of the DNN based on difference information between a result of processing an optical signal corresponding to the representative frame in the first group by using the OOTF and the DNN and the labeled signal, and determine the second specification of the DNN based on difference information between a result of processing an optical signal corresponding to the representative frame in the second group by using the OOTF and the DNN and the labeled signal.
[0048] The processor can be further configured to execute the one or more instructions to transmit information indicating whether a DNN-based tone mapping process is required to the display device, based on a difference between a maximum value of luminance that the display device is capable of displaying and a threshold luminance value being less than or equal to a predetermined value.
[0049] The processor can be further configured to execute the one or more instructions to transmit information indicating that a DNN-based tone mapping process is not required to the display device, based on a difference between a maximum value of luminance that the display device is capable of displaying and a threshold luminance value being less than or equal to a predetermined value.
[0050] The processor can be further configured to execute the one or more instructions to receive performance information from the display device, determine one specification capable of being used for tone mapping of a light signal corresponding to the first digital image, among specifications of a plurality of DNNs, based on the received performance information, and transmit AI meta information indicating the determined one specification of the DNN to the display device.
[0051] The processor can be further configured to execute the one or more instructions to determine a restriction condition of the DNN based on pixel values of the first digital image, and the restriction condition can include at least one of a minimum number of layers included in the DNN, a minimum size of a filter kernel used in at least one layer, or a minimum number of filter kernels used in at least one layer.
[0052] According to another aspect of the disclosure, an image display method includes obtaining encoding data of a first digital image and artificial intelligence (AI) meta information indicating a specification of a deep neural network (DNN), obtaining a second digital image corresponding to the first digital image by decoding the encoding data, obtaining a light signal converted from the second digital image according to a predetermined electro-optical transfer function (EOTF), and obtaining a display signal by processing the light signal using an optical-optical transfer function (OOTF) and a high dynamic range (HDR) DNN set according to the AI meta information.
[0053] According to another aspect of the disclosure, an image display method includes obtaining encoding data of a first digital image and artificial intelligence (AI) meta information indicating a specification of a deep neural network (DNN), obtaining a second digital image corresponding to the first digital image by decoding the encoding data, obtaining a light signal converted from the second digital image according to a predetermined electro-optical transfer function (EOTF), and obtaining a display signal by processing the light signal using an optical-optical transfer function (OOTF) and a high dynamic range (HDR) DNN set according to the AI meta information.
[0054] According to another aspect of the disclosure, a method of providing meta information includes determining a specification of a deep neural network (DNN) based on difference information between a result of processing a light signal corresponding to a first digital image by using an optical optical transfer function (OOTF) and the DNN and a labeled signal, and transmitting artificial intelligence (AI) meta information indicating the determined specification of the DNN to a display device.
[0055] Inventive Modes
[0056] Various types of changes or modifications can be made to the disclosed embodiments and the specific embodiments shown in the drawings and described in detail above. However, it should be understood that the specific embodiments do not limit the disclosure to specific forms, but include every modification, equivalent or alternative form within the spirit and technical scope of the disclosure.
[0057] In the description of the embodiments, detailed description of related known features can be omitted in order to obscure the gist of the disclosure. In addition, numbers (e.g., first and second) used in the description of the embodiments are only identifiers for distinguishing one element from another element.
[0058] When it is described that one component is "connected" or "linked" to another component, it should be understood that one component can be directly connected to another component, or can be connected or linked to another component via a third component therebetween, unless there is a description to the contrary, even if direct connection is possible.
[0059] In addition, regarding components such as "... unit" and "... module" used in the specification, two or more components can be combined into a single component, or a single component can be divided into two or more components according to a subdivided function. In addition, each component to be described below can additionally perform part or all of a function that another component is configured to perform, in addition to its main function, and part of the main function of each component can be exclusively performed by another component.
[0060] Throughout the disclosure, expressions such as "at least one of A, B, and / or C" indicate only A, only B, only C, both A and B, both A and C, both B and C, all of A, B, and C, or a variation thereof.
[0061] In addition, in the specification, the term "light signal" indicates a signal having a linear luminance value. The linear luminance value can exist two-dimensionally (i.e., in a horizontal direction and a vertical direction). The luminance value of the "light signal" can be represented by a floating point. The "light signal" can include scene light collected by a camera sensor or display light output from a display. Since the "light signal" corresponds to light existing in a natural condition, the "light signal" has a linear luminance value.
[0062] In addition, in the present specification, the term "display signal" indicates a signal that is to be displayed, which is tone-mapped from a "light signal" based on artificial intelligence (AI) and / or an optical-optical transfer function (OOTF). The "display signal" can be referred to as a display light. The "display signal" is represented as an image through a display.
[0063] In addition, in the present specification, the term "digital image" indicates data having non-linear luminance values. The non-linear luminance values exist two-dimensionally, that is, in a horizontal direction and a vertical direction. The luminance values of the "digital image" can be represented by a fixed point. In the present specification, the luminance values of the "digital image" can be referred to as pixel values. Since the "digital image" is converted from a light signal according to the visual characteristics of a human being, the "digital image" has non-linear luminance values that are different from light existing in a natural condition.
[0064] In addition, the "light signal", the "display signal", and the "digital image" can include at least one frame. In this context, a "frame" includes a luminance value at one point in time among luminance values over time.
[0065] In addition, the luminance values of the "light signal", the "display signal", and the "digital image" can be represented as RGB values or luminance values.
[0066] In addition, in the present specification, the term "opto-electric transfer function" (OETF) is a function that defines a relationship between luminance values of a light signal and luminance values of a digital image. A digital image can be obtained by converting luminance values of a light signal according to an OETF. The OETF can convert luminance values included in a narrow range, which have a relatively small size, among luminance values of a light signal, to luminance values of a wide range, and convert luminance values included in a wide range, which have a relatively large size, among luminance values of a light signal, to luminance values of a narrow range. The OETF can convert a light signal to a digital image suitable for the cognitive visual characteristics of a human being, thereby enabling optimal bits to be allocated in quantization of the digital image. That is, in a digital image converted from a light signal according to an OETF, a larger number of bits can be allocated to a region corresponding to a dark region of the light signal, and a smaller number of bits can be allocated to a region corresponding to a bright region of the light signal.
[0067] In addition, in the present specification, the term "electro-optical transfer function" (EOTF) is a function that defines a relationship between luminance values of a digital image and luminance values of a light signal, and can have an inverse relationship with the OETF. A light signal can be obtained by converting luminance values of a digital image according to an EOTF.
[0068] Also, in the present specification, the term "optical-optical transfer function" (OOTF) is a function that defines a relationship between a luminance value of any one optical signal and a luminance value of another optical signal. Another optical signal can be obtained by converting a luminance value of any one optical signal according to the OOTF.
[0069] Also, in the present specification, the term "tone mapping" indicates an operation of converting an optical signal into a display signal according to an OOTF and / or an AI.
[0070] Also, in the present specification, the term "deep neural network" (DNN) is a representative example of an artificial neural network model that simulates a brain neural network, and is not limited to an artificial neural network model using a specific algorithm.
[0071] Also, in the present specification, the term "structure of a DNN" indicates at least one of a number of layers constituting a DNN, a type of layer, a size of a filter kernel used in at least one layer, or a number of filter kernels used in at least one layer.
[0072] Also, in the present specification, the term "parameter of a DNN" is a value used in a calculation operation of each layer constituting a DNN, and can include, for example, at least one of a weight to be used when an input value is applied to a specific formula or a bias value to be added to or subtracted from a result value of a specific formula. The "parameter" can be expressed in a matrix form. Furthermore, the "parameter" is a value optimized as a result of training, and can be updated by separate training data as the case can be. The "parameter" can be determined through a training operation of a DNN using training data after a structure of the DNN is determined.
[0073] Also, in the present specification, the term "specification of a DNN" indicates at least one of a structure or a parameter of a DNN. For example, in the present specification, the expression "determining a specification of a DNN" indicates determining a structure of a DNN, determining a parameter of a DNN, or determining a structure and a parameter of a DNN.
[0074] Also, in the present specification, the term "high dynamic range (HDR) DNN" is a DNN to be used for tone mapping of an optical signal, and is set to have a specification of a DNN determined through one or more embodiments described below.
[0075] Also, in the specification, the term "setting an HDR DNN (or an OOTF)" can indicate storing an HDR DNN (or an OOTF) having a specification indicated by AI meta information (or OOTF meta information), modifying a previously stored HDR DNN (or OOTF) having an arbitrary specification so that the previously stored HDR DNN (or OOTF) has a specification indicated by AI meta information (or OOTF meta information), or generating an HDR DNN (or an OOTF) having a specification indicated by AI meta information (or OOTF meta information). In other words, the term "setting an HDR DNN (or an OOTF)" can indicate various types of operations that enable a display device to use an HDR DNN (or an OOTF) having a specification indicated by AI meta information (or OOTF meta information).
[0076] Hereinafter, embodiments will be described in detail.
[0077] Figure 1 An image providing method and a display method according to an embodiment are illustrated.
[0078] Referring to Figure 1 A first digital image having a non-linear luminance value is obtained by applying an OETF 201 to a first optical signal having a linear luminance value. Although Figure 1 An image providing apparatus 200 is illustrated as converting a first optical signal according to an OETF 201, but the application of the OETF 201 can be implemented by a camera sensor, and in this case, the image providing apparatus 200 obtains a first digital image generated as a result of the application of the OETF 201.
[0079] The image providing apparatus 200 performs an encoding operation 202 and an image analysis 203 on the first digital image, and transmits AI display data including encoded data and meta information as a result of the encoding operation 202 and the image analysis 203 to a display apparatus 1600. The meta information includes information to be used by the display apparatus 1600 in tone mapping 1603.
[0080] According to an embodiment, the tone mapping 1603 uses an OOTF and an HDR DNN, and the image providing apparatus 200 transmits OOTF meta information used by the display apparatus 1600 to set the OOTF to the display apparatus 1600, and transmits AI meta information used by the display apparatus 1600 to set the HDR DNN to the display apparatus 1600. Since the meta information is derived as a result of the analysis of the first digital image, the display apparatus 1600 can display an image of excellent quality through the tone mapping 1603 based on the meta information.
[0081] The encoding operation 202 of the image providing apparatus 200 can include generating prediction data by predicting the first digital image, generating residual data corresponding to a difference between the first digital image and the prediction data, transforming the residual data of the spatial domain component into the residual data of the frequency domain component, quantizing the residual data transformed into the frequency domain component, entropy-encoding the quantized residual data, and the like. The encoding operation 202 can be implemented by using one of image compression schemes using frequency transform, such as Moving Picture Experts Group 2 (MPEG-2), H.264 Advanced Video Coding (AVC), MPEG-4, High Efficiency Video Coding (HEVC), VC-1, VP8, VP9, Alliance for Open Media Video (AV1), and the like.
[0082] The encoded data can be transmitted in the form of a bitstream. The encoded data can include data obtained based on the pixel values of the first digital image, for example, the residual data corresponding to the difference between the first digital image and the prediction data. In addition, the encoded data includes a plurality of pieces of information used in the encoding operation of the first digital image. For example, the encoded data can include prediction mode information, motion information, quantization parameter-related information, and the like, used for encoding the first digital image. The encoded data can be generated according to the rules (e.g., syntax) of the image compression scheme used in the encoding operation 202, among the image compression schemes using frequency transform, such as MPEG-2, H.264 AVC, MPEG-4, HEVC, VC-1, VP8, VP9, AV1, and the like.
[0083] The meta information can be transmitted in the form of a bitstream by being included in the encoded data. According to an implementation example, the meta information can be transmitted in the form of a frame or a packet by being separated from the encoded data. The encoded data and the meta information can be transmitted through the same network or different networks. Although Figure 1 Although both the meta information and the encoded data are shown as being transmitted from the image providing apparatus 200 to the display apparatus 1600, according to an implementation example, the meta information and the encoded data can be transmitted from different apparatuses to the display apparatus 1600, respectively.
[0084] The display apparatus 1600 having received the AI meta information performs a decoding operation 1601 on the encoded data to restore the second digital image having the non-linear luminance values. Here, the decoding operation 1601 can include generating quantized residual data by entropy-decoding the encoded data, inverse-quantizing the quantized residual data, transforming the residual data of the frequency domain component into the residual data of the spatial domain component, generating prediction data, obtaining the second digital image by using the prediction data and the residual data, and the like. The decoding operation 1601 can be implemented by an image decompression method corresponding to one of the image compression schemes using frequency transform, such as MPEG-2, H.264 AVC, MPEG-4, HEVC, VC-1, VP8, VP9, AV1, and the like.
[0085] The display apparatus 1600 obtains a second light signal converted from the second digital image according to the previously determined EOTF 1602. The second light signal includes linear luminance values. The EOTF 1602 and the OETF 201 can have inverse relationships with each other.
[0086] The display apparatus 1600 obtains a display signal having linear luminance values by applying a tone mapping 1603 based on the meta information to the second light signal. The display signal is output on a screen of the display apparatus 1600.
[0087] Because the meta information is derived as a result of analysis of the first digital image, the display apparatus 1600 can set an OOTF and an HDR DNN optimized for the first digital image based on the meta information, and display an image of excellent quality by performing the tone mapping 1603 based on the set OOTF and HDR DNN.
[0088] In the present disclosure, the tone mapping 1603 of the second light signal is performed based on a DNN. The image providing apparatus 200 determines which specification of DNN to use to perform the tone mapping 1603 of the second light signal in order to maximize improvement in the quality of the image to be displayed. In addition, the image providing apparatus 200 transmits meta information (in particular, AI meta information) indicating the specification of the DNN to the display apparatus 1600, so that the display apparatus 1600 performs the tone mapping 1603 based on the DNN. That is, by providing the AI meta information from the image providing apparatus 200 to the display apparatus 1600, a viewer can watch an image having a wide luminance range and luminance values improved according to a context. It should be understood that embodiments are not limited to the display apparatus 1600, and include an image processing apparatus that decodes and processes (including tone mapping) an image signal to be output to, for example, a display.
[0089] Hereinafter, referring to Figures 2 to 22 The configuration and operation of the image providing apparatus 200 and the configuration and operation of the display apparatus 1600 are described in detail.
[0090] Figure 2 is a block diagram of an image providing apparatus 200 according to an embodiment.
[0091] Referring to Figure 2 The image providing apparatus 200 according to an embodiment can include an image processor 210 and a transmitter 230. The image processor 210 can include an encoder 212 and an image analyzer 214. The transmitter 230 can include a data processor 232 and a communication interface 234.
[0092] Although Figure 2It is shown that the image processor 210 and the transmitter 230 are separate, but the image processor 210 and the transmitter 230 can be implemented by a single processor. In this case, the single processor can be implemented by a dedicated processor or by a combination of a general-purpose processor such as an application processor (AP), a central processing unit (CPU), or a graphics processing unit (GPU) and software (S / W). In addition, the dedicated processor can be implemented with a memory for implementing the embodiments or with a memory processor for using an external memory.
[0093] In addition, the image processor 210 and the transmitter 230 can be implemented by a plurality of processors. In this case, the plurality of processors can be implemented by a combination of a plurality of general-purpose processors such as an AP, a CPU, and a GPU and S / W.
[0094] The encoder 212 can encode the first digital image according to an image compression scheme based on a frequency transform. As a result of encoding the first digital image, encoded data is generated and transmitted to the data processor 232.
[0095] The encoded data can include data obtained based on pixel values of the first digital image, for example, residual data corresponding to a difference between the first digital image and prediction data. In addition, the encoded data includes information used when the first digital image is encoded. For example, the encoded data can include prediction mode information, motion information, quantization parameter-related information, etc. used to encode the first digital image. In addition, as described below, the encoded data can include at least one of OOTF meta information or AI meta information.
[0096] The image analyzer 214 can analyze the first digital image to generate meta information for tone mapping in the display device 1600. The meta information can be transmitted to the data processor 232.
[0097] In particular, the image analyzer 214 includes an OOTF determiner 215 and a DNN determiner 216. The OOTF determiner 215 and the DNN determiner 216 determine a specification of an OOTF and a specification of a DNN, respectively, for tone mapping.
[0098] The OOTF determiner 215 can determine the specification of the OOTF through various methods. As one method, the OOTF determiner 215 can determine an OOTF having a specification corresponding to a characteristic of the first digital image among previously stored OOTFs having different specifications. Here, the characteristic of the first digital image can include a distribution, a bias, a variance, a histogram, etc. of the first digital image. As another method, the OOTF determiner 215 can process the first digital image or a first optical signal corresponding to the first digital image by using a previously trained DNN, and obtain an OOTF having a certain specification based on a processing result. As another method, the OOTF determiner 215 can obtain an OOTF having a specification determined by a manager.
[0099] When the specification of the OOTF is determined, the OOTF determiner 215 can generate OOTF meta information. The OOTF meta information is used for the display device 1600 to set the OOTF.
[0100] Figure 3 is a graph illustrating an OOTF.
[0101] In Figure 3 , the horizontal axis indicates a luminance value of an optical signal before tone mapping, and the vertical axis indicates a luminance value of an optical signal after tone mapping. The OOTF is used to convert an optical signal having a wide range of luminance values into an optical signal having a relatively narrow range of luminance values.
[0102] The OOTF can include a Bezier curve 300. The Bezier curve 300 includes an inflection point 310 and one or more anchor points 330, in which the Bezier curve 300 is linearly changed from an origin to the inflection point 310 and is nonlinearly changed from the inflection point 310. That is, the Bezier curve 300 can include a first-order curve from the origin to the inflection point 310 and a multi-order curve from the inflection point 310.
[0103] The anchor points 330 can indicate inflection points of a curve, and in the Bezier curve 300, the number of the anchor points 330 can be 1 or more.
[0104] The OOTF meta information indicates a specification of the OOTF, and can include at least one of information indicating a position of the inflection point 310, information indicating a position of the anchor points 330, or information indicating a number of the anchor points 330. In this context, the information indicating the position of the inflection point 310 can include an x-axis value K S and a y-axis value K F of the inflection point 310. In addition, the information indicating the position of the anchor points 330 can include real values indicating the positions of the anchor points 330.
[0105] Referring back to Figure 2The OOTF having the specification determined by the OOTF determiner 215 can be provided to the DNN determiner 216. The DNN determiner 216 determines a specification of a DNN to be used for tone mapping of a second light signal based on the first digital image and the OOTF. The DNN can include a plurality of layers, and each layer can be a convolution layer, an activation layer, a normalization layer, or a pooling layer.
[0106] The determination of the specification of the DNN by the DNN determiner 216 indicates determination of a structure of the DNN and / or parameters of the DNN. The structure of the DNN can be specified by the number of layers, the type of layers, the size of a filter kernel used in at least one layer, and the number of filter kernels used in at least one layer. The filter kernel can be used for convolution processing of input data in the convolution layer. In addition, the parameters of the DNN can include at least one of a weight or a bias value to be used when input data is used in a layer. For example, the parameters of the DNN can include a weight of a filter kernel to be used when input data is convolution-processed in a convolution layer. Output data can be determined by a multiplication operation and an addition operation between the weight of the filter kernel and a sample value of the input data. The convolution operation in the convolution layer can be performed as in the related art.
[0107] The DNN determiner 216 determines the specification of the DNN while continuously changing the specification of the DNN used for effective tone mapping for the second light signal.
[0108] Hereinafter, the following will be described with reference to Figures 4 to 7 A specific method of determining the specification of the DNN to be used for tone mapping, which is performed by the DNN determiner 216, is described.
[0109] Figure 4 A method of determining the specification of the DNN performed by the image providing apparatus 200 according to an embodiment is illustrated.
[0110] The first light signal 410 corresponding to the first digital image is converted according to the OOTF 415. Here, the OOTF 415 is determined by the OOTF determiner 215.
[0111] When the first light signal 410 corresponding to the first digital image is not stored in the image providing apparatus 200, or based on the first light signal 410 corresponding to the first digital image is not stored in the image providing apparatus 200, the DNN determiner 216 converts the first digital image into the first light signal 410 according to the EOTF, and converts the first light signal 410 according to the OOTF 415.
[0112] The first light signal 410 is processed by the DNN 420 of the previously determined specification. The display signal 430 is obtained by adding the processing result of the OOTF 415 to the output result of the DNN 420.
[0113] The display signal 430 is compared with the previously generated annotation signal 440, and the specification of the DNN 420 is changed according to a difference between the display signal 430 and the annotation signal 440. In this regard, the difference between the display signal 430 and the annotation signal 440 can be calculated as at least one of an L1 norm value, an L2 norm value, a structural similarity (SSIM) value, a peak signal-to-noise ratio-human visual system (PSNR-HVS) value, a multi-scale SSIM (MS-SSIM) value, a variance of illumination factor (VIF) value, or a video multi-method assessment fusion (VMAF) value.
[0114] The DNN determiner 216 can determine the difference between the display signal 430 and the annotation signal 440 while continuously changing the specification of the DNN 420, and determine the specification of the DNN 420 capable of minimizing the respective difference.
[0115] According to an embodiment, the DNN determiner 216 can determine the difference between the display signal 430 and the annotation signal 440 while continuously changing the parameters of the DNN 420 in the DNN 420 of a fixed structure, and determine the parameters of the DNN 420 capable of minimizing the respective difference. In this case, the DNN determiner 216 can determine different parameters for the DNN 420 of various structures. For example, the DNN determiner 216 can determine the parameters of the DNN 420 capable of minimizing the difference between the display signal 430 and the annotation signal 440 for the DNN 420 of a structure "a", and determine the parameters of the DNN 420 capable of minimizing the difference between the display signal 430 and the annotation signal 440 for the DNN 420 of a structure "b" different from the structure "a". Various parameters of the DNN 420 of various structures are determined to consider the performance of the display device 1600 outputting the display signal 430. This is described below with reference to FIGS. 6 to 8. Figure 8
[0116] The annotation signal 440 can be generated based on a result of processing the first light signal 410 according to the OOTF 415, and, for example, a manager or a user can change the luminance value of the signal converted from the first light signal 410 according to the OOTF 415 while monitoring the light signal converted from the first light signal 410 according to the OOTF 415 through a display. The annotation signal 440 can be obtained as a result of the luminance value change. In particular, when the light signal converted according to the OOTF 415 is displayed on the display and there is a portion having a low luminance value difficult to recognize, or based on the light signal converted according to the OOTF 415 being displayed on the display and there being a portion having a low luminance value difficult to recognize, the luminance value of the corresponding portion can be increased to generate the annotation signal 440 easily recognizable in general.
[0117] The method of determining the annotation signal 440 will now be described in detail. The annotation signal 440 can be determined based on various types of displays. Since the performance of the display apparatus 1600 can vary, the annotation signal 440 is determined by considering various types of displays. Accordingly, the annotation signal 440 can be determined for each type of display, thereby determining the specification of the DNN 420 for each annotation signal 440.
[0118] For example, the manager or the user can change the luminance value of the signal converted from the first light signal 410 according to the OOTF 415 while monitoring the light signal converted from the first light signal 410 according to the OOTF 415 through the display "A". Accordingly, the annotation signal 440 corresponding to the display "A" is determined. In addition, the manager or the user can change the luminance value of the signal converted from the first light signal 410 according to the OOTF 415 while monitoring the light signal converted from the first light signal 410 according to the OOTF 415 through the display "B". Accordingly, the annotation signal 440 corresponding to the display "B" is determined.
[0119] The DNN determiner 216 can determine the difference between the display signal 430 and the annotation signal 440 corresponding to the display "A" while changing the specification of the DNN 420, and determine the specification of the DNN 420 capable of minimizing the corresponding difference. In addition, the DNN determiner 216 can determine the difference between the display signal 430 and the annotation signal 440 corresponding to the display "B" while changing the specification of the DNN 420, and determine the specification of the DNN 420 capable of minimizing the corresponding difference.
[0120] The display for determining the annotation signal can display different ranges of luminance values. For example, the display "A" can display a range of luminance values of 0.001 nit to 800 nit, and the display "B" can display a range of luminance values of 0.001 nit to 1000 nit.
[0121] As described below with reference to Figure 8 When determining the specification of the DNN 420 for various types of displays, the DNN determiner 216 can identify the performance of the display apparatus 1600 outputting the display signal, and transmit AI meta information indicating the specification of the DNN 420 determined based on a display having a similar performance to the identified performance to the display apparatus 1600.
[0122] As described above, the OOTF is used to convert any one luminance value before tone mapping 1:1 into another luminance value, but because the luminance value of the light signal located in the surrounding environment is not considered in the 1:1 conversion, the improvement in the quality of the image is limited. Accordingly, according to an embodiment, by determining the best quality of the labeled signal 440, and then determining the specifications of the DNN 420 capable of generating a display signal 430 similar to the labeled signal 440, not only the tone mapping of the 1:1 conversion scheme can be performed, but also the AI-based tone mapping considering the luminance value of the light signal located in the surrounding environment can be performed.
[0123] Figure 5 A method of determining specifications of a DNN 520 performed by the image providing apparatus 200 according to another embodiment is illustrated.
[0124] Referring to Figure 5 , the first light signal 510 corresponding to the first digital image is converted according to the OOTF 515. Here, the specifications of the OOTF 515 are determined by the OOTF determiner 215. When the first light signal 510 corresponding to the first digital image is not stored in the image providing apparatus 200, or based on the first light signal 510 corresponding to the first digital image is not stored in the image providing apparatus 200, the DNN determiner 216 converts the first digital image according to the EOTF.
[0125] In addition, the first light signal 510 is converted into a first intermediate image according to the OETF 550. The first intermediate image is processed by the DNN 520 of the previously determined specifications. A second intermediate image is obtained as a processing result of the DNN 520. The first intermediate image can be the first digital image, and according to an implementation example, the conversion operation of the OETF 550 can be omitted, and the first digital image can be input to the DNN 520.
[0126] The second intermediate image is converted into a light signal according to the EOTF 560, and a display signal 530 is obtained by adding the signal converted according to the EOTF 560 and the signal converted according to the OOTF 515. The display signal 530 is compared with the previously generated labeled signal 540, and the specifications of the DNN 520 are changed or determined according to the difference between the display signal 530 and the labeled signal 540.
[0127] The DNN determiner 216 can determine the difference between the display signal 530 and the labeled signal 540 while continuously changing the specifications of the DNN 520, and determine the specifications of the DNN 520 capable of minimizing the corresponding difference.
[0128] When comparing Figure 4 and Figure 5In the DNN specification determination method shown in FIG. 6, the first light signal 610 having a linear luminance value is processed by the DNN 620 in FIG. 6, but the first intermediate image having a non-linear luminance value is processed by the DNN 520 in FIG. 5. Figure 4 In the DNN specification determination method shown in FIG. 6, the first light signal 610 having a linear luminance value is processed by the DNN 620 in FIG. 6, but the first intermediate image having a non-linear luminance value is processed by the DNN 520 in FIG. 5. Figure 5
[0129] As described above, the annotation signal 540 can be determined based on various types of displays, and in this case, various specifications of the DNN 520 suitable for various types of displays can be determined. In addition, when the specification of the DNN 520 is determined, the DNN 520 of a fixed structure can be utilized to determine the parameters of the DNN 520 that minimize the difference between the display signal 530 and the annotation signal 540.
[0130] Figure 6 A method of determining a specification of a DNN 620 performed by the image providing apparatus 200 according to another embodiment is shown.
[0131] Referring to FIG. 6, Figure 6 , the first light signal 610 corresponding to the first digital image is processed according to the OOTF 615, and a display signal 630 is obtained by processing the result of the processing by the DNN 620 of the previously determined specification. When the first light signal 610 corresponding to the first digital image is not stored in the image providing apparatus 200, the DNN determiner 216 converts the first digital image into the first light signal 610 according to the EOTF.
[0132] According to an embodiment, OOTF meta information can also be input to the DNN 620 together with the light signal converted from the first light signal 610 according to the OOTF 615. The OOTF meta information is determined according to the characteristics of the first digital image. Accordingly, when the light signal is processed, the DNN 620 can process the light signal according to the characteristics of the first digital image by considering the input OOTF meta information together.
[0133] The display signal 630 is compared with the annotation signal 640, and the specification of the DNN 620 is changed or determined according to the difference between the display signal 630 and the annotation signal 640. The DNN determiner 216 can determine the difference between the display signal 630 and the annotation signal 640 while continuously changing the specification of the DNN 620, and determine the specification of the DNN 620 that can minimize the corresponding difference.
[0134] As described above, the annotation signal 640 can be determined based on various types of displays, and in this case, DNNs 620 of various specifications suitable for various types of displays can be determined. In addition, when the specifications of the DNNs 620 are determined, the parameters of the DNNs 620 that minimize the difference between the display signal 630 and the annotation signal 640 can be determined with the DNNs 620 of fixed structures.
[0135] Figure 7 A method of determining a specification of a DNN 720 performed by the image providing apparatus 200 according to another embodiment is illustrated.
[0136] Referring to Figure 7 , the first optical signal 710 corresponding to the first digital image is processed by the DNN 720 of the previously determined specification. Both the first optical signal 710 and the OOTF meta information can be input to the DNN 720. The signal output from the DNN 720 is processed according to the OOTF 715, and the display signal 730 is obtained as a processing result. When the first optical signal 710 corresponding to the first digital image is not stored in the image providing apparatus 200, or based on the first optical signal 710 corresponding to the first digital image is not stored in the image providing apparatus 200, the DNN determiner 216 converts the first digital image to the first optical signal 710 according to the EOTF.
[0137] The display signal 730 is compared with the annotation signal 740, and the specification of the DNN 720 is changed or determined according to the difference between the display signal 730 and the annotation signal 740. The DNN determiner 216 can determine the difference between the display signal 730 and the annotation signal 740 while continuously changing the specification of the DNN 720, and determine the specification of the DNN 720 that can minimize the corresponding difference.
[0138] As described above, the annotation signal 740 can be determined based on various types of displays, and in this case, DNNs 720 of various specifications suitable for various types of displays can be determined. In addition, when the specifications of the DNNs 720 are determined, or based on the specifications of the DNNs 720 are determined, the parameters of the DNNs 720 that minimize the difference between the display signal 730 and the annotation signal 740 can be determined with the DNNs 720 of fixed structures.
[0139] When the specification of the DNN is determined, or based on the specification of the DNN being determined, the DNN determiner 216 can set a constraint condition of the DNN according to a characteristic of the first digital image checked from the pixel value of the first digital image. The constraint condition of the DNN can include at least one of a minimum number of layers included in the DNN, a maximum number of layers included in the DNN, a minimum size of a filter kernel used in at least one layer, a maximum size of a filter kernel used in at least one layer, a minimum number of filter kernels used in at least one layer, or a maximum number of filter kernels used in at least one layer. The characteristic of the first digital image can be determined by a maximum luminance value, an average luminance value, a variance of luminance values, or a luminance value corresponding to a percentile of a specific value of the first digital image.
[0140] When the constraint condition is set, the DNN determiner 216 can determine a DNN of a specification that minimizes a difference between a display signal and a labeled signal within a range satisfying the constraint condition. In other words, when a minimum number of layers included in the DNN is determined to be 3, the DNN determiner 216 can determine a specification of the DNN to include three or more layers as a DNN for tone mapping.
[0141] When a range of luminance values of the first digital image is large or a distribution thereof is complex, the DNN determiner 216 can determine at least one of a minimum number of layers included in the DNN, a minimum size of a filter kernel used in at least one layer, or a minimum number of filter kernels used in at least one layer to be greater than when a range of luminance values of the first digital image is small or a distribution thereof is simple.
[0142] For example, when a difference between an average luminance value and a maximum luminance value of the first digital image is greater than or equal to a previously determined value, the DNN determiner 216 can determine a minimum size of a filter kernel to be 5x5 and a minimum number of layers to be 5, and when the difference between the average luminance value and the maximum luminance value of the first digital image is less than the previously determined value, the DNN determiner 216 can determine the minimum size of the filter kernel to be 3x3 and the minimum number of layers to be 3.
[0143] As another example, when a variance of luminance values of the first digital image is greater than or equal to a previously determined value, the DNN determiner 216 can determine a minimum size of a filter kernel to be 5x5 and a minimum number of layers to be 5. Otherwise, when the variance of luminance values of the first digital image is less than the previously determined value, the DNN determiner 216 can determine the minimum size of the filter kernel to be 3x3 and the minimum number of layers to be 3.
[0144] As another example, when a difference between a luminance value corresponding to an a-th percentile (where a is a rational number) in the first digital image and an average luminance value of the first digital image is greater than or equal to a previously determined value, the DNN determiner 216 can determine a minimum size of a filter kernel as 5x5 and a minimum number of layers as 5. Otherwise, when the difference between the luminance value corresponding to the a-th percentile in the first digital image and the average luminance value of the first digital image is less than the previously determined value, the DNN determiner 216 can determine the minimum size of the filter kernel as 3x3 and the minimum number of layers as 3. The luminance value corresponding to the a-th percentile indicates a luminance value when the number of luminance values less than the luminance value corresponding to the percentile exists for a of the entire luminance values.
[0145] Referring back to Figure 2 When the specification of the DNN for tone mapping is determined, the DNN determiner 216 generates AI meta information indicating the determined specification of the DNN. For example, the AI meta information can include information on at least one of the number of layers, the type of layer, the number of filter kernels used in at least one layer, the size of the filter kernel used in at least one layer, or the weight or bias value of the filter kernel used in at least one layer.
[0146] When a transmission request for the first digital image is received from the display device 1600, or based on the transmission request for the first digital image being received from the display device 1600, the DNN determiner 216 transmits, to the display device 1600 through the transmitter 230, AI meta information indicating the DNN specification determined with respect to the first digital image.
[0147] As described above, the DNN determiner 216 can determine a plurality of DNN specifications for tone mapping of a light signal corresponding to the first digital image. In this case, the DNN determiner 216 can select any one DNN specification from among the plurality of DNN specifications in response to the transmission request for the first digital image, and transmit AI meta information indicating the selected DNN specification to the transmitter 230. Herein, the plurality of DNN specifications can be different from each other. When any one DNN specification is selected from among the plurality of DNN specifications, the DNN determiner 216 can consider the performance of the display device 1600 that has requested the first digital image.
[0148] Figure 8 is a table indicating various specifications of DNNs determined by the DNN determiner 216.
[0149] As Figure 8As illustrated, the DNNs of various specifications determined by the DNN determiner 216 can be classified according to the type of the display used to determine the specifications of the DNNs. For example, based on the display "A", the specifications of a K1 DNN and the specifications of a K2 DNN are determined, where the structure of the K1 DNN includes 4 layers and 20 filter kernels, and the structure of the K2 DNN includes 2 layers and 6 filter kernels. That is, the K1 DNN and the K2 DNN are determined based on the same type of display, but have different structures.
[0150] When the display device 1600 requests the image providing device 200 to transmit the first digital image, the display device 1600 can transmit the performance information of the display device 1600 to the image providing device 200. The performance information of the display device 1600 is information from which the performance of the display device 1600 is confirmed, and can include information about, for example, the manufacturer and model of the display device 1600.
[0151] When the performance of the display device 1600 is confirmed, the DNN determiner 216 can select a DNN specification determined based on a display having a function similar to the function of the display device 1600 from among the plurality of DNN specifications, and transmit AI meta information indicating the selected DNN specification to the data processor 232. Specifically, when the performance of the display device 1600 corresponds to the performance of the display "A", the DNN determiner 216 can transmit A1 meta information indicating the specifications of the K1 DNN or the K2 DNN to the data processor 232. Here, the performance of the display device 1600 corresponding to the performance of the display "A" can mean that the range of luminance values that the display device 1600 can represent is greater than or equal to the range of luminance values that the display "A" can represent.
[0152] Also, the DNN determiner 216 can transmit AI meta information indicating a specification of a DNN implementable by the display device 1600 having a structure of a K1 DNN or a K2 DNN to the data processor 232. Due to a computational load of a DNN including a large number of layers or using a large number of kernels, the DNN can not be operable in the low-performance display device 1600. In this case, even if AI meta information of a DNN including a large number of layers or using a large number of kernels is transmitted to the display device 1600, the display device 1600 can not implement the DNN confirmed from the AI meta information, and thus, the display device 1600 can not perform AI-based tone mapping. Accordingly, the DNN determiner 216 checks a performance of the display device 1600, selects a specification of an operable DNN based on the checked performance of the display device 1600, and provides AI meta information indicating the selected DNN specification to the data processor 232. Here, the performance of the display device 1600 checked by the DNN determiner 216 can include a performance related to at least one of a calculation speed and a calculation amount of the display device 1600, such as a processing speed of a CPU and a memory size. For example, when a display device 1600 corresponding to an "A" display cannot operate a DNN including more than 2 layers, the DNN determiner 216 transmits A1 meta information indicating a specification of a K2 DNN to the data processor 232.
[0153] According to an embodiment, the AI meta information can include information indicating whether a DNN-based tone mapping process is necessary. The information indicating whether the DNN-based tone mapping process is necessary can include a flag. The DNN determiner 216 can determine whether the display device 1600 needs or will perform the DNN-based tone mapping by considering a performance of the display device 1600.
[0154] For example, when a difference between a maximum luminance value representable by the display device 1600 and a threshold value is a certain value or more, the DNN determiner 216 can determine that the DNN-based tone mapping process is necessary. Otherwise, when the difference between the maximum luminance value representable by the display device 1600 and the threshold value is less than the certain value, the DNN determiner 216 can determine that the DNN-based tone mapping process is not necessary. Here, the threshold value can be a maximum luminance value representable by a main display for image analysis.
[0155] When a difference between a maximum luminance value of the main display and a maximum luminance value of the display device 1600 is not large, the DNN determiner 216 can determine that the DNN-based tone mapping process is not necessary, and the display device 1600 can perform only the OOTF-based tone mapping process according to the AI meta information including information indicating that the DNN-based tone mapping process is not necessary.
[0156] The reason why the maximum luminance value representable by the main display is compared with the maximum luminance value representable by the display apparatus 1600 is because when the manager or the user determines the OOTF having the best specification while watching the light signal according to the preset OOTF tone mapping by using the main display, the display apparatus 1600 having similar performance to the main display can reproduce an image of excellent quality only with the tone mapping based on the OOTF.
[0157] Referring back to Figure 2 , the data processor 232 obtains display data having a specific format by processing at least one of the encoded data or the meta information. The AI display data obtained by the data processor 232 is described below with reference to Figure 10 and Figure 12 .
[0158] The communication interface 234 transmits the AI display data to the display apparatus 1600 through a network. In this context, the network can include a wired network and / or a wireless network.
[0159] According to an embodiment, the AI display data obtained as a result of the processing of the data processor 232 can be stored in a data storage medium including a magnetic medium such as a hard disk, a floppy disk, and a magnetic tape, an optical recording medium such as a compact disk read only memory (CD-ROM) and a digital versatile disk (DVD), a magneto-optical medium such as an optical floppy disk, or the like.
[0160] According to an embodiment, the image providing apparatus 200 can transmit only the meta information and the meta information (e.g., AI meta information or AI meta information and OOTF meta information) among the encoded image data to the display apparatus 1600. In this case, the encoder 212 shown in Figure 2 may be omitted from the image providing apparatus 200. The display apparatus 1600 can receive the AI meta information from the image providing apparatus 200 and receive the encoded data of the first digital image from another apparatus (e.g., a server). In addition, the display apparatus 1600 can obtain a display signal by performing AI and OOTF-based tone mapping on a second light signal corresponding to a second digital image.
[0161] Figure 9 Frames constituting the first digital image 900 are shown.
[0162] As described above, the DNN determiner 216 determines the DNN specification for tone mapping based on the first digital image 900, and as Figure 9As illustrated, when the first digital image 900 includes a plurality of frames, the DNN determiner 216 can determine a DNN specification for each frame. Accordingly, a DNN specification for a first frame can be different from a DNN specification for a second frame. The DNN determiner 216 can process a first light signal corresponding to the first frame according to the OOTF and the DNN, and determine a DNN specification according to a difference between a display signal obtained as a result of the processing and a labeled signal. Thereafter, the DNN determiner 216 can process a first light signal corresponding to the second frame according to the OOTF and the DNN, and determine a DNN specification according to a difference between a display signal obtained as a result of the processing and a labeled signal.
[0163] According to an embodiment, the DNN determiner 216 can divide frames included in the first digital image 900 into a plurality of groups, and determine a DNN specification for each group. The DNN determiner 216 can divide frames included in the first digital image 900 into a first group 901 including frames t0 to ta-1, a second group 902 including frames ta to tb-1, and a third group 903 including frames tb to tn according to characteristics of the frames. In addition, the DNN determiner 216 can select a representative frame from each of the first group 901, the second group 902, and the third group 903, and determine a DNN specification corresponding to each group of the selected representative frame. That is, the DNN determiner 216 can process a first light signal corresponding to a representative frame of the first group 901 according to the OOTF and the DNN, and determine a DNN specification of the first group 901 according to a difference between a display signal obtained as a result of the processing and a labeled signal. In addition, the DNN determiner 216 can process a first light signal corresponding to a representative frame of the second group 902 according to the OOTF and the DNN, and determine a DNN specification of the second group 902 according to a difference between a display signal obtained as a result of the processing and a labeled signal. In addition, the DNN determiner 216 can process a first light signal corresponding to a representative frame of the third group 903 according to the OOTF and the DNN, and determine a DNN specification of the third group 903 according to a difference between a display signal obtained as a result of the processing and a labeled signal.
[0164] The DNN determiner 216 can classify frames having similar characteristics into the same group. Whether frames have similar characteristics can be determined based on a variance of luminance values of the frames and / or a histogram similarity of the luminance values. For example, frames whose variance of luminance values or histogram similarity belong to a certain range can be determined as the same group.
[0165] The DNN determiner 216 can classify frames in which a scene change occurs or initial frames to previous frames of frames in which a next scene change occurs into the same group.
[0166] Optionally, the DNN determiner 216 can determine a plurality of groups each including a previously determined (i.e., predetermined) number of frames continuous in time.
[0167] Referring to Figure 9 Although the frames included in the first group 901, the second group 902, and the third group 903 are continuous in time, the frames included in each group can not be continuous in time. For example, the first frame, the third frame, etc. can be determined as the first group, the second frame, the fifth frame, etc. can be determined as the second group, and the fourth frame, the sixth frame, etc. can be determined as the third group.
[0168] When the DNN specification is determined in the frame unit or the group unit of the first digital image, AI meta information indicating each of the determined DNN specifications is transmitted to the display device 1600 through the transmitter 230.
[0169] When it is necessary or will be transmitted to indicate various specifications of AI meta information by determining the DNN specification in the frame unit or the group unit of the first digital image, the DNN determiner 216 can generate AI meta information indicating the specification of the first DNN required or used for tone mapping. In addition, when AI meta information indicating the specification of the DNN after the first DNN is generated, the generated AI meta information can include difference information from the specification of the previous DNN. For example, when the first DNN includes 3 convolution layers and the second DNN includes 2 convolution layers, the AI meta information indicating the specification of the first DNN can include information indicating that the first DNN includes 3 convolution layers, and the AI meta information indicating the specification of the second DNN can include information indicating that one layer will be omitted from the first DNN.
[0170] Hereinafter, AI display data including encoded data and meta information is described in detail.
[0171] Figure 10 AI display data 1000 according to an embodiment is illustrated.
[0172] The AI display data 1000 including a single file can include AI meta information 1012 and encoded data 1032. Here, the AI display data 1000 can be included in a video file in a specific container format. The specific container format can be MPEG-4 Part 14 (MP4), Audio Video Interleave (AVI), Matroska Video (MKV), Flash Live Video (FLV), etc. The video file can include a metadata box 1010 and a media data box 1030.
[0173] Metadata box 1010 includes information about encoded data 1032 included in media data box 1030. For example, metadata box 1010 may include information about the type of the first digital image, the type of codec used to encode the first digital image, the playback time of the first digital image, etc. Additionally, metadata box 1010 may include AI metadata 1012. AI metadata 1012 may be encoded according to an encoding scheme provided in a specific container format and stored in metadata box 1010. Media data box 1030 may include encoded data 1032 generated according to the syntax of a specific image compression scheme. OOTF metadata may be included in metadata box 1010 along with AI metadata 1012, or may be included in media data box 1030.
[0174] AI metadata 1012 may include AI metadata for the first digital image, AI metadata for a group of frames, and AI metadata for a single frame. When a DNN of the same specification is determined for all frames included in the first digital image, the AI metadata for the group of frames and the AI metadata for a single frame can be omitted from the metadata box 1010. Optionally, when the specification of the DNN is determined for each group of frames of the first digital image, the AI metadata for the first digital image and the AI metadata for a single frame can be omitted from the metadata box 1010.
[0175] Figure 11 It shows that it includes Figure 10 The structure of AI metadata 1012 in the AI display data 1000 shown.
[0176] exist Figure 11 In the AI metadata 1012, AI_HDR_DNN_flag 1100 indicates whether DNN-based tone mapping processing is required. When AI_HDR_DNN_flag 1100 indicates that DNN-based tone mapping processing is required, information such as AI_HDR_num_layers 1105, AI_HDR_out_channel 1111, AI_HDR_in_channel 1112, and AI_HDR_filter_size 1113 can be included in the AI metadata 1012. Otherwise, when AI_HDR_DNN_flag 1100 indicates that DNN-based tone mapping is not required, information such as AI_HDR_num_layers 1105, AI_HDR_out_channel 1111, AI_HDR_in_channel 1112, and AI_HDR_filter_size 1113 may not be included in the AI metadata 1012.
[0177] The AI_HDR_num_layers 1105 indicates the number of layers included in the DNN for tone mapping.
[0178] In addition, the AI_HDR_out_channel 1111, the AI_HDR_in_channel 1112, the AI_HDR_filter_size 1113, the AI_HDR_weights 1114, and the AI_HDR_bias 1115 indicate the specifications of the first layer included in the DNN. Specifically, the AI_HDR_out_channel 1111 indicates the number of channels of data output from the first layer, and the AI_HDR_in_channel 1112 indicates the number of channels of data input to the first layer. In addition, the AI_HDR_filter_size 1113 indicates the size of a filter kernel used in the first layer, the AI_HDR_weights 1114 indicates the weight of the filter kernel used in the first layer, and the AI_HDR_bias 1115 indicates a bias value to be added to or subtracted from a result value of a specific formula in the first layer.
[0179] In addition, the AI_HDR_out_channel 1121, the AI_HDR_in_channel 1122, the AI_HDR_filter_size 1123, the AI_HDR_weights 1124, and the AI_HDR_bias 1125 indicate the specifications of the second layer included in the DNN. Specifically, the AI_HDR_out_channel 1121 indicates the number of channels of data output from the second layer, and the AI_HDR_in_channel 1122 indicates the number of channels of data input to the second layer. In addition, the AI_HDR_filter_size 1123 indicates the size of a filter kernel used in the second layer, the AI_HDR_weights 1124 indicates the weight of the filter kernel used in the second layer, and the AI_HDR_bias 1125 indicates a bias value to be added to or subtracted from a result value of a specific formula in the second layer.
[0180] In Figure 11According to the AI_HDR_num_layers 1105, AI_HDR_out_channel 1111, AI_HDR_out_channel 1121, AI_HDR_in_channel 1112, AI_HDR_in_channel 1122, AI_HDR_filter_size 1113, and AI_HDR_filter_size 1123, the structure of the DNN can be determined, and according to AI_HDR_weights 1114, AI_HDR_weights 1124, AI_HDR_bias 1115, and AI_HDR_bias 1125, the parameters of the DNN can be determined.
[0181] The AI_HDR_out_channel, AI_HDR_in_channel, AI_HDR_filter_size, AI_HDR_weights, and AI_HDR_bias indicating the specifications of each layer can exist as many as the number of layers confirmed from the AI_HDR_num_layers.
[0182] Figure 12 An AI display data 1200 according to another embodiment is illustrated.
[0183] Referring to Figure 12 The AI meta information 1234 can be included in the encoded data 1232. The video file can include a metadata box 1210 and a media data box 1230, and when the AI meta information 1234 is included in the encoded data 1232, the metadata box 1210 can not include the AI meta information 1234. The OOTF meta information can be included in the encoded data 1232 together with the AI meta information 1234, or in the metadata box 1210.
[0184] The media data box 1230 can include the encoded data 1232 including the AI meta information 1234. The AI meta information 1234 can be encoded according to a video codec used to encode the first digital image.
[0185] Since the AI meta information 1234 is included in the encoded data 1232, the AI meta information 1234 can be decoded according to the decoding order of the encoded data 1232.
[0186] The coded data 1232 includes video unit data (e.g., video parameter set) including information related to all frames included in the first digital image, frame group unit data (e.g., sequence parameter set) including information related to frames included in a group, frame unit data (e.g., picture parameter set) including information related to a single frame, etc. When the DNN of the same specification is determined for all frames in the first digital image, the AI meta information can be included in the video unit data. Alternatively, when the specification of the DNN is determined in a group unit, AI meta information indicating the specification of the DNN corresponding to each group can be included in the data of each frame group unit or in the frame unit data corresponding to the first frame of each group. When the specification of the DNN is determined in a group unit, the AI meta information corresponding to each group can include identification information (e.g., picture order count) of a frame using the corresponding AI meta information. This can be useful when the frames included in each group are not continuous in time.
[0187] When the specification of the DNN is determined in a frame unit, AI meta information indicating the specification of the DNN corresponding to each frame can be included in the data of each frame unit.
[0188] Figure 13 The structure of the AI meta information 1234 included in the AI display data 1200 according to an embodiment is illustrated. Figure 12 The structure of the AI meta information 1234 included in the AI display data 1200 according to an embodiment is illustrated.
[0189] As described above, because the coded data 1232 is generated according to the rules (e.g., syntax) of the image compression scheme using the frequency transform, the AI meta information 1234 can also be included in the coded data 1232 according to the syntax.
[0190] The AI meta information 1234 can be included in a video parameter set, a sequence parameter set, or a picture parameter set. Alternatively, the AI meta information 1234 can be included in a supplemental enhancement information (SEI) message. The SEI message includes additional information other than information required to restore the second digital image (e.g., prediction mode information, motion vector information, etc.). The SEI message includes a single network abstraction layer (NAL) unit, and can be transmitted in a frame group unit or a frame unit.
[0191] Referring to FIG. 12, Figure 13, AI_HDR_DNN_flag 1301 is included in the AI meta information 1234. The AI HDR DNN flag 1301 indicates whether DNN-based tone mapping processing is required (or whether DNN-based tone mapping processing will be performed). When the AI HDR DNN flag 1301 indicates that DNN-based tone mapping processing is required, AI HDR num layers 1303, AI HDR in channel[i] 1304, AI HDR out channel[i] 1305, AI HDR filter width[i] 1306, AI HDR filter height[i] 1307, AI HDR bias[i][j] 1308, and AI HDR weight[i][j][k][l] 1309 are included in the AI meta information 1234.
[0192] Otherwise, when the AI HDR DNN flag 1301 indicates that DNN-based tone mapping processing is not required (or will not be performed), AI HDR num layers 1303, AI HDR in channel[i] 1304, AI HDR out channel[i] 1305, AI HDR filter width[i] 1306, AI HDR filter height[i] 1307, AI HDR bias[i][j] 1308, and AI HDR weight[i][j][k][l] 1309 are not included in the AI meta information 1234.
[0193] The AI HDR num layers 1303 indicates the number of layers included in the DNN for tone mapping. In addition, the AI HDR out channel[i] 1305 indicates the number of channels of output data from the i-th layer, and the AI HDR in channel[i] 1304 indicates the number of channels of input data to the i-th layer. In addition, the AI HDR filter width[i] 1306 and the AI HDR filter height[i] 1307 indicate the width and height of a filter kernel used in the i-th layer, respectively.
[0194] In addition, AI_HDR_bias[i][j] 1308 indicates a bias value to be added to or subtracted from a result value of a specific formula for the output data of the jth channel of the ith layer, and AI_HDR_weight[i][j][k][l] 1309 indicates a weight of the lth sample in a filter kernel associated with the output data of the jth channel of the ith layer and the input data of the kth channel.
[0195] The parsing performed by the display device 1600 according to an embodiment is described below. Figure 13 The method of AI meta information 1234 illustrated.
[0196] Figure 14 is a flowchart of an image providing method according to an embodiment.
[0197] Referring to Figure 14 In operation S1410, the image providing device 200 determines a specification of a DNN corresponding to the first digital image. Specifically, the image providing device 200 determines a specification of a DNN to be used for tone mapping based on the first light signal corresponding to the first digital image. As described above, when the first digital image includes a plurality of frames, the image providing device 200 can determine a DNN specification for each frame or each group.
[0198] The image providing device 200 can determine a specification of an OOTF corresponding to the first digital image. When the first digital image includes a plurality of frames, the image providing device 200 can determine a specification of an OOTF for each frame, for each block divided from a frame, or for each group of frames. The same specification of an OOTF can be determined for all frames included in the first digital image.
[0199] In operation S1420, the image providing device 200 receives a transmission request for the first digital image from the display device 1600. The image providing device 200 can communicate with the display device 1600 through a wired / wireless network (e.g., the Internet).
[0200] In operation S1430, the image providing device 200 encodes the first digital image. The image providing device 200 can encode the first digital image through an image compression scheme based on frequency transformation.
[0201] In operation S1440, the image providing device 200 transmits encoded data of the first digital image and AI meta information indicating the specification of the DNN determined in operation S1410 to the display device 1600. The image providing device 200 can also transmit OOTF meta information to the display device 1600 together with the encoded data and the AI meta information.
[0202] As described above, the AI meta information can include information indicating whether DNN-based tone mapping is required. When it is determined that DNN-based tone mapping is not required, the image providing apparatus 200 generates AI meta information including information indicating that DNN-based tone mapping is not required. Otherwise, when it is determined that DNN-based tone mapping is required, the image providing apparatus 200 generates AI meta information including information indicating that DNN-based tone mapping is required. When the AI meta information includes information indicating that DNN-based tone mapping is not required, information indicating the specification of the DNN determined in operation S1410 can not be included in the AI meta information.
[0203] Figure 15 is a flowchart of an image providing method according to another embodiment.
[0204] Referring to Figure 15 In operation S1510, the image providing apparatus 200 determines the specifications of a plurality of DNNs corresponding to the first digital image. The plurality of DNNs having various specifications can have different structures and / or different parameters. For example, a first DNN can include 4 layers, and a second DNN can include 3 layers. As another example, both the first DNN and the second DNN can include 4 convolution layers, in which the number of filter kernels used in the convolution layers of the first DNN is 3, and the number of filter kernels used in the convolution layers of the second DNN is 4. The specifications of the plurality of DNNs can be determined based on different types of displays, respectively. That is, when the annotation signals are determined for each of the different types of displays, various specifications of the DNNs capable of generating a display signal having the smallest difference from the annotation signal can be determined. Alternatively, various specifications of DNNs having different structures can be determined based on any one type of display.
[0205] When the first digital image includes a plurality of frames, the image providing apparatus 200 can determine various specifications of the DNNs in frame units or group units, respectively.
[0206] The image providing apparatus 200 can determine the specification of an OOTF corresponding to the first digital image. When the first digital image includes a plurality of frames, the image providing apparatus 200 can determine the specification of the OOTF for each frame, for each block divided from the frame, or for each group. The OOTF of the same specification can be determined for all frames included in the first digital image.
[0207] The image providing apparatus 200 receives a transmission request for the first digital image and performance information of the display apparatus 1600 from the display apparatus 1600 at operation S1520. The image providing apparatus 200 can communicate with the display apparatus 1600 through a wired / wireless network (e.g., the Internet). The performance information of the display apparatus 1600 is information from which the performance of the display is confirmed, and can include, for example, manufacturer information and model information of the display apparatus 1600.
[0208] The image providing apparatus 200 selects a DNN specification determined based on a display having a similar performance to the performance of the display apparatus 1600 from among specifications of a plurality of DNNs by considering the performance of the display apparatus 1600 at operation S1530. When a plurality of DNN specifications are determined based on a display having a similar performance to the performance of the display apparatus 1600, the image providing apparatus 200 selects a DNN specification having a structure that can be implemented by the display apparatus 1600 from among the plurality of DNN specifications.
[0209] The image providing apparatus 200 encodes the first digital image at operation S1540. The image providing apparatus 200 can encode the first digital image through an image compression scheme based on frequency transformation.
[0210] The image providing apparatus 200 transmits encoded data of the first digital image and AI meta information indicating the specification of the DNN selected at operation S1530 to the display apparatus 1600 at operation S1550. The image providing apparatus 200 can also transmit OOTF meta information together with the encoded data and the AI meta information to the display apparatus 1600.
[0211] Figure 16 is a block diagram of a display apparatus 1600 according to an embodiment.
[0212] Referring to Figure 16 The display apparatus 1600 according to an embodiment can include a receiver 1610, an image processor 1630, and a display 1650. The receiver 1610 can include a communication interface 1612, a parser 1614, and an output unit 1616, and the image processor 1630 can include a decoder 1632 and a converter 1634.
[0213] Although Figure 16 The receiver 1610 is shown as being separate from the image processor 1630, but the receiver 1610 and the image processor 1630 can be implemented by a single processor. In this case, the single processor can be implemented by a dedicated (or proprietary) processor or by a combination of a general-purpose processor such as an AP, a CPU, or a GPU and S / W. Also, the dedicated processor can be implemented with a memory for implementing the embodiments or with a memory processor using an external memory.
[0214] In addition, the receiver 1610 and the image processor 1630 can be implemented by a plurality of processors. In this case, the plurality of processors can be implemented by a combination of a plurality of general-purpose processors such as an AP, a CPU, and a GPU and S / W.
[0215] The display 1650 can include various types of displays capable of outputting a display signal, such as a liquid crystal display (LCD), a light emitting diode (LED) display, an organic light emitting diode (OLED) display, and a quantum dot light emitting diode (QLED) display.
[0216] Although Figure 16 Although it is shown that the display apparatus 1600 includes all of the receiver 1610, the image processor 1630, and the display 1650, according to an embodiment, the display apparatus 1600 can include only the receiver 1610 and the image processor 1630, and the display apparatus 1600 (e.g., an image processing device) can transmit the tone-mapped display signal to a separate display.
[0217] The receiver 1610 receives and parses AI display data, and transmits encoded data and meta information to the image processor 1630, respectively.
[0218] In particular, the communication interface 1612 receives AI display data through a network. The AI display data includes encoded data and meta information. The meta information includes OOTF meta information and AI meta information. The encoded data and the meta information can be received through a homogeneous network or a heterogeneous network. At least one of the AI meta information or the OOTF meta information can be included in the encoded data. According to an embodiment, the communication interface 1612 can receive the meta information from the image providing apparatus 200, and receive the encoded data from another device (e.g., a server). Alternatively, the communication interface 1612 can receive the AI meta information from the image providing apparatus 200, and receive the encoded data and the OOTF meta information from another device (e.g., a server).
[0219] According to an embodiment, the communication interface 1612 can transmit a transmission request message for a first digital image to the image providing apparatus 200 to receive AI display data. In this case, the communication interface 1612 can further transmit performance information of the display apparatus 1600 to the image providing apparatus 200. The performance information of the display apparatus 1600 can include manufacturer information and model information of the display apparatus 1600.
[0220] The parser 1614 receives the AI display data received through the communication interface 1612 and parses the AI display data to separate the encoded data and the meta information. For example, a header of data obtained by the communication interface 1612 is read to identify whether the data is the encoded data or the meta information. Here, as an example, the parser 1614 separates the encoded data from the meta information based on the header of the data received through the communication interface 1612 and transmits the encoded data and the meta information to the output unit 1616, and the output unit 1616 transmits the encoded data and the meta information to the decoder 1632 and the converter 1634, respectively. In this case, the parser 1614 can check which codec (e.g., MPEG-2, H.264, MPEG-4, HEVC, VC-1, VP8, VP9, AV1, or the like) is used to generate the encoded data. The parser 1614 can transmit the corresponding information to the decoder 1632 through the output unit 1616 so that the encoded data is processed using the checked codec.
[0221] When both the AI meta information and the OOTF meta information are included in the encoded data, the parser 1614 can transmit the encoded data including the AI meta information and the OOTF meta information to the decoder 1632.
[0222] As Figure 10 and Figure 11As shown, when the AI meta information 1012 is included in the metadata box 1010 and the encoded data 1032 is included in the media data box 1030, the parser 1614 can extract the AI meta information 1012 included in the metadata box 1010 and transmit the AI meta information 1012 to the converter 1634, and extract the encoded data 1032 included in the media data box 1030 and transmit the encoded data 1032 to the decoder 1632. Specifically, the parser 1614 extracts the AI_HDR_DNN_flag 1100, the AI_HDR_num_layers 1105, the AI_HDR_in_channel 1112, the AI_HDR_in_channel 1122, the AI_HDR_out_channel 1111, the AI_HDR_out_channel 1121, the AI_HDR_filter_size 1113, the AI_HDR_filter_size 1123, the AI_HDR_bias 1115, the AI_HDR_bias 1125, the AI_HDR_weights 1114, and the AI_HDR_weights 1124, and provides the AI_HDR_DNN_flag 1100, the AI_HDR_num_layers 1105, the AI_HDR_in_channel 1112, the AI_HDR_in_channel 1122, the AI_HDR_out_channel 1111, the AI_HDR_out_channel 1121, the AI_HDR_filter_size 1113, the AI_HDR_filter_size 1123, the AI_HDR_bias 1115, the AI_HDR_bias 1125, the AI_HDR_weights 1114, and the AI_HDR_weights 1124 to the converter 1634.
[0223] In addition, as shown in FIG. 12, when the AI meta information 1234 is included in the encoded data 1232, the parser 1614 can extract the encoded data 1232 included in the media data box 1230 and transmit the encoded data 1232 to the decoder 1632. Figure 12
[0224] According to an embodiment, the AI display data parsed by the parser 1614 can be obtained from a data storage medium including a magnetic medium such as a hard disk, a floppy disk, and a magnetic tape, an optical recording medium such as a CD-ROM and a DVD, a magneto-optical medium such as an optical floppy disk, or the like.
[0225] The decoder 1632 recovers a second digital image corresponding to the first digital image based on the encoded data. The decoder 1632 recovers the second digital image according to an image decompression scheme based on a frequency transform. The second digital image obtained by the decoder 1632 is provided to the converter 1634.
[0226] When the encoded data includes the AI meta information and / or the OOTF meta information, the decoder 1632 decodes the AI meta information and / or the OOTF meta information included in the encoded data, and provides the decoded AI meta information and / or the OOTF meta information to the converter 1634.
[0227] When the encoded data includes the AI meta information, a method of parsing the AI meta information performed by the decoder 1632 can be performed as described with reference to Figure 13
[0228] The decoder 1632 extracts the AI_HDR_DNN_flag 1301 included in the encoded data. The AI_HDR_DNN_flag 1301 indicates whether DNN-based tone mapping is required. When the AI_HDR_DNN_flag 1301 indicates that DNN-based tone mapping is not required, the decoder 1632 stops parsing the AI meta information, and provides information indicating that DNN-based tone mapping is not required or will not be performed to the converter 1634.
[0229] When the AI_HDR_DNN_flag 1301 indicates that DNN-based tone mapping is required, the decoder 1632 extracts the AI_HDR_num_layers 1303 from the encoded data. The AI_HDR_num_layers 1303 indicates the number of layers included in the DNN used for tone mapping.
[0230] The decoder 1632 extracts the AI_HDR_in_channel[i] 1304, the AI_HDR_out_channel[i] 1305, the AI_HDR_filter_width[i] 1306, and the AI_HDR_filter_height[i] 1307 as many as the number of layers included in the DNN. The AI_HDR_out_channel[i] 1305 indicates the number of channels of output data from the i-th layer, and the AI_HDR_in_channel[i] 1304 indicates the number of channels of input data to the i-th layer. In addition, the AI_HDR_filter_width[i] 1306 and the AI_HDR_filter_height[i] 1307 indicate the width and the height of a filter kernel used in the i-th layer, respectively.
[0231] Thereafter, the decoder 1632 extracts AI_HDR_bias[i][j] 1308 as many times as the number of output channels of the i-th layer included in the DNN. AI_HDR_bias[i][j] 1308 indicates a bias value to be added to or subtracted from a result value of a specific formula to be applied to output data of the j-th channel of the i-th layer.
[0232] The decoder 1632 extracts AI_HDR_weight[i][j][k][l] 1309 as many times as the number of input channels of the i-th layer x the number of output channels of the i-th layer x the width of a filter kernel used in the i-th layer x the height of a filter kernel used in the i-th layer. AI_HDR_weight[i][j][k][l] 1309 indicates a weight of the l-th sample in a filter kernel associated with output data of the j-th channel of the i-th layer and input data of the k-th channel.
[0233] The decoder 1632 provides AI_HDR_in_channel[i] 1304, AI_HDR_out_channel[i] 1305, AI_HDR_filter_width[i] 1306, AI_HDR_filter_height[i] 1307, AI_HDR_bias[i][j] 1308, and AI_HDR_weight[i][j][k][l] 1309 extracted from the encoded data to the converter 1634 as AI meta information.
[0234] As described above, when the AI meta information and / or the OOTF meta information are included in the SEI message of the encoded data, the decoder 1632 can transmit the SEI message to the converter 1634, and the converter 1634 can obtain the AI meta information and / or the OOTF meta information from the SEI message. For example, size information of the SEI message can be stored in a header of the encoded data, and the decoder 1632 can check the size of the SEI message from the header, extract the SEI message of the checked size from the encoded data, and transmit the SEI message to the converter 1634.
[0235] The operation of parsing the AI meta information from the SEI message by the converter 1634 is the same as or similar to the operation of parsing the AI meta information from the encoded data by the decoder 1632, and thus a redundant description thereof is omitted here.
[0236] The converter 1634 sets the OOTF based on the meta information (specifically, the OOTF meta information), and sets the HDR DNN based on the AI meta information. The OOTF can have a specification indicated by the OOTF meta information, and the HDR DNN can have a specification indicated by the AI meta information. In addition, the converter 1634 performs tone mapping on a second optical signal corresponding to the second digital image by using the OOTF and the HDR DNN, thereby obtaining a display signal.
[0237] When the second digital image includes a plurality of frames, the converter 1634 can obtain a plurality of pieces of AI meta information corresponding to the respective frames. In addition, the converter 1634 can independently set the HDR DNN for the respective frames based on the plurality of pieces of AI meta information. In this case, the specification of the HDR DNN set for any one frame can be different from the specification of the HDR DNN set for another frame.
[0238] Alternatively, when the second digital image includes a plurality of frames, the converter 1634 can obtain AI meta information corresponding to each frame group. In addition, the converter 1634 can independently set the HDR DNN for the respective frame groups based on the plurality of pieces of AI meta information. In this case, the specification of the HDR DNN set for any one frame group can be different from the specification of the HDR DNN set for another frame group. The AI meta information corresponding to each frame group can include identification information (e.g., a picture order count or an index number) of frames to which the AI meta information is applied. This can be useful when the frames included in each group are not continuous in time.
[0239] According to an implementation example, the converter 1634 can divide the frames included in the second digital image into a plurality of groups according to the characteristics of the frames, and set the HDR DNN for each group by using the AI meta information sequentially provided from the image providing apparatus 200. In this case, the AI meta information can not include identification information of frames to which the AI meta information is applied, but the converter 1634 must divide the frames into groups based on the same standard as in the image providing apparatus 200.
[0240] When the second digital image includes a plurality of frames, the converter 1634 can obtain AI meta information corresponding to all of the frames. In addition, the converter 1634 can set the HDR DNN for all of the frames based on the AI meta information.
[0241] Hereinafter, the operation of the converter 1634 will be described with reference to the following drawings. Figures 17 to 20 The operation of the converter 1634 after setting the OOTF and the HDR DNN based on the OOTF meta information and the AI meta information will be described.
[0242] Figure 17 The operation of the converter 1634 after setting the OOTF and the HDR DNN based on the OOTF meta information and the AI meta information will be described.
[0243] Referring toFigure 17 The second optical signal 1710 corresponding to the second digital image is converted according to the OOTF 1715. Here, the OOTF 1715 is set based on the OOTF meta information. The second digital image is converted into the second optical signal 1710 according to the EOTF.
[0244] The second optical signal 1710 is processed by the HDR DNN 1720 set based on the AI meta information. The display signal 1730 is obtained by adding the result processed by the OOTF 1715 to the output result of the HDR DNN 1720.
[0245] Figure 18 A tone mapping operation performed by the display apparatus 1600 according to another embodiment is illustrated.
[0246] Referring to Figure 18 The second optical signal 1810 corresponding to the second digital image is converted according to the OOTF 1815. The OOTF 1815 is set based on the OOTF meta information.
[0247] In addition, the second optical signal 1810 is converted into a first intermediate image according to the OETF 1850. The first intermediate image is processed by the HDR DNN 1820 set based on the AI meta information. As a result of the processing of the HDR DNN 1820, a second intermediate image is obtained. The first intermediate image can be the second digital image. In this case, the conversion operation of the OETF 1850 can be omitted, and the second digital image obtained in the decoding operation can be processed by the HDR DNN 1820.
[0248] The second intermediate image is converted into an optical signal according to the EOTF 1860, and the display signal 1830 is obtained by adding the signal converted according to the EOTF 1860 to the signal converted according to the OOTF 1815.
[0249] Figure 19 A tone mapping operation performed by the display apparatus 1600 according to another embodiment is illustrated.
[0250] Referring to Figure 19 The second optical signal 1910 corresponding to the second digital image is processed according to the OOTF 1915, and the optical signal converted from the second optical signal 1910 according to the OOTF 1915 is processed by the HDR DNN 1920 set based on the AI meta information, thereby obtaining the display signal 1930.
[0251] According to an embodiment, OOTF meta information can also be input to the HDR DNN 1920. The OOTF meta information is determined according to characteristics of the first digital image corresponding to the second digital image, and thus the HDR DNN 1920 can also consider the input OOTF meta information to process the light signal.
[0252] Figure 20 A tone mapping operation performed by the display apparatus 1600 according to another embodiment is illustrated.
[0253] Referring to Figure 20 The second light signal 2010 corresponding to the second digital image is processed by the HDR DNN 2020 set based on the AI meta information. When the second light signal 2010 is input to the HDR DNN 2020, OOTF meta information can also be input to the HDR DNN 2020. The signal output from the HDR DNN 2020 is processed according to the OOTF 2015, and as a result of the processing, a display signal 2030 is obtained.
[0254] Figure 21 is a flowchart of a display method according to an embodiment.
[0255] Referring to Figure 21 In operation S2110, the display apparatus 1600 obtains encoding data of a first digital image and AI meta information. The display apparatus 1600 can also obtain OOTF meta information.
[0256] In operation S2120, the display apparatus 1600 obtains a second digital image by decoding the encoding data. When the AI meta information is included in the encoding data, the display apparatus 1600 can obtain the AI meta information by decoding the encoding data.
[0257] In operation S2130, the display apparatus 1600 obtains a second light signal converted from the second digital image according to an EOTF.
[0258] In operation S2140, the display apparatus 1600 determines whether DNN-based tone mapping processing is needed based on the AI meta information.
[0259] In operation S2150, when it is determined that the DNN-based tone mapping process is needed, the display device 1600 sets an HDR DNN for tone mapping based on the AI meta information. According to an embodiment, when the AI meta information is obtained in units of respective frames included in the second digital image, the display device 1600 can set an HDR DNN for tone mapping for each frame. Alternatively, when the AI meta information is obtained in units of groups of frames included in the second digital image, the display device 1600 can set an HDR DNN for tone mapping for each group. Alternatively, when the AI meta information is obtained in units of all frames included in the second digital image, the display device 1600 can set a single HDR DNN for tone mapping.
[0260] In operation S2160, the display device 1600 obtains a display signal by processing the second optical signal obtained in operation S2130 through an OOTF set based on the OOTF meta information and an HDR DNN set based on the AI meta information. The display signal is output as an image on the display 1650.
[0261] In operation S2170, when it is determined that the DNN-based tone mapping process is not needed, the display device 1600 obtains a display signal by processing the second optical signal through an OOTF set based on the OOTF meta information. The display signal is output as an image on the display 1650.
[0262] Figure 22 is a flowchart of a display method according to another embodiment.
[0263] Referring to Figure 22 In operation S2210, the display device 1600 obtains encoding data of frames of the first digital image, first AI meta information, and second AI meta information.
[0264] In operation S2220, the display device 1600 obtains frames of the second digital image by decoding the encoding data.
[0265] In operation S2230, the display device 1600 sets a first HDR DNN based on the first AI meta information and a second HDR DNN based on the second AI meta information.
[0266] In operation S2240, the display device 1600 obtains a display signal by processing a second optical signal corresponding to frames of a first group among the frames of the second digital image through the first HDR DNN and an OOTF. Further, in operation S2250, the display device 1600 obtains a display signal by processing a second optical signal corresponding to frames of a second group among the frames of the second digital image through the second HDR DNN and the OOTF.
[0267] The first AI meta information and the second AI meta information can include identification numbers of frames to which the first AI meta information and the second AI meta information are respectively applied. When the frames included in each group are continuous in time, the first AI meta information and the second AI meta information can respectively include identification numbers of first frames and last frames to which the first AI meta information and the second AI meta information are applied.
[0268] The display signal corresponding to the frames of the first group and the display signal corresponding to the frames of the second group are displayed as an image on the display 1650.
[0269] The above-described embodiments can be written or implemented as a computer executable program or code, and the program or code can be stored in a medium. Figure 10 The above-described embodiments can be written or implemented as a computer executable program or code, and the program or code can be stored in a medium. Figure 22 The above-described embodiments can be written or implemented as a computer executable program or code, and the program or code can be stored in a medium. Figure 12 The above-described embodiments can be written or implemented as a computer executable program or code, and the program or code can be stored in a medium.
[0270] The above-described embodiments can be written or implemented as a computer executable program or code, and the program or code can be stored in a medium.
[0271] The medium can continuously store the computer executable program or temporarily store the computer executable program for execution or download. In addition, the medium can include various recording devices or storage devices in the form of a single hardware or a combination of several hardware, and the medium is not limited to a medium directly connected to a specific computer system, but can be distributed over a network. Examples of the medium can include magnetic media (such as a hard disk, a floppy disk, and a magnetic tape), optical recording media (such as a CD-ROM and a DVD), magneto-optical media (such as a floptical disk, a ROM, a RAM, and a flash memory) configured to store program instructions. In addition, examples of other media can include an application store for distributing applications, other sites for supplying or distributing various S / W, and recording media and storage media managed by a server, etc.
[0272] The image providing apparatus and the image providing method thereof and the display apparatus and the display method thereof according to the embodiments can improve the quality of an image to be displayed through AI-based tone mapping.
[0273] However, the effects achieved by the image providing apparatus and the image providing method thereof and the display apparatus and the display method thereof according to the embodiments are not limited to the above description, and other effects not described can be clearly understood by those skilled in the art to which the present disclosure pertains.
[0274] Although the technical idea of the disclosure has been described in detail with reference to the embodiments, the technical idea of the disclosure is not limited to the above-described embodiments, and various modifications and changes can be made by those skilled in the art within the scope of the technical idea of the disclosure.
Claims
1. A display device comprising: a memory storing one or more instructions; and a processor configured to execute the stored one or more instructions to: obtain encoding data of a first digital image and artificial intelligence (AI) meta information indicating a specification of a deep neural network (DNN), obtain a second digital image corresponding to the first digital image by decoding the encoding data, obtain an optical signal converted from the second digital image according to a predetermined electro-optical transfer function (EOTF), convert the optical signal by using an optical-optical transfer function (OOTF), input the optical signal to a high dynamic range (HDR) DNN set according to the AI meta information, and obtain a display signal by adding a signal converted from the optical signal according to the OOTF to an output signal of the HDR DNN. 2.The display device of claim 1, wherein: the second digital image comprises a plurality of frames; and the processor is further configured to execute the stored one or more instructions to: obtain first AI meta information for frames in a first group among the plurality of frames and second AI meta information for frames in a second group among the plurality of frames, and independently set the HDR DNN for the frames in the first group according to the first AI meta information and independently set the HDR DNN for the frames in the second group according to the second AI meta information. 3.The display device of claim 2, wherein: the first AI meta information comprises first identification information of the frames to which the first AI meta information is applied; and the second AI meta information comprises second identification information of the frames to which the second AI meta information is applied. 4.An image providing device comprising: a memory storing one or more instructions; and a processor configured to execute the stored one or more instructions to: convert an optical signal corresponding to a first digital image by using an optical-optical transfer function (OOTF), input a first intermediate image converted from the optical signal according to an optical-electrical transfer function (OETF) to a high dynamic range (HDR) DNN of a predetermined specification, convert a second intermediate image output from the HDR DNN according to an electro-optical transfer function (EOTF), obtain a display signal by adding a signal converted from the optical signal according to the OOTF to a signal converted from the second intermediate image according to the EOTF, determine a specification of the DNN based on difference information between the display signal and a labeled signal, encode the first digital image, and transmit encoding data of the first digital image and artificial intelligence (AI) meta information indicating the determined specification of the DNN to a display device.
5. The image providing apparatus as claimed in claim 4, wherein The labeled signal is predetermined based on a signal converted from the optical signal according to the OOTF. 6.The image providing device of claim 4, wherein: the first digital image comprises a plurality of frames; and the processor is further configured to execute the stored one or more instructions to independently determine a first specification of the DNN for frames in a first group among the plurality of frames and a second specification of the DNN for frames in a second group among the plurality of frames.
7. The image providing apparatus as claimed in claim 6, wherein The processor is further configured to execute the stored one or more instructions to: determine a representative frame in a first group and a representative frame in a second group among the plurality of frames, determine a first specification of the DNN based on difference information between a result of processing a light signal corresponding to the representative frame in the first group by using the OOTF and the DNN and the labeled signal, and determine a second specification of the DNN based on difference information between a result of processing a light signal corresponding to the representative frame in the second group by using the OOTF and the DNN and the labeled signal.
8. The image providing apparatus as claimed in claim 4, wherein The processor is further configured to execute the stored one or more instructions to transmit information indicating whether DNN-based tone mapping processing is required to the display device according to a luminance value that the display device can display.
9. The image providing apparatus as claimed in claim 8, wherein The processor is further configured to execute the stored one or more instructions to transmit information indicating that DNN-based tone mapping processing is not required to the display device based on a difference between a maximum value of luminance that the display device can display and a threshold luminance value being less than or equal to a predetermined value. 10.A method of displaying an image performed by a display device, the method comprising: obtaining encoded data of a first digital image and artificial intelligence (AI) meta information indicating a specification of a deep neural network (DNN), obtaining a second digital image corresponding to the first digital image by decoding the encoded data, obtaining a light signal converted from the second digital image according to a predetermined electro-optical transfer function (EOTF), converting the light signal by using an optical-optical transfer function (OOTF), inputting the light signal to a high dynamic range (HDR) DNN set according to the AI meta information, and obtaining a display signal by adding a signal converted from the light signal according to the OOTF to an output signal of the HDR DNN. 11.A non-transitory computer-readable recording medium having stored therein a computer-readable program for executing a method of displaying an image performed by a display device according to claim 10.
Citation Information
Patent Citations
Methods, systems and apparatus for electro-optical and opto-electrical conversion of images and video
CN107431825A
HDR image representations using neural network mappings
WO2019199701A1