Electro-optical transfer function conversion and signal legitimization
By determining the backward shaping function for the target display in an electronic processor, the efficiency and accuracy of electro-optical transfer function conversion, signal legalization and backward shaping function in high-dynamic range video signal processing is solved, and a more efficient and accurate video rendering effect is achieved.
Patent Information
- Application Number
- CN202080054727.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-07-30
- Filing Date
- 2020-07-27
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2040-07-27
AI Technical Summary
When processing high dynamic range video signals, the prior art has problems with the efficiency and accuracy of electro-optical transfer function conversion, signal legalization and backward shaping functions, resulting in poor video rendering effect.
The backward shaping function for rendering video on the target display is determined by an electronic processor by determining the sample pixel set from the received video data, converting the sample pixel set to the electro-optical transfer function of the second color space according to the electro-optical transfer function of the first color space, and determining the backward shaping function by repeatedly applying and adjusting the sample backward shaping function to minimize the difference between predicted pixel values.
Improves the speed, efficiency and accuracy of video signal conversion, and improves the effects of HDR-TV image rendering and signal processing.
Smart Images

Figure CN114175647B_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 62 / 880,266, filed Jul. 30, 2019, and European Patent Application No. 19189052.4, filed Jul. 30, 2019, each of which is hereby incorporated by reference in its entirety. Background of the Invention Field of the Invention
[0004] This application generally relates to video signal conversion for high - dynamic range video (HDR).
[0005] Description of the Related Art
[0006] As used herein, the term "dynamic range (DR)" can refer to the ability of the human visual system (HVS) to perceive the intensity range (e.g., luminance, brightness) in an image, e.g., from the darkest black (dark tone) to the brightest white (highlight). In this sense, DR is related to "scene - referred" intensity. DR can also refer to the ability of a display device to adequately or approximately render an intensity range of a particular breadth. In this sense, DR is related to "display - referred" intensity. Unless a specific meaning is explicitly specified at any point in the description herein, it should be inferred that the term can be used in either sense, e.g., interchangeably.
[0007] As used herein, the term "high - dynamic range (HDR)" refers to a DR breadth that spans approximately 14 to 15 orders of magnitude of the human visual system (HVS). In practice, the DR breadth that humans can simultaneously perceive may be slightly truncated relative to HDR. As used herein, the term "enhanced dynamic range (EDR) or visual dynamic range (VDR)" can be related to this DR either individually or interchangeably: the DR that can be perceived by the human visual system (HVS) including eye movement within a scene or image, allowing some light adaptation changes over the scene or image. As used herein, EDR can refer to a DR that spans 5 to 6 orders of magnitude. Thus, while EDR may be slightly narrower than HDR relative to a true scene reference, EDR can represent a wide DR breadth and can also be referred to as HDR.
[0008] In fact, an image includes one or more color components (e.g., luminance Y and chrominance Cb and Cr), where each color component is represented with a precision of n bits per pixel (e.g., n = 8). Using linear light luminance encoding, an image where n ≤ 8 (e.g., a 24-bit color JPEG image) is considered a standard dynamic range image, while an image where n > 8 can be considered an enhanced dynamic range image.
[0009] The reference electro-optical transfer function (EOTF) of a given display characterizes the relationship between the color values (e.g., luminance) of an input video signal and the output screen color values (e.g., screen luminance) produced by the display. For example, ITU-R BT.1886 defines the reference EOTF for flat panel displays based on the measured characteristics of cathode ray tubes (CRTs). Given a video stream, information about its EOTF is typically embedded as metadata in the bitstream. As used herein, the term "metadata" refers to any auxiliary information that is transmitted as part of an encoded bitstream and aids the decoder in rendering the decoded image. Such metadata can include, but is not limited to, color space or gamut information, reference display parameters, and auxiliary signal parameters as described herein. In this document, BT.1886, Rec.2020, BT.2100, etc. refer to sets of definitions for various aspects of HDR video promulgated by the International Telecommunication Union (ITU).
[0010] Most consumer desktop displays currently support a luminance of 200 to 300 cd / m 2 or nits. Most consumer HDTVs range from 300 to 500 nits, where some models reach 1000 nits (cd / m 2 ). Thus, such traditional displays represent a lower dynamic range (LDR) associated with HDR or EDR, also known as standard dynamic range (SDR). As the availability of HDR content increases due to the development of both capture devices (e.g., cameras) and HDR displays (e.g., Dolby Laboratories' PRM-4200 professional reference monitor), HDR content can be color graded and displayed on HDR displays that support a higher dynamic range (e.g., from 1,000 nits to 5,000 nits or higher). Such displays can be defined using an alternative EOTF that supports high luminance capabilities (e.g., 0 to 10,000 nits). An example of such an EOTF is defined in SMPTE ST 2084:2014, "High Dynamic Range EOTF of Mastering Reference Displays". Generally, and without limitation, the methods of the present disclosure relate to any dynamic range.
[0011] As used herein, the term "forward shaping" refers to the process of mapping (or quantizing) an image from its original bit depth and coding format (e.g., gamma or SMPTE 2084) to an image of a lower or the same bit depth and a different coding format, which process allows the use of coding methods (such as AVC, HEVC, etc.) to improve compression. At the receiver, after decompressing the shaped signal, the receiver may apply an inverse shaping function to restore the signal to its original high dynamic range. The receiver may receive the backward shaping function in the form of a look-up table (LUT) or parameters, e.g., as coefficients of a piecewise polynomial approximation of the function.
[0012] The methods described in this section are methods that can be pursued, but not necessarily methods that have been previously conceived or pursued. Thus, unless otherwise specified, no method described in this section should be considered to be prior art merely by virtue of its inclusion in this section. Similarly, unless otherwise indicated, problems identified with respect to one or more methods should not be considered to be identified in any prior art based on this section. Summary of the Invention
[0013] Aspects of the present disclosure relate to systems and methods for improved electro-optical transfer function conversion, signal legitimization, and backward shaping functions.
[0014] In an exemplary aspect of the present disclosure, a device is provided. The device includes an electronic processor. The device is configured to determine a backward shaping function for rendering video on a target display.
[0015] The electronic processor is configured to determine a set of sample pixels from received video data; define a first set of sample pixels from the set of sample pixels according to a first electro-optical transfer function in a first color representation of a first color space; convert the first set of sample pixels to a second electro-optical transfer function in the first color representation of the first color space via a mapping function, thereby generating a second set of sample pixels according to the second electro-optical transfer function from the first set of sample pixels; convert the first set of sample pixels and the second set of sample pixels from the first color representation to a second color representation of the first color space; and determine the backward shaping function based on the converted first set of sample pixels and the converted second set of sample pixels. The electronic processor is configured to determine the backward shaping function by repeatedly applying and adjusting a sample backward shaping function to minimize the difference between predicted pixel values obtained by applying the sample backward shaping function to pixels in the converted first set of sample pixels and pixels in the converted second set of sample pixels.
[0016] In another exemplary aspect of the present disclosure, the device may be implemented as a method or implemented with the method, the method being for converting a signal and / or a non-transitory computer-readable medium storing instructions that, when executed by a processor of a computer, cause the computer to perform operations.
[0017] Aspects of the present disclosure may provide improvements in conversion speed, conversion efficiency, conversion accuracy, etc. In this way, aspects of the present disclosure provide conversion of images and provide improvements in technical fields such as at least HDR-TV image rendering, signal processing, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings, together with the following detailed description, are incorporated in and form a part of the specification, and are used to further illustrate aspects and explain the various principles and advantages of those aspects. Like reference numerals in the drawings refer to the same or functionally similar elements in the various views.
[0019] Figure 1A is an exemplary process of a video transmission pipeline according to aspects of the present disclosure.
[0020] Figure 1B is an exemplary process of content adaptive shaping according to aspects of the present disclosure.
[0021] Figure 2A is a graph illustrating the output of an example prediction algorithm according to aspects of the present disclosure.
[0022] Figure 2B is a graph illustrating the output of an example prediction algorithm according to aspects of the present disclosure.
[0023] Figure 2C is a graph illustrating the output of an example prediction algorithm according to aspects of the present disclosure.
[0024] Figure 3A is a process flow diagram of an exemplary backward shaping function for determining a process based on a prediction function generated according to aspects of the present disclosure.
[0025] Figure 3B is a process flow diagram of an exemplary backward shaping function for determining a process based on a prediction function generated according to aspects of the present disclosure.
[0026] Figure 3C is a process flow diagram of an exemplary backward shaping function for determining a process based on a prediction function generated according to aspects of the present disclosure.
[0027] Figure 3DIt is a process flow diagram for determining an exemplary backward shaping function of an imaging process based on a prediction function generated according to various aspects of the present disclosure.
[0028] Figure 4 It is a flowchart illustrating an exemplary method for determining a backward shaping function according to various aspects of the present disclosure.
[0029] Figure 5A It is a graph of an exemplary piecewise equation for signal legitimization according to various aspects of the present disclosure.
[0030] Figure 5B It is a graph of an exemplary approximate sigmoid curve for signal legitimization according to various aspects of the present disclosure.
[0031] Figure 6 It is capable of implementing according to various aspects of the present disclosure Figure 4 of an exemplary processing device for a process.
[0032] Those skilled in the art should understand that the elements in the drawings are illustrated for clarity and are not necessarily drawn to scale. For example, the dimensions of some elements in the drawings may be enlarged relative to other elements to help improve the understanding of various aspects of the present disclosure.
[0033] Where appropriate, device and method components have been represented by symbols in the drawings, showing only those specific details relevant to understanding the various aspects of the present disclosure, so as not to obscure the present disclosure with details that are obvious to those of ordinary skill in the art who benefit from the description herein. Detailed Description
[0034] Overview
[0035] This overview presents a basic description of some aspects of the present disclosure. It should be noted that this overview is not an extensive or exhaustive summary of the aspects of the present disclosure. Additionally, it should be noted that this overview is not intended to be understood as identifying any particularly important aspects or elements of the present disclosure, nor is it intended to specifically delineate any scope of the aspects, nor to generally depict the present disclosure. This overview only presents some concepts related to exemplary aspects in a compressed and simplified format and should be understood as merely a conceptual prelude to the more detailed description of the following aspects. Note that although separate aspects are discussed herein, any combination of the aspects and / or partial aspects discussed herein may be combined.
[0036] The techniques described herein can be used to minimize the requirements for memory bandwidth, data rate, and / or computational complexity in video applications, which may include the display of video content and / or the streaming of video content between a (multiple) video stream server and a (multiple) video stream client.
[0037] Video applications as described herein can refer to any one or more of the following: video display applications, virtual reality (VR) applications, augmented reality (AR) applications, automotive entertainment applications, telepresence applications, display applications, etc. Example video content can include, but is not limited to, any one or more of the following: audio-visual programs, movies, video programs, TV broadcasts, computer games, AR content, VR content, automotive entertainment content, etc.
[0038] Example video stream clients can include, but are not limited to, any one or more of the following: display devices, computing devices with near-eye displays, head-mounted displays (HMDs), mobile devices, wearable display devices, set-top boxes with displays such as TVs, video monitors, etc.
[0039] As used herein, a "video stream server" can refer to one or more upstream devices that prepare omnidirectional video content and stream the omnidirectional video content to one or more video stream clients for rendering at least a portion of the omnidirectional video content (e.g., a portion corresponding to a user's field of view or viewport, etc.) on one or more displays. The display on which the omnidirectional video content is rendered can be part of the one or more video stream clients or can operate in conjunction with the one or more video stream clients.
[0040] Example video stream servers can include, but are not limited to, any one of the following: cloud-based video stream servers located remotely from the video stream client(s), local video stream servers connected to the video stream client(s) via a local wired or wireless network, VR devices, AR devices, automotive entertainment devices, digital media devices, digital media receivers, set-top boxes, game consoles (e.g., Xbox (TM)), general purpose personal computers, tablet computers, dedicated digital media receivers such as Apple TV (TM) or Roku (TM) boxes, etc.
[0041] The present disclosure and its aspects can be implemented in various forms, including: hardware or circuitry controlled by a computer-implemented method, computer program products, computer systems and networks, user interfaces and application programming interfaces; and hardware-implemented methods, signal processing circuits, memory arrays, application specific integrated circuits, field programmable gate arrays, etc. The foregoing summary is only intended to give a general idea of the various aspects of the present disclosure and does not limit the scope of the present disclosure in any way.
[0042] In some aspects of the present disclosure, the mechanisms described herein form part of a media processing system, which includes, but is not limited to, any one or more of the following: cloud-based servers, mobile devices, virtual reality systems, augmented reality systems, head-up display devices, head-mounted display devices, CAVE systems, wall-sized displays, video game devices, display devices, media players, media servers, media production systems, camera systems, home-based systems, communication devices, video processing systems, video codec systems, production studio systems, streaming servers, cloud-based content service systems, handheld devices, game consoles, televisions, theater displays, laptop computers, notebook computers, tablet computers, cellular wireless telephones, e-book readers, point-of-sale terminals, desktop computers, computer workstations, computer servers, computer kiosks, or various other types of terminals and media processing units. For ease of description, some or all of the example systems presented herein are illustrated with a single example of each of their component parts. Some examples may not describe or illustrate all of the components of the system. Other aspects of the present disclosure may include more or fewer of each of the illustrated components, may combine some components, or may include additional or alternative components.
[0043] Example video transmission processing pipeline
[0044] Figure 1A An example process of video transmission pipeline 100A is depicted, which shows various stages from video capture to video content display. A sequence of video frames 102 is captured or generated using image generation block 105. The video frames 102 can be digitally captured (e.g., by a digital camera) or generated by a computer (e.g., using computer animation) to provide video data 107. Alternatively, the video frames 102 can be captured on film by a film camera. The film is converted to a digital format to provide video data 107. During the production stage 110, the video data 107 is edited to provide a video production stream 112.
[0045] The video data 112 of the production stream is then provided to a processor at block 115 for post-production editing. The post-production editing at block 115 can include adjusting or modifying the color or brightness in a specific region of the image to enhance the image quality or achieve a specific look of the image according to the creative intent of the video creator. This is sometimes referred to as "color timing" or "color grading". Other edits (e.g., scene selection and sequencing, image cropping, adding computer-generated visual effects, etc.) can be performed at block 115 to produce a final version 117 of the work for distribution. During the post-production editing 115, the video image is viewed on a reference monitor 125.
[0046] After post-production 115, the video data of the final work 117 can be transmitted to an encoding box 120 for downstream transmission to decoding and playback devices such as televisions, set-top boxes, and movie theaters. In some aspects, the encoding box 120 can include audio and video encoders for generating an encoded bitstream 122, such as those defined by ATSC, DVB, DVD, Blu-ray, and other transmission formats. In a receiver, the encoded bitstream 122 is decoded by a decoding unit 130 to generate a decoded signal 132 that represents the same or a closely approximate version of the signal 117. The receiver can be attached to a target display 140, which can have characteristics that are completely different from those of the reference display 125. In such a case, a display management box 135 can be used to map the dynamic range of the decoded signal 132 to the characteristics of the target display 140 by generating a display mapping signal 137.
[0047] Additionally, optionally, or alternatively, the encoded bitstream 122 is further encoded with image metadata, which includes but is not limited to backward reshaping data, which can be used by a downstream decoder to perform backward reshaping on the signal 117 in order to generate a backward reshaped image that is the same or similar to a target HDR image that can be optimized for rendering on an HDR display device. In some aspects of the present disclosure, one or more conversion tools implementing inverse tone mapping, inverse display management, etc. can be used to generate the target HDR image from the signal 117.
[0048] In some aspects of the present disclosure, the target HDR image can be directly generated from the video data 112 at the post-production edit 115. During the post-production edit 115, the target HDR image is viewed on a second reference display (not shown) that supports high dynamic range by the same or a different colorist who is performing post-production edit operations on the target HDR image.
[0049] Signal shaping
[0050] Currently, many digital interfaces for video transmission (such as serial digital interface (SDI)) are limited to 12 bits per component per pixel. Additionally, many compression standards (such as H.264 (or AVC) and H.265 (or HEVC)) are limited to 10 bits per component per pixel. Therefore, within existing infrastructure and compression standards, efficient encoding and / or quantization are needed to support HDR content with a dynamic range from approximately 0.001 to 10,000 cd / m 2 (or nits).
[0051] As used herein, the term "PQ" refers to Perceptual Quantizer. The human visual system responds to increasing light levels in a highly non-linear manner. The ability of a human to observe a stimulus is affected by factors such as the luminance of the stimulus, the size of the stimulus, the spatial frequency that makes up the stimulus, and the luminance level to which the eye is adapted at a particular moment when viewing the stimulus. In aspects of the present disclosure, a perceptual quantizer function maps linear input gray levels to output gray levels that better match the contrast sensitivity threshold in the human visual system. An example of a PQ mapping function is described in SMPTE ST 2084:2014, "High Dynamic Range EOTF of Mastering Reference Displays", which is incorporated herein by reference in its entirety, where, for each luminance level (i.e., stimulus level) given a fixed stimulus size, the minimum visible contrast step at that luminance level is selected according to the most sensitive adaptation level and the most sensitive spatial frequency (according to the HVS model). For example, compared to a traditional gamma curve that represents the response curve of a physical cathode ray tube (CRT) device and coincidentally may be very similar to the way the human visual system responds, the PQ curve mimics the true visual response of the human visual system using a relatively simple functional model.
[0052] Figure 1B Depicts an example process 100B of content adaptive shaping according to aspects of the present disclosure. As compared with Figure 1A Items given the same reference numeral may refer to the same element. Given an input frame 117, the forward shaping block 150 analyzes the input constraints and the encoding constraints and generates a codeword mapping function that maps the input frame 117 to a re-quantized output frame 152. For example, according to certain EOTFs, the input 117 may be gamma-encoded or PQ-encoded. In some aspects of the present disclosure, metadata may be used to convey information about the shaping process to downstream devices such as decoders. After encoding 120, the frames within decoding 130 may be processed by a backward shaping function that converts the frame 122 back to the EOTF domain (e.g., gamma or PQ) for further downstream processing such as the display management process 135 discussed previously.
[0053] As mentioned above, the backward shaping function is ideally configured such that the resulting backward-shaped image is the same as or approximate to a target (e.g., HDR) image that can be optimized for rendering on a display device. In other words, the quality of the image produced on the display device depends on the accuracy of the backward shaping function.
[0054] Backward Shaping Optimization - Minimum Mean Square Error Predictor
[0055] In the following section, the optimization of the backward shaping function is described. In Figure 1A the system of, a luminance signal channel predictor is constructed using cumulative distribution function (CDF) matching. As described in U.S. Patent No. 10,264,287, an SDR CDF is constructed based on an SDR histogram generated from the distribution of SDR codewords in one or more SDR images, which is incorporated herein by reference in its entirety. Similarly, an HDR CDF is constructed based on an HDR histogram generated from the distribution of HDR codewords in one or more HDR images corresponding to one or more SDR images. Then, a histogram transfer function is generated based on the SDR CDF and the HDR CDF. Then, the histogram transfer function can be used to determine the backward shaping metadata in order to determine the backward shaping function.
[0056] When chrominance channel information interferes with the luminance channel, errors may occur in the CDF. Therefore, a minimum mean square error (MMSE) predictor can be used to minimize the prediction error (in terms of mean square error or MSE) for each luminance range.
[0057] Although the MMSE predictor is a solution in terms of MSE, the predictor does not guarantee a monotonic non-decreasing property. To avoid any artifacts resulting from the non-monotonic non-decreasing property, monotonic non-decreasing is enforced via CDF matching of the MMSE predictor. Final curve smoothing is also applied to ensure that the curve is smooth.
[0058] First, the source signal is defined as s ij and the reference signal is defined as r ij (where i is the pixel position at frame index j) and the bit depths of the source signal and the reference signal are defined as B s and B r , respectively. For each source signal bin b, by taking the set of source signals having values in bin b (denoted as Φ b,j ), the average value of the reference signals mapped from the same source signal bin is obtained. Non-limitingly, as an example, the number of bins can be set to the total number of codewords in the signal (e.g., 2 Bs ); however, in other embodiments, a smaller number of bins can be selected to reduce computational complexity. It is known that
[0059] Φ b,j ={i|s ii ==b}, (1)
[0060] The average value is expressed as:
[0061]
[0062] The mapping t b,j =f jMMSE (b) is the MMSE predictor. Figure 2A An example of the resulting MMSE prediction (200A) is shown.
[0063] Backward shaping optimization - monotonically non - decreasing
[0064] If the MMSE predictor is used alone, the mapping is not monotonically non - decreasing. As Figure 2A shown, the mapping values of some bins with larger bin indices are smaller than those from smaller bin indices. In other words, in a non - monotonically non - decreasing curve, for two bins it can be observed that:
[0065]
[0066] To avoid artifacts (resulting from the above property), the monotonically non - decreasing (MND) curve should have the following property for all bins:
[0067]
[0068] As described in U.S. Patent No. 10,264,278, by utilizing the cumulative density function (CDF) based on the SDR histogram and the CDF based on the HDR histogram, CDF matching generates the backward shaping function (or BLUT). In an embodiment, to construct the CDF, the SDR histogram of the source signal can still be used; however, the HDR CDF can be constructed by using the MMSE prediction function of Equation (2). For example, given the SDR histogram, the MMSE prediction function is used to map each element of the SDR histogram to the HDR histogram to determine the histogram transfer function. Given the two histograms, the construction of the CDF and the rest of the CDF matching algorithm are still similar to U.S. Patent Serial No. 10,264,278. The output BLUT from this CDF matching algorithm will also satisfy the MND property of Equation (4).
[0069] Using the final mapping table After applying MND (applied to the Figure 2A curve shown in Figure 2B ), the resulting curve is shown in the graph 200B of
[0070] Backward shaping optimization - curve smoothing
[0071] CDF matching can ensure MND; however, the mapping function needs to be smooth enough so that it can be approximated by an 8 - segment second - order polynomial. Therefore, a smoothing filter is applied to the curve as follows.
[0072] First, the codewords b (b u and b lThe upper and lower limits of the moving average, where n is half of the overall filter taps.
[0073] b u = min(b + n, b max + 1) (5)
[0074] b l = max(b - n, b min - 1) (6)
[0075] Then apply the first moving average.
[0076]
[0077] Then apply the second moving average.
[0078]
[0079] Note that for a 10-bit signal, the value of n can be 4. As Figure 2C shown, the resulting curve shown in graph 200C is smoother than the Figure 2B previous curve, making it easier to model the resulting curve using an 8-segment 2nd order polynomial.
[0080] The MMSE predictor can be applied to the forward shaping path or the backward shaping path. The MMSE predictor can also be applied to various EOTFs, not just HLG to PQ.
[0081] Backward Shaping Optimization - BLUT Similarity Weighted Smoothing
[0082] The luminance backward shaping function for each video frame needs to be smoothed in the time domain to prevent sudden and unexpected intensity changes / visual flicker in consecutive video frames within a scene. Scene Cut Aware Backward Lookup Table (BLUT) smoothing can be used to reduce flicker. However, due to the defects of the automatic scene cut detector, visual flicker problems may still occur in automatically detected scene cut instances. Therefore, a smoothing mechanism that is not affected by false scene cut detections is needed. Since the differences between adjacent BLUTs already indicate different content, the process described herein averages BLUTs with similar shapes and excludes BLUTs with different trends. In other words, the BLUT similarity weighted average is used to solve this temporal stability problem.
[0083] Define T j as the unsmoothed BLUT at frame j, the normalized HDR codeword value at the b-th SDR luminance codeword of frame j N S is the total number of SDR codewords. Consider using for smoothing T ja symmetric window of M frames on each side (a total of 2M + 1 frames) of the central frame j([j - M, j + M]), let be the smoothed output BLUT of the frame at j.
[0084] The BLUT similarity is measured according to the normalized squared difference at each SDR codeword b of each j-th frame, where the integer m ∈ [j - M, j + M]. The BLUT similarity of the BLUT of the m-th frame at the b-th codeword relative to the central frame at j is defined as The resulting definition is:
[0085]
[0086] Note that
[0087] a content-dependent weighting factor is used as a multiplier for the BLUT similarity and is determined for each b-th codeword by the SDR image histogram of the j-th frame as follows:
[0088]
[0089] In the above example, the logarithm of the histogram is taken so that the range of the histogram values is less than the range of the actual number of pixels of each codeword. Adding 1 when taking the logarithm ensures that the weighting factor of any histogram remains finite.
[0090] To smooth the BLUT of the j-th frame, the weights of the BLUTs of each m-th frame are used, where m comes from its temporal neighborhood [j - M, j + M]. For example, the weights can be calculated as an exponential term or a Gaussian term based on both the histogram and the BLUT difference. The weight w of the BLUT of the m-th frame used to smooth the BLUT of the j-th frame j,m is calculated as:
[0091]
[0092] where γ is a constant determined empirically to make the smoothing work properly (e.g., 130). Thus, the weights are specific to a frame and are the same for all codewords in that frame.
[0093] Using the weight w j,m as a multiplier for each such that m ∈ [j - M, j + M], the smoothed BLUT of the central frame j is calculated as
[0094]
[0095] In some aspects, twelve frames (M = 12) are used for BLUT smoothing. The resulting BLUT for the j-th frame is used for the next curve fitting process.
[0096] Multivariate Multiple Regression (MMR) Optimization
[0097] In video processing, a multi-color channel MMR predictor can be implemented that allows prediction of an input signal in a first dynamic range using a corresponding enhanced dynamic range signal in its second dynamic range and a multivariate MMR operator (e.g., the predictor described in U.S. Patent No. 8,811,490, which is incorporated herein by reference in its entirety). The resulting data can be used to determine a backward shaping function. The process of selecting prediction parameters (MMR coefficients) for the MMR predictor is described below. After collecting color mapping pairs from each pixel or as a 3D mapping table, the MMR coefficients are solved by the least squares method. The mapping from the source image to the reference image is represented as follows:
[0098]
[0099] where and represent the source values of the plane Y, C0, and C1 of the k-th entry in the mapping table, respectively, and represent the reference mapping values of the plane Y, C0, and C1 of the k-th entry in the mapping table, respectively. Note that K is the total number of entries in the mapping table.
[0100] Two vectors are constructed using the average reference chromaticity values.
[0101]
[0102]
[0103] A matrix is constructed using the source values.
[0104]
[0105] where p k T = [1 s0 Y s0 C0 s0 C1 s0 Y ·s0 C0 s0 Y ·s0 C1 s0 Y ·s0 C0 ·s0 C1 ...] contains all terms supported by the MMR predictor.
[0106] Calculate the MMR by solving the following optimization problem:
[0107]
[0108]
[0109] where x C0 and x C0 are the MMR coefficients of C0 and C1, respectively.
[0110] denotes:
[0111] A = S T S(19)
[0112]
[0113]
[0114] The MMR coefficients can be calculated by solving a linear problem:
[0115]
[0116]
[0117] Problems may occur when the A matrix is close to being singular (the above linear problem is ill-conditioned). To obtain a stable solution to the ill-conditioning problem, the following method can be applied.
[0118] Multivariate multiple regression optimization - Gaussian elimination
[0119] The MMR coefficients can be solved by:
[0120]
[0121]
[0122] However, calculating the inverse matrix of A can be time-consuming. One solution is to apply Gaussian elimination.
[0123] During Gaussian elimination, the A matrix is transformed into its upper triangular form. Then back substitution is applied to solve for the MMR coefficients. When some (multiple) rows of the A matrix are close to a linear combination of some other rows, the A matrix is close to being singular. This means that the corresponding (multiple) MMR terms are close to being linearly related to some other MMR terms. Removing these terms will make the problem better conditioned, resulting in a more stable solution.
[0124] Matrix A, vector b C0 and b C1 are represented as follows, where P is the total number of MMR terms.
[0125]
[0126]
[0127]
[0128] For ease of description, the following process is described with respect to C0. It should be noted that C1 can be processed similarly. The following elimination is described with reference to the following matrix:
[0129]
[0130] As described in more detail in Table 1, if |a m,m | ≤ ε, where ε is a small threshold, then row m and column m will be ignored, which is equivalent to removing the m-th MMR term (e.g., the set ) and the m-th equation from the system of equations, otherwise the solver continues to eliminate the remaining part of the row. In some aspects, the predetermined threshold ε is approximately 1e-6. Then back substitution is applied to solve Equation (22) above in order to calculate the MMR coefficients using the predetermined threshold ε. Thus, the linear problem A·x C0 = b C0 and A·x C1 = b C1 are solved using valid MMR terms, and the coefficients of the invalid terms will be zero. During the Gaussian elimination process described above, the linear dependencies of the MMR terms are removed, making the solution relatively stable. An example process of the pseudocode is depicted in Table 1.
[0131] Table 1: Example Pseudocode of a Stable Gaussian Solver
[0132]
[0133]
[0134] EOTF Conversion - Full Data Point Optimization
[0135] Figure 3A is Process Diagram 300A that illustrates the EOTF conversion process implemented by Controller 600, which will be described in more detail below with reference to Figure 6 and is described in conjunction with Figure 4 Process Diagram 300A below Figure 4FIG. 400 is a flow chart of a method for determining a backward shaping function for video conversion, specifically, conversion from a first EOTF to another EOTF. The following paragraphs will describe an example EOTF conversion from a Hybrid Log Gamma (HLG) signal to a Perceptual Quantizer (PQ) signal, specifically, from HLG Rec. 2020 at 1000 nits to PQ Rec. 2020 at 1000 nits. However, it should be understood that the system is not necessarily limited to conversions between these specific types of signals.
[0136] First, at block 410, the controller 600 determines a first set of sample points based on the composite data (e.g., received video data) from the color grid. Specifically, a set of sample points (pixels) is collected from the composite data (e.g., Figure 1A video 117), denoted as Φ in Figure 3A . In this document, the terms "sample point" and "sample pixel" may be used interchangeably to refer to the same thing. The set of sample points Φ is defined by first constructing a 1D sampling array q i with M samples and the i-th point in the normalized domain (i ∈ [0, M - 1]), as shown below, where i represents the pixel position.
[0137]
[0138] Then, using the 1D array q i 3D sample points are constructed in 3D space, denoted as the 3D array q ijk below, where j and k are the frame index and depth of the pixel, respectively.
[0139] q ijk = (q i , q j , q k ) (27)
[0140] Thus, the set of sample points Φ is the sample points collected in the 3D space {q ijk}.
[0141] Returning to Figure 4 , at block 415, the controller 600 defines a first set of sample points from the set of sample points according to a first electro-optical transfer function in a first color representation of a first color space. For example, at block 302 ( Figure 3A ), according to the electro-optical transfer function, the set of sample points Φ is considered (or defined as) HLG at 1000 nits in the Rec. 2020 color space (RGB), denoted as Φ HLG,RGB,R2020 . Note that the values in Φ remain unchanged here. In Figure 4At the block 420, the processor 600 converts a first set of sample points according to a first electro-optical transfer function to a second electro-optical transfer function in a first color representation in a first color space via a mapping function, thereby generating a second set of sample points according to the second electro-optical transfer function. For example, as shown in block 304 of Figure 3A , the set of sample points Φ HLG,RGB,R2020 is converted to Rec.2020PQ 1000 nits RGB points via ITU-R BT.2100, and the result is represented as Φ PQ,RGB,R2020 (block 306). Figure 3A As shown in block 304 of Figure 3A , the set of sample points Φ HLG,RGB,R2020 is converted to Rec.2020PQ 1000 nits RGB points via ITU-R BT.2100, and the result is represented as Φ PQ,RGB,R2020 PQ,RGB,R2020 (block 306).
[0142] In some aspects, at block 425, the controller 600 converts a first set of sample pixels (points) according to a first electro-optical transfer function and a second set of sample pixels (points) according to a second electro-optical transfer function to a second color representation in a first color space. For example, in this example, to obtain the backward shaping function, the sample points are converted from an RGB color representation to a YCbCr color representation in the same color space Rec.2020. Here, the processed sets of sample pixels (points) Φ HLG,RGB,R2020 and Φ PQ,RGB,R2020 are both converted from the first color representation RGB to the second color representation YCbCr in the same color space Rec.2020 (blocks 308 and 310 of Figure 3A respectively). For the set of sample pixels (points) Φ HLG,RGB,R2020 , the Rec.2020HLG YCbCr pixels (points) are defined as
[0143]
[0144]
[0145] and the entire set of sample pixels (points) is defined as Φ HLG,YCbCr,R2020 (block 312). HLG,RGB,R2020 and Φ PQ,RGB,R2020 are both converted from the first color representation RGB to the second color representation YCbCr in the same color space Rec.2020 (blocks 308 and 310 of Figure 3A respectively). For the set of sample pixels (points) Φ HLG,RGB,R2020 , the Rec.2020HLG YCbCr pixels (points) are defined as
[0143]
[0144]
[0145] and the entire set of sample pixels (points) is defined as Φ HLG,YCbCr,R2020 (block 312). Figure 3A For the set of sample pixels (points) Φ HLG,RGB,R2020 , the Rec.2020HLG YCbCr pixels (points) are defined as HLG,RGB,R2020
[0143]
[0144]
[0145] and the entire set of sample pixels (points) is defined as Φ HLG,YCbCr,R2020
[0143]
[0144]
[0145] and the entire set of sample pixels (points) is defined as Φ HLG,YCbCr,R2020 HLG,YCbCr,R2020 (block 312).
[0145] For the processed set of sample points Φ PQ,RGB,R2020 , the Rec.2020PQ YCbCr points are defined as
[0146]
[0147] Figure 3A and the entire set of sample points is defined as Φ PQ,YCbCr,R2020 PQ,RGB,R2020
[0146]
[0147] Figure 3A and the entire set of sample points is defined as Φ PQ,YCbCr,R2020
[0146]
[0147] Figure 3A PQ,YCbCr,R2020 ( Figure 3A block 314 of Figure 3A ).
[0148] The backward shaping function formula is defined as follows:
[0149]
[0150] where, is the predicted PQ value of each HLG pixel.
[0151] Returning to Figure 4, at block 430, the controller 600 determines a backward shaping function based on a set of converted first sample pixels (points) according to a first electro-optical transfer function and a set of converted second sample pixels (points) according to a second electro-optical transfer function. In this example, to find the backward shaping function formula, the following optimization problem (block 316) is solved.
[0152]
[0153] The optimization equation (31) can be solved iteratively by repeatedly applying and adjusting the sample backward shaping function to minimize the difference between the result of the sample backward shaping function from equation (30) (the predicted PQ value obtained by applying the sample backward shaping function to the pixels in the set of converted first sample pixels) and the pixels in the set of second sample points (29) according to the second electro-optical transfer function. As pointed out above, according to some aspects of the present disclosure, the backward shaping function can be a polynomial function. Thus, the method explained above enables approximate conversion between the HLG system and the PQ system (e.g., at the encoder side) without performing a full conversion. Using the above process, it is found that there may be some prediction errors.
[0154] EOTF Conversion - Common Data Point Optimization
[0155] The methods and processes described below provide solutions for improving the accuracy of the backward shaping function determination described above by leveraging a small range within the Rec. 2020 color space. Figure 3B is a process diagram 300B that illustrates a modified EOTF conversion process implemented by the controller 600 ( Figure 6 ). It should be noted that process diagram 300B includes steps / blocks similar to those in process diagram 300A and are correspondingly labeled the same (specifically, blocks 302, 304, 306, 308, 310, 312, 314, and 316).
[0156] In Figure 3B the example illustrated, after determining the set of sample points from the received source data, generating the first set of sample points according to the first electro-optical transfer function of the first color space further includes, at block 318, generating a third set of sample points according to a third electro-optical transfer function in the first color representation of the second color space (here, PQ 1000 nits Rec. 709 RGB) from the set of data points, and generating the first set of sample points according to the first electro-optical transfer function based on the third set of sample points according to the third electro-optical transfer function, wherein the second color space is smaller than the first color space. In Figure 3B the example illustrated, the third electro-optical transfer function in the first color representation of the second color space is represented as Φ PQ,RGB,R709As mentioned above, the second color space is smaller than the first color space. The second color space can be determined or selected based on the visual content of the received data. For example, in this example, the data can be a natural scene. Therefore, Rec.709 is selected because although it is smaller than color spaces of higher definition standards such as Rec.2020, it includes most of the colors required for natural scenes. By using a smaller color space that includes the most commonly used colors in a particular scene, when using a predictor to approximately convert pixel values, non-linearity can be reduced and prediction error can be reduced.
[0157] At block 320, the controller 600 converts the third electro-optical transfer function in the first color representation of the second color space to the container of the first electro-optical transfer function of the first color representation of the first color space, thereby generating a first set of sample pixels (points) according to the first electro-optical transfer function of block 302. In this example, the container in the first color representation of the first color space defined by the first set of sample points is Rec.2020 HLG RGB, and the resulting set of sample pixels (points) is denoted as Φ HLG,RGB,R2020 ( Figure 3B of block 302). Then, the signal at block 302 is processed similarly to the corresponding blocks (blocks 304, 306, 308, 310, 312, 314, and 316) of Figure 3A method 300A of
[0158] In some aspects of the present disclosure, in order to include data in a wider color space, the controller 600 can interpolate (described below) the third set of sample points according to the third electro-optical transfer function of the second color space and the fourth set of sample points according to the fourth electro-optical transfer function of the first color representation of the third color space when defining the first set of sample points according to the first electro-optical transfer function. Thus, the resulting first set of sample points according to the first electro-optical transfer function includes a weighted combination of the third set of sample points according to the third electro-optical transfer function of the second color space and the fourth set of sample points according to the fourth electro-optical transfer function of the third color space. It should be noted that interpolation includes converting both the third set of sample points and the fourth set of sample points to a common electro-optical transfer function of a common color space.
[0159] For example, in this example, sample points can be interpolated from Rec.709 and Rec.2020. Figure 3C is illustrated by the controller 600( Figure 6) is implemented. It should be noted that process diagram 300C includes steps / boxes that are similar to and are labeled the same as those in process diagram 300A (specifically, blocks 302, 304, 306, 308, 310, 312, 314, and 316). It should also be noted that the processes performed at blocks 322 and 324 (and 326 and 328) are the same as those in Figure 3B The processes performed at blocks 318 and 320 of method 300B are similar.
[0160] At block 322, the controller 600 defines a set of sample points Φ, such as in the Rec.709PQ 1000 nit RGB color space, as Figure 3B The frame 318 is similar to Φ1 PQ,RGB,R709 (a third set of sample points according to a third electro-optical transfer function of the second color space). Then, at block 324, the controller 600 converts the third set of sample points Φ to the Rec. 2020 HLG container as Φ1 HLG,RGB,R2020 Then, at block 326, a copy of the original set of sample points Φ of the video data is defined as Φ2, such as in the Rec.2020PQ 1000 nit RGB color space. PQ,RGB,R2020 (a fourth set of sample points according to a fourth electro-optical transfer function of a third color space). Then, at block 328, the set Φ2 PQ,RGB,R2020 Converted to Rec.2020HLG container as Φ2 HLG,RGB,R2020 (e.g., a container of a first color representation (RGB) converted to a common color space Rec. 2020 of a third sample set). At block 330, the controller 600 weights and combines the data points in all color channels, such as
[0161]
[0162] Then, the resulting HLG set Φ from the above equation (32) HLG,RGB,R2020 (box 302) is converted to Rec.2020PQ 1000 nit RGB points (box 304), and the resulting set is denoted as Φ PQ,RGB,R2020 (Block 306). Then, the set Φ HLG,RGB,R2020 and Φ PQ,RGB,R2020 The second color representation YCbCr is converted to the same color space Rec. 2020 at blocks 308 and 310 , respectively, and at block 316 , the resulting set is used to compute a backward shaping function.
[0163] signal legitimation
[0164] To further improve the method described above, a signal legitimization function / process can be implemented. For example, as described below, a signal legitimization function configured to modify an input to conform to a predetermined range can be applied to the first set of sample points Φ. Signal legitimization is to correct out-of-range input signals so that they are within the desired legal range. A pipeline (e.g., Figure 1A pipeline 100A) can introduce out-of-range signals into the video data during processing, which may result in undesired artifacts in the final video signal. As described below, in some aspects of the present disclosure, the signal legitimization function implements hard clipping. In some aspects of the present disclosure, the signal legitimization function is a piecewise linear function. In some aspects of the present disclosure, the signal legitimization function is an S-shaped curve.
[0165] Signal Legitimization - Input Signal Legitimization
[0166] Input signal legitimization can be implemented using a hard clipping method (clipping signals outside the desired range). Although it is simple to implement, the final visual product may not be sufficient. To address this issue, soft clipping or a gradual transition can be applied near the boundaries of the legal range.
[0167] One method is to apply piecewise linear legitimization. Piecewise linear legitimization maintains a linear relationship between the input and the legitimized signal in the middle range and applies compression to signals near the legal / illegal boundary. First, the input range is defined as [x L , x H , the pivot point is defined as [x p1 , x p2 , and the legitimization function is defined as f L pwl (), and the corresponding legitimized value is:
[0168]
[0169]
[0170]
[0171]
[0172] The piecewise equation can be expressed as
[0173]
[0174] Figure 5A is the graph 500A of the above piecewise equation, where the input range is [x L , x H = [-0.2, 1.2], and the pivot point is [x p1 , xp2 ] = [0.20.8]. As can be seen from graph 500A, there are first order discontinuities at pivot points 502A and 502B, which may cause global model problems. This problem can be solved by approximating the piecewise linearity with an S-shaped curve, which can be characterized by the following equation:
[0175]
[0176] The above variables a1, a2, a3 and a4 represent four parameter models. Using the given segment model f L pwl (x), the parameters can be calculated via nonlinear optimization. Figure 5B It shows that by giving the segmentation parameter [x L , x H ] = [-0.2, 1.2] and [x p1 , x p2 ]=[0.20.8] is a graph 500B of an approximate S-shaped curve (with the following parameters) obtained.
[0177] a1=-0.0645 (39)
[0178] a2=1.0645 (40)
[0179] a3=0.5000 (41)
[0180] a4=1.6007 (42)
[0181] EOTF conversion-signal legalization
[0182] Using the above techniques, Figure 3A Method 300A (and respectively Figure 3B and Figure 3C Methods 300B and 300C) may be further modified to incorporate signal legalization. Figure 3D The modified EOTF conversion process (using signal legalization) implemented by the controller 500 (FIG. 5) is illustrated. Similar to the EOTF conversion described above, a set of sample points Φ is collected (as described above with respect to equations (1) and (2)). This set of sample points Φ is defined as the illegal input signal (block 332). Next, the controller 600 performs the legalization function (f of equations (12) and (13) above, respectively) by L pwl (x) or f L sgm (x) is applied to each sample point q i (Block 334) to construct the corresponding legalization set, thereby creating the legalization value q i L(Frame 336).
[0183]
[0184] The above 1D array is used to construct 3D sample points in 3D space using the following formula.
[0185]
[0186] The collected legalized points {q L ijk} are represented as the set Φ L .
[0187] To obtain the backward shaping function, at block 308, the sample points in the set Φ are defined in Rec.2020 HLG YCbCr points as
[0188]
[0189] and the transformed set is represented as Φ in,YCbCr,R2020 .
[0190] At block 310, the sample points in the legalized set Φ L are defined in Rec.2020 PQ YCbCr at block 310 and are represented as
[0191]
[0192] and the transformed legalized set is represented as Φ lg,YCbCr,R2020 .
[0193] Then, the backward shaping function (block 316) is calculated from the input illegal signal Φ in,YCbCr,R2020 to the legal signal Φ lg,YCbCr,R2020 . Similar to block 316 described above regarding Figure 3A the backward shaping function formula is defined as follows:
[0194]
[0195] where, is the predicted value.
[0196] To obtain the backward shaping function formula, the following optimization problem is solved (block 316).
[0197]
[0198] Example hardware device
[0199] Figure 6is a block diagram of a controller 600 according to some aspects of the present disclosure. The controller 600 can be the device described above for generating a backward warping function for rendering video on a target display. The controller 600 includes an electronic processor 605, a memory 610, and an input / output interface 615. The electronic processor 605 can be configured to, for example, execute the methods described with reference to Figure 4 . The electronic processor 605 obtains and provides information (e.g., from the memory 610 and / or the input / output interface 615), and processes the information by executing one or more software instructions or modules that can be stored in a random access memory (“RAM”) region of the memory 610, for example, or a read-only memory (“ROM”) of the memory 610 or another non-transitory computer-readable medium (not shown). The software can include firmware, one or more applications, program data, filters, rules, one or more program modules, and other executable instructions. The electronic processor 605 can include multiple cores or individual processing units. The electronic processor 605 is configured to fetch and execute software and the like related to the control processes and methods described herein from the memory 610.
[0200] The memory 610 can include one or more non-transitory computer-readable media and includes a program storage area and a data storage area. As described herein, the program storage area and the data storage area can include a combination of different types of memories. The memory 610 can take the form of any non-transitory computer-readable medium.
[0201] The input / output interface 615 is configured to receive inputs and provide system outputs. The input / output interface 615 obtains information and signals from both devices internal and external to the controller 600 (e.g., a video data source of the post-production 115( Figure 1A )) and provides information and signals to the devices (e.g., via one or more wired and / or wireless connections). The controller 600 can include or be configured to act as an encoder, a decoder, or both.
[0202] Equivalents, Extensions, Alternatives, and Miscellaneous
[0203] In the foregoing specification, specific aspects of the present disclosure have been described. However, those of ordinary skill in the art should understand that various modifications and changes can be made without departing from the scope of the present disclosure set forth in the following claims. Accordingly, the specification and drawings should be regarded as illustrative rather than restrictive in nature, and all such modifications are intended to be included within the scope of the present disclosure's teachings.
[0204] Regarding the processes, systems, methods, heuristics, etc. described herein, it should be understood that although the steps of such processes, etc. have been described as occurring in a particular, ordered sequence, such processes may be performed using the described steps in an order different from that described herein. It should also be understood that certain steps may be performed simultaneously, other steps may be added, or some steps described herein may be omitted. In other words, the process descriptions herein are provided for purposes of illustration of certain aspects and should in no way be construed as limiting the claims.
[0205] In addition, in this document, relational terms such as first and second, top and bottom, etc. may be used solely to distinguish one entity or action from another entity or action, and do not necessarily require or imply any actual such relationship or order between such entities or actions. The terms “comprises / comprising,” “has / having,” “includes / including,” “contains / containing,” or any other variation thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises, has, includes, or contains a list of elements does not include only those elements but may include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without further limitation, an element preceded by “comprises...a,” “has...a,” “includes...a,” or “contains...a” does not preclude the presence of additional, identical elements in the process, method, article, or apparatus that comprises, has, includes, or contains the element. Unless expressly stated otherwise herein, the terms “a” and “an” are defined as one or more. The terms “substantially,” “essentially,” “approximately,” “about,” or any other form thereof are defined as being close to what is understood by a person of ordinary skill in the art, and in the context of the present disclosure, the terms may be defined as within 10%, within 5%, within 1%, or within 0.5%. As used herein, the term “coupled” is defined as connected, but not necessarily directly, and not necessarily mechanically. An apparatus or structure that is “configured” in a certain way is at least configured in that way, but may also be configured in ways not listed.
[0206] It should be understood that some aspects of the present disclosure may include one or more general or special-purpose processors (or "processing devices"), such as microprocessors, digital signal processors, custom processors, and field programmable gate arrays (FPGAs), as well as uniquely stored program instructions (including both software and firmware), and the uniquely stored program instructions control one or more processors to implement some, most, or all of the functions of the methods and / or apparatuses described herein in combination with certain non-processor circuits. Alternatively, some or all of the functions may be implemented by a state machine without stored program instructions, or in one or more application specific integrated circuits (ASICs), where each function or some combination of certain functions is implemented as custom logic. Of course, a combination of these two methods may be used.
[0207] In addition, the present disclosure may be implemented as a computer-readable storage medium having computer-readable code stored thereon for programming a computer (e.g., including a processor) to perform the methods described and claimed herein. Examples of such computer-readable storage media include, but are not limited to, hard disks, CD-ROMs, optical storage devices, magnetic storage devices, ROM (read-only memory), PROM (programmable read-only memory), EPROM (erasable programmable read-only memory), EEPROM (electrically erasable programmable read-only memory), and flash memory. Further, it is expected that, despite the likely significant efforts and many design choices motivated by, for example, available time, current technology, and economic considerations, those skilled in the art will readily produce such software instructions and programs as well as ICs with minimal experimentation when guided by the concepts and principles disclosed herein.
[0208] All terms used in the claims are intended to be given the broadest reasonable interpretation and ordinary meaning as understood by those who know the technology described herein, unless an express contrary indication appears herein. In particular, the use of singular articles such as "a", "the", "said", etc. should be understood to recite one or more of the indicated elements unless the claim recites an express contrary limitation.
[0209] Aspects of the present disclosure may adopt any one or more of the following exemplary configurations:
[0210] (1) An apparatus for generating high-dynamic range video data and including a memory and an electronic processor. The electronic processor is configured to: determine a set of sample points based on synthesized data; define a first set of sample points from the set of sample points according to a first electro-optical transfer function of a first color space; convert the first set of sample points according to the first electro-optical transfer function to a second electro-optical transfer function via a mapping function; generate a second set of sample points according to the second electro-optical transfer function; and determine a backward shaping function based on the first set of sample points according to the first electro-optical transfer function and the second set of sample points according to the second electro-optical transfer function.
[0211] (2) The apparatus according to (1), wherein the electronic processor is configured to determine the backward shaping function by repeatedly applying and adjusting a sample backward shaping function to minimize a difference between a result of the sample backward shaping function and the second set of sample points according to the second electro-optical transfer function.
[0212] (3) The apparatus according to (1) or (2), wherein converting the first set of sample points according to the first electro-optical transfer function to the second electro-optical transfer function via the mapping function includes applying a signal legalization function to the first set of sample points according to the first electro-optical transfer function.
[0213] (4) The apparatus according to any one of (1) to (3), wherein the first electro-optical transfer function is hybrid log gamma.
[0214] (5) The apparatus according to any one of (1) to (4), wherein the second electro-optical transfer function is a perceptual quantizer.
[0215] (6) The apparatus according to any one of (1) to (5), wherein defining the first set of sample points according to the first electro-optical transfer function includes: generating a third set of sample points according to a third electro-optical transfer function of a second color space from the set of data points; and generating the first set of sample points according to the first electro-optical transfer function based on the third set of sample points according to the third electro-optical transfer function of the second color space, wherein the second color space is smaller than the first color space.
[0216] (7) The device according to (6), wherein defining the first set of sample points according to the first electro-optical transfer function includes: interpolating the third set of sample pixels according to the third electro-optical transfer function in the second color space and the fourth set of sample pixels according to the fourth electro-optical transfer function in the third color space, such that the first set of sample pixels according to the first electro-optical transfer function includes a weighted combination of the third set of sample points according to the third electro-optical transfer function in the second color space and the fourth set of sample points according to the fourth electro-optical transfer function in the third color space, wherein the interpolation includes converting the third set of sample points and the fourth set of sample points to a common electro-optical transfer function in a common color space.
[0217] (8) The device according to any one of (1) to (7), wherein the electronic processor is further configured to determine the backward shaping function based on backward shaping function data from a least mean square error predictor.
[0218] (9) The device according to (8), wherein a plurality of parameters of the least mean square error predictor are determined based on a multi-channel multivariate regression model.
[0219] (10) The device according to any one of (1) to (9), wherein the electronic processor is further configured to determine the backward shaping function based on a smooth equal-weight backward lookup table.
[0220] (11) The device according to any one of (1) to (10), wherein the device is an encoder.
[0221] (12) A method for converting a signal corresponding to a first electro-optical transfer function into a signal corresponding to a second electro-optical transfer function, the method comprising: determining a set of sample points according to synthetic data; defining a first set of sample points from the set of sample points according to the first electro-optical transfer function in a first color space; converting the first set of sample points according to the first electro-optical transfer function to a second electro-optical transfer function via a mapping function; generating a second set of sample points according to the second electro-optical transfer function; and determining a backward shaping function based on the first set of sample points according to the first electro-optical transfer function and the second set of sample points according to the second electro-optical transfer function.
[0222] (13) The method according to (12), wherein determining the backward shaping function includes repeatedly applying and adjusting a sample backward shaping function to minimize the difference between the result of the sample backward shaping function and the second set of sample points according to the second electro-optical transfer function.
[0223] (14) The method according to (12) or (13), wherein converting the first set of sample points according to the first electro-optical transfer function to the second electro-optical transfer function via the mapping function includes applying a signal legitimization function to the first set of sample points according to the first electro-optical transfer function.
[0224] (15) The method according to (14), wherein the signal legitimization function performs hard clipping.
[0225] (16) The method according to (14), wherein the signal legitimization function is a piecewise linear function.
[0226] (17) The method according to (14), wherein the signal legitimization function is an S-curve.
[0227] (18) The method according to any one of (12) to (17), wherein the first electro-optical transfer function is a hybrid logarithmic gamma.
[0228] (19) The method according to any one of (12) to (18), wherein the second electro-optical transfer function is a perceptual quantizer.
[0229] (20) The method according to any one of (12) to (19), wherein defining the first set of sample points according to the first electro-optical transfer function includes: generating a third set of sample points of a third electro-optical transfer function according to a second color space from the set of data points; and generating the first set of sample points according to the first electro-optical transfer function based on the third set of sample points of the third electro-optical transfer function according to the second color space, wherein the second color space is smaller than the first color space.
[0230] (21) The method according to (20), wherein defining the first set of sample points according to the first electro-optical transfer function includes: interpolating the third set of sample points of the third electro-optical transfer function according to the second color space and the fourth set of sample points of a fourth electro-optical transfer function according to a third color space, such that the first set of sample points according to the first electro-optical transfer function includes a weighted combination of the third set of sample points of the third electro-optical transfer function according to the second color space and the fourth set of sample points of the fourth electro-optical transfer function according to the third color space, wherein the interpolation includes converting the third set of sample points and the fourth set of sample points to a common electro-optical transfer function in a common color space.
[0231] (22) The method according to any one of (12) to (21), wherein the backward shaping function is further determined based on backward shaping function data from a least mean square error predictor.
[0232] (23) The method as described in (22), wherein a plurality of parameters of the least mean square error predictor are determined based on a multi-channel multivariate regression model.
[0233] (24) The method as described in (23), wherein calculating the solution of the multi-channel multivariate regression (MMR) model includes using Gaussian elimination to reduce the ill-conditioned condition in the MMR model.
[0234] (25) The method as described in any one of (12) to (24), wherein determining the backward shaping function is based on a smooth equal-weight backward lookup table.
[0235] (26) A non-transitory computer-readable medium storing instructions that, when executed by a processor of a computer, cause the computer to perform the method as described in any one of (12) to (25).
[0236] Aspects of the present invention can be understood from the following enumerated example embodiments (EEE):
[0237] 1. A device for generating high dynamic range video data and comprising:
[0238] A memory; and
[0239] An electronic processor configured to:
[0240] Determine a set of sample points according to the synthesized data;
[0241] Define a first set of sample points from the set of sample points according to a first electro-optical transfer function in a first color space;
[0242] Convert the first set of sample points according to the first electro-optical transfer function to a second electro-optical transfer function via a mapping function, thereby generating a second set of sample points according to the second electro-optical transfer function; and
[0243] Determine a backward shaping function based on the first set of sample points according to the first electro-optical transfer function and the second set of sample points according to the second electro-optical transfer function.
[0244] 2. The device as described in EEE 1, wherein the electronic processor is configured to determine the backward shaping function by repeatedly applying and adjusting a sample backward shaping function to minimize the difference between the result of the sample backward shaping function and the second set of sample points according to the second electro-optical transfer function.
[0245] 3. The device as described in EEE 1 or 2, wherein converting the first set of sample points according to the first electro-optical transfer function to the second electro-optical transfer function via the mapping function includes applying a signal legitimization function to the first set of sample points according to the first electro-optical transfer function, the signal legitimization function being configured to modify the input to conform to a predetermined range.
[0246] 4. The device as described in any one of EEE 1 to 3, wherein the first electro-optical transfer function is hybrid logarithmic gamma.
[0247] 5. The device as described in any one of EEE 1 to 4, wherein the second electro-optical transfer function is a perceptual quantizer.
[0248] 6. The device as described in any one of EEE 1 to 5, wherein defining the first set of sample points according to the first electro-optical transfer function includes:
[0249] generating a third set of sample points of a third electro-optical transfer function according to a second color space from the set of data points, and
[0250] generating the first set of sample points according to the first electro-optical transfer function based on the third set of sample points of the third electro-optical transfer function according to the second color space, wherein the second color space is smaller than the first color space.
[0251] 7. The device as described in EEE 6, wherein defining the first set of sample points according to the first electro-optical transfer function includes: interpolating the third set of sample points of the third electro-optical transfer function according to the second color space and a fourth set of sample points of a fourth electro-optical transfer function according to a third color space, such that the first set of sample pixels according to the first electro-optical transfer function includes a weighted combination of the third set of sample points of the third electro-optical transfer function according to the second color space and the fourth set of sample points of the fourth electro-optical transfer function according to the third color space, wherein the interpolation includes converting the third set of sample points and the fourth set of sample points to a common electro-optical transfer function in a common color space.
[0252] 8. The device as described in any one of EEE 1 to 7, wherein the electronic processor is further configured to determine the backward shaping function based on backward shaping function data from a least mean square error predictor.
[0253] 9. The device as described in any one of EEE 1 to 8, wherein the device is an encoder.
[0254] 10. A method for converting a signal corresponding to a first electro-optical transfer function into a signal corresponding to a second electro-optical transfer function, the method comprising:
[0255] Determining a set of sample points based on the synthesized data;
[0256] Defining a first set of sample points from the set of sample points according to a first electro-optical transfer function of a first color space;
[0257] Converting the first set of sample points according to the first electro-optical transfer function to the second electro-optical transfer function via a mapping function, thereby generating a second set of sample points according to the second electro-optical transfer function; and
[0258] Determining a backward shaping function based on the first set of sample points according to the first electro-optical transfer function and the second set of sample points according to the second electro-optical transfer function.
[0259] 11. The method according to EEE 10, wherein determining the backward shaping function includes repeatedly applying and adjusting a sample backward shaping function to minimize the difference between the result of the sample backward shaping function and the second set of sample points according to the second electro-optical transfer function.
[0260] 12. The method according to EEE 10 or 11, wherein converting the first set of sample points according to the first electro-optical transfer function to the second electro-optical transfer function via the mapping function includes applying a signal legalization function to the first set of sample points according to the first electro-optical transfer function, the signal legalization function being configured to modify the input to conform to a predetermined range.
[0261] 13. The method according to EEE 12, wherein the signal legalization function implements hard clipping.
[0262] 14. The method according to EEE 12 or 13, wherein the signal legalization function is a piecewise linear function.
[0263] 15. The method according to EEE 12 or 13, wherein the signal legalization function is an S-shaped curve.
[0264] 16. The method according to any one of EEE 10 to 15, wherein the first electro-optical transfer function is a hybrid logarithmic gamma.
[0265] 17. The method according to any one of EEE 10 to 16, wherein the second electro-optical transfer function is a perceptual quantizer.
[0266] 18. The method according to any one of EEE 10 to 17, wherein defining the first set of sample points according to the first electro-optical transfer function includes:
[0267] generating a third set of sample points of a third electro-optical transfer function according to a second color space from the set of data points, and
[0268] generating the first set of sample points according to the first electro-optical transfer function based on the third set of sample points of the third electro-optical transfer function according to the second color space, wherein the second color space is smaller than the first color space.
[0269] 19. The method according to EEE 18, wherein defining the first set of sample points according to the first electro-optical transfer function includes: interpolating the third set of sample points of the third electro-optical transfer function according to the second color space and a fourth set of sample points of a fourth electro-optical transfer function according to a third color space, such that the first set of sample pixels according to the first electro-optical transfer function includes a weighted combination of the third set of sample points of the third electro-optical transfer function according to the second color space and the fourth set of sample points of the fourth electro-optical transfer function according to the third color space, wherein the interpolation includes converting the third set of sample points and the fourth set of sample points to a common electro-optical transfer function in a common color space.
[0270] 20. The method according to any one of EEE 10 to 19, wherein the backward shaping function is further determined based on backward shaping function data from a least mean square error predictor.
[0271] 21. A non-transitory computer-readable medium storing instructions that, when executed by a processor of a computer, cause the computer to perform the method according to any one of EEE 10 to 20.
Claims
1. An apparatus for determining a backward shaping function, the apparatus comprising: an electronic processor configured to: determine a set of sample pixels from received video data; define a first set of sample pixels from the set of sample pixels according to a first electro-optical transfer function in a first color representation of a first color space; convert the first set of sample pixels according to the first electro-optical transfer function to a second electro-optical transfer function in the first color representation of the first color space via a mapping function, thereby generating a second set of sample pixels according to the second electro-optical transfer function from the first set of sample pixels; convert the first set of sample pixels and the second set of sample pixels from the first color representation to a second color representation of the first color space; determine a backward shaping function based on the converted first set of sample pixels and the converted second set of sample pixels; wherein the electronic processor is configured to determine the backward shaping function by repeatedly applying and adjusting a sample backward shaping function to minimize the difference between predicted pixel values obtained by applying the sample backward shaping function to pixels in the converted first set of sample pixels and pixels in the converted second set of sample pixels.
2. The device according to claim 1, wherein, The received video data includes one or more first images in a first dynamic range, and the second set of sample pixels belongs to one or more second images in a second dynamic range, the first dynamic range being lower than the second dynamic range, and wherein the electronic processor is further configured to: - determine a first cumulative density function based on a first histogram generated from a first codeword distribution in the one or more first images, - determine a second cumulative density function based on a second histogram generated from a second codeword distribution in the one or more second images, and - determine a histogram transfer function based on the first cumulative density function and the second cumulative density function for determining the backward shaping function.
3. The device according to claim 2, wherein, The electronic processor is further configured to determine a predicted value by applying a predictor for minimizing the mean square error.
4. The device according to claim 3, wherein, The electronic processor is configured to use the predictor to map each codeword from the first codeword distribution to the second codeword distribution to determine the histogram transfer function.
5. The device according to claim 3, wherein A plurality of parameters of the minimum mean square error predictor are determined based on a multi-channel multivariate regression model.
6. The device according to any one of claims 2 to 4, wherein, The backward shaping function is a luminance backward shaping function.
7. The device according to claim 5, wherein, The backward shaping function is a chrominance backward shaping function.
8. The device according to claim 1, wherein, The electronic processor is further configured to determine the backward shaping function based on a smooth equal-weight backward lookup table.
9. The device according to claim 1, wherein, The electronic processor is configured to determine the set of sample pixels as a three-dimensional pixel array q of sample pixels ijk , where i indicates the pixel position of the corresponding one-dimensional pixel array q having M samples i , and where j and k are the frame index and depth of the pixel 10. The device according to claim 1, wherein, Converting the first set of sample pixels according to the first electro-optical transfer function to the second electro-optical transfer function via the mapping function includes applying a signal legalization function to the first set of sample pixels for forcing the range of the first set of sample pixels to be within a predetermined range.
11. The device according to claim 10, wherein, The signal legalization function is one of the following functions: - a clipping function that includes clipping sample pixels in the first set of sample pixels outside the predetermined range - Piecewise linear function, or - S-curve function.
12. The device according to claim 1, wherein, The first electro-optical transfer function is a hybrid logarithmic gamma.
13. The device according to claim 1, wherein, The second electro-optical transfer function is a perceptual quantizer.
14. The device according to claim 1, wherein, Defining the first set of sample pixels according to the first electro-optical transfer function includes: Generating a third set of sample pixels of a third electro-optical transfer function in the first color representation according to a second color space from the set of sample pixels, and Generating the first set of sample pixels according to the first electro-optical transfer function based on the third set of sample pixels of the third electro-optical transfer function according to the second color space, wherein the second color space is smaller than the first color space.
15. The device according to claim 14, wherein, Defining the first set of sample pixels according to the first electro-optical transfer function includes: interpolating the third set of sample pixels of the third electro-optical transfer function in the first color representation according to the second color space and the fourth set of sample pixels of the fourth electro-optical transfer function in the first color representation according to a third color space, such that the first set of sample pixels according to the first electro-optical transfer function includes a weighted combination of the third set of sample pixels of the third electro-optical transfer function according to the second color space and the fourth set of sample pixels of the fourth electro-optical transfer function according to the third color space.
16. The device according to claim 14, wherein, The electronic processor is configured to convert the third electro-optical transfer function in the first color representation of the second color space to a container of the first electro-optical transfer function in the first color representation of the first color space.
17. The device according to claim 15, wherein, The electronic processor is configured to convert the fourth electro-optical transfer function of the third color space to a container of the first electro-optical transfer function in the first color representation of the first color space.
18. The device according to claim 1, wherein The device is an encoder or a decoder.
19. A method for determining a backward shaping function, the method comprising: Determining a set of sample pixels from received video data; Defining a first set of sample pixels from the set of sample pixels according to a first electro-optical transfer function in a first color representation of a first color space; Converting the first set of sample pixels according to the first electro-optical transfer function to a second electro-optical transfer function in the first color representation of the first color space via a mapping function, thereby generating a second set of sample pixels according to the second electro-optical transfer function from the first set of sample pixels; Converting the first set of sample pixels and the second set of sample pixels from the first color representation to a second color representation of the first color space, and Determining a backward shaping function based on the converted first set of sample pixels and the converted second set of sample pixels, wherein determining the backward shaping function includes repeatedly applying and adjusting a sample backward shaping function to minimize the difference between the predicted pixel values obtained by applying the sample backward shaping function to the pixels in the converted first set of sample pixels and the pixels in the converted second set of sample pixels.
20. The method according to claim 19, wherein, Converting the first set of sample pixels according to the first electro-optical transfer function to the second electro-optical transfer function via the mapping function includes applying a signal legitimization function to the first set of sample pixels for forcing the range of the first set of sample pixels within a predetermined range.
21. The method according to claim 20, wherein, The signal legitimization function is one of the following functions: - A clipping function that includes clipping sample pixels in the first set of sample pixels outside the predetermined range, - A piecewise linear function, or - An S-curve function.
22. The method according to any one of claims 19 to 21, wherein, The first electro-optical transfer function is a hybrid logarithmic gamma.
23. The method according to any one of claims 19 to 21, wherein The second electro-optical transfer function is a perceptual quantizer.
24. The method according to any one of claims 19 to 21, wherein, Defining the first set of sample pixels according to the first electro-optical transfer function includes: Generating a third set of sample pixels of a third electro-optical transfer function in a first color representation according to a second color space from the set of sample pixels, and Generating the first set of sample pixels according to the first electro-optical transfer function based on the third set of sample pixels of the third electro-optical transfer function in the second color space, wherein the second color space is smaller than the first color space.
25. The method according to claim 24, wherein, Defining the first set of sample pixels according to the first electro-optical transfer function includes: interpolating the third set of sample pixels of the third electro-optical transfer function in a first color representation according to the second color space and a fourth set of sample pixels of a fourth electro-optical transfer function in the first color representation according to a third color space such that the first set of sample pixels according to the first electro-optical transfer function includes a weighted combination of the third set of sample pixels of the third electro-optical transfer function in the second color space and the fourth set of sample pixels of the fourth electro-optical transfer function in the third color space.
26. The method according to claim 24, further comprising converting a third electro-optical transfer function of a first color representation of the second color space to a container of a first electro-optical transfer function in a first color representation of the first color space.
27. The method according to claim 25, further comprising converting a fourth electro-optical transfer function of the third color space to a container of a first electro-optical transfer function in a first color representation of the first color space.
28. The method according to any one of claims 19 to 21, wherein The predicted value of the sample post-shaping function is determined by applying a predictor for minimizing the mean square error.
29. A computer-readable medium storing instructions that, when executed by a processor of a computer, cause the computer to perform the method according to any one of claims 19 to 28.
30. A computer program product having one or more computer programs that, when executed by a processor of a computer, cause the computer to perform the method according to any one of claims 19 to 28.
Citation Information
Patent Citations
Motion vector derivation method, moving picture coding method and moving picture decoding method
US10264278B2
Inverse luma / chroma mappings with histogram transfer and approximation
US10264287B2
Multiple color channel multiple regression predictor
US8811490B2
Display method and display device
US20180048845A1
Inverse luma / chroma mappings with histogram transfer and approximation
US20180098094A1