High dynamic range video format with low dynamic range compatibility
The described video format with a recovery map track adapts HDR content for both LDR and HDR displays, addressing inefficiencies in existing formats by preserving image quality and reducing storage needs.
Patent Information
- Application Number
- JP2025505597
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-05-31
- Filing Date
- 2023-10-30
- Publication Date
- 2025-11-07
AI Technical Summary
Existing video formats struggle to efficiently display high dynamic range (HDR) content on both low dynamic range (LDR) and HDR display devices without losing visual quality or requiring specialized hardware, and current HDR-to-LDR conversion methods often result in reduced image quality due to global tone mapping.
A video format that includes a recovery map track encoding luminance differences between HDR and LDR frames, allowing devices to adapt the dynamic range of the video to match their display capabilities by applying recovery map gains, enabling both LDR and HDR display without specialized hardware.
Enables efficient storage and high-quality display of HDR or LDR videos on a wide range of devices, preserving local contrast and detail, and reducing storage requirements, while allowing dynamic range adjustment based on display capabilities.
Smart Images

Figure 2025536502000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 421,155, filed October 31, 2022, and entitled "HIGH DYNAMIC RANGE IMAGE FORMAT WITH LOW DYNAMIC RANGE COMPATIBILITY," and to U.S. Provisional Patent Application No. 63 / 439,271, filed January 16, 2023, and entitled "HIGH DYNAMIC RANGE IMAGE AND VIDEO FORMATS WITH LOW DYNAMIC RANGE COMPATIBILITY," and to International Application No. PCT / US2023 / 023998, filed May 31, 2023, and entitled "HIGH DYNAMIC RANGE IMAGE FORMAT WITH LOW DYNAMIC RANGE COMPATIBILITY," the entire contents of all of which are incorporated herein by reference. [Background technology]
[0002] Users of devices such as smartphones or other digital camera devices capture and store large numbers of photos and videos in their image libraries. High dynamic range (HDR) images and videos can be captured by some cameras, which offer greater dynamic range and more realistic image quality than low dynamic range (LDR) images and videos captured by many other (e.g., older) cameras. HDR images and videos are best viewed on display devices capable of displaying the full dynamic range of HDR images and videos.
[0003] The background statement provided herein is intended to generally provide a context for the disclosure. The work of the inventors identified herein, to the extent that it is described in this Background section, and aspects of the present specification that may not qualify as prior art at the time of filing, are not admitted expressly or impliedly as prior art to the present disclosure. Summary of the Invention
[0004]
[0003] Embodiments described herein relate to methods, devices, and computer-readable media for providing an HDR video format with low dynamic range compatibility. In some embodiments, the computer-implemented method includes acquiring a first video including a plurality of first frames, each depicting a respective scene, the first frames having a first dynamic range; acquiring a second video including a plurality of second frames, each depicting a respective scene of a corresponding one of the first frames, the second frames having a second dynamic range different from the first dynamic range; and generating a recovery map track. Generating the recovery map track includes, for each first one of the first frames and a corresponding second one of the second frames, generating a recovery map frame based on the first frame and the corresponding second frame, the recovery map frame encoding a luminance difference between a portion of the first frame and a corresponding portion of the corresponding second frame. The first video and the recovery map track are provided in a video container, the video container being readable to display a derived video based on applying the recovery map track to the first video, the derived video including a plurality of derived frames having a dynamic range different from the first dynamic range.
[0005] Various implementations of this method are described. In some implementations, the luminance difference is scaled by a range scaling factor including a ratio of a maximum luminance of the first frame to a maximum luminance of the corresponding second frame. In some implementations, the method further includes generating a metadata track including respective metadata frames associated with corresponding recovery map frames, and providing the first video and recovery map tracks in a video container includes providing the metadata tracks in the video container. In some implementations, the luminance difference is scaled by a range scaling factor including a ratio of a maximum luminance of the first frame to a maximum luminance of the corresponding second frame, and each metadata frame in the timed metadata track includes a range scaling factor associated with an associated recovery map frame in the recovery map track.
[0006] In some implementations, the second dynamic range is greater than the first dynamic range and the dynamic range of the derived frame is greater than the first dynamic range. In some implementations, the video container is readable to display the first video by a first display device capable of displaying the first dynamic range and to display the derived video by a second display device capable of displaying a dynamic range greater than the first dynamic range. In some implementations, obtaining the first video includes performing range compression on the second video.
[0007] In some embodiments, generating each recovery map frame includes encoding a luminance gain such that applying the luminance gain to the luminance of each pixel in the first frame results in a corresponding pixel in the corresponding second frame. In some embodiments, generating each recovery map frame includes encoding the luminance gain in logarithmic space, where the value of the recovery map frame is proportional to the difference in logarithms of luminance divided by the logarithm of a range scaling factor. In some embodiments, generating each recovery map frame is based on recovery(x,y)=log(pixel_gain(x,y)) / log(range scaling factor), where recovery(x,y) is the recovery map frame for pixel location (x,y) in the corresponding second frame, and pixel_gain(x,y) is the ratio of luminance at location (x,y) in the corresponding second frame relative to the first frame.
[0008] In some implementations, generating each recovery map frame includes encoding the recovery map frame into a bilateral grid. In some implementations, providing a video container with a recovery map track includes encoding the recovery map track to have a different resolution than the first video and the same aspect ratio as the first video. In some implementations, the method further includes obtaining a video container, determining to display the second video by a second display device, scaling a luminance of a plurality of pixels of a first frame of the first video within the frame container based on a particular luminance output of the second display device and based on the corresponding recovery map frame to obtain a derived frame, and displaying the derived frame by the second display device as an output frame having a different dynamic range than the first frame. In some implementations, the method further includes determining a maximum luminance display capability of the second display device, and scaling the luminance of the plurality of pixels includes increasing a luminance of highlights in the first video to a luminance level that is equal to or less than the maximum luminance display capability.
[0009] In some implementations, the second dynamic range is lower than the first dynamic range and the dynamic range of the derived frame is lower than the first dynamic range. In some implementations, the video container is readable to display the first video by a first display device capable of displaying the first dynamic range and to display the derived video by a second display device capable of displaying only a dynamic range lower than the first dynamic range. In some implementations, the method includes performing range compression on the first video to obtain the second video.
[0010] In some embodiments, generating a recovery map track includes generating a plurality of recovery map tracks, and providing the recovery map track in the video container includes providing a plurality of recovery map tracks in the video container, each of the plurality of recovery map tracks encoding a luminance difference between a portion of a first frame and a corresponding portion of a corresponding second frame, and each of the plurality of recovery map tracks readable from the video container to display a respective derived video, each respective derived video having one or more characteristics different from each other. In some embodiments, the one or more characteristics of the respective derived videos different from each other include dynamic range, a larger dynamic range provided from applying a first recovery map track of the plurality of recovery map tracks and a lower dynamic range provided from applying a second recovery map track.
[0011] In some implementations, a computer-implemented method includes obtaining a portion of a video container, the portion including: a first video including a plurality of first frames, each depicting a respective scene, the first frames having a first dynamic range; and a recovery map track including a plurality of recovery map frames, each recovery map frame corresponding to a respective one of the first frames. Each recovery map frame encodes a luminance gain of pixels of the corresponding first frame scaled by an associated range scaling factor including a ratio of a maximum luminance of the corresponding first frame to a maximum luminance of a corresponding second frame, the corresponding second frame depicting a respective scene of the first frame and having a second dynamic range different from the first dynamic range. The method includes determining whether to display the first video or one of a derived video including a plurality of derived frames having a dynamic range different from the first dynamic range. In response to determining to display the first video, the method includes causing at least a portion of the first video to be displayed by a display device, and in response to determining to display the derived video, the method includes applying a gain of a recovery map frame of the recovery map track to a luminance of a pixel of a corresponding first frame to determine a corresponding pixel value of each corresponding derived frame of the derived video, and causing at least a portion of the derived video to be displayed by a display device.
[0012] Various implementations of this method are described. In some implementations, the video container includes a metadata track including respective range scaling coefficients associated with recovery map frames of the recovery map track, each of the respective range scaling coefficients being used to scale a luminance gain of a pixel of a corresponding first frame, the luminance gain being provided by the associated recovery map frame. In some implementations, the display device is capable of displaying a display dynamic range different from the first dynamic range, and applying the gain of the recovery map includes adapting a luminance of pixel values of the corresponding derived frame to a display dynamic range of the display device. In some implementations, the dynamic range of the derived frame has a maximum luminance that is the lesser of a maximum luminance of the display device and a maximum luminance of a second dynamic range of a second frame used to generate the recovery map frame.
[0013] In some implementations, applying a recovery map gain in response to determining to display the derived video includes scaling the luminance of the first frame based on a particular luminance output of the display device and based on the corresponding recovery map frame. In some implementations, scaling the luminance of the first frame is performed based on derived_frame(x,y) = first_frame(x,y) + log(display_factor) * recovery(x,y), where derived_frame(x,y) is a log-space version of the corresponding derived frame, first_frame(x,y) is a log-space version of the first frame recovered from the frame container, display_factor is the minimum and maximum display luminance of the second display device, recovery(x,y) is the recovery map for pixel location (x,y) of the first frame, and the range scaling factor is the ratio of the maximum luminance of the first frame to the maximum luminance of the corresponding second frame.
[0014] In some implementations, each recovery map frame is encoded with a bilateral grid, and the method further includes decoding the recovery map frame from the bilateral grid. In some implementations, the second dynamic range is lower than the first dynamic range, and the dynamic range of the derived frame is lower than the first dynamic range. In some implementations, the method further includes decoding additional information from the video container in response to determining to display the derived video, the additional information including a range scaling factor. In some implementations, the method further includes decoding the recovery map track from blocks included in a frame container separate from the first video. In some implementations, determining whether to display one of the first video or the derived video is performed by a server device, the determination being based on a display dynamic range of a display device included in or coupled to the client device, causing at least a portion of the first video to be displayed by the display device includes streaming the first video from the server device to the client device so that the client device displays the first video, and causing at least a portion of the derived video to be displayed by the display device includes streaming the derived video from the server device to the client device so that the client device displays the derivative video.In some implementations, obtaining a portion of the video container includes receiving a first video and a recovery map track by a client device from a server device, determining whether to display one of the first video or the derived video is performed by the client device, the determination being based on a display dynamic range of a display device included in or coupled to the client device, causing at least a portion of the first video to be displayed by the display device is performed by the client device, and applying a gain of the recovery map frame and causing at least a portion of the derived video to be displayed by the display device is performed by the client device.
[0015] In some implementations, a computer-implemented method includes obtaining a video container, the video container including: a first video including a plurality of first frames, each of the first frames depicting a respective scene, the first frames having a first dynamic range; and a recovery map track including a plurality of recovery map frames, each recovery map frame corresponding to a respective one of the first frames, each recovery map frame encoding a luminance gain of pixels of a corresponding first frame scaled by a range scaling factor including a ratio of a maximum luminance of the corresponding first frame to a maximum luminance of a corresponding second frame, the corresponding second frame depicting a respective scene of the first frame and having a second dynamic range different from the first dynamic range. The method includes determining to display a derived video including a plurality of derived frames having a dynamic range different from the first dynamic range; and, in response to determining to display the derived video, applying gains of recovery map frames of the recovery map track to luminances of pixels of corresponding first frames to determine corresponding pixel values of each of the corresponding derived frames of the derived video; and displaying at least a portion of the derived video by a display device.
[0016] In some implementations, a system includes a processor and a memory coupled to the processor, the memory storing instructions that, when executed by the processor, cause the processor to perform operations. The operations include obtaining a video container including a first frame video including a plurality of first frames, the first frames having a first dynamic range, and a recovery map track including a plurality of recovery map frames, each recovery map frame corresponding to a respective one of the first frames and each recovery map frame encoding a luminance gain of a pixel of the corresponding first frame. The operations include determining whether to display one of the first video or a derivative video including a plurality of derived frames corresponding to the first frames of the depicted subject matter and having a second dynamic range different from the first dynamic range. In response to determining to display the first video, at least a portion of the first video is displayed by a first display device. In response to determining to display the derived video, the method applies a gain of the recovery map frame of the recovery map track to a luminance of a pixel of the corresponding first frame to determine a corresponding pixel value of each corresponding derived frame of the derived video, and causes at least a portion of the derived video to be displayed by the second display device, wherein applying the gain includes scaling a luminance of the corresponding first frame based on a particular luminance output of the second display device and based on the recovery map frame.
[0017] In some implementations, a computer-implemented method includes obtaining a video container, the video container including: a first video including a plurality of first frames, each depicting a respective scene, the first frames having a first dynamic range; and a recovery map track including a plurality of recovery map frames, each corresponding to a respective one of the first frames, each recovery map frame encoding a luminance gain for pixels of the corresponding first frame. The method includes determining to display a derived video including a plurality of derived frames corresponding to the first frames of the depicted subject matter, the derived video having a second dynamic range different from the first dynamic range. In response to determining to display the derived video, the method applies gains of the recovery map frames to luminances of pixels of the corresponding first frames to determine corresponding pixel values for each of the corresponding derived frames of the derived video, and displays at least a portion of the derived video by a display device. Applying the gains includes scaling the luminance of the first frames based on a specific luminance output of the second display device and based on the recovery map frames.
[0018] Some implementations may include a computing device including a processor and a memory coupled to the processor, the memory may store instructions that, when executed by the processor, cause the processor to perform operations including one or more of the operations and / or features of any of the methods described above.
[0019] Some embodiments include a non-transitory computer-readable medium having stored thereon instructions that, when executed by a processor, cause the processor to perform operations that may be similar to the operations and / or features of any of the methods, systems, and / or computing devices described above. [Brief explanation of the drawings]
[0020] [Figure 1] FIG. 1 is a block diagram of an exemplary network environment that may be used in one or more implementations described herein. [Figure 2] 1 is a flow diagram illustrating an example method for encoding video into a backward-compatible high dynamic range video format, according to some implementations. [Figure 3] 1 is a flow diagram illustrating an example method for generating a recovery map track based on HDR and LDR videos, according to some implementations. [Figure 4] 1 is a flow diagram illustrating an example method for encoding a recovery map into a bilateral grid, according to some implementations. [Figure 5] 1 is a flow diagram illustrating an example method for decoding video into a backward-compatible high dynamic range video format and displaying HDR video, according to some implementations. [Figure 6] 1 is a diagrammatic representation of an exemplary video container that may be used to provide a backward-compatible high dynamic range video format, according to some implementations. [Figure 7] 1 is a diagram of an exemplary image showing high dynamic range, according to some implementations. [Figure 8] 1A-1C are diagrams of exemplary images depicting low dynamic range, according to some implementations. [Figure 9] 1A-1C are diagrams of exemplary images depicting low dynamic range, according to some implementations. [Figure 10] 1A-1C are diagrams of exemplary images depicting low dynamic range, according to some implementations. [Figure 11] FIG. 1 is a block diagram of an example computing device that can be used to implement one or more features described herein. DETAILED DESCRIPTION OF THE INVENTION
[0021] This disclosure relates to a backward-compatible HDR video format. The video format provides a container that can capture and display both low dynamic range (LDR) video and high dynamic range (HDR) video. The video format can be used to display an LDR version of the video by an LDR display device and an HDR version of the video by an HDR display device.
[0022] In some implementations, a video container implementing a video format is created. A video with a lower dynamic range (e.g., an LDR video) and a corresponding video with a larger (e.g., higher) dynamic range (e.g., an HDR video) are obtained. In some examples, the HDR video is captured by a camera, and the LDR video is created from the HDR video, for example, using tone mapping or other processes. The LDR video can be in a standard video format, for example, using a standard video codec. A recovery map track is generated based on the HDR and LDR videos, including recovery map frames, each of which encodes a luminance difference (e.g., gain) between a portion (e.g., a pixel or pixel value) of an associated frame of the LDR video and a corresponding portion (e.g., a pixel or pixel value) of a corresponding frame of the HDR video. In some implementations, the difference is scaled by a range scaling factor associated with the recovery map frame, the range scaling factor including a ratio of the maximum luminance of a frame of the HDR video to the maximum luminance of the corresponding frame of the LDR video. In some exemplary implementations, each recovery map frame can be a scalar function that encodes luminance gain in logarithmic space, with the recovery map value being proportional to the logarithmic difference in luminance divided by the logarithm of the range scaling factor. In some implementations, the recovery map track can be encoded into a data structure or data representation that can reduce the storage space required to store the recovery map track. The recovery map track and LDR video are provided in a video container that can be stored, for example, as a new HDR video format.
[0023] The video container can be read and processed by a device to display output video that is LDR video or output HDR video (e.g., derived video). A device that reads and displays video stored in the described video formats can use an LDR display device (e.g., a display screen) that can display LDR video, e.g., that has a display dynamic range that can only display up to the low dynamic range of LDR video and cannot display the greater dynamic range of HDR video. For such an LDR display device, the device accesses the LDR video in the video container and ignores the recovery map track. Because the LDR video is in a standard format, the device can easily display the video. Thus, in some implementations, the LDR video can still be displayed by a device that does not implement code to detect or apply the recovery map track. If the accessing device can use a display device that can display HDR video (e.g., that can display a luminance range at its output that is greater than the maximum luminance of the LDR video), the device accesses the LDR video and recovery map in the video container and applies the luminance gain encoded in the recovery map to pixels of the LDR video to determine each corresponding pixel value of the output HDR video to be displayed. In some examples, the device scales pixel brightness of the LDR video based on the brightness output of the HDR display device (e.g., the maximum brightness output of the display device) and the brightness gain stored in the recovery map. The output HDR video is displayed on the HDR display device with a high dynamic range. In some implementations, the HDR video and the recovery map track can be stored in a video container, and the recovery map track can be applied to the HDR video to obtain an LDR video with a lower dynamic range, which has high visual quality because it is locally tone-mapped from the HDR video via the recovery map.
[0024] The described features provide several technical advantages, including enabling efficient storage and high-quality display of high-dynamic-range video or lower-dynamic-range (e.g., standard) video by a wide range of devices and using a single video format. For example, video formats are provided that enable any device to display video with a dynamic range appropriate to its display capabilities. Furthermore, video displayed from these formats does not lose visual quality (particularly its local contrast, e.g., detail) due to conversion from one dynamic range to another. In various implementations, the output HDR video generated from the video container can be a lossless, exact version of the original HDR video used to create the video container, or the output HDR video can be a version that is (visually) very similar to the original HDR video due to a recovery map that stores luminance information.
[0025] Described features may include using a recovery map to provide an output HDR video derived from an LDR video. The recovery map may include values based on a luminance gain ratio scaled by a range scaling factor, which is the ratio of the maximum luminance of the original HDR video frame to the corresponding LDR video frame. For example, the recovery map indicates the relative amount to scale each pixel, and the range scaling factor indicates the specific amount of scaling to be performed, which provides context for the relative values in the recovery map (e.g., expanding the range of amounts by which pixels can be scaled for a particular image). The range scaling factor is advantageous in that it allows for efficient, accurate, and compact specification of recovery map values in a normalized range, which can then be used to efficiently adjust for a particular output dynamic range. Using the range scaling factor efficiently uses all bits used to store the recovery map, thereby enabling a wider variety of videos to be accurately encoded.
[0026] Additionally, in some embodiments or cases, when displaying the output HDR video, pixel luminance of the LDR video frames is also scaled based on a display factor. The display factor can be based on the luminance output of the HDR display device displaying the output video. This allows HDR video with any dynamic range beyond LDR to be generated from the container based on the display capabilities of the display device. The dynamic range of the output HDR video is not constrained by any standard high dynamic range or the dynamic range of the original HDR video. Because HDR display devices can vary significantly in terms of the luminance at which they can display images, an advantage of this feature is that the video can be displayed with high-quality rendition on display devices of any dynamic range. Furthermore, the described features enable modification of the dynamic range of the output HDR video (e.g., beyond LDR) in real time based on user input or other conditions, for example, to reduce user visual strain or fatigue, or for other uses. Display scaling allows the output HDR video to have any dynamic range beyond that of the LDR video.
[0027] In contrast, conventional techniques may encode HDR images for displays with a specific, fixed dynamic range or maximum brightness. Such techniques cannot take advantage of HDR display devices with a dynamic range or brightness greater than the encoded dynamic range. Furthermore, such conventional techniques may produce low-quality images if the HDR display device is unable to display as high a dynamic range as encoded in the original HDR image. For example, techniques such as roll-off curves may be required to reduce the dynamic range of an HDR image for display on an HDR display device, which often undesirably reduces the local contrast of the image and reduces the visual quality of the image.
[0028] In further contrast, traditional HDR video technologies may require specialized hardware to decode data provided in a specific format. Devices lacking the necessary hardware may be unable to decode the video at all without specialized software that may be underpowered. Some devices may also lack displays capable of rendering content in HDR. Furthermore, HDR video typically requires tone mapping to convert the video to a lower dynamic range (e.g., standard dynamic range, or SDR), a common use case for supporting legacy software not designed to work with HDR content or for displaying video on displays that do not adequately support HDR content. Existing HDR-to-LDR conversion technologies may not preserve all of the desired artistic intent in the video because such conversion takes the form of a single global tone mapping, lacking the information necessary to do so reliably for a variety of content. For example, the hybrid-log-gamma (HLG) transfer function, while designed to be somewhat backward compatible, can actually cause noticeable degradation when rendered by a 10-bit HDR hardware decoder and displayed in an 8-bit SDR environment. Additionally, both HLG and Perceptual Quantization (PQ) transfer functions provide global tone mapping and do not account for lighting differences across a particular scene. Some formats may have limitations in capturing local tone mapping differences within a frame. As a result, some parts of a scene may take better advantage of the higher dynamic range than others, potentially resulting in a loss of detail or fidelity.
[0029] The described features provide a sharable, backward-compatible mechanism for creating and rendering video containing high dynamic range (HDR) content beyond formats such as 10-bit HDR. The described video format enables next-generation consumer HDR content and, in some implementations, may also enable support for professional video workflows without requiring dedicated hardware codecs for decoding / encoding. One or more implementations of the described HDR format also address issues with current HDR transfer functions, such as HLG and PQ, which assume global tone mapping. In contrast, the described features enable local tone mapping, which is superior at preserving detail in scenes with varying levels of brightness. Thus, the described HDR format can produce HDR-rendered content with higher fidelity and quality compared to traditional HDR formats.
[0030] Additionally, the described features can reduce the storage requirements of the video format. For example, the described video format allows a single LDR video to be stored with a recovery map track that requires less storage, thus saving significant storage space over formats that store multiple videos. For example, some implementations include compressing the recovery map track, and in some of these examples, the recovery map track can be encoded into a bilateral grid. Such a grid can store the recovery map track with significantly reduced storage space requirements and enable the recovery map track to provide an output HDR video with little loss in visual quality compared to the original HDR video.
[0031] A technical effect of one or more described implementations is that a device can display video having a dynamic range that better corresponds to the dynamic range of a particular output display device being used to display the video, compared to conventional systems. For example, such conventional systems may provide video that does not have a dynamic range that corresponds to the dynamic range of the display device, resulting in reduced video display quality. Features described herein can reduce such drawbacks, for example, by providing a recovery map track in the video format and / or scaling the video output to better configure the video dynamic range for a particular display device. Additionally, a technical effect of one or more described implementations is that a device consumes fewer computational resources to obtain results. For example, a technical effect of the described techniques is reduced consumption of system processing and / or storage resources, compared to conventional systems that do not provide one or more of the described techniques or features. For example, conventional systems may need to store and / or provide full LDR video and full HDR video to display video with an appropriate dynamic range for a particular display device, requiring additional storage and communication bandwidth resources, compared to the described techniques. In another example, conventional systems may rely on storing only HDR video and then tone mapping the HDR video to a lower dynamic range for LDR display, which can degrade the visual quality of the HDR video when presented on an LDR display for a wide variety of HDR videos.
[0032] As referred to herein, dynamic range relates to the ratio between the brightest and darkest parts of a scene. The dynamic range of an LDR format typically does not exceed a certain low range, such as a standard dynamic range (SDR) file in a low-range color space / color profile. Similarly, the displayed dynamic range output of an LDR display device is low. As referred to herein, an image with a larger (or higher) dynamic range than an LDR image (e.g., JPEG) or LDR video is considered an HDR image or HDR video, and a display device capable of displaying that larger dynamic range is considered an HDR display device. An HDR image or HDR video can store pixel values that span a wider tonal range than an LDR. In some examples, an HDR image or HDR video can more accurately represent the dynamic range of a real-world scene and / or may have a lower dynamic range that is greater than the dynamic range of an LDR image or LDR video (e.g., HDR images and videos may generally exceed 8 bits per color channel).
[0033] The use of "logarithm" in this description refers to a particular base of logarithm, and can be any base number, for example, all logarithms herein can be base 2, or base 10, etc.
[0034] In addition to the description herein, users may be provided with controls that allow them to make choices about both whether and when the systems, programs, or features described herein enable the collection of user information (e.g., images from the user's library, social networks, social actions or activities, occupation, user preferences, or the user's current location, the user's messages, or the user's device characteristics), as well as whether content or communications are sent from the server to the user. Furthermore, certain data may be processed in one or more ways so that personally identifiable information is removed before it is stored or used. For example, a user's identification information may be processed so that personally identifiable information about the user cannot be determined, or if location information is obtained (e.g., to the city, zip code, or state level), the user's geographic location may be generalized so that the user's specific location cannot be determined. Thus, users may control what information is collected about them, how that information is used, and what information is provided to them.
[0035] 1 shows a block diagram of an exemplary network environment 100 that may be used in some implementations described herein. In some implementations, network environment 100 includes one or more server systems, e.g., server system 102 in the example of FIG. 1, and multiple client devices, e.g., client devices 120-126, each associated with a respective user, U1-U4. Server system 102 and each of client devices 120-126 may be configured to communicate with network 130.
[0036] Server system 102 can include server device 104 and video database 110. In some implementations, server device 104 can provide video application 106a. In FIG. 1 and the remaining figures, a letter following a reference number, e.g., "106a," denotes a reference to the element with that specific reference number. A reference number in text without a following letter, e.g., "106," denotes a general reference to an embodiment of the element with that reference number.
[0037] Video database 110 may be stored on a storage device that is part of server system 102. In some embodiments, video database 110 may be implemented using a relational database, a key-value structure, or other types of database structures. In some embodiments, video database 110 may include multiple partitions, each partition corresponding to a respective video library for each of users 1-4. For example, as seen in FIG. 1, video database 110 may include a first video library (Library 1, 108a) for user 1 and other video libraries (Library 2, ..., Library n) for various other users. While FIG. 1 depicts a single video database 110, it should be understood that video database 110 may be implemented as a distributed database, e.g., across multiple database servers. Furthermore, while FIG. 1 depicts multiple partitions, one for each user, in some embodiments, each video library may be implemented as a separate database.
[0038] Video library 108a may store multiple videos associated with user 1, metadata associated with multiple images, and one or more other database fields stored in association with the multiple videos. Access permissions for video library 108a may be restricted to allow user 1 to control how videos and other data in video library 108a may be accessed, for example, by video application 106, by other applications, and / or by one or more other users. Server system 102 may be configured to implement access permissions such that a particular user's video data is accessible only if authorized by the user.
[0039] A video, as referred to herein, can include multiple digital frames (images included in the video) having pixels with one or more pixel values (e.g., color values, brightness values, etc.). A video can be a sequence of images or frames that can optionally include audio, or can be a dynamic image (e.g., an animation, animated GIF, cinemagraph in which some images include movement and other images are static, etc.). A video, as used herein, can be understood as any of the above. In some implementations, a video can include a single image or frame.
[0040] Network environment 100 may include one or more client devices, e.g., client devices 120, 122, 124, and 126, which may communicate with each other and / or with server system 102 via network 130. Network 130 may be any type of communication network, including one or more of the Internet, a local area network (LAN), a wireless network, a switched or hubbed connection, etc. In some implementations, network 130 may include peer-to-peer communication between devices, for example, using a peer-to-peer wireless protocol (e.g., Bluetooth, Wi-Fi Direct, etc.). An example of peer-to-peer communication between two client devices 120 and 122 is indicated by arrow 132.
[0041] In various implementations, users 1, 2, 3, and 4 may communicate with server system 102 and / or with each other using respective client devices 120, 122, 124, and 126. In some examples, users 1, 2, 3, and 4 can interact with each other via applications executing on their respective client devices and / or server system 102 and / or via network services, such as a social networking service or other type of network, implemented on server system 102. For example, each client device 120, 122, 124, and 126 can communicate data with one or more server systems, such as server system 102.
[0042] In some implementations, server system 102 may provide appropriate data to client devices so that each client device can receive communicated or shared content uploaded to server system 102 and / or network services. In some examples, users 1-4 may interact via image / video sharing, voice or video conferencing, audio, video, or text chat, or other communication modes or applications.
[0043] The network services implemented by the server system 102 may include systems that enable users to perform various communications, form links and associations, upload and post shared content such as images, video, text, audio, and other types of content, and / or perform other functions. For example, a client device may display received data, such as content posts, that are transmitted or streamed to the client device and originate from different client devices (or directly from different client devices) via a server and / or network service, or that originate from a server system and / or network service. In some implementations, client devices may communicate directly with each other, for example, using peer-to-peer communication between client devices as described above. In some implementations, a "user" may include one or more programs or virtual entities as well as people interfacing with a system or network.
[0044] In some implementations, any of client devices 120, 122, 124, and / or 126 may host one or more applications. For example, as shown in FIG. 1, client device 120 may host video application 106b. Client devices 122-126 may also host similar applications. Video application 106a may be implemented using hardware and / or software on client device 120. In different implementations, video application 106a may be, for example, a standalone client application running on any of client devices 120-124, or may work in conjunction with video application 106b hosted on server system 102.
[0045] The video application 106 may provide various features related to video (and images) that are implemented with the user's permission. For example, such features may include one or more of capturing video using a camera, modifying the video, determining image / video quality (e.g., based on factors such as face size, blur, number of faces, image composition, lighting, exposure, etc.), storing the video in a video library 108, encoding and decoding the images and videos into any of a variety of image and video formats (including those described herein), providing a user interface for viewing the displayed video or other image-based production or compilation, etc.
[0046] Client device 120 may include video library 108b for user 1, which may be a standalone video library. In some implementations, video library 108b may be available in combination with video library 108a on server system 102. For example, with the user's permission, video library 108a and video library 108b may be synchronized over network 130. In some implementations, video library 108 may include multiple videos related to user 1, such as videos captured by the user (e.g., using the camera of client device 120 or other device), videos shared with user 1 (e.g., from the respective video libraries of other users 2-4), videos downloaded by user 1 (e.g., from a website, messaging application, etc.), screenshots, and other videos or images. In some implementations, video library 108b on client device 120 may include a subset of the videos in video library 108a on server system 102. For example, such an implementation may be advantageous if a limited amount of storage space is available on client device 120.
[0047] In various implementations, client device 120 and / or server system 102 may include other applications (not shown), which may be applications that provide various types of functionality. User interfaces on client devices 120, 122, 124, and / or 126 may enable the display of user content and other content, including images, videos, image-based works, data, and other content, as well as communications, privacy settings, notifications, and other data. Such user interfaces may be displayed using software on the client device, software on the server device, and / or a combination of client and server software running on server device 104, e.g., application software or client software communicating with server system 102. User interfaces may be displayed by a display device of the client or server device, e.g., a touchscreen or other display screen, a projector, etc. In some implementations, application programs running on the server system may communicate with client devices to receive user input at the client devices and output data, such as visual data, audio data, etc., at the client devices.
[0048] For ease of explanation, FIG. 1 shows one block for server system 102, server device 104, and video database 110, and four blocks for client devices 120, 122, 124, and 126. Server blocks 102, 104, and 110 may represent multiple systems, server devices, and network databases, and the blocks may be provided in different configurations than shown. For example, server system 102 may represent multiple server systems that can communicate with other server systems over network 130. In some implementations, server system 102 may include, for example, a cloud-hosted server. In some examples, video database 110 may be stored in a storage device separate from server device 104 and provided within server system block(s) that can communicate with server device 104 and other server systems over network 130.
[0049] Also, there may be any number of client devices. Each client device may be any type of electronic device, such as a desktop computer, a laptop computer, a portable or mobile device, a mobile phone, a smartphone, a tablet computer, a television, a television set-top box or entertainment device, a wearable device (e.g., display glasses or goggles, a wristwatch, a headset, an armband, jewelry, etc.), a personal digital assistant (PDA), a media player, a gaming device, etc. In some implementations, network environment 100 may not have all of the components shown and / or may have other elements, including other types of elements, instead of or in addition to the elements described herein.
[0050] Other implementations of the features described herein may use any type of system and / or service. For example, other networked services (e.g., connected to the Internet) may be used instead of or in addition to a social networking service. Any type of electronic device may utilize the features described herein. Some implementations may provide one or more features described herein on one or more client or server devices that are disconnected from or intermittently connected to a computer network. In some examples, a client device that includes or is connected to a display device may display posts of content stored on a storage device local to the client device, e.g., content previously received over a communications network.
[0051] FIG. 2 is a flow diagram illustrating an exemplary method 200 for encoding video in a backward-compatible high dynamic range video format, e.g., an HDR video format with low dynamic range (LDR) compatibility, according to some implementations. In some implementations, method 200 may be performed, for example, on server system 102 as shown in FIG. 1. In some implementations, some or all of method 200 may be implemented on one or more client devices, such as client devices 120, 122, 124, or 126 of FIG. 1, one or more server devices, such as server device 104 of FIG. 1, and / or both the server device(s) and the client device(s). In the illustrated example, an implementing system includes one or more digital processors or processing circuits (“processors”) and one or more storage devices (e.g., databases or other storage). In some implementations, different components of one or more servers and / or clients may perform different blocks or other portions of method 200. In some examples, devices are described as performing blocks of method 200. Some implementations may have one or more blocks of method 200 performed by one or more other devices (e.g., other client devices or server devices) that may send results or data to the first device.
[0052] Various types of devices can perform the encoding of video data and / or generation of recovery map data and video containers described in method 200. For example, in various implementations, such encoding and generation can be performed by a camera on a mobile device, a server device in the cloud or on a network (e.g., after receiving a video stream from a camera or other image / video capture device), a desktop computer, dedicated video hardware in the device, software implemented by a CPU or GPU, etc.
[0053] In some implementations, method 200, or portions thereof, may be initiated automatically by the system. For example, the method (or portions thereof) may be executed periodically or may be executed based on one or more specific events or conditions, such as a client device launching video application 106, the application's image capture device capturing new video, the device receiving video over the network, uploading new video to server system 102, a predetermined period of time having elapsed since the last execution of method 200, and / or the occurrence of one or more other conditions that may be specified in settings read by the method.
[0054] User permission may be obtained to use user data in method 200 (blocks 210-220). For example, the user data for which permission is obtained may include images or videos stored on a client device (e.g., any of client devices 120-126) and / or a server device, image / video metadata, user data associated with the use of a video application, image-based production, etc. The user is provided with the option to selectively provide permission to access all user data, any user data, or none of the user data. If the user's permission is insufficient for particular user data, method 200 may be performed without using the user data's permission, for example, using other data (e.g., images or videos not associated with the user).
[0055] Method 200 may begin at block 210. In block 210, high dynamic range (HDR) video is acquired by a device ("original HDR video"). The HDR video includes multiple frames, each depicting a respective scene. Each frame can be considered an image in this disclosure. In some implementations, the HDR video is provided in any of several standard formats for high dynamic range video. Examples of some HDR standards are HDR10, HDR10+, DolbyVision, Hybrid Log-Gamma (HLG), etc. In some examples, the HDR video may be captured by an image capture device, such as a camera on a client device or another device. In additional examples, the HDR video may be received over a network from a different device or retrieved from storage accessible by the device. In some implementations, the HDR video may be generated based on any of several image or video generation techniques, such as ray tracing. In some implementations, HDR video may be compressed, e.g., having frames that are independently encoded and frames that are predicted based on other frames, so that every complete frame displayed from the video does not need to be stored in an independently decodable manner. The techniques herein can optionally be used in combination with video encoding / decoding schemes, e.g., recovery map frames can be generated from raw video frames or decoded video frames.
[0056] In block 212, low dynamic range (LDR) video is acquired by the device (e.g., "original LDR video"). The LDR video includes multiple frames depicting the same scene and / or subject matter as corresponding frames of the HDR video. For example, frames of the LDR video correspond to frames of the HDR video in the depicted subject matter (and / or the LDR frames can have the same position or index from the beginning of the video as the corresponding HDR frames). Each frame can be considered an image in this disclosure. LDR video frames have a lower dynamic range than HDR video frames. In some implementations, the LDR video frames can have the dynamic range of a standard LDR video format or can have any dynamic range lower than the HDR video frames. For example, the LDR video can be provided in a standard video format, such as a format encoded with a bit depth of 8 bits or greater and having a Rec. 709 / sRGB color gamut, HEVC encoded at 8 or 10 bits, or VP8 / VP9 / AV1 encoded at 8 or 10 bits. In some examples, LDR video has a dynamic range provided by 8 bits per pixel per channel, and HDR video has a dynamic range that exceeds that of LDR video (e.g., a bit depth of typically 10 bits or more, a wide gamut color space, and an HDR-oriented transfer function). In some implementations, LDR video may be compressed, e.g., may have frames that are independently encoded and frames that are predicted based on other frames, so that every complete frame displayed from the video does not need to be stored in an independently decodable manner. The techniques herein can optionally be used in combination with a video encoding / decoding scheme; for example, recovery map frames can be generated from either raw video frames or decoded video frames.
[0057] In various implementations, the HDR video of block 210 can be acquired before the LDR video is acquired or can be acquired after the LDR video is acquired. For example, in some implementations, the HDR video is acquired and the LDR video is derived from the HDR video. In some examples, the LDR video can be generated based on the HDR video using local tone mapping or other processes. Local tone mapping, for example, modifies each pixel (or pixel region) according to local characteristics of the image frame (rather than according to a global tone mapping function that modifies all pixels in the same way). Any of a variety of local tone mapping techniques can be used to reduce the tonal values in the HDR frame to be appropriate for a corresponding LDR frame with a lower dynamic range. For example, local Laplacian filtering and / or other techniques may be used in the tone mapping process. In some implementations, different global tone curves (or digital gains) can be used in different regions of the frame. In some implementations, different tone mapping techniques can be used depending on the type of image data, for example, different techniques for still images and video images. In other examples, exposure fusion techniques can be used to generate an LDR frame from a corresponding HDR frame and / or from multiple captured or generated frames, e.g., combining portions of multiple frames captured or synthetically generated at different exposure levels into a combined LDR frame. In some exemplary implementations, a combination of tone mapping and exposure fusion techniques can be used to generate each LDR frame (e.g., using Laplacian tone mapping pyramids from different frames with different exposure levels, providing a weighted blend of the pyramids, and collapsing the blended pyramid).
[0058] In other examples, the LDR video may be obtained from other sources. For example, the HDR video may be captured by a camera device, and the LDR video may be captured by the same camera device, e.g., before or after the capture of the HDR video, or at least partially simultaneously. In some examples, the camera device may capture the HDR video and the LDR video based on camera settings, e.g., according to a dynamic range indicated in the camera settings or a user-selected setting. In some implementations, each captured video frame may be processed to provide an SDR output frame for the SDR video and an HDR output frame for the HDR video. Block 212 may be followed by block 214.
[0059] In block 214, a recovery map track is generated based on the HDR video and the LDR video. The recovery map track includes a sequence of recovery map frames; for example, the track can be considered a video. Each recovery map frame encodes the difference in luminance between a portion (e.g., pixels) of a respective HDR video frame and a corresponding portion (e.g., pixels) of a corresponding LDR video frame. Thus, there can be a respective recovery map frame in the recovery map track associated with each frame of the HDR video and the LDR video, and the frames of the HDR video and the LDR video correspond to each other (e.g., depict the same scene or subject). In other words, each recovery map frame is associated with a pair of corresponding frames of the HDR video and the LDR video (e.g., recovery map frame 1 is associated with HDR frame 1 and LDR frame 1, recovery map frame 2 is associated with HDR frame 2 and LDR frame 2, etc.). Each recovery map frame is used to convert the luminance of the associated LDR video frame to the luminance of the corresponding HDR video frame. For example, a respective brightness gain and a respective range scaling factor can be encoded in each recovery map frame (and / or the range scaling factor can be included in a separate metadata track, as described below), such that applying the brightness gain to the brightness value of an individual pixel in the associated LDR frame results in the same (or similar, e.g., nearly the same) pixel as the corresponding pixel in the corresponding HDR frame. An example of generating a recovery map track is described in more detail below with reference to FIG. 3. Block 214 may be followed by block 216.
[0060] In block 216, the recovery map track can be encoded into a recovery element. In various implementations, the recovery element can be a track, data structure, algorithm, neural network, etc. that includes or provides recovery map information for the recovery map frames. For example, the recovery element can be a track whose frames are compressed into a standard format having the same aspect ratio as the frames of the LDR video. The frames of the recovery element can have any resolution, e.g., the same resolution as the resolution of the LDR video, or a different resolution. For example, the recovery element can have a lower resolution than the LDR video, e.g., 1 / 4 the resolution of the LDR images. In some examples, the recovery map track can be encoded using H.264 / AVC, H.265 / HEVC, VP9, AV1, or other video codecs.
[0061] In some examples, the recovery map track can be encoded as 8-bit video. In various implementations, the recovery map track can be encoded in a single channel or multiple channels (e.g., RGB channels). In some examples, the values of each recovery map frame can be encoded as 8-bit unsigned integer values, each representing a recovery value and stored in one pixel of the recovery map frame. In some exemplary implementations, the encoding can result in a continuous representation between -2.0 and 2.0, which can be compressed via, for example, H.264 / AVC, H.265 / HEVC, VP9, AV1, or other codecs. In some implementations, the encoding can include one bit to indicate whether the map value is positive or negative, with the remaining bits interpreted as the magnitude of the value within the range of recovery map values (e.g., -1 to +1). In some implementations, such as this example, the magnitude representation can exceed 1.0 and -1.0 because some pixels may require more attenuation or gain than is represented by the associated range scaling factor to accurately recover or represent the luminance range of the original HDR video.
[0062] In single-channel encoding, a channel can specify adjustments to make to a target frame based on luminance (lightness). Thus, the chromaticity or hue of the target frame is not altered by applying the recovery map associated with that frame. Single-channel encoding can preserve the existing chromaticity or hue of the target frame (e.g., preserve the RGB ratio) and typically requires less storage space than multi-channel encoding. In some implementations, recovery elements can be encoded as multi-channel values. Multi-channel encoding can allow for adjustments to the color (e.g., chromaticity or hue) of the target frame while applying the recovery map to the target frame. Multi-channel encoding can compensate for color loss that may have occurred in an LDR frame, such as color rolloff at high lightness. This allows for HDR luminance recovery as well as accurate color recovery (e.g., resaturating the sky to a blue hue instead of only increasing lightness, which makes the sky appear white). In some implementations, a multi-channel encoded recovery map can be used to recover a wider color gamut of the target frame (or compensate for color gamut loss). For example, the ITU-R recommended BT.2020 color gamut can be recovered from LDR video in the standard RGB (sRGB) / BT.709 color gamut.
[0063] In some implementations, different bit depths, such as a bit depth greater than 8 bits, can be used for the recovery map values of the recovery map frame. For example, in some implementations, 8-bit depth may be the minimum bit depth that enables HDR representation in combination with an 8-bit single-channel gain map, and lower bit depths may not provide enough information to provide HDR video content without artifacts such as banding. In some implementations, the recovery elements can be encoded as floating-point values. In some implementations, the recovery elements can be a data structure, algorithm, neural network, etc. that encodes the values of the recovery map frame. For example, the recovery map tracks can be encoded with the weights of a neural network used as the recovery elements. Block 216 can be followed by block 218.
[0064] In block 218, the LDR video and recovery elements are provided to a video container that can be stored as a new HDR video format. In some implementations, the LDR video can be considered a base video in an image container. The video container can be read by a device to display the LDR video or a corresponding HDR video generated based on the LDR video and the recovery elements. In some implementations, the video container can be a standard video container, such as an ISOBMFF / MP4 file, a WebM container, Matroska, etc.
[0065] In some implementations, the recovery map track can be quantized to the precision of the video container in which the recovery map track is placed. For example, in some implementations where LDR video frames are encoded in a video encoding format that uses 8-bit unsigned integer values, the values of the recovery map frame can be quantized to an 8-bit format for storage. In some examples, each value can represent a recovery value and is stored in one pixel of the recovery map frame. In some examples, each encoded value can be defined as: encode(x,y)=recovery(x,y)*63.75+127.5 This encoding results in a representation ranging from -2.0 to 2.0, with each value quantized to, for example, one of 256 possible encoded values. In other implementations, other ranges and quantizations can be used, for example, -1.0 to 1.0, or other ranges.
[0066] The video container may include additional information related to the display of the output HDR video based on the contents of the video container. For example, the additional information may include metadata provided with the video container, the recovery element, and / or the LDR video. The metadata may encode information regarding how the LDR video and / or the HDR video derived therefrom is presented on a display device. Such metadata may include, for example, the version of the recovery map format used in the container, the range scaling factor(s) used in the recovery map track, the storage method (e.g., whether the recovery map track is encoded in a bilateral grid or other data structure or otherwise compressed by a specified compression technique), the resolution of the recovery map track, guide weights for bilateral grid storage, and / or other recovery map or image / video characteristics.
[0067] In some examples, an LDR video may include metadata, such as a metadata data structure or directory, that defines the order and characteristics of items (e.g., files) within a video container, and each file within the container may have a corresponding media item in the data structure. The media item may describe the location of an associated file within the video container and basic characteristics of the associated file. For example, a container element may be encoded into the metadata of the LDR video, which defines the format version and the data structure or directory of media items within the container. In some examples, the metadata is stored according to a data model that allows the LDR video to be read by devices or applications that do not support and are unable to read metadata. In some examples, some metadata may be stored in the video container as a timed metadata track (e.g., video) that includes frames corresponding to frames in the recovery map track. Some examples of video containers that can be used are described below with reference to FIG. 13. Block 218 may be followed by block 220.
[0068] In block 220, the video container can be stored in storage, for example, for access by one or more devices. For example, the video container can be transmitted over a network to one or more server systems and made accessible to multiple client devices that can access the server systems.
[0069] In some implementations, the server device can store the video container, including additional information such as metadata, in cloud storage (or other storage). The server can store the LDR video in the container in any format. In response to the server device receiving a request for video from a client device, the server device can determine the particular video data (e.g., video stream, LDR, or HDR) to process for the requesting client device based on system settings, user preferences, client device characteristics such as the type of client device (e.g., mobile client device, desktop computer client device), and display device characteristics of the client device such as dynamic range and resolution. For example, for some client devices, the server device can provide the video container to the client device, and the client device can decode the video container to obtain video (base video or derived video) for display with an appropriate dynamic range on the client's display device. In some implementations, the server device can extract and decode the video data (and recovery map track, if necessary) from the video container, generate a single output video stream (e.g., base video stream or derived video stream) according to the techniques described herein (e.g., as in FIG. 5), and provide the single determined video stream to the client device.
[0070] In some implementations (e.g., for some types of client devices), a server device can provide multiple video streams to a client device from a video container stored on the server, including a base video stream and a recovery map track (and, in some implementations, a timed metadata track). In some implementations or cases, for example, if the base video stream is LDR, the client device can determine whether it can display the high dynamic range of HDR video, and if so, can utilize the stream at the client device to create a derived output HDR video, such as that described with reference to FIG. 5. In some implementations, the client device can apply the recovery map track to the base video and create the derived output video using its processing hardware, e.g., a CPU, GPU, or other processor. In some implementations, the client device can include a video decoder, e.g., a hardware video decoder, that can efficiently apply the recovery map track to the base video stream and provide the derived output video for display by the client device.
[0071] In some implementations, for example, for some types of client devices where the server device has information describing client device characteristics (such as the dynamic range of the client device's display device), the server device can similarly provide multiple video streams to the client device (in the case of client device processing as described above) if the server device determines that the client device has a display device with a dynamic range appropriate for displaying the output video generated by the two streams. If the server device determines that the client device does not have such an appropriate display device, the server device can provide the client device with a single video stream, for example, an LDR video stream that is base video data in a video container (or, in some implementations, an LDR video stream produced by the server from a base HDR video stream in a video container using, for example, a recovery map track in the video container, tone mapping techniques, etc.).
[0072] In some implementations, a client device can receive a video container from a server device or other device or source and can locally provide output video suitable for display from the video container using techniques described herein (e.g., with reference to FIG. 5). For example, the client device outputs the output video based on the dynamic range of its display device as the base video in the container or as a derived video created by applying the recovery map track in the container to the base video.
[0073] In some implementations, multiple recovery map tracks (e.g., recovery elements) can be included in a video container. Each recovery map track can include different values to provide one or more different characteristics to the output video based on that recovery map track, as described herein. For example, a particular one of the multiple recovery map tracks can be selected and applied to provide output video having characteristics based on the selected recovery map track, which may be more suitable for a particular use or application than other recovery map track(s) in the video container. For example, some or all of the multiple recovery map tracks can be associated with specific different display device characteristics of a target output display device. In some examples, a user (or device configuration) may use different recovery map tracks to display SDR or HDR video on display devices with different dynamic ranges (e.g., different peak brightness) or different color gamuts, or when using different video decoders.
[0074] In some implementations, a respective set of associated range scaling coefficients (e.g., a respective timed metadata track, as in FIG. 6) may be stored in the container for each of these recovery map tracks. In some implementations, one or more of the recovery map tracks may be associated with a stored indication of particular display device characteristics of the target display device, as described below with respect to FIG. 5, allowing selection of one of the recovery map tracks based on the particular target display device being used for display. In some implementations, a single range scaling coefficient may be used with multiple recovery maps, e.g., a single range scaling coefficient for all frames in the LDR video and the SDR video (as described below with respect to block 310 of FIG. 3) and / or a single set of range scaling coefficients may be used with multiple recovery map tracks.
[0075] In some implementations, additional metadata indicating the intended usage of the recovery map tracks can be stored in the video container. In some implementations in which multiple recovery map tracks are stored in the video container, such metadata can include an indication of the intended usage for each of the multiple recovery map tracks. For example, if recovery map track 1 is intended for mapping to LDR video, a map, table, or index with a key for track 1 can indicate a value that is a list of specified output formats associated with the recovery map. In some examples, the metadata can specify an indication of the output range (e.g., "LDR" or "HDR," or HDR can be specified as a multiplier of the LDR range). In some examples, the metadata can specify a bit depth, e.g., "output metadata: bit_depth X," where X can be 8, 10, 12, etc. In some implementations, the specified bit_depth may be a minimum bit depth; for example, "LDR" may indicate an 8-bit minimum for output via an opto-electronic transfer function (OETF) and a 12-bit minimum for linear / extended range output, while "HDR" may indicate a 10-bit minimum for output via an OETF and a 16-bit minimum for linear / extended range output. This type of metadata may also be used for multi-channel recovery maps, in which the color space may vary and the metadata may include an indication of the intended output color space. For example, the metadata may be specified as "output metadata: bit_depth X, color_space," where color_space may be, for example, Rec. 709, Rec. 2020, etc. In some implementations, for a single-channel recovery map track, the output color space may be the same as the primary (base) video.
[0076] 3 is a flow diagram illustrating an example method 300 for generating a recovery map track based on an HDR video and an LDR video, according to some implementations. For example, method 300 may be implemented for block 214 of method 200 of FIG. 2, or may be performed in other implementations. In some implementations, the LDR video and the (original) HDR video may be obtained as described with reference to FIG. 2.
[0077] Method 300 may begin at block 302, where a pair of corresponding frames of an LDR video and an HDR video are selected for processing. For example, the corresponding frames may depict the same scene or subject matter and / or have the same position or index from the beginning of the video. Block 302 may be followed by block 304.
[0078] In block 304, a linear luminance can be determined for the selected LDR frame. In some implementations, the original LDR frame (e.g., captured in the LDR video of block 212 of FIG. 2) can be a nonlinear or gamma-encoded frame, and a linear version of the LDR frame can be generated, for example, by converting the primary image color space of the nonlinear LDR frame to a linear version. For example, a color space with a standard RGB (sRGB) transfer function is converted to a linear color space that preserves the color primaries of sRGB.
[0079] For a linear LDR frame, a linear luminance can be determined. For example, the luminance (Y) function can be defined as: Y ldr (x,y)=primary_color_profile_to_luminance(LDR(x,y)) where Yldr is the linear luminance of the low dynamic range image defined in the range of 0.0 to 1.0, and primary_color_profile_to_luminance is a function that converts the primary colors of the image (LDR frame) to a linear luminance value Y of each pixel of the LDR frame at coordinates (x, y). Block 304 may be followed by block 306.
[0080] In block 306, a linear luminance is determined for the selected HDR frame. In some implementations, the original HDR frame (e.g., obtained in the HDR video of block 210 of FIG. 2) can be a nonlinear or three-channel encoded image (e.g., perceptual quantizer (PQ) encoded or hybrid log-gamma (HLG) encoded), and the three-channel linear version of the HDR frame can be generated, for example, by converting the nonlinear HDR frame to a linear version. In other implementations, any other color space and color profile can be used.
[0081] For a linear HDR frame, a linear luminance can be determined. For example, the luminance (Y) function can be defined as: Y hdr (x,y)=primary_color_profile_to_luminance(HDR(x,y)) where Yhdr is the linear luminance of the low dynamic range image defined in the range 0.0 to range scaling factor, and primary_color_profile_to_luminance is a function that converts the primary colors of the image (HDR frame) to a linear luminance value Y of each pixel of the HDR frame at coordinates (x,y). Block 306 may be followed by block 308.
[0082] In block 308, a pixel gain function is determined based on the linear luminance determined in blocks 304 and 306. The pixel gain function is defined as the ratio between the Yhdr function and the Yldr function. For example: pixel_gain(x,y)=Yhdr(x,y) / Yldr(x,y)
[0083] where Yhdr is the linear luminance of the high dynamic range image defined in the range 0.0 to the range scaling factor, and primary_color_profile_to_luminance is a function that converts the primary colors of the image (HDR frame) to a linear luminance value Y of each pixel of the HDR frame at coordinates (x,y).
[0084] Since zero is a valid luminance value, it is possible for Yhdr or Yldr to be zero, which leads to potential problems in the above equations or when determining logarithms as described below. In some implementations, the case where Yhdr and / or Yldr are zero can be handled by, for example, defining the pixel_gain function as 1, or by adjusting the calculation to avoid this case using a sufficiently small additional factor (e.g., adding or clamping to a small factor), which is represented, for example, as epsilon (ε) in the following equations: pixel_gain(x,y)=(Yhdr(x,y)+ε) / (Yldr(x,y)+ε)
[0085] In some implementations, different values of epsilon (ε) may be used in the numerator and denominator of the above equation. Block 308 may be followed by block 310.
[0086] In block 310, a range scaling factor is determined based on the maximum luminance of the selected HDR frame to the maximum luminance of the selected LDR frame. For example, the range scaling factor may be the ratio of the maximum luminance of the HDR frame to the maximum luminance of the LDR frame. The range scaling factor is also referred to herein as a range compression factor and / or a range expansion factor, which is the multiplicative inverse of the range compression factor. For example, if an LDR frame is determined from an HDR frame, the range scaling factor may be derived from the amount by which the total HDR range is compressed to generate the LDR frame. For example, this factor may indicate the amount by which the luminance of the highlights of the HDR frame is reduced to map the HDR frame to the LDR dynamic range, or the amount by which the luminance of the shadows of the HDR frame is increased to map the HDR frame to the LDR dynamic range. The range scaling factor may be a linear value that can be multiplied with the total LDR range to obtain a total HDR range in linear space. In some examples, if the range scaling factor is 3, the shadows of the HDR frame are boosted three times, or the highlights of the HDR frame are reduced three times (or a blend of such shadow boosting and highlight reduction is performed). In some implementations, the range scaling factors may be defined in other ways (e.g., by the camera application or other application, a user, or other content creator) to provide specific visual effects that may change the appearance of the output image or frame derived from the range scaling factors.
[0087] In some implementations, a “universal” range scaling factor may be determined or selected in block 310. A single universal range scaling factor may be associated with a subset of recovery map frames in a recovery map track, or with all recovery map frames. If the current recovery map frame (determined in block 312 below) is associated with a universal scaling factor (e.g., if a subset is used, the current recovery map frame is within the subset of frames), the universal scaling factor is selected for use in block 310 if the universal range scaling factor was previously determined. In some examples, the universal range scaling factor may have been determined at any time in method 300. In some implementations, if the universal range scaling factor has not yet been determined, it may be determined in block 310. In some exemplary implementations, such a universal range scaling factor may be determined as the ratio between the maximum HDR luminance of all frames in the HDR video divided by the maximum LDR luminance of all frames in the LDR video. In some examples, the universal range scaling factor can be associated with a subset of the plurality of recovery map frames, which in turn is associated with at least a set of LDR frames of the LDR video (or HDR frames of the HDR video) having a threshold pixel or content similarity, determined, for example, by a pixel similarity measurement technique. In some implementations using the universal scaling factor, the dynamic range may be reduced to represent all possible luminances in the HDR video. Block 310 may be followed by block 312.
[0088] In block 312, a recovery map frame is determined that encodes a pixel gain function scaled by an associated range scaling factor. Thus, the recovery map frame is based on two linear frames containing the luminance of the desired HDR frame (Yhdr) and the luminance of the LDR frame (Yldr) via a pixel gain function. In some implementations, the recovery map frame is a scalar function that encodes normalized pixel gain in logarithmic space, scaled by an associated range scaling factor (e.g., multiplied by the inverse of the range scaling factor). In some examples, the recovery map frame value is proportional to the difference between the logarithmic HDR and LDR luminances divided by the logarithm of the range scaling factor. For example, the recovery map frame can be defined as follows: recovery(x,y)=log(pixel_gain(x,y)) / log(range_scaling_factor)
[0089] In some exemplary implementations, recovery(x,y) may tend to be in the range of −1 to +1. For display on an HDR display device, values below zero make the pixel (in an LDR frame) darker, and values above 0 make the pixel brighter. In some examples of the effect of the recovery map range, for example, to emphasize highlights when displaying a frame on an HDR display (such as the example described with reference to FIG. 5), these values may typically be in the range ∼[0..1] (e.g., applying the recovery map holds shadows constant and emphasizes highlights). Thus, the brightest areas of the frame may have values close to 1, and darker areas of the frame may typically have values close to 0. In some implementations, for pushing (darkening) shadows when displaying on an HDR display device, these values may typically be in the range ∼[−1..0] (e.g., applying the recovery map holds highlights stable and darkens shadows again). In some implementations, these values can cover the entire range [-1..1] to provide a combination of enhancing highlights and re-darkening shadows when displayed on an HDR display device.
[0090] In some implementations of the above, the frames are converted to logarithmic space for ease of manipulation, and processing is performed in logarithmic space. In some implementations, processing can be performed in linear space via exponentiation and exponential interpolation, which is mathematically equivalent. In some implementations, processing can be performed in linear space without exponentiation, e.g., simply interpolating / extrapolating in linear space.
[0091] When a recovery map in a recovery map frame is applied to an LDR frame for display, the associated range scaling factor generally indicates the amount by which to increase pixel brightness, indicating that some pixels may be made brighter or darker. The recovery map indicates the relative amount to scale each pixel (or pixel region), and the range scaling factor indicates the specific amount of scaling to be performed. The range scaling factor provides normalization and context for the relative values in the recovery map. The range scaling factor allows for efficient use of all bits used to store the recovery map, regardless of the amount of scaling applied via the recovery map. This means that a greater variety of frames can be accurately encoded. For example, all maps may contain values in a range, e.g., -1 to 1, and with the values given context by the range scaling factor, this range of values can represent maps that scale equally well to a larger total range (e.g., 2 or 8). For example, one frame in the described format may have a range scaling factor of 2, while a second frame may have a range scaling factor of 8. Without a range scaling factor, the maximum / minimum absolute gain values are determined within the map. If this value were, for example, 2, then a second frame with range scale 8 could not be represented. If this value were, for example, 8, then 2 bits of information would be wasted for every pixel in the recovery map when storing a frame with range scale 2. In such implementations without a range scaling factor, it may be more likely to have (undesirable) visual differences in the displayed output HDR frame compared to the original HDR frame used for encoding (e.g., a severe form of such differences may be banding).
[0092] In some implementations, when pixel_gain is 0.0, the recovery function can be defined as -2.0, which is the maximum representable attenuation. The recovery function can be outside the range of -1.0 to +1.0 because one or more regions or locations within a frame may require greater scaling (attenuation or gain) than is represented by the range scaling factor to recover the dynamic range of the original HDR frame. The range of -2.0 to +2.0 may be sufficient to accommodate such greater scaling (attenuation or gain). Block 312 may be followed by block 314.
[0093] At block 314, it may be determined whether there is another recovery map frame of the recovery map track to determine. For example, this block may have a positive result if a recovery map frame has not been generated for one or more frames of the video. If there is another recovery map frame to determine, the method may continue to block 302 and select a new pair of corresponding frames of the LDR video and the HDR video for generating a recovery map frame. If there are no more recovery map frames to determine (e.g., if all frames of the LDR video and the HDR video have associated recovery map frames or another condition occurs that stops determining recovery map frames), block 314 may be followed by block 316.
[0094] At block 316, the recovery map frames forming the recovery map track are encoded into a compressed format. This may allow for a reduction in the storage space required to store the recovery map track. The compression format may be a variety of different compressions, formats, etc. In some implementations, the recovery map frames are encoded into a bilateral grid, as described in more detail below with respect to FIG. 4 (e.g., each recovery map frame may be encoded as a recovery map within a respective bilateral grid, with pixel_gain as described above with respect to blocks 308-310). In some implementations, the recovery map frames may be compressed or otherwise reduce storage requirements based on one or more additional or alternative compression techniques. In some implementations, the recovery map frames may be compressed into a low-resolution recovery frame having a lower resolution than the recovery map frame determined at block 312. This recovery frame may occupy less storage space than a full-resolution recovery map frame. For example, H.264 / AVC, H.265 / HEVC, VP9, AV1, or other codecs may be used to encode / decode recovery map frames in various implementations.
[0095] In some implementations, block 316 may be performed after block 312 for the particular recovery map frame being processed, and block 314 may be performed after block 316 in those implementations. In some implementations, blocks 304 and 306 may be performed in parallel. In some implementations, block 316 may not be performed, and the recovery map track may directly store the recovery map frame(s). In some implementations, method 300 may be performed in parallel for multiple pairs of LDR and HDR video frames to speed up the generation of the recovery map.
[0096] Figure 4 is a flow diagram illustrating an example method 400 for encoding recovery map frames of a recovery map track into a bilateral grid, according to some implementations. For example, method 400 can be implemented in blocks 312 and 316 of method 300 of Figure 3, where one or more recovery map frames are determined, encoded, and / or stored in a compressed format based on the luminance of HDR and LDR video frames, and in the implementation of Figure 4, the compressed format is a bilateral grid. The bilateral grid is a three-dimensional data structure of grid cells mapped to pixels of the LDR frame based on guide weights, as described below.
[0097] Method 400 may begin at block 402, where a three-dimensional data structure of grid cells is defined. The three-dimensional data structure has a width, height, and depth in several cells, shown as size WxHxD, where the width and height correspond to the width and height of the LDR frame currently being processed, and the depth corresponds to the number of layers of WxH cells. A cell in the grid is defined as a vector element of (width, height) and length D and maps to multiple pixels in the LDR frame. Thus, the number of cells in the grid can be much less than the number of pixels in the LDR frame. In some examples, the width and height can be approximately 3% of the LDR frame size, and the depth can be 16. For example, a 1920x1080 LDR frame can have a grid of size on the order of 64x36x16. Block 402 may be followed by block 404.
[0098] At block 404, a set of guide weights is defined. In some implementations, the set of guide weights may be defined as a multi-element vector of floating-point values, which can be used to generate and decode a bilateral grid. The guide weights represent the weight of each element of an input pixel of an LDR frame, LDR(x,y), for determining the corresponding depth D of the cell that maps to that input pixel. For example, the set of guide weights may be defined as a five-element vector: guide_weights={weight_r,weight_g,weight_b,weight_min,weight_max} where weight_r, weight_g, and weight_b refer to the weights of the color components (red, green, or blue) of the respective input pixels, and weight_min and weight_max refer to the weights of the minimum component value of the r, g, and b values, and the weights of the maximum component value of the r, g, and b values, respectively (e.g., min(r,g,b) and max(r,g,b)).
[0099] In one example, the LDR frame is in Rec. 709 color space, and the guide weights may be set to the following values, for example, to represent the balance between the luminance of the LDR(x,y) pixel and the minimum and maximum component values of that pixel: guide_weights={0.1495,0.2935,0.057,0.125,0.375}
[0100] In this example, these guide weight values weight the grid depth lookup by 50% for pixel intensities (r, g, b values), 12.5% for minimum component values, and 37.5% for maximum pixel component values. In some implementations, these values sum to 1.0, e.g., all grid cells can always be looked up (as opposed to values summing less than 1.0), and pixel depths cannot overflow beyond the final depth level of the grid (as opposed to values summing greater than 1.0). Guide weight values can be determined heuristically, for example, by having a human visually evaluate the results. In some implementations, a different number of guide weights can be used, e.g., only three guide weights for pixel intensities (r, g, b values).
[0101] Block 404 may be followed by block 406 .
[0102] In block 406, the grid cells are mapped to pixels of the LDR frame based on the guide weights, which are used to determine the depth D corresponding to the cell that maps to that input pixel.
[0103] In some implementations, the width (x) and height (y) of a grid cell are mapped to a pixel (x,y) in the LDR frame. For example, a pixel x,y in the LDR frame can be converted to the x,y coordinates of a grid cell as follows: grid_x=ldr_x / ldr_W*(grid_W-1) grid_y=ldr_y / ldr_H*(grid_H-1) where grid_x and grid_y are the x and y positions of the grid cell, ldr_x and ldr_y are the x and y coordinates of the pixel in the LDR frame that is mapped to the grid cell, ldr_W and ldr_H are the total pixel width and height of the LDR frame, and grid_W and grid_H are the total width and height cell dimensions of the bilateral grid, respectively.
[0104] In some implementations, for the depth of a grid cell mapped to an LDR frame, the following formula can be used: z_idx(x,y)=(D-1) * (guide_weights ·{r,g,b,min(r,g,b),max(r,g,b)}) where D is the total cell depth dimension of the grid, and z_idx(x,y) is the depth of the bilateral grid of the cell mapped to the LDR frame pixel (x,y). This formula determines, for a particular pixel of the LDR frame, the dot product of a vector of guide weights and a vector of corresponding values, and then scales the result by the size of the depth dimension in the grid to indicate the depth of the corresponding cell for that particular pixel. For example, the guide weight multiplication result can be a value between 0.0 and 1.0 (inclusive), and this result is scaled to various depth levels (buckets) of the grid. For example, if D is 16, a multiplication value of 0.0 results in a mapping to bucket 0, a value of 0.5 results in a mapping to bucket 7, and a value of 1.0 results in a mapping to bucket 15 (the last bucket). Block 406 may be followed by block 408.
[0105] At block 408, the recovery map frame values are determined using the solution of a set of linear equations. For example, in some implementations, the input to the bilateral grid is defined using the following equations: pixel_gain(x,y)=bilateral_grid_linear(x,y,z_idx(x,y)) where pixel_gain(x,y) is the value at pixel (x,y) of the pixel_gain function described above with respect to method 300 of FIG. 3. Bilateral_grid_linear is the bilateral grid content, representing the pixel gain in linear space. The coordinates x, y, and z_idx are coordinates of the bilateral grid, which are different from the x, y coordinates of the LDR frame. Via z_idx, the guide weights inform, for each pixel of the LDR frame, the depth at which its pixel_gain is located in the bilateral grid, thereby informing how to configure the linear set of equations to solve in block 410. Block 408 may be followed by block 410.
[0106] In some implementations, the encoding is defined as the solution to a set of linear equations that minimizes the following for each grid cell: (pixel_gain(x,y)*Yldr(x,y)-Yhdr(x,y)) 2
[0107] The Yldr and Yhdr values are the luminance values of the LDR and HDR frames, respectively, at the (x, y) pixel location of the LDR frame. The pixel_gain(x, y) value that minimizes the value of the above equation is solved based on when each pixel is looked up in the depth vector as shown in the above equation, which is predetermined based on the guide weights.
[0108] The resulting value is defined as follows: For{x,y,z}over{[0,W),[0,H),[0,D)}:bilateral_grid_linear(x,y,z)
[0109] This definition specifies that the bilateral grid is defined from 0 to the total grid width for x, from 0 to the total grid height for y, and from 0 to the total grid depth for z. For example, if the bilateral grid is 64x36x16, then the grid values are defined as 0 to 63 (inclusive) for x, 0 to 35 (inclusive) for y, and 0 to 15 (inclusive) for z. Block 408 may be followed by block 410.
[0110] In block 410, the resolved values are transformed and stored in a bilateral grid. For example, the resolved values can be transformed as follows: bilateral_grid(x,y,z)=log(pixel_gain(x,y)) / log(range_scaling_factor) or bilateral_grid(x,y,z)=log(bilateral_grid_linear) / log(range_scaling_factor) Bilateral_grid is the content of the bilateral grid, representing pixel gains in logarithmic space. These values are stored as the encoded recovery map frame. As shown in these equations, the normalized pixel gains between the LDR and HDR frames are encoded in logarithmic space and scaled by the range scaling factor, as described above with respect to block 310 of FIG. 3. The encoded recovery map frame can be similarly determined for each set of corresponding LDR and HDR frames of the video.
[0111] When decoding a recovery map frame from a bilateral grid (as described in more detail with respect to FIG. 5 ) to display an HDR frame with a larger dynamic range than the corresponding LDR frame, each recovery map frame value is extracted from the bilateral grid. In the following equation, r, g, and b are the respective component values of pixel LDR(x, y) of the LDR frame. z_idx(x,y)=(D-1)*(guide_weights ·{r,g,b,min(r,g,b),max(r,g,b)}) recovery(x,y)=bilateral_grid(x,y,z_idx(x,y)) where recovery(x,y) is the extracted recovery map frame. In some implementations, z_idx, as determined in the first equation, can be rounded to the nearest integer before being input to bilateral_grid(x,y,z) in the second equation.
[0112] FIG. 5 is a flow diagram illustrating an example method 500 for decoding video into a backward-compatible high dynamic range video format and displaying HDR video according to some implementations. In some implementations, method 500 may be performed on, for example, server system 102 as shown in FIG. 1. In some implementations, some or all of method 500 may be implemented on one or more client devices, such as client devices 120, 122, 124, or 126 of FIG. 1, one or more server devices, such as server device 104 of FIG. 1, and / or both the server device(s) and the client device(s). In the described example, an implementing system includes one or more digital processors or processing circuits (“processors”) and one or more storage devices (e.g., databases or other storage). In some implementations, different components of one or more servers and / or clients may perform different blocks or other portions of method 500. In some examples, devices are described as performing blocks of method 500. Some implementations may have one or more blocks of method 500 performed by one or more other devices (e.g., other client devices or server devices) that may send results or data to the first device.
[0113] In some implementations, method 500, or portions thereof, may be initiated automatically by the system. For example, the method (or portions thereof) may be executed periodically or may be executed based on one or more specific events or conditions, such as a client device launching image application 106, a client device capturing or receiving a new image (including video), uploading a new image (including video) to server system 102, a predetermined period of time having elapsed since the last execution of method 500, and / or one or more other conditions occurring that may be specified in settings read or received by the system executing the method.
[0114] Method 500 can be performed by a device different from the device that created the video container. For example, a first device can create a video container and store it in a storage location (e.g., a server device) accessible by a second device different from the first device. The second device can access the video container and display the video therefrom according to method 500. In some implementations, a single device can create a video container as described in method 200 and decode and display the video container as described in method 500. In some implementations, the first device can decode a video stream (or video file) from a video container according to method 500 and provide the decoded video stream or file to a second device to display the video stream or file.
[0115] As described above with reference to FIG. 2, in some embodiments, a first device (e.g., a server device in the cloud) can decode two video streams (e.g., a base video stream and a derived video stream based on the base video stream and the recovery map frame) from a video container and transmit the two video streams to a second device, which selects and displays one of the video streams as determined by the second device, for example, based on the display device characteristics of the second device.
[0116] In some implementations, a first device (e.g., a server device in the cloud) can decode multiple video streams from a video container (e.g., the base video stream and the recovery map track (and in some implementations, the timed metadata track, or metadata) can be included in different streams) and send the video streams to a second device, which (based on its display device characteristics) selects the base video stream to display directly (e.g., ignores the recovery map track) or applies the recovery map track to the base video stream to generate a derived video stream that is displayed by the second device. In some examples, the application of the recovery map track to the base video stream can be performed by a hardware decoder and / or by the CPU and / or GPU of the second device. In various implementations, a dedicated hardware codec can be used to accelerate the application of the recovery map track to video frames.
[0117] User permission may be obtained to use user data in method 500 (blocks 510-534). For example, the user data for which permission is obtained may include images stored on a client device (e.g., any of client devices 120-126) and / or a server device, image metadata, user data associated with the use of an image application or image-based production, etc. The user may be provided with options to selectively provide permission to access all user data, any user data, or none of the user data. If the user's permission is insufficient for particular user data, method 500 may be performed without using the user data's permission, for example, using other data (e.g., images not associated with the user).
[0118] Method 500 may begin at block 510. In block 510, a video container is obtained that includes an LDR video and a recovery map track. The recovery map track includes a plurality of recovery map frames, as described above. For example, the video container may be a container created by method 200 of FIG. 2 that generates a recovery map track from HDR video and LDR video, where the HDR video has a larger dynamic range than the LDR video. The recovery map track, as described above, may be provided in the container as a recovery element in some implementations. The video container may have files for a standardized format supported and recognized by a device executing method 500, such as ISOBMFF and / or MP4. Other examples of usable formats may include WebM and Matroska, and / or Extensible Metadata Platform (XMP)-encoded data (e.g., which may be stored in an ISOBMFF container and used to store metadata such as information about the format version and items in the container). In some implementations, the recovery map track may be stored as a secondary video track according to a standard (e.g., ISOBMFF), or the recovery map track may be stored as a different type of auxiliary video stored in a data structure box, such as an rdat box. Block 510 may be followed by block 512.
[0119] At block 512, the LDR video within the video container is extracted. In some example implementations, this extraction may include decoding the LDR video from its standard file format into a format for use in display and / or processing. The decoded value from each encoded recovery map frame may be determined by the following formula, where r is the range and n is the number of bits in each recovery map value: recovery(x,y)=(encode(x,y)-(2^n-1) / 2) / ((2^n-1) / 2r)
[0120] In some example implementations, with reference to the example quantized encode map (encode(x,y)) described above for block 218 of FIG. 2 (n=8 and r=2), the decoded values from each encoded recovery map for a frame can be determined as follows: recovery(x,y)=(encode(x,y)-127.5) / 63.75 This provides a scale and offset for a range of recovery map values from -2 to 2, stored as 8 bits. In another example, if the values stored in the recovery map frame are from -1 to 1, the decoded value can be determined as follows: recovery(x,y)=(encode(x,y)-127.5) / 127.5
[0121] Other implementations may use other ranges, scales, offsets, bits of storage, etc. Block 512 may be followed by block 514.
[0122] At block 514, it is determined whether to display the output video, which is HDR video derived from the extracted LDR video (e.g., derived video), or the output video, which is LDR video (e.g., base video of the video container). The output video is displayed on at least one target display device associated with or in communication with the device performing method 500 and accessing the video container. The target display device may be any suitable output display device, such as a display screen, touch screen, projector, display goggles or glasses, etc. The displayed output HDR video may have any dynamic range greater than the dynamic range of the HDR video. For example, up to the dynamic range of the original HDR video used in creating the video container (e.g., at block 210 of FIG. 2 ), including the LDR video in the container.
[0123] The decision to display the output HDR video can be based on one or more characteristics of the target display device that will display the output video. For example, the dynamic range of the target display device can determine whether to display HDR video or LDR video as the output video. In some examples, if the target display device can display a dynamic range greater than the dynamic range of the LDR video in the video container, it is determined to display the HDR video. If the target display device can only display a low dynamic range of the LDR video in the container, it is determined not to display the HDR video.
[0124] In some implementations or cases, a decision may be made not to display HDR video even if the target display device is capable of accommodating a larger dynamic range than the LDR video. For example, user settings or preferences may dictate that a video with a lower dynamic range be displayed, or the displayed video will be displayed as an LDR video visually compared to other LDR videos, etc.
[0125] If it is determined that the HDR video will not be displayed at block 514, the method proceeds to block 516, where the LDR video extracted from the video container is displayed by the target display device. Typically, the LDR video is displayed by an LDR-compatible target output device that cannot display a dynamic range greater than that of the LDR video. The LDR video can be output directly by the target display device. In this case, the recovery map track and any metadata related to the HDR video format in the video container are ignored. In some implementations, the LDR video is provided in a standard format and is readable and displayable by any device capable of reading the standard format.
[0126] If it is determined at block 514 that the HDR video is to be displayed as the output video, the method proceeds to block 518, where the recovery element in the video container is extracted. Metadata related to displaying the output HDR video based on the recovery element may also be obtained from the video container. For example, the recovery element may be or include a recovery map track as described above, which may include and / or be associated with metadata as described above. Metadata may also be included in the container and / or in the HDR video from the video container. For example, the obtained metadata may include respective range scaling coefficients associated with recovery map frames in the recovery map track. In some implementations, these range scaling coefficients may be included in a timed metadata track that includes frames, each frame including a respective range scaling coefficient associated with a corresponding recovery map frame in the recovery map track.
[0127] In some implementations, multiple recovery map tracks (e.g., recovery elements) may be included in the video container. In some of these implementations, one of the multiple recovery elements may be selected in block 518 for processing below. In some implementations, the selected recovery element may be specified by user input or user setting, or in some implementations, may be selected automatically (e.g., without current user input) by the device executing method 500 or a portion thereof, based, for example, on one or more characteristics of the target output display device or other output display component, where such characteristics may be obtained via an operating system call or other available source. For example, in some implementations, one or more of the recovery elements may be associated with an indication of particular display device characteristics stored in the video container (or otherwise accessible to the device executing block 518). If it is determined that the target display device has one or more particular characteristics (e.g., dynamic range, peak brightness, color gamut, etc.), block 518 may select a particular recovery element associated with those characteristics. If multiple recovery elements are associated with different sets of range scaling coefficients (e.g., different timed metadata tracks as in FIG. 13), the associated set of range scaling coefficients is also selected and retrieved. Block 518 may be followed by block 520.
[0128] In block 520, the recovery map track may optionally be decoded from the recovery element. For example, the recovery map track may be in a compressed format or other encoded format within the recovery element and may be decompressed or decoded. For example, in some implementations, the recovery map track may be encoded in a bilateral grid as described above, and the recovery map track is decoded from the bilateral grid. Block 520 may be followed by block 522.
[0129] At block 524, it is determined whether to scale the luminance of the extracted LDR video for the output HDR video. In some implementations, the scaling is based on the display output capabilities of the target display device. For example, in some implementations or cases, the generation of the output HDR video is based in part on a particular luminance that the target display device can display. In some implementations or cases, this particular luminance may be the maximum luminance that the target display device can display, such that the maximum luminance of the display device limits the luminance of the output HDR video. In some implementations or cases, it is determined not to scale the luminance of the LDR video, e.g., the generation of the output HDR video does not take into account the luminance capabilities of the target display device.
[0130] If it is determined at block 524 not to scale the luminance of the output video based on the capabilities of the target display device, the method proceeds to block 526, where the luminance gain of each recovery map frame of the recovery map track is applied to pixels of associated frames of the LDR video, and the pixels of these respective frames are scaled based on the associated range scaling factor of the recovery map frame to determine corresponding pixels of corresponding frames of the output HDR video. For example, the luminance gain and scaling are applied to the luminance of pixels of associated frames of the LDR video to determine each corresponding pixel value of corresponding frames of the output HDR video that are displayed. This causes the output HDR video to have the same dynamic range as the original HDR video used to create the video container.
[0131] In some example implementations, each frame of LDR video and the associated display coefficients may be transformed into logarithmic space, and the following formula may be used: HDR*(x,y)=LDR*(x,y)+log(range_scaling_factor)*recovery(x,y) where HDR*(x,y) is pixel (x,y) of the recovered HDR frame (output frame of the output video) in logarithmic space, LDR*(x,y) is pixel (x,y) of the logarithmic space version of the extracted LDR frame (e.g., log(LDR(x,y))), log(range_scaling_factor) is the range scaling factor in the associated logarithmic space, and recovery(x,y) is the recovery map value at pixel (x,y) of the associated recovery map frame. For example, recovery(x,y) is the normalized pixel gain in logarithmic space.
[0132] In some examples, when a per-pixel gain, e.g., recovery(x,y)*range scaling factor, is applied, an interpolation may be provided between the LDR and HDR versions of each frame of video. This is essentially an interpolation between the original and range-compressed versions of the video, and the amount of interpolation is dynamic, both on a per-pixel and global basis, because it depends on the range scaling factor. In some examples, this can be considered an extrapolation when the recovery map value is less than 0 or greater than 1.
[0133] In some implementations, the LDR video from the video container is nonlinearly encoded, e.g., gamma encoded, and the primary image color space of the frames of the nonlinear LDR video is converted to a linear version before the LDR video is converted to LDR*(x,y) in logarithmic space in the above equation. For example, a color space with a standard RGB (sRGB) transfer function can be converted to a linear color space that preserves the sRGB primaries, similar to the manner described above for block 302 of FIG. 3. In some implementations, the LDR video in the container can be linearly encoded, and no such conversion from the nonlinear color space is used.
[0134] Applying a luminance gain to the luminance of individual pixels of the LDR video results in pixels that correspond to the original (e.g., originally encoded) HDR video. If lossless storage of the LDR video and recovery map track is provided, the output HDR video can have the same pixel values as the original HDR video (e.g., if the recovery map track has the same resolution as the LDR video). Otherwise, the resulting output HDR video from block 532 can closely approximate the original HDR video, and typically, a user cannot visually perceive the difference when viewing the output video.
[0135] In some implementations, the output video may be converted from logarithmic space to an output color space for display. In some implementations, the color space of the output video may differ from the color space of the LDR video extracted from the video container. Block 526 may be followed by block 534, described below.
[0136] If at block 524 it is determined to scale the luminance of the output video based on the output of the target display device, the method proceeds to block 528. At block 528, the maximum luminance output of the target display device is determined. The target display device may be an HDR display device, and these types of devices can display maximum luminance values that may vary based on the model, manufacturer, etc. of the display device. In some implementations, the maximum luminance output of the target display device may be obtained via an operating system call or other source. Block 528 may be followed by block 530, described below.
[0137] At block 530, a display factor is determined. The display factor is used to scale the dynamic range of the associated frame of the displayed output HDR video, as described below. In some implementations, the display factor is determined to be less than or equal to the minimum luminance output of the target display device and less than or equal to the range scaling factor of the associated frame. For example, this can be stated as follows: Display coefficient ≦ min (maximum display brightness, range scaling coefficient) This determination can limit the luminance scaling of the output video by the maximum display luminance of the target display device. This prevents scaling the output video to a luminance range that is present in the original HDR video from which the video container was created (as indicated by the range scaling factor) but that is greater than the maximum luminance output range of the target display device. In some implementations, the maximum display luminance may be a dynamic characteristic of the target display device and may be adjustable, for example, by user settings, an application program, etc.
[0138] In some implementations, the display factor can be set to a luminance that is greater than the associated range scaling factor and less than or equal to the maximum display luminance. For example, this can be done if the target display device is capable of outputting a dynamic range greater than that present in the original HDR video (which may be indicated by the range scaling factor). In some examples, a user (via a user setting or selection input) can instruct the output video to be brighter.
[0139] In some implementations, the display factor may be set to a luminance that is less than the maximum display luminance of the target display device and the maximum luminance of the original HDR video. For example, the display factor may be set by user input, user preference or setting, etc., based on the particular display luminance of the display device indicated by one or more display conditions, and may be set to cause the output of a lower luminance level, which may be desirable in some cases or applications, as described below. Block 530 may be followed by block 532, described below.
[0140] In block 532, the luminance gain of each recovery map frame in the recovery map track is applied to pixels of the associated frame of the LDR video, and the luminance of the pixels of the associated frame of the LDR video is scaled based on a display factor to determine corresponding pixels of the corresponding frame of the output HDR video. For example, the luminance gain encoded in the recovery map frame is applied to the luminance of pixels of the associated LDR frame to determine each corresponding pixel value of the corresponding frame of the output HDR video to be displayed. These pixel values are scaled based on the display factor associated with the LDR frame determined in block 530, for example, based on the specific luminance output of the target display device (which may be the maximum luminance output of the target display device). In some implementations, this allows the luminance of highlights in the LDR video to be increased to a level that the display device can display, up to a maximum level based on the dynamic range of the original HDR video, or the luminance of shadows in the LDR video to be decreased to a lower limit based on the dynamic range of the original HDR video, up to a level that the display device can display.
[0141] In some example implementations, each frame of LDR video and the associated display coefficients may be transformed into logarithmic space, and the following formula may be used: HDR*(x,y)=LDR*(x,y)+log(display_factor)*recovery(x,y) where HDR*(x,y) is pixel (x,y) of the recovered HDR frame (output frame of the output video) in logarithmic space, LDR*(x,y) is pixel (x,y) of the logarithmic space version of the extracted LDR frame (e.g., log(LDR(x,y))), log(display_factor) is the associated display factor in logarithmic space, and recovery(x,y) is the recovery map value at pixel (x,y) of the associated recovery map frame. For example, recovery(x,y) is the normalized pixel gain in logarithmic space.
[0142] In some examples, when a per-pixel gain, e.g., recovery_map(x,y) * display factor, is applied, an interpolation may be provided between the LDR and HDR versions of each frame of video. This is essentially an interpolation between the original and range-compressed versions of the video, and the amount of interpolation is dynamic, both on a per-pixel and global basis, because it depends on the display factor. In some examples, this can be considered an extrapolation when the recovery map value is less than 0 or greater than 1.
[0143] In some implementations, the LDR video from the video container is nonlinearly encoded, e.g., gamma encoded, and the primary image color space of the frames of the nonlinear LDR video is converted to a linear version before the LDR video is converted to LDR*(x,y) in logarithmic space in the above equation. For example, a color space with a standard RGB (sRGB) transfer function can be converted to a linear color space that preserves the sRGB primaries, similar to the manner described above for block 302 of FIG. 3. In some implementations, the LDR video in the container can be linearly encoded, and no such conversion from the nonlinear color space is used.
[0144] Applying a luminance gain to the luminance of individual pixels of the LDR video results in pixels that correspond to the original (e.g., originally encoded) HDR video. If lossless storage of the LDR video and recovery map track is provided, the output HDR video can have the same pixel values as the original HDR video (e.g., if the recovery map track has the same resolution as the LDR video). Otherwise, the resulting output HDR video from block 532 can approximate the original HDR video, and typically, a user cannot visually perceive the difference when viewing the output video.
[0145] In some implementations, the output video may be converted from logarithmic space to an output color space for display. In some implementations, the color space of the output video may differ from the color space of the LDR video extracted from the video container. Block 532 may be followed by block 534.
[0146] In block 534, the output HDR video determined in block 526 or block 532 is displayed by a target display device. The output HDR video is displayed by an HDR-compatible display device capable of displaying a dynamic range greater than that of the LDR video. The output HDR video generated from the video container may be the same as or similar to the original HDR video used to create the video container, e.g., nearly the same and visually indistinguishable from the original HDR video, because the recovery map stores luminance information.
[0147] In some implementations, the output HDR video (obtained in block 526 or 532) may be suitable for display by a target display device, for example, as a video having linear image frames or another output HDR video format. For example, the output HDR video may have been processed by the display coefficients of block 532 and may not need to be further modified for display, such that the output HDR video resulting from applying the recovery map track is directly rendered for display.
[0148] In some implementations or cases, the output HDR video may be converted to a different format for display by the target display device, e.g., converted to a video format that may be more suitable for display by the target display device than the output HDR video. For example, the output HDR video from block 526 (or in some implementations from block 532) may be converted to a luminance range that is more suitable for the capabilities of the target display (which may be in addition to using display factors as in block 532, in some implementations). In some implementations, the output HDR video may be converted to standard video having a standard video format, such as HLG / PQ or other standard HDR format. In some of these examples, common tone mapping techniques may be applied to such standard video for display on the target display device.
[0149] In some implementations or cases, the output HDR video generated for a target display device via the display coefficients of block 532 may provide higher quality video display on the target device than output HDR video converted and generated via a general tone mapping technique without being further converted to a standard format and processed for display using a general tone mapping technique. For example, local tone mapping provided via a recovery map may provide higher quality video, e.g., greater detail, than general or global tone mapping. Furthermore, there may be ambiguities and differences in the conversion of HDR video to some standard formats using a global tone mapping technique based on the specifications and implementations of different devices.
[0150] In some implementations, techniques can be used to output frames to a display device in a linear format, for example, an extended range format where 0 represents black, 1 represents SDR white, and values greater than 1 represent HDR lightness. Such techniques can avoid problems with global tone mapping techniques for handling the display capabilities of a target display device. One or more techniques described herein can address the display capabilities of the display device by applying a recovery map via display coefficients as described above, so that the extended range values do not exceed the display capabilities of the display device.
[0151] As described above, the device can scale the pixel brightness of the LDR video based on the maximum brightness output of the HDR display device and based on the brightness gain stored in the recovery map. In some implementations, for example, in many photography and videography applications, scaling the LDR video can be used to generate an HDR output video that reproduces the full dynamic range of the original HDR video used to create the video container (if the target display device is capable of displaying that full range). In some implementations, scaling the LDR video can set any dynamic range for the output HDR video that exceeds the range of the LDR video. In some implementations, the dynamic range may also be less than the range of the original HDR video. For example, the brightness of the output video can be scaled based on the capabilities of the target display device or other criteria, or can be scaled lower than the original HDR video and / or lower than the maximum capabilities of the target display device, for example, to reduce the displayed brightness in the video that may be visually uncomfortable or fatiguing to the viewer. For example, a fatigue reduction algorithm and one or more light sensors in communication with a device implementing the algorithm can be used to detect one or more of various factors, such as ambient light around the display device, time of day, etc., to determine a reduction in the dynamic range or maximum brightness of the output video. For example, the reduced brightness level of the output video can be obtained by scaling the recovery map frame value and / or display coefficients used in determining the HDR video, as described in some embodiments herein, such as scaling the range_scaling_factor and / or the display coefficients in the exemplary equations above. In another example, a device may determine that its battery power level is below a threshold level and that battery power should be conserved (e.g., by avoiding high display brightness), in which case it causes the device to display the HDR video at a brightness below that of the original HDR video.
[0152] In some implementations, the luminance of the LDR video can be scaled larger (e.g., higher) than the dynamic range of the original HDR video, for example, if the target display device has a larger dynamic range than the original HDR video. The dynamic range of the output HDR video is not constrained to the dynamic range of the original HDR video or to any particular high dynamic range. In some examples, the maximum value of the dynamic range of the output HDR video can be lower than the maximum dynamic range of the original HDR video and / or lower than the maximum dynamic range of the target output device.
[0153] In some implementations, the display of the output HDR video can be scaled incrementally from a lower brightness level to the full (maximum) brightness level of the output HDR video (the maximum brightness level determined as described above). For example, the incremental scaling can be performed over a specific period of time, such as 2-4 seconds (which may be user-configurable in some implementations). The incremental scaling can avoid user discomfort that may result from a sudden, large increase in brightness when displaying a higher brightness HDR video after displaying a lower brightness (e.g., displaying an LDR video). The incremental scaling of an LDR video to its HDR version can be performed, for example, by scaling the recovery map frame values and / or display factors used in determining the HDR video, as described herein, for example, in multiple steps, to gradually brighten the LDR video to the HDR version of the video. For example, the range_scaling_factor and / or display_factor can be scaled according to the above formula. In other implementations, multiple scaling operations can be performed sequentially to obtain intermediate images with gradually increasing dynamic ranges over a period of time. Other HDR video formats that use the above-described global tone mapping techniques to scale LDR video to HDR video instead of the local tone mapping described for HDR formats herein may experience reduced display quality when performing such gradual scaling or may not support gradual scaling.
[0154] In some examples, this gradual scaling of the HDR video brightness can be performed or triggered in response to specific display conditions. Such conditions can include, for example, when an output HDR video is displayed by a display device immediately after displaying an LDR video (or when an output HDR video frame is displayed immediately after displaying an LDR video frame). In some examples, a grid of LDR videos and / or images can be displayed, and a user can select one of these videos (or a frame thereof) to fill the screen with a display of an HDR version of the selected video or frame (where the HDR video can be determined based on recovery elements from the associated video container, as described above). To avoid a sudden, large increase in brightness when changing from an LDR video display to an HDR video display, the HDR video can be initially displayed at the lower brightness level of the LDR video and gradually updated (e.g., gradually displaying video content with a higher dynamic range and incrementally increasing to its maximum brightness level, e.g., with fade-in and fade-out animations). In another example, a display condition that triggers the gradual scaling can include overall screen brightness being set at a low level (e.g., based on low ambient light levels around the device), as described herein. For example, the overall screen brightness may have a brightness level below a certain threshold level to trigger a gradual scaling of HDR brightness.
[0155] In some cases or conditions, such as when one or more HDR videos are being displayed and are replaced by the output HDR video, the output HDR video may be immediately displayed at its full brightness without gradual scaling. For example, when a second HDR video is displayed following the display of a first HDR video, the second HDR video may be immediately displayed at its full brightness because the user is already accustomed to viewing the increased brightness of the first HDR video.
[0156] In some implementations, after displaying the output video (e.g., either LDR or HDR output video), the device providing the display of the output video can receive user input from a user (or from another source, e.g., another device) instructing it to modify the display of the output video, e.g., increase or decrease the dynamic range of the display up to the maximum dynamic range of the display device and / or the original HDR video. For example, in response to receiving user input to modify the display, the device can modify the display coefficients (described above) according to the user input, thereby changing the dynamic range of the displayed output video. In some examples, if the user decreases the dynamic range, the display coefficients can be decreased by a corresponding amount, resulting in the output video being displayed with a lower dynamic range. This display change can be provided in real time in response to the user input.
[0157] In some implementations, after displaying the output video, the device providing the display of the output video may receive input, e.g., from a user or another source, e.g., another device, directing the output video to be modified or edited. In some implementations, the recovery map track may be discarded after such modification, e.g., it may no longer be properly applied to the LDR video. In some implementations, the recovery map track may be saved after such modification. In some implementations, the recovery map track may be updated based on modifications to the output video to reflect a modified version of the LDR video.
[0158] In some examples, when modifications are performed by a user on an HDR display and / or via an HDR-aware editing application program, the characteristics of the modifications are known over a larger dynamic range, and recovery map frames can be determined based on the edited video frames. For example, in some implementations, when a video frame is loaded into an image editing interface, the associated recovery map frame can be decoded into a full-resolution image and presented to the editor as an alpha channel, other channel, or data associated with the video frame. For example, the recovery map frame can be displayed in the editor as an additional channel in the list of channels for the video frame. One or more corresponding frames of both the LDR video and the HDR output video can be displayed in the editor to provide visualization of the modifications to both versions of the video. When a user makes edits to the video using the editor's tools, corresponding edits are made to the recovery map frame and the corresponding LDR video frame (base frame). In some examples, the user can decide whether to edit the video in RGB or RGBA (red, green, blue, alpha). In some implementations, a user can manually edit a recovery map frame (e.g., the alpha channel) and see the changes in the displayed HDR frame while the displayed corresponding LDR frame remains the same.
[0159] In some implementations, an overall dynamic range control can be provided in the image editing interface, allowing the associated range scaling factors to be adjusted higher or lower. Editing operations on the image, such as cropping and resampling, can update the alpha channel in the same manner as the RGB channels. At the time of saving the video, the LDR video can be saved back to the video container. The saved LDR video may have been edited in the interface, or may be generated from the edited HDR video via local tone mapping, or may be a new LDR video generated based on edited HDR video frame(s) scaled (e.g., divided) by the associated recovery map frame(s) at each pixel. The edited recovery map track can be saved back to the video container. In various implementations, if such encoding is used, the recovery map track can be processed back to a bilateral grid. Alternatively, the video and recovery map track can be saved losslessly, with larger storage requirements, for example, if additional editing of the video is performed.
[0160] In some embodiments, a video editor or any other program can convert the LDR video and recovery map track in the described image formats into standard HDR video, such as 10-bit HDR10 or HDR10+, HLG, or Dolby Vision video, or video in a different HDR format.
[0161] In various implementations, various blocks of methods 200, 300, 400, and / or 500 may be combined, divided into multiple blocks, performed in parallel, or performed asynchronously. In some implementations, one or more blocks of these methods may not be performed or may be performed in a different order than shown in these figures. For example, in various implementations, blocks 210 and 212 of FIG. 2 and / or blocks 304 and 306 of FIG. 3 may be performed in a different order or at least partially in parallel. Methods 200, 300, 400, and / or 500, or portions thereof, may be repeated any number of times using additional inputs. For example, in some implementations, method 200 may be performed when one or more new videos are received by the device performing the method and / or when they are stored in a user's image library.
[0162] In some implementations, the roles of the HDR video and the LDR video described with reference to Figures 2-5 may be reversed. For example, the original HDR video may be stored in a video container (e.g., as a base video or primary video) instead of an original LDR image having a lower dynamic range than the HDR video. A recovery map track indicating how to determine a derived output LDR video from the base HDR video for an LDR display device capable of displaying only a dynamic range lower than that of the HDR video may be determined and stored in the video container, as described above. The recovery map track may include recovery map frames having gain and / or scaling factors based on the base HDR video frames and the corresponding original LDR video frames; in some implementations, the original LDR video frames may be tone-mapped frames (or other range transformations of the HDR video), as described above. In some examples, the base video may be 10-bit HDR video, and each frame of the recovery map track may be generated from the base video using local tone mapping and based on the LDR video having an 8-bit LDR bit depth.
[0163] Some of the described implementations may enable applications that record HDR video to obtain and output higher quality LDR video based on local tone mapping of the HDR video via a recovery map track, without the limitations of current technologies that provide global or general tone mapping (e.g., HLG / PQ transfer functions). The local tone mapping provided by using the recovery map techniques described herein may enable higher fidelity conversion of HDR video to LDR video, for applications such as video sharing or uploading video to services or applications that only support LDR video, without applying global or general tone mapping.
[0164] In some implementations, a brightness control may be provided for displaying HDR video, such as the HDR video described herein. The HDR brightness control may be adjusted by a user of the device to specify the maximum brightness (e.g., lightness) at which the HDR video is displayed by the display device. In some implementations, the maximum brightness of the HDR video may be determined as the maximum value of the average brightness of the displayed pixels of the HDR video, because such an average may indicate the brightness of the video as perceived by the user. In other implementations, the maximum brightness may be determined as the maximum value of the brightest pixel of the HDR video, as the maximum value of one or more specific pixels of the HDR video (or the maximum value of the average values of such specific pixels), and / or based on one or more other characteristics of the HDR video. If the brightness control is set to a value less than the maximum displayable brightness of the HDR video, the device reduces the maximum brightness of the output HDR video to a lower brightness level based on the brightness control. This allows a user to adjust the maximum brightness of the displayed HDR video if the user prefers a lower maximum brightness level, for example, to avoid discomfort or fatigue due to glare from the high-brightness HDR video and / or to reduce the contrast between the high-brightness HDR video and other content on the display screen outside of the HDR video (e.g., LDR video, user interface elements, etc., which may be displayed at lower brightness). In some cases, lowering the maximum brightness level of the HDR video can reduce battery power consumption by the device providing the display due to the lower brightness of the display.
[0165] For example, the HDR brightness control may be implemented as a global or system brightness control provided by the device (e.g., provided by an operating system running on the device) that controls HDR video brightness for applications running on the device and displaying HDR video. In some examples, the brightness control may apply to all such applications running on the device, all applications of a particular type running on the device (e.g., video viewing and video editing applications), and / or specific applications specified by the user. This allows a user to specify a single brightness adjustment for all (or many) HDR videos displayed by the device, for example, without having to adjust the brightness for each individual displayed HDR video. Such a system brightness control can provide display consistency across all applications running on the device and can also provide the user with more control over the maximum brightness of the displayed HDR video. In some implementations, the HDR brightness control may be implemented together with other system controls on the device, for example, as an accessibility control, a user preference, a user setting, etc. The value set by the system brightness control can be used to adjust the output of the HDR video used to display that video, in addition to any other scaling factors, such as a display factor based on the maximum display device brightness described above, individual brightness settings for the video, etc.
[0166] In some examples, the system's brightness control may be implemented as a slider or value that a user can adjust to indicate, for example, a maximum brightness value or a percentage of the maximum displayable brightness of the HDR video by the display device (e.g., where the maximum brightness may be determined as the maximum average brightness of the pixels in the HDR video, the maximum value of the brightest or specific pixel in the HDR video, or a maximum value based on one or more other video characteristics). For example, if the brightness control is set to 70% by the user, the HDR video may be displayed at 70% of its maximum brightness (e.g., if the maximum brightness of the display device is used to determine the maximum brightness, as described above, it will be displayed at 70%). In some implementations, the system's brightness control may allow a user to specify different maximum HDR video brightness values for different screen brightnesses. For example, if the entire screen is displaying content at a higher screen brightness, the maximum HDR video brightness may be higher, such as 90% or 100%, because there is less contrast between the brightness of the entire screen and the brightness of the HDR video. At low screen brightness, a lower maximum HDR video brightness, such as 50% or 60%, can reduce the contrast between the overall screen brightness and the brightness of the HDR video. In some implementations, a system brightness control can be provided (e.g., via a slider or other interface control) as a setting related to the maximum SDR video brightness. For example, this setting can set the maximum HDR brightness value to a value N times the maximum LDR video brightness, where N has a maximum configurable value based on the maximum brightness of the display device used to display the video. In some examples, a setting value of 1 can indicate the maximum brightness of the LDR (base) video, and a setting that is the maximum or maximum setting of the brightness control indicates display at the maximum brightness of the display device.
[0167] In some implementations, these variable settings for HDR video brightness can be input via controls in a user interface. For example, a user can define a curve with a variable maximum HDR brightness relative to the current screen brightness. In some implementations, a plurality of predefined HDR brightness configurations can be provided in the user interface, and the user can select one of these configurations for use with the device. For example, each configuration can provide a different set of maximum HDR brightness values associated with a particular range of screen brightness values. In some implementations, a toggle control can be provided that, when selected by the user, removes HDR brightness so that the HDR video is displayed at LDR brightness. For example, when the toggle is selected by the user, an associated recovery map or recovery element (described herein) can be ignored, and the base video is displayed.
[0168] 6 is a diagrammatic representation of an exemplary video container 600 that may be used to provide video in a backward-compatible high dynamic range video format, e.g., to store backward-compatible high dynamic range video, according to some implementations. For example, video container 600 may be used in any of the implementations described above with respect to FIGS. 2-5. Other types of video containers may be used in various other implementations.
[0169] In this example, the video container provides data according to the MPEG-4 video format (or other ISO-based format). The video container 600 contains several blocks or units of data, such as "boxes" or "atoms," which can be identified by a type identifier and a length. These boxes can include an "ftyp" 602, which contains information about the MP4 format variant used by the file, such as the file's encoding type, compatibility, and / or intended use. A "moov" box 604 can define the timescale, duration, and indicate characteristics of the primary video data within the container (described below). A "meta" box 606 can contain metadata for the primary video data, including metadata for a top-level box, rdat 610, associated with recovery map data, such as its offset in bytes and length, as described below.
[0170] The video container 600 includes an "mdat" box 608 that contains the primary video data (e.g., base video) of the container. This can be a conventional video track within the container. For example, the primary video data can be original LDR video data, which can be converted to HDR video data with the process described above. In some other implementations, the primary video data can be original HDR video data, which can be converted to SDR video as described above. Video output devices and players that do not support the use of recovery maps to obtain output video according to the described video formats can still read and play the primary video track when loading the file.
[0171] The video container 600 includes an "rdat" box 610. In some implementations, the payload of the rdat box may be a file (e.g., an MP4 file within the MP4 file 600) that includes several boxes associated with the recovery map track. For example, the rdat box 610 may include an "ftyp" box 612 that is similar to the ftyp box 602 but applies to the rdat box 610. The "moov" box 614 may define the timescale, duration, and display characteristics of the recovery map track within the rdat box 610. The "meta" box 616 may include static metadata that applies to all frames of the recovery map track, such as a version number and / or other static metadata of the recovery map definition within the rdat box 610. For example, this box may include a "hdlr" type set to "mdta" to indicate the structure and an mdta key to the ilst entry for the recovery map track static metadata.
[0172] The rdat box 610 includes an "mdat" box 618 that can contain a recovery map track, which is the payload of the rdat box. In some implementations, the mdat box can also contain a temporal metadata track. The recovery map track can be video that includes recovery map frames, as described herein, that can be applied to associated frames of primary video data in the mdat box 608 to obtain a derived output video. In some implementations, the recovery map track can be stored at a different resolution than the primary video track and has the same aspect ratio and orientation as the primary video track.
[0173] The time-specific metadata track may be video including a respective set of gain map rendering parameters determined based on each recovery map frame, where the gain map rendering parameters include instructions or parameters used to apply associated recovery map values to corresponding frames of the primary video to obtain an output video, e.g., to fully recover the high dynamic range of the original video. For example, the gain map rendering parameters may include respective range scaling coefficients associated with the recovery map frames of the recovery map track, as described herein, and / or other parameters that may be used in some implementations, such as one or more particular epsilon values that may be desirable in a transformation operation (described above with reference to FIG. 3), dynamic metadata used in some alternative implementations that may be applied to sub-portions of the video, or other parameters that may be used as input to a transformation operation in various implementations.
[0174] In some implementations, the recovery map track and the timed metadata track are stored in separate top-level boxes to prevent video playback software and devices that do not recognize the described backward-compatible video format from unintentionally selecting the recovery map track as the primary video track.
[0175] In some implementations, multiple recovery map tracks may be included in the mdat box 618, and in some implementations, a respective associated timed metadata track may be stored in the mdat box for each of these recovery map tracks. As previously described, multiple recovery map tracks allow for selecting one of these tracks for use and providing output video with characteristics specific to the selected recovery map track. In some implementations, as previously described, a single timed metadata track may be associated with and usable for multiple recovery map tracks stored in a video container. In some implementations, one or more recovery map tracks in a video container may be associated with one respective timed metadata track, and multiple other recovery map tracks in a video container may be associated with a single timed metadata track. In some implementations, metadata may be stored in the mdat box 618 (and / or other box(es) in the rdat box 610) that identifies the intended use of the single recovery map track or multiple recovery map tracks provided in the mdat box 618, similar to that described above with respect to FIG. 2. In some examples, such intended use may include specifying a range of particular display characteristics suitable for displaying video generated based on a particular recovery map track, such as a particular dynamic range, maximum brightness, color gamut, bit depth, color space, etc.
[0176] 7 is an illustration of an exemplary approximation of a video frame 700 having a high dynamic range (the full dynamic range of an HDR image cannot be presented in this representation, e.g., an LDR depiction of an HDR scene). For example, HDR frame 700 may have been captured by a camera capable of capturing HDR video and images. Frame 700 shows image regions 702, 704, and 706 in detail, and the greater dynamic range allows these regions to be rendered in a way that appears natural to a human observer.
[0177] 8 shows an example representation of an LDR frame 800 having a lower dynamic range than HDR frame 700, such as the low dynamic range of a standard JPEG image. LDR frame 800 depicts the same scene as images 700 and 800. For example, frame 800 may be an LDR frame of a video captured by a camera. In this example, sky image region 802 is properly exposed, but foreground image regions 804 and 806 are in shadow and not fully exposed due to the lack of range compression in the frame, resulting in these regions being overly dark and losing visual detail.
[0178] 9 is another example of frame 900, which, like LDR frame 800, has a lower dynamic range than HDR frame 700, depicting the same scene as frames 700 and 800. For example, frame 900 may be an LDR frame captured by a camera. In this example, due to the lack of range compression within the frame, sky image region 902 is overexposed and therefore its details are blown out, while foreground image regions 904 and 906 are properly exposed and contain adequate detail.
[0179] FIG. 10 is an example of a range-compressed frame 1000. For example, frame 1000 can be locally tone mapped from HDR frame 700. Local tone mapping can be performed using any of a variety of tone mapping techniques. For example, the shadow areas of frame 800 can have increased luminance while preserving the contrast of image edges. Thus, while frame 1000 has a lower dynamic range than frame 700 (although not perceptible in these figures), it can preserve more frame detail than frame 800 of FIG. 8 and frame 900 of FIG. 9. A range-compressed image such as frame 1000 can be used as an LDR frame in the video formats described herein and can be included in an image container described as an LDR version of the provided HDR frame.
[0180] 11 is a block diagram of an example device 1100 that may be used to implement one or more features described herein. In one example, device 1100 may be used to implement a client device, such as any of client devices 120-126 shown in FIG. 1. Alternatively, device 1100 may implement a server device, such as server device 104. In some implementations, device 1100 may be used to implement a client device, a server device, or both a client device and a server device. Device 1100 may be any suitable computer system, server, or other electronic or hardware device, as described above.
[0181] One or more methods described herein can operate in several environments and platforms, for example, as a standalone computer program that can run on any type of computing device, as a web application having a web page, a program running on a web browser, or as a mobile application (“app”) running on a mobile computing device (e.g., a mobile phone, a smartphone, a tablet computer, a wearable device (e.g., a wristwatch, an armband, jewelry, headwear, virtual reality goggles or glasses, augmented reality goggles or glasses, a head-mounted display, etc.), a laptop computer, etc.). In one example, a client / server architecture can be used, e.g., a mobile computing device (as a client device) sends user input data to a server device and receives final output data from the server for output (e.g., display). In another example, all computations can be performed within a mobile app (and / or other apps) on the mobile computing device. In another example, computations can be split between the mobile computing device and one or more server devices.
[0182] In some implementations, device 1100 includes a processor 1102, memory 1104, and an input / output (I / O) interface 1106. Processor 1102 may be one or more processors and / or processing circuits that execute program code and control basic operations of device 1100. A "processor" includes any suitable hardware system, mechanism, or component that processes data, signals, or other information. A processor may include a general-purpose central processing unit (CPU) having one or more cores (e.g., in a single-core, dual-core, or multi-core configuration), multiple processing units (e.g., in a multiprocessor configuration), a graphics processing unit (GPU), a field programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a complex programmable logic device (CPLD), dedicated circuits for achieving functionality (e.g., one or more hardware image decoders and / or video decoders), dedicated processors for performing neural network model-based processing, neural circuits, systems with processors optimized for matrix calculations (e.g., matrix multiplication), or other systems. In some embodiments, the processor 1102 may include one or more coprocessors that perform neural network processing. In some embodiments, the processor 1102 may be a processor that processes data to produce a probabilistic output; for example, the output produced by the processor 1102 may be inaccurate or accurate within a range from an expected output. The processing need not be limited to a particular geographic location or have time limitations. For example, the processor may perform its functions in "real time," "offline," "batch mode," etc. Portions of the processing may be performed by different (or the same) processing systems at different times and in different locations. A computer may be any processor in communication with a memory.
[0183] Memory 1104 is typically provided in device 1100 for access by processor 1102 and may be any suitable processor-readable storage medium, such as random access memory (RAM), read-only memory (ROM), electrically erasable read-only memory (EEPROM), flash memory, etc., located separately from and / or integrated with processor 1102, suitable for storing instructions for execution by the processor. Memory 1104 may store software operated on server device 1100 by processor 1102, including an operating system 1108, a video application 1110 (which may be, for example, video application 106 of FIG. 1), other applications 1112, and application data 1114. Other applications 1112 may include applications such as a data display engine, a web host engine, a map application, a video display engine, a notification engine, a social networking engine, a media display application, a communication application, a web host engine or application, a media sharing application, etc. In some implementations, video application 1110 may include instructions that enable processor 1102 to perform the functions described herein, e.g., some or all of the methods of Figures 2-5. In some implementations, images and video stored in the formats described herein (including recovery maps, recovery map tracks, metadata, metadata tracks, etc.) may be stored as application data 1114 or other data in memory 1104 and / or on other storage devices of one or more other devices in communication with device 1100. In some examples, video application 1110, or other applications stored in memory 1104, may include image encoding / video encoding and container creation module(s) (e.g., performing the methods of Figures 2-4) and / or image decoding / decoding module(s) (e.g., performing the method of Figure 5).Alternatively, such modules may be combined into fewer or a single module or application.
[0184] Any software in memory 1104 may alternatively be stored on any other suitable storage location or computer-readable medium. Additionally, memory 1104 (and / or other connected storage device(s)) may store one or more messages, one or more taxonomies, electronic encyclopedias, dictionaries, digital maps, thesauri, knowledge bases, message data, grammars, user preferences, and / or other instructions and data used in the features described herein. Memory 1104 and any other type of storage (such as magnetic disk, optical disk, magnetic tape, or other tangible medium) may be considered "storage" or "storage device."
[0185] The I / O interface 1106 may provide functionality that allows the server device 1100 to interface with other systems and devices. The interfaced devices may be included as part of the device 1100 or may be separate and in communication with the computing device 1100. For example, network communication devices, storage devices (e.g., memory and / or databases), and input / output devices may communicate through the I / O interface 1106. In some implementations, the I / O interface may connect to interface devices such as input devices (keyboards, pointing devices, touchscreens, microphones, cameras, scanners, sensors, etc.) and / or output devices (display devices, speaker devices, printers, motors, etc.).
[0186] Some examples of interfaced devices that may be connected to I / O interface 1106 may include one or more display devices 1120 that can be used to display content, such as images, videos, and / or user interfaces of the applications described herein. Display device 1120 may be connected to device 1100 via a local connection (e.g., a display bus) and / or via a network connection and may be any suitable display device. Display device 1120 may include any suitable display device, such as an LCD, LED, or plasma display screen, a CRT, a television, a monitor, a touchscreen, a 3-D display screen, or other visual display device. Display device 1120 may also function as an input device, such as a touchscreen input device. For example, display device 1120 may be a flat display screen provided on a mobile device, multiple display screens provided on a glasses or headset device, or a monitor screen of a computing device.
[0187] The I / O interface 1106 can interface to other input and output devices. Some examples include one or more cameras that can capture images and video and / or detect gestures. Some implementations can provide microphones for capturing audio (e.g., as part of captured video, voice commands, etc.), radar or other sensors for detecting gestures, audio speaker devices for outputting audio, or other input and output devices.
[0188] For ease of explanation, FIG. 11 illustrates one block for each of the processor 1102, memory 1104, I / O interface 1106, and software blocks 1108-1114. These blocks may represent one or more processors or processing circuits, operating systems, memories, I / O interfaces, applications, and / or software modules. In other implementations, the device 1100 may not have all of the components illustrated, and / or may have other elements, including other types of elements, instead of or in addition to the components illustrated herein. Although some components are described as performing the blocks and operations described in some implementations herein, any suitable component or combination of components of the environment 100, the device 1100, a similar system, or any suitable processor(s) associated with such a system may perform the described blocks and operations.
[0189] The methods described herein can be implemented by computer program instructions or code executable by a computer. For example, the code can be executed by one or more digital processors (e.g., microprocessors or other processing circuits) and stored in a computer program product, including a non-transitory computer-readable medium (e.g., storage medium), such as a magnetic, optical, electromagnetic, or semiconductor storage medium including semiconductor or solid-state memory, magnetic tape, removable computer diskette, random access memory (RAM), read-only memory (ROM), flash memory, disk, optical disk, solid-state memory drive, etc. The program instructions can also be contained in electronic signals and provided as electronic signals, for example, in the form of software as a service (SaaS) delivered from a server(s) (e.g., a distributed system and / or cloud computing system). Alternatively, one or more methods can be implemented in hardware (e.g., logic gates), or a combination of hardware and software. Exemplary hardware can be a programmable processor (e.g., a field programmable gate array (FPGA), complex programmable logic device), general-purpose processor, graphics processor, application-specific integrated circuit (ASIC), etc. One or more of the methods may be implemented as part of or a component of an application executing on the system, or as an application or software running in conjunction with other applications and the operating system.
[0190] Although the description is given with respect to specific embodiments thereof, these specific embodiments are merely exemplary and not limiting, and the concepts illustrated in the examples may be applied to other examples and embodiments.
[0191] In addition to the above, a user may be provided with controls that allow the user to choose both whether and when the systems, programs, or features described herein may enable the collection of user information (e.g., information regarding the user's social network, social actions, or activities, occupation, user preferences, or the user's current location) and whether content or communications are sent from the server to the user. Furthermore, certain data may be processed in one or more ways such that personally identifiable information is removed before it is stored or used. For example, a user's identifying information may be processed so that personally identifiable information about the user cannot be determined, or if location information is obtained (e.g., to the city, zip code, or state level), the user's geographic location may be generalized so that the user's specific location cannot be determined. Thus, a user may control what information is collected about them, how that information is used, and what information is provided to them.
[0192] It should be noted that the functional blocks, operations, features, methods, devices, and systems described in this disclosure may be combined or divided into different combinations of systems, devices, and functional blocks, as would be known to one of ordinary skill in the art. Any suitable programming language and programming techniques may be used to implement the routines of a particular embodiment. Different programming techniques, e.g., procedural or object-oriented, may be used. The routines may be executed on a single processing device or on multiple processors. While steps, operations, or computations may be presented in a particular order, the order may be changed in different particular embodiments. In some embodiments, multiple steps or operations shown herein as sequential may be performed simultaneously.
Claims
1. 1. A computer-implemented method comprising: acquiring a first video comprising a plurality of first frames, each of the first frames depicting a respective scene, the first frames having a first dynamic range, the method further comprising: acquiring a second video including a plurality of second frames, each of the second frames depicting the respective scene of a corresponding one of the first frames, the second frames having a second dynamic range different from the first dynamic range, the method further comprising: generating a recovery map track, wherein generating the recovery map track includes, for each first frame of the first frames and a corresponding second frame of the second frames: generating a recovery map frame based on the first frame and the corresponding second frame, the recovery map frame encoding a luminance difference between a portion of the first frame and a corresponding portion of the corresponding second frame, the method further comprising:
11. A method comprising: providing the first video and the recovery map track in a video container, the video container being readable to display a derived video based on applying the recovery map track to the first video, the derived video including a plurality of derived frames having a dynamic range different from the first dynamic range.
2. The method of claim 1 , wherein the luminance difference is scaled by a range scaling factor comprising a ratio of a maximum luminance of the first frame to a maximum luminance of the corresponding second frame.
3. 10. The method of claim 1, further comprising generating a metadata track including each metadata frame associated with a corresponding recovery map frame, and wherein providing the video container with the first video and the recovery map track comprises providing the metadata track with the video container.
4. the luminance difference is scaled by a range scaling factor comprising a ratio of a maximum luminance of the corresponding second frame to a maximum luminance of the corresponding first frame; The method of claim 3 , wherein each said metadata frame in said timed metadata track includes a respective range scaling factor associated with an associated recovery map frame in said recovery map track.
5. 10. The method of any preceding claim, wherein the second dynamic range is greater than the first dynamic range, and the dynamic range of the derived frame is greater than the first dynamic range.
6. The video container displaying the first video by a first display device capable of displaying the first dynamic range; displaying the derived video by a second display device capable of displaying a dynamic range greater than the first dynamic range. The method of claim 5, wherein the data is readable as follows:
7. The method of claim 5 or 6, wherein obtaining the first video includes performing range compression on the second video.
8. 10. The method of claim 1, wherein generating each recovery map frame comprises encoding a luminance gain such that applying the luminance gain to the luminance of an individual pixel of the first frame results in a corresponding pixel in the corresponding second frame.
9. Generating each recovery map frame involves: recovery(x,y)=log(pixel_gain(x,y)) / log(range scaling factor), 9. The method of claim 8, wherein recovery(x, y) is the recovery map frame at pixel location (x, y) of the corresponding second frame, and pixel_gain(x, y) is the ratio of luminance at the location (x, y) of the corresponding second frame to the first frame.
10. 10. The method of any preceding claim, wherein generating each recovery map frame comprises encoding the recovery map frame into a bilateral grid.
11. 10. The method of claim 1, wherein providing the recovery map track in the video container comprises encoding the recovery map track to have a different resolution than the first video and to have the same aspect ratio as the first video.
12. obtaining the video container; determining to display the second video by the second display device; scaling luminance of a plurality of pixels of the first frame of the first video within the frame container based on a particular luminance output of the second display device and based on the corresponding recovery map frame to obtain the derived frame; 10. The method of any preceding claim, further comprising: after said scaling, causing said derived frame to be displayed by said second display device as an output frame having a different dynamic range than said first frame.
13. determining a maximum brightness display capability of the second display device; 13. The method of claim 12, wherein scaling the plurality of pixel intensities includes increasing the intensity of highlights in the first video to an intensity level that is less than or equal to the maximum intensity display capability.
14. The method of claim 1 , wherein the second dynamic range is lower than the first dynamic range, and the dynamic range of the derived frame is lower than the first dynamic range.
15. The video container displaying the first video by a first display device capable of displaying the first dynamic range; displaying the derived video on a second display device capable of displaying a dynamic range lower than the first dynamic range; 15. The method of claim 1 or 14, wherein the data is readable as follows:
16. The method of claim 14 or 15, wherein obtaining the second video includes performing range compression on the first video.
17. generating the recovery map track includes generating a plurality of recovery map tracks, and providing the recovery map track in the video container includes providing the plurality of recovery map tracks in the video container; 10. A method according to any preceding claim, wherein each of the plurality of recovery map tracks encodes a difference in luminance between a portion of the first frame and a corresponding portion of the corresponding second frame, and wherein each of the plurality of recovery map tracks can be read from the video container to display a respective derived video, each respective derived video having one or more characteristics different from each other.
18. 18. The method of claim 17, wherein the one or more characteristics of the respective derived videos that differ from one another include dynamic range, a greater dynamic range being provided from applying a first recovery map track of the plurality of recovery map tracks and a lower dynamic range being provided from applying a second dynamic range of the plurality of recovery map tracks.
19. 1. A computer-implemented method comprising: obtaining a portion of a video container, the portion comprising: a first video including a plurality of first frames, each of the first frames depicting a respective scene, the first frames having a first dynamic range, the portion further comprising: a recovery map track including a plurality of recovery map frames, each recovery map frame corresponding to a respective one of the first frames, each recovery map frame encoding a luminance gain of pixels of the corresponding first frame scaled by an associated range scaling factor including a ratio of a maximum luminance of the corresponding first frame to a maximum luminance of a corresponding second frame, the corresponding second frame depicting the respective scene of the first frames and having a second dynamic range different from the first dynamic range; The method further comprises: determining whether to display one of the first video or a derived video including a plurality of derived frames having a dynamic range different from the first dynamic range; In response to determining to display the first video, displaying at least a portion of the first video by a display device; In response to determining to display the derivative video, applying the gains of the recovery map frames of the recovery map track to luminances of pixels of the corresponding first frames to determine corresponding pixel values of each of the corresponding derived frames of the derived video; and causing at least a portion of the derivative video to be displayed by the display device.
20. 20. The method of claim 19, wherein the video container includes a metadata track including respective range scaling coefficients associated with the recovery map frames of the recovery map track, each of the respective range scaling coefficients being used to scale the luminance gain of pixels of the corresponding first frame, the luminance gain being provided by an associated recovery map frame.
21. 21. The method of claim 19, wherein the display device is capable of displaying a display dynamic range different from the first dynamic range, and applying the gain of the recovery map includes adapting the luminance of the pixel values of the corresponding derived frame to the display dynamic range of the display device.
22. The dynamic range of the derived frame is: the maximum brightness of the display device; a maximum luminance of the second dynamic range of the second frame used to generate the recovery map frame; and The method of any of claims 19 to 21, having a maximum brightness which is the smaller of
23. applying the gain of the recovery map in response to determining to display the derivative video includes: The method of any of claims 19 to 22, comprising scaling the luminance of the first frame based on a particular luminance output of the display device and based on the corresponding recovery map frame.
24. The scaling of the luminance of the first frame comprises: derived_frame(x,y)=first_frame(x,y)+log(display_factor)*recovery(x,y), 24. The method of claim 23, wherein derived_frame(x,y) is a logarithmic space version of the corresponding derived frame, first_frame(x,y) is a logarithmic space version of the first frame recovered from the frame container, display_factor is a minimum range scaling factor and a maximum display luminance of a second display device, recovery(x,y) is the recovery map for pixel location (x,y) of the first frame, and the range scaling factor is a ratio of a maximum luminance of the first frame to a maximum luminance of the corresponding second frame.
25. 10. The method of any preceding claim, wherein each recovery map frame is encoded with a bilateral grid, and further comprising decoding the recovery map frame from the bilateral grid.
26. 20. The method of claim 19, wherein the second dynamic range is lower than the first dynamic range, and the dynamic range of the derived frame is lower than the first dynamic range.
27. A method according to any one of claims 19 to 26, further comprising, in response to determining to display the derivative video, decoding additional information from the video container, the additional information including the range scaling factor.
28. The method of any of claims 19 to 27, further comprising decoding the recovery map track from blocks contained in the frame container separate from the first video.
29. determining whether to display one of the first video or the derivative video is performed by a server device, the determination being based on a display dynamic range of the display device included in or coupled to the client device; causing the display device to display at least the portion of the first video includes streaming the first video from the server device to the client device such that the client device displays the first video; 29. The method of claim 19, wherein causing at least a portion of the derivative video to be displayed by the display device includes streaming the derivative video from the server device to the client device such that the client device displays the derivative video.
30. obtaining the portion of the video container includes receiving, by a client device, the first video and the recovery map track from a server device; determining whether to display one of the first video or the derivative video is performed by a client device, the determination being based on a display dynamic range of the display device included in or coupled to the client device; causing the display device to display at least the portion of the first video is performed by the client device; 29. The method of claim 19, wherein applying the gain of the recovery map frame and causing at least a portion of the derived video to be displayed by the display device is performed by the client device.
31. 1. A system comprising: a processor; a memory coupled to the processor, the memory storing instructions that, when executed by the processor, cause the processor to perform operations, the operations including: obtaining a video container, the video container comprising: a first frame video including a plurality of first frames, the first frames having a first dynamic range, the video container further comprising: a recovery map track including a plurality of recovery map frames, each recovery map frame corresponding to a respective one of the first frames, each recovery map frame encoding a luminance gain of a pixel of the corresponding first frame, the operations further comprising: determining whether to display one of the first video or a derivative video including a plurality of derivative frames corresponding to the first frames of the depicted subject matter and having a second dynamic range different from the first dynamic range; In response to determining to display the first video, displaying at least a portion of the first video by a first display device; In response to determining to display the derivative video, applying the gains of the recovery map frames of the recovery map track to luminances of pixels of the corresponding first frames to determine corresponding pixel values of each of the corresponding derived frames of the derived video, wherein applying the gains includes scaling the luminances of the corresponding first frames based on a particular luminance output of the second display device and based on the recovery map frames, the operations further comprising: causing at least a portion of the derivative video to be displayed by the second display device.
Citation Information
Patent Citations
Encoding, Decoding, and Representing High Dynamic Range Images
JP2007534238A
Apparatus and method for converting the dynamic range of an image
JP2014532195A
Specifying visual dynamic range coding operations and parameters
US20140341305A1
Backward-compatible HDR codecs with temporal scalability
US20170295382A1
Backward-compatible video capture and distribution
US20190182487A1
Cited By
Single channel encoding into a multi-channel container and subsequent image compression
JP2025531920A
Single-channel encoding to a multi-channel container and subsequent image compression
JP7837471B2