Stereoscopic High Dynamic Range Video
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- DOLBY LABORATORIES LICENSING CORP
- Filing Date
- 2023-04-19
- Publication Date
- 2026-04-27
AI Technical Summary
Current technologies face challenges in efficiently encoding and decoding stereoscopic high dynamic range (HDR) video, particularly in achieving optimal compression and image quality while supporting higher dynamic ranges than standard dynamic range (SDR).
The proposed solution involves a method for stereoscopic HDR video encoding and decoding that includes merging left and right views, applying reshaping operations, and generating composer metadata to optimize compression and image quality. This method utilizes specific encoding and decoding processes, including forward and inverse reshaping, to ensure compatibility with both HDR and SDR displays.
This approach enables efficient encoding and decoding of stereoscopic HDR video, improving compression efficiency and image quality while supporting higher dynamic ranges, thereby enhancing the immersive experience in stereo HDR video applications.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] [Related Applications] This application claims priority to U.S. Provisional Application No. 63 / 338,781, filed May 5, 2022, which is incorporated by reference in its entirety.
[0002] [Technical field] The present invention relates generally to images, and more particularly to stereoscopic transmission of high dynamic range video. [Background technology]
[0003] As used herein, the term "dynamic range (DR)" may relate to the ability of the human visual system (HVS) to perceive a range of intensities (e.g., luminance, luma) in an image, e.g., from darkest gray (black) to brightest white (highlight). In this scene, DR relates to "scene-referred" intensity. DR may also relate to the ability of a display device to properly or approximately render an intensity range of a particular width. In this scene, DR relates to "display-referred" intensity. At any point in the description herein, unless it is explicitly specified that a particular scene has a particular importance, it should be presumed that the terms may be used synonymously for either scene, e.g.,
[0004] As used herein, the term "high dynamic range (HDR)" refers to a DR width that spans approximately 14-15 times the magnitude of the human visual system (HVS) or greater. In fact, the DR that humans can simultaneously perceive in a wide range of intensity ranges may be omitted in some manner in relation to HDR. As used herein, the terms "enhanced dynamic range (EDR)" or "visual dynamic range (VDR)" refer individually or synonymously to the DR perceivable in a scene or image by the human visual system (HVS), including eye movements, allowing any light adaptation to change across the scene or image.
[0005] In practice, an image includes one or more color components (e.g., luma Y and chroma Cb and Cr), with each color component represented with n bits of precision per pixel (e.g., n=8). For example, using gamma luminance coding, an image with n≦8 (e.g., a color 24-bit JPEG image) may be considered to be a standard dynamic range image, while an image with n≧10 may be considered to be an extended dynamic range image. EDR and HDR images may be stored and distributed using high-definition (e.g., 16-bit) floating-point formats such as the OpenEXR file format developed by Industrial Light and Magic.
[0006] Most consumer desktop displays currently range from 200 to 300 cd / m 2 or nits of brightness. Most consumer HDTVs are in the 300-500 nits range, with newer models capable of 1000 nits (cd / m 2). Such conventional displays exhibit lower dynamic range (LDR) characteristics, also referred to as standard dynamic range (SDR) as opposed to HDR or EDR. As the availability of HDR content increases due to advances in both capture devices (e.g., cameras) and HDR displays (e.g., Dolby Laboratories' PRM-4200 Professional Reference Monitor), HDR content may be color graded and displayed on HDR displays that support a higher dynamic range (e.g., 1000 nits to 5000 nits or higher). In general, but not by way of limitation, the methods of the present disclosure relate to higher dynamic range than SDR.
[0007] As used herein, the term "display management" refers to processing performed at a receiver to render an image for a target display. For example, such processing may include, but is not limited to, tone mapping, color gamut mapping, color management, frame rate conversion, etc.
[0008] High Dynamic Range (HDR) content generation and playback is now widespread, as HDR technology provides more realistic and lifelike images than previous formats. Stereo HDR combined with stereoscopic video is expected to provide a more immersive experience. As the inventors recognize, improved techniques for stereoscopic transmission and display of stereoscopic HDR video are being developed to improve upon existing coding schemes.
[0009] The approaches described in this section are approaches that could be pursued, but not necessarily approaches that have been previously conceived or pursued. Thus, unless otherwise noted, it should not be assumed that any of the approaches described in this section qualify as prior art merely by virtue of their inclusion in this section. Similarly, problems identified with one or more of the approaches should not be assumed to have been recognized in any prior art under this section unless specifically indicated. [Brief description of the drawings]
[0010] Embodiments of the present invention are illustrated by way of example, and not by way of limitation, in the accompanying drawings, in which like reference symbols represent similar elements and in which:
[0011] [Figure 1A] 1 illustrates a stereoscopic HDR video encoding process according to a first embodiment of the present invention.
[0012] [Figure 1B] 1 illustrates a stereoscopic HDR video decoding process according to a first embodiment of the present invention.
[0013] [Figure 2A] 1 illustrates a stereoscopic HDR video encoding and decoding process according to a second embodiment of the present invention. [Figure 2B] 1 illustrates a stereoscopic HDR video encoding and decoding process according to a second embodiment of the present invention.
[0014] [Figure 2C] 13 illustrates a stereoscopic HDR video decoding process according to a third embodiment of the present invention.
[0015] [Diagram 3] 1 illustrates merging of left and right stereoscopic views before reshaping in accordance with an exemplary embodiment of the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0016] Herein, a method for stereoscopic HDR video encoding and decoding is described. Throughout the following detailed description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present invention. However, it will be apparent that the present invention may be practiced without some of these specific details. In other instances, well-known structures and devices are not described in exhaustive detail in order to avoid occluding, obscuring, or obscuring the present invention.
[0017] <Summary> Exemplary embodiments described herein relate to a method for a stereoscopic HDR video pipeline. In an embodiment, a processor includes: receiving a first view and a second view of a scene in a first codeword representation; applying a merge function to merge the first view and the second view into an input merged view of the first codeword representation; applying a reshaping operation to the input merged view to generate a reshaped merged view of a second codeword representation and composer metadata, the composer metadata enabling a composer function operating on the reshaped merged view to generate an output that approximates the input merged view of the first codeword representation; applying a partitioning function to the reshaped merged view to generate a first reshaped view and a second reshaped view in the second codeword representation; encoding the first reshaped view and the second reshaped view to generate a coded bitstream; The coded bitstream and the composer metadata are combined to generate a coded output.
[0018] In a second embodiment, the processor receiving a bitstream including the coded reshaped data in a second codeword representation and the composer metadata; Demultiplexing the bitstream to extract the coded reshaped data and composer metadata; Decoding the coded reshaped data to generate a reshaped merged view of the second codeword representation; applying a composer function to the reshaped merged view to generate an output merged view of the first codeword representation based on the composer metadata; A first output view and a second output view of the first codeword representation are generated based on the output merged view.
[0019] Stereo HDR Video Architecture 1 shows an example embodiment of an encoding pipeline for stereoscopic video. As used herein, the term "metadata" refers to any auxiliary information transmitted as part of a coded bitstream or sequence that assists a decoder to render a decoded image. Such metadata may include, but is not limited to, color space or gamut information, reference display parameters, and auxiliary signal parameters, as described herein.
[0020] As used herein, "forward reshaping" refers to the process of sample-to-sample or codeword-to-codeword mapping of a digital image from its original bit depth and original codeword distribution or representation (e.g., gamma, PQ, HLG, etc.) to an image of the same or different bit depth and a different codeword distribution or representation. Reshaping allows for improved compression ratio or improved image quality at a fixed bit rate. For example, and without limitation, reshaping may be applied to 10-bit or 12-bit PQ coded HDR video to improve coding efficiency in 10-bit video coding architectures. At the receiver, after decompressing the received signal (which may or may not have been reshaped), the receiver may apply an "inverse (or backward) reshaping function" to restore the signal to its original codeword distribution and / or achieve a higher dynamic range.
[0021] As shown in FIG. 1, the source content includes a left view, a right view, and metadata (102) to aid in proper display management (DM). For example, such metadata may include "L1 metadata," indicating minimum, mean, and maximum luminance values associated with an input frame or image. L1 metadata may be calculated by converting RGB data to luma-chroma format (e.g., YCbCr) and calculating the minimum, mean (average), and maximum values in the Y plane, or may be calculated directly in RGB space. For example, in one embodiment, L1Min indicates the minimum of the PQ-encoded min(RGB) values of the image while taking into account active areas (e.g., by excluding gray or black bars, letterbox bars, etc.). min(RGB) indicates the minimum of the color component values {R, G, B} of a pixel. The values of L1Mid and L1Max may also be calculated in the same manner by replacing the min() function with average() and max() functions. For example, L1Mid indicates the average of the PQ encoded max(RGB) values of the image, and L1Max indicates the maximum of the PQ encoded max(RGB) values of the image. In some embodiments, the L1 metadata may be normalized to [0, 1].
[0022] Given two views, the merging unit (105) merges the left and right views. Merging can be done horizontally or vertically, as will be explained in more detail later (depending on letterbox and pillarbox detection). Merging packs the left view video (having dimensions HxW) and the right view video (having dimensions HxW) together into a single frame. This merging is performed to optimize the efficiency of the reshaper (110). The packed images are passed through the reshaper (110) to generate a reshaped video and corresponding composer metadata (112). The reshaped bitstream is passed through a video codec (125) (e.g., HEVC, VVC, etc.) for compression. The multiplexer (130) multiplexes the coded video bitstream (127) with the composer metadata (112) and DM metadata (102) to generate an output bitstream (132).
[0023] The output of the reshaper (110) is a merged reshaped signal, which in a particular embodiment is split again in spatial splitter 115 into left and right views and goes through another frame packing process (120), this time to optimize the video encoding (125). For example, depending on the required encoding format, frame packing options include: Full resolution of each view in a side-by-side packing format with coded frame resolution of Hx2W. Horizontally downsample each view by dimension Hx(W / 2) and pack them side by side (SbS) to the final image resolution HxW. Vertically downsample each view with dimensions (H / 2)xW and pack them top-and-bottom (TaB) to the final image resolution HxW.
[0024] In one embodiment, if the output format of the spatial merging unit 105 matches the frame packing format used for video encoding, the split unit (115) and the frame packing unit (120) can be removed from the pipeline.
[0025] Note that if half resolution (either horizontal or vertical downsampling) is selected, resolution reformatting can also be performed before left and right view merging to reduce the reshaping computations by half. Note that merging (105) and frame packing (120) can have different formats. For example, left and right view merging can merge the views using TaB to facilitate reshaping, while frame packing (120) can split left and right and re-merge as SbS for packing to improve coding efficiency or other concerns.
[0026] The composer coefficients (112) are generated from a single merged frame to ensure that the same reshaping function is applied to both views.
[0027] FIG. 1B shows an example embodiment of the decoding process. The decoder side performs the reverse operation. After demultiplexing (135), a suitable video decoder (140) (matching the encoder 125) decodes the bitstream to obtain a reshaped signal. Then, the composer (145) applies the composer metadata (112) to the decoder signal to reconstruct the HDR signal. Finally, the splitter 150 regenerates the two views. The display metadata 102 can be passed directly to a display management process (not shown).
[0028] FIG. 2A shows another embodiment of the encoding process by utilizing a scalable video encoder (125B) (e.g., using HEVC temporal scalability). On the encoder side, for a given left / right pair, we spatially merge them into a single frame and perform reshaping. Composer coefficients are generated from this single merged frame. Then, in block 115, we split this merged frame into left and right reshaped images. The left view is coded as the first temporal sublayer and the right view is coded as the second temporal sublayer.
[0029] On the decoder side, as shown in FIG. 2B, after demultiplexing (135), the scalable video decoder (140B) generates two temporal sub-layers that are passed to the composer (145) to generate a reconstructed HDR image. Then, in unit 150B, the left and right views are temporally split for subsequent post-processing and display.
[0030] Other decoder implementations are possible, such as outputting the scalable video decoder (140B) to two buffers, each processed by its own composer, thus having two composers (145L, 145R) running at half speed. This alternative is shown in Figure 2C.
[0031] In alternative embodiments, a multi-view encoder (such as a multi-view HEVC encoder) can be used. In such a scenario, instead of using a scalable video encoder (125B) in FIG. 2B, a multi-view video encoder can be substituted. Similarly, for the decoder, the scalable video decoder (140B) in FIG. 2C can be substituted with a multi-view video decoder.
[0032] Stereo Reshaper As mentioned before, the input to the reshaper is the merged representation of the left and right views. One straightforward way to perform the reshaping and obtain the composer coefficients is to run the reshaper on each view separately. However, since the reshaping operation is content-dependent, this approach produces different composer coefficients for each view. To ensure that both views share the same forward and backward reshaping functions, it is recommended that the reshaper input takes the merged left and right views as one image. This ensures that the statistics of both views are taken into account together and the reshaping function is the same for both views. First, various compatibility scenarios are verified.
[0033] Non Backwards compatibility (NBC) In the context of this disclosure, backward compatibility refers to whether a decoder needs to be able to display an SDR version of the incoming content. That is, the incoming input will be displayed even if the decoder cannot apply the opposite or inverse reshaping. In a non-backwards compatible system, the decoder needs to apply the reshaped metadata (112) to generate a displayable HDR output.
[0034] In one embodiment, the NBC codec core takes in a single image and calculates block-based statistics and techniques (see, for example, references [1-2]) to determine the number of codewords required for each luminance range. Using that information, a forward lookup table (LUT) can be constructed to reshape the input HDR signal to a lower bit depth in the reshaped domain. The reshaper also outputs composer coefficients, which allow the decoder to construct a reverse reshaping LUT to reconstruct the HDR signal. To reuse this single-view functionality for stereo signals, the left and right views are merged together and sent to a single-view NBC codec, which outputs a reshaped stereo video and corresponding metadata. Note that while both side-by-side and top-and-bottom merging are feasible, a typical system may have letterbox and pillarbox detectors to exclude these non-textured dark regions from the reshaping operation. To allow existing detectors to work without modification, it is recommended to have a side-by-side merged format if the video has letterbox (black bars above and below). If the video has pillarbox (black bars on the left and right) detection, then the top-and-bottom merge format is recommended, so at the output of Merge 105, the output image will either be two views arranged side-by-side or two views arranged above and below.
[0035] After reshaping, in unit 115, the reshaped merged video can be split again into reshaped left and right views, depending on the choice of the next stage video encoder (125).
[0036] Single Layer Backwards Compatibility (SLBC) When coding HDR video in the SLBC format, the first design priority is to preserve the fidelity of the HDR content. Therefore, the reshaping is optimized to preserve as much of the original HDR view as possible while an SDR version is created. There are two types of SLBC codecs, depending on whether the SDR video version is available during the reshaping process. These are described separately below.
[0037] A reference SDR is available. If both reference SDR and HDR images of the same view are available, the single-view SLBC reshaping algorithm can take one HDR frame and one SDR frame as input and generate a reshaped SDR and the corresponding composer coefficients. If the device does not have HDR playback, it can still play back the SDR base layer. Alternatively, an HDR decoder can apply the composer metadata to reconstruct the HDR signal.
[0038] The single-view SLBC algorithm performs CDF matching by constructing a mapping curve that matches the histograms between the HDR and SDR versions of the input image (see, for example, References [3-4]). The algorithm also constructs a dynamic 3D mapping table (d3DMT) by scanning the chroma color components to resolve the appropriate composer-related metadata (see References [5]). The HDR of the left view and the HDR of the right view can be merged into one HDR image, as shown in Figure 3. The HDR of the left view and the HDR of the right view can be merged into one HDR image, as shown in Figure 3. The forward reshaping can then be reused to output a reshaped SDR signal. The corresponding composer coefficients are also output.
[0039] As in the NBC case, the SLBC pipeline can have a letterbox / pillarbox detector. Thus, depending on the box type, a control signal (302) can be used to control whether the merging is performed in SbS or TaB format. As before, the reshaped merged SDR (307) signal can be split into left and right views after the reshaper, depending on the video codec selection of the next stage.
[0040] The reference SDR is unavailable. If no reference SDR is available, the reshaper operates only on the HDR input, producing a reshaped SDR output and corresponding composer coefficients. As in the NBC case, the left view HDR and right view HDR can be merged as inputs. The merged HDR image can be passed to a no-reference SLBC codec (see, for example, [6-7]) to output a merged reshaped SDR and composer coefficients. As before, the reshaped merged SDR image can be split into left and right views after the reshaper, depending on the choice of video codec for the next stage.
[0041] Single Layer inverse Display mapping (SLiDM) For certain HDR profiles (such as Dolby Vision Profile 8.4), the original content is in SDR format. The decoder can reconstruct the HDR version, but with reshaping optimized to preserve the original SDR content. For such formats, the encoder can apply inverse mapping techniques (e.g. inverse tone mapping or inverse display management) to generate the HDR signal (reference [4]). If both SDR and HDR versions are available (reference [4]), the merging of the left and right views is the same as the process described for the SLBC codec (see Figure 3). If no HDR signal is available (reference [8]) and a static mapping from SDR to HDR is selected, the available SDR signal is transmitted as is, so there is no need to include units 105 and 115. The reshaping coefficients are simply multiplexed together with the coded stereoscopic signal.
[0042] For example, Dolby Vision Profile 8.4, which converts HLG to PQ, does not require access to the video content, since this mode works as an EOTF conversion. The composer coefficients are static for the entire video sequence. Other profiles allow for the creation of HDR versions from SDR using dynamic composer coefficients. The coefficients are created according to the content and characteristics. In this case, it is necessary to merge the left and right views and perform a content analysis to generate the appropriate metadata 112.
[0043] Stereo Coding Details As mentioned above, the video encoder 125 can be either a single-layer encoder, a scalable encoder, or a multi-view encoder. These options are described in detail in this section. Without loss of generality, additional details are described using HEVC (Reference [9]) as an example. However, the same ideas can be applied to other codecs such as AVC, AV1, VVC, etc. In this description, it is assumed that the left view is the base view and the right view is the extended view. HEVC specifies three methods to support stereo 3D encoding:
[0044] The first method is to use the frame packing arrangement SEI message of single-layer HEVC. The syntax definition of the syntax parameter frame_packing_arrangement_type supports SbS (side-by-side), TaB (top-bottom), and TI (temporal interleaving) as shown in Table D.8 of Reference [9].
[0045] In SbS, the spatial resolution of the original video is reduced by half horizontally. In TaB, the spatial resolution of the original image is reduced by half vertically. The decoder performs an up-conversion process to restore the decoded video to its original resolution.
[0046] When using temporal interleaving, the frames maintain their original resolution, but the video frame rate is doubled, and stereo 3D is possible by interleaving the two views. In the decoder, the left view is extracted from the even frames and the right view from the odd frames, or vice versa. In practice, for 8kx2kx96Hz content, such an approach requires the use of HEVC Main10, Level6.2.
[0047] Temporal Scalability The second option is to support stereoscopic video coding in single-layer HEVC using temporal sublayers. For example, in such a scenario, the left view is coded on even frames and the right view on odd frames. Such a solution also supports backward compatibility for the single-view case, i.e. the decoder can only decode the left view.
[0048] In one embodiment, assume that the maximum value of TemporalId is set to maxTemporalId (maxTemporalId>0). The right view is coded with a TemporalId equal to maxTemporalId, and the left view is coded with a TemporalId smaller than maxTemporalId.
[0049] Specifically, in a sequence parameter set (SPS), the syntax sps_max_sub_layers_minus1 is signaled. The signal sps_max_sub_layers_minus1 plus 1 specifies the maximum number of temporal sub-layers that may be present in each coded video sequence (CVS) that references the SPS. In one embodiment, maxTemporalId may be set equal to the value of sps_max_sub_layers_minus1.
[0050] In network abstraction layer (NAL) unit headers, the TemporalId is signaled by the syntax nuh_temporal_id_plus1, which is defined as follows: nuh_temporal_id_plus1 minus 1 specifies the temporal identifier of the NAL unit. The value of nuh_temporal_id_plus1 is set equal to TemporalId+1.
[0051] Thus, for the right view, TemporalId is set equal to maxTemporalId. For the left view, TemporalId is set less than maxTemporalId and greater than or equal to 0. This setting supports backward compatibility by extracting bitstreams with TemporalId less than MaxTemporalId.
[0052] To code time-interleaved video more efficiently, the following encoder settings are proposed: 1) In inter-view prediction, if a decoded picture of the left view is used to predict a picture of the right view, the reference picture of the left view needs to be marked as a long-term reference picture. In particular, in line with the MV-HEVC design, only decoded pictures of the left view with the same time instance should be used as long-term reference pictures of the right view. 2) QP Adaptation: In traditional temporal scalability, a high QP is generally used for the best temporal sublayer. In this application, the best temporal sublayer is used to code the right view, so the QP rule should be adjusted to have the best stereoscopic quality.
[0053] The main reason behind 1) is the picture order count (POC) setting in the proposed embodiment. 1) POC: By definition, at the same time instance, the left and right views should have the same POC. In the proposed solution, POC needs to be assigned based on temporal interleaving, and the meaning of POC between the left and right views no longer has a true temporal meaning. 2) Impact of POC on coding: In HEVC, in merge or advance motion vector prediction (AVMP) modes, a list of candidates is created from the motion information of spatially or temporally neighboring predicted blocks. In this process, the motion vectors (MVs) from neighboring blocks are temporally scaled using POC. Since POC is now fake, in order to keep coding efficiency, it is necessary to mark inter-view reference pictures as long-term reference pictures and disable scaling of MVs associated with long-term reference pictures.
[0054] For example, to associate metadata with left and right views in HDR coding, nuh_temporal_id_plus1 needs to be checked to ensure that the left and right views are distinct, and POC order and POC difference need to be used to associate left and right views with the same time instance.
[0055] For example, if POC1 is assigned to a left view, it needs to find an associated right view. The right view POC2 needs to have the following properties: 1) POC2 is larger than POC1 and has the smallest POC gap with POC1. 2) nuh_temporal_id_plus1 is the same as maxTemporalId+1.
[0056] If temporal scalability is to be supported, the proposed solution would need to be adapted to allow for the indication of additional temporal sublayers within the left and right views. In this case, it may be much easier to use the frame packing alignment SEI message and the MV-HEVC solution to avoid confusing the intent.
[0057] Such a case for temporal scalability might look like this: For example, most of a film is made at 24fps, but some scenes are made at 96fps. Not all decoders can support 96fps, so you might need a compatible 24 and an additional 96 for compatible decoders. You could also set a base frame rate of 24fps and use temporal scalability only for 96fps scenes. In practice, for 4kx2kx192fps such an implementation requires HEVC Main10 level 6.1.
[0058] MV-HEVC HEVC has two extensions to support 3D video: MV-HEVC and 3D-HEVC. MV-HEVC, the multiview extension, allows efficient coding of multiple camera views and associated auxiliary pictures by reusing a single-layer decoder without modifying the block-level processing modules. This enables inter-view prediction with only high-level syntax (HLS) changes. 3D-HEVC, the 3D video extension, targets coded representation of both multiple views and associated depth maps. This involves modifications to the low-level coding tool modules. Currently, MV-HEVC only supports Multiview Main Profile. In practice, to support HDR, it is necessary to define Multiview Main 10 Profile or Stereo Main 10 Profile for stereoscopic video.
[0059] 360-degree video (3DoF) support The above solution can also be extended to support 360-degree video. A standardized compression scheme for 360-degree video involves first projecting the video onto a 2D plane and then applying a compression scheme for 2D video. A typical delivery workflow is as follows (Reference
[10] ): A multi-camera array captures videos, and then image stitching is applied to obtain a spherical video. Spherical video is "unfolded" onto a 2D plane using a projection such as equirectangular (ERP) or cubemap. Video encoding using HEVC, AVC, VVC, etc., followed by packaging and delivery Receive, unpackage and decode the 2D video on the receiving side Project a 2D plane onto a sphere by specifying a specific viewpoint Rendering on the display So, from a coding and delivery point of view, 2D 360-degree video is no different from 2D flat video. For stereoscopic delivery of 360-degree HDR video, the solutions above can be applied. References 1. US 10,419,762, “Content-adaptive perceptual quantizer for high dynamic range images.” 2. US 10,032,262, “Block-based content-adaptive reshaping for high dynamic range images.” 3. US 10701375, “Encoding and decoding reversible production-quality single-layer video signals.” 4. US 10,264,287, “Inverse Luma / Chroma mappings with histogram transfer and approximation.” 5. US 11,277,627, “High-fidelity full reference and high-efficiency reduced reference encoding in end-to-end single-layer backward compatible encoding Pipepeline.” 6. US Patent Application Publication 2022-0046245-A1, “Interpolation of reshaping functions.” 7. WIPO PCT Patent Application Publication WO 2021 / 216607, “Reshaping functions for HDR imaging with continuity and reversibility constraints.” 8. US Patent Application Ser. No. 17 / 630,901, “Electro-optical transfer function conversion and signal legalization,” filed on 27 Jan 2022, GM. Su, et al. 9. ITU-T Rec. H.265, “High efficiency video coding,” ITU, version 08 / 2021. 10. Ref: “Recent trends and challenges in 360-degree video compression”, Y. Ye, ICME 2018, Hot3D talk. Each of these references is incorporated by reference in its entirety.
[0060] Exemplary Computer System Implementation The embodiments of the present invention may be implemented by a computer system, a system configured in electronic circuits and components, an integrated circuit (IC) device such as a microcontroller, a field programmable gate array (FPGA) or other configurable or programmable logic device (PLD), a discrete time or digital signal processor (DSP), an application specific IC (ASIC), and / or an apparatus including one or more of such systems, devices, or components. The computer and / or IC may execute, control, or perform instructions related to image transformation as described herein. The computer and / or IC may calculate any of the various parameters or values related to stereoscopic HDR video processing as described herein. The image and video embodiments may be implemented in hardware, software, firmware, and various combinations thereof.
[0061] A particular implementation of the present invention includes a computer processor executing software instructions that cause the processor to perform the method of the present invention. For example, one or more processors in a display, an encoder, a set-top box, a transcoder, etc. may perform the method related to stereoscopic HDR video processing described above by executing software instructions in a program memory accessible to the processor. The present invention may be provided in the form of a program product. The program product may include any tangible non-transitory medium that carries a set of computer-readable signals that include instructions that, when executed by a data processor, cause the data processor to perform the method of the present invention. A program product according to the present invention may be in any of a variety of tangible forms. The program product may include a physical medium, such as, for example, a magnetic data storage medium including a floppy disk, a hard disk drive, an optical data storage medium including a CD-ROM, a DVD, an electronic data storage medium including a ROM, a flash RAM, etc. The computer-readable signals on the program product may be optically compressed or encrypted.
[0062] Although components (e.g., software modules, processors, components, devices, circuits, etc.) have been referred to above, unless otherwise indicated, references to those components (including references to "means") should be interpreted to include equivalents of those components, any components that perform the functions of the described components (e.g., are functionally equivalent), and components that are not structurally equivalent to the disclosed structures that perform the functions in the illustrated exemplary embodiments of the present invention.
[0063] <Equivalents, Extensions, Alternatives and Miscellaneous> Exemplary embodiments related to stereoscopic HDR video processing are described. In the above specification, embodiments of the invention have been described with reference to numerous specific details that may vary from implementation to implementation. Thus, the sole and exclusive indication of what the invention is, and what the applicant intends the invention to be, is set forth in the claims issued in a particular form hereby, including any subsequent amendments. Any definitions expressly set forth herein for terms contained in such claims shall govern the meaning of such terms as used in the claims. Thus, any limitations, elements, features, advantages, or attributes not expressly set forth in the claims should not limit the scope of the claims in any manner. The specification and drawings should therefore be considered in an illustrative, rather than restrictive, sense.
Claims
1. A method for encoding stereoscopic high dynamic range (HDR) video, wherein the method is: The steps include receiving the first and second views of the scene in the first codeword representation, A step of applying a merge function to merge the first view and the second view into a first merged view of the first codeword representation, wherein the merge function uses a first frame packing format. A step of applying a reshaping process to the first merge view to generate a reshaped merge view and composer metadata for a second codeword representation, wherein the composer metadata enables a composer function operating on the reshaped merge view to generate an output that approximates the first merge view of the first codeword representation. The steps include: applying a splitting function to the reshaped merge view to generate a first reshaped view and a second reshaped view of the second codeword representation; The steps include merging the first reshaped view and the second reshaped view into a second merged view using a different second frame packing format, The steps include encoding the second merge view to generate a coded bitstream, The steps include: combining the coded bitstream and the composer metadata to generate coded output; A method that includes this.
2. The method according to claim 1, wherein the first merged view includes a side-by-side arrangement of the first view and the second view, or a top-and-bottom arrangement of the first view and the second view.
3. The method according to claim 2, wherein a side-by-side arrangement is applied when the first view and the second view are in letterbox format, and a top-and-bottom arrangement is applied when the first view and the second view are in pillar format.
4. The step of encoding the second merge view is: The steps include: frame packing the second merge view into a single frame according to the aforementioned different second frame packing format; The steps include compressing the single frame using a video coder and Supplemental Enhancement Information (SEI) messaging indicating the different second frame packing format, The method according to claim 1, including the method described in claim 1.
5. The method according to claim 1, wherein the step of encoding the second merge view includes the step of using time interleaving.
6. The method according to claim 1, wherein the step of encoding the second merged view includes the step of using scalable video coding, multiview coding, or 3D coding.
7. The encoding by the aforementioned scalable video coding is The steps include setting a TemporalId equal to maxTemporalId for the second reshaped view, A step of setting a TemporalId for the first reshaped view that is less than maxTemporalId but greater than or equal to 0, The method according to claim 6, including the method described in claim 6.
8. The aforementioned reshaping process is The steps include receiving the first SDR view and the second SDR view of the aforementioned scene, The steps include: applying the merge function to merge the first SDR view and the second SDR view into an SDR merge view; The steps include: applying the reshaping process using both the SDR merge view and the first merge view to generate a reshaped merge view of the second codeword representation and the composer metadata; The method according to claim 1, including the method described in claim 1.
9. A device including a processor and configured to perform the method according to claim 1.
10. A non-temporary computer-readable storage medium storing computer-executable instructions that cause one or more processors to perform the method described in claim 1.
11. The first frame packing format is selected to optimize the efficiency of the reshaping process, The second frame packing format is selected to optimize the encoding. The method according to claim 1.