Frame Rate Scalable Video Coding

Frame rate scalable video coding addresses the limitations of HDR content distribution by allowing flexible frame rate and shutter angle adjustments, ensuring compatibility and quality across different devices and formats.

JP7824458B2Active Publication Date: 2026-03-04DOLBY LABORATORIES LICENSING CORP
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2026-03-04

AI Technical Summary

Technical Problem

Current video distribution technologies are limited by playback device capabilities, particularly for high dynamic range (HDR) content, which is often restricted to 4K resolution and 60 frames per second, necessitating separate content versions for different devices and lacking scalable distribution methods.

Method used

The implementation of frame rate scalable video coding techniques that allow encoding and decoding processes to adjust frame rates and shutter angles, enabling flexible rendering of video content across various devices and formats, including HDR content, by using metadata and encoding strategies to support variable frame rates and shutter angles within a fixed container.

Benefits of technology

Enables backward-compatible distribution of HDR content at varying resolutions and frame rates, maintaining artistic intent while simplifying content creation and distribution, and optimizing playback on diverse devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007824458000045
    Figure 0007824458000045
  • Figure 0007824458000046
    Figure 0007824458000046
  • Figure 0007824458000047
    Figure 0007824458000047
Patent Text Reader

Abstract

To provide a method for frame-rate scalable video coding and a non-transitory processor-readable medium.SOLUTION: In a method for frame rate scalability, support is provided for input and output video sequences with a variable frame rate and variable shutter angle across scenes, or for input video sequences with a fixed input frame rate and input shutter angle, but a decoder is allowed to generate a video output at a different output frame rate and shutter angle than corresponding input values. A decoder is allowed to decode more computationally-efficiently a specific backward compatible target frame rate and shutter angle among those allowed.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority from U.S. Provisional Application No. 62 / 816,521, filed March 11, 2019, U.S. Provisional Application No. 62 / 850,985, filed May 21, 2019, U.S. Provisional Application No. 62 / 883,195, filed August 6, 2019, and U.S. Provisional Application No. 62 / 904,744, filed September 24, 2019, each of which is incorporated by reference in its entirety.

[0002] This document relates generally to images. More particularly, an embodiment of the present invention relates to frame rate scalable video coding. [Background technology]

[0003] As used herein, the term "dynamic range" (DR) may relate to the ability of the human visual system (HVS) to perceive a range of intensities (e.g., luminance, luma) in an image, e.g., from darkest gray (black) to brightest white (highlight). In this sense, DR relates to "scene-referred" intensity. DR may also relate to the ability of a display device to properly or approximately render an intensity range of a particular width. In this sense, DR relates to "display-referred" intensity. Unless explicitly designated to have a particular meaning at any point in the description herein, it should be inferred that terms may be used in either sense, e.g., interchangeably.

[0004] As used herein, the term high dynamic range (HDR) refers to a DR width that spans 14-15 orders of magnitude of the human visual system (HVS). In practice, the DR that humans can simultaneously perceive across a wide range of intensities can be somewhat truncated in terms of HDR.

[0005] In practice, an image contains one or more color components (e.g., luma Y and chroma Cb and Cr), each represented with n bits of precision per pixel (e.g., n=8). Using linear luminance encoding, images with n≦8 (e.g., color 24-bit JPEG images) are considered standard dynamic range (SDR) images, while images with n>8 are considered enhanced dynamic range images. HDR images can also be stored and distributed using high-precision (e.g., 16-bit) floating-point formats, such as the OpenEXR file format developed by Industrial Light and Magic.

[0006] Currently, the distribution of video high dynamic range content, such as Dolby Vision from Dolby Labs or HDR10 on Blu-ray, is limited to 4K resolution (e.g., 4096 × 2160 or 3840 × 2160) and 60 frames per second (fps) due to the capabilities of many playback devices. It is expected that in future versions, content up to 8K resolution (e.g., 7680 × 4320) and 120 fps may be available for distribution and playback. To simplify the HDR playback content ecosystem, such as Dolby Vision, it is desirable for future content types to be compatible with existing playback devices. Ideally, content producers should be able to adopt and distribute future HDR technologies without having to derive and distribute special versions of content compatible with existing HDR devices (e.g., HDR10 or Dolby Vision). As recognized by the present inventors, improved techniques for scalable distribution of video content, particularly HDR content, are desirable.

[0007] The approaches described in this section are approaches that could be pursued, but not necessarily approaches that have been previously conceived or pursued. Thus, unless otherwise noted, it should not be assumed that any of the approaches described in this section qualify as prior art merely by virtue of their inclusion in this section. Likewise, it should not be inferred that problems identified with one or more approaches have been recognized in any prior art based on this section, unless otherwise noted. [Brief explanation of the drawings]

[0008] Embodiments of the present invention are illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings, in which like reference numerals refer to similar elements and in which: [Figure 1] 1 illustrates an exemplary process of a video delivery pipeline. [Figure 2] 1 illustrates an exemplary process for combining consecutive original frames to render a target frame rate at a target shutter angle, according to one embodiment of the present invention. [Figure 3] 1 illustrates an exemplary display of an input sequence with a variable input frame rate and a variable shuttle angle within a container with a fixed frame rate, according to one embodiment of the present invention. [Figure 4] 10 shows an exemplary display for temporal scalability at various frame rates and shutter angles with backward compatibility according to one embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0009] Exemplary embodiments relating to frame rate scalability for video coding are described herein. In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present invention. It will be apparent, however, that the present invention may be practiced without these specific details. In other instances, well-known structures and devices are not described in exhaustive detail in order to avoid unnecessarily occluding, obscuring, or obfuscating the present invention.

[0010] overview

[0003] In one embodiment, a system including a processor receives an encoded bitstream including encoded video frames, one or more of the encoded frames being encoded at a first frame rate and a first shutter angle. The processor receives a first flag indicating a group of encoded frames to be decoded at a second frame rate and a second shutter angle, accesses from the encoded bitstream values ​​of the second frame rate and the second shutter angle for the group of encoded frames, and generates a decoded frame at the second frame rate and the second shutter angle based on the group of encoded frames, the first frame rate, the first shutter angle, the second frame rate, and the second shutter angle.

[0011] In a second embodiment, there is provided a decoder having a processor, receiving an encoded bitstream including a group of encoded video frames, wherein all of the encoded video frames in the encoded bitstream are encoded at a first frame rate; receiving a number of N combined frames; receiving a value for a baseline frame rate; accessing a group of N consecutive encoded frames, wherein an i-th encoded frame in the group of N consecutive encoded frames represents an average of the first i-th input video frames encoded at the encoder at a baseline frame rate and the i-th shutter angle, based on a first shutter angle and a first frame rate, where i=1, 2, ...N; accessing from the encoded bitstream or from user-entered values ​​for the second frame rate and the second shutter angle to decode a group of N consecutive encoded frames at the second frame rate and the second shutter angle; generating a decoded frame at a second frame rate and a second shutter angle based on the group of N consecutive encoded frames, the first frame rate, the first shutter angle, the second frame rate, and the second shutter angle.

[0012] In a third embodiment, the coded video stream structure is: a coded picture section containing a coding of a sequence of video pictures; A signaling section, a shutter interval time scale parameter indicating the number of time units that elapse in one second; a shutter interval clock tick parameter indicating the number of time units of a clock running at the frequency of the shutter interval time scale parameter, where the shutter interval divided by the shutter interval time scale parameter indicates an exposure duration value; a shutter interval duration flag indicating whether exposure duration information is fixed for all temporal sub-layers in the coded picture section; If the shutter interval duration flag indicates that the exposure duration information is fixed, then the coded version of the sequence of video pictures for all temporal sub-rays in the coded picture section is decoded by calculating exposure duration values ​​based on the shutter interval time scale parameter and the shutter interval clock tick parameter; otherwise, The signaling section includes one or more arrays of sub-layer parameters, and values ​​in the one or more arrays of sub-layer parameters combined with the shutter interval time scale parameter are used to calculate a corresponding sub-layer exposure duration value for each sub-layer to display a decoded version of the temporal sub-layers of the sequence of video pictures.

[0013] Video Delivery Processing Pipeline Example Figure 1 illustrates an exemplary process of a conventional video distribution pipeline (100), showing various stages from video capture to video content display. A series of video frames (102) are captured or generated using an image generation block (105). The video frames (102) may be digitally captured (e.g., by a digital camera) or computer-generated (e.g., using computer animation) to provide video data (107). Alternatively, the video frames (102) may be captured on film by a film camera. The film is converted to a digital format to provide the video data (107). In a production phase (110), the video data (107) is edited to provide a video production stream (112).

[0014] The video data in the production stream (112) is then provided to a processor in block (115) for post-production editing. Post-production editing in block (115) can involve adjusting or modifying the color or brightness in specific areas of the image to enhance image quality or achieve a particular look for the image according to the video creator's creative intent. This is known as "color timing" or "color grading." Other editing (e.g., scene selection and sequencing, image cropping, addition of computer-generated visual special effects, judder or blur control, frame rate control, etc.) is performed in block (115), and the video image is displayed on a reference display (125) during the final version of the production for distribution. Following post-production (115), the video data in the final production (117) may be delivered to an encoding block (120) for downstream delivery to decoding and playback devices such as television sets, set-top boxes, movie theaters, etc. In some embodiments, the encoding block (120) may include audio and video encoders, such as those defined by ATSC, DVB, DVD, Blu-Ray, and other distribution formats, to generate an encoded bitstream (122). At the receiver, the encoded bitstream (122) is decoded by a decoding unit (130) to generate a decoded signal (132) that represents an identical or approximation of the signal (117). The receiver may be attached to a target display (140), which may have characteristics that are quite different from the reference display (125). In that case, a display management block (135) may be used to map the dynamic range of the decoded signal (132) to the characteristics of the target display (140) by generating a display map signal (137).

[0015] Scalable Coding Scalable coding is already part of many video coding standards, such as MPEG-2, AVC, and HEVC. In embodiments of the present invention, scalable coding is extended to improve performance and flexibility, especially as it relates to very high resolution HDR content.

[0016] As used herein, the term "shutter angle" refers to an adjustable shutter setting that controls the percentage of time the film is exposed to light during each frame interval. For example, in one embodiment:

number

[0017] The term originates from traditional mechanical rotating shutters, but modern digital cameras also allow electronically adjustable shutters. Cinematographers can use the shutter angle to control the amount of motion blur or judder recorded in each frame. Note that alternative terms such as "exposure time," "exposure duration," "shutter interval," and "shutter speed" may be used instead of "exposure time." Similarly, the term "frame duration" may be used instead of "frame interval." Alternatively, "frame interval" may be replaced with "1 / frame rate." Exposure time values ​​are typically less than or equal to the frame duration. For example, a shutter angle of 180 degrees indicates an exposure time that is half the frame duration. In some situations, the exposure time may be longer than the frame duration of the encoded video, for example, if the encoding frame rate is 120 fps and the frame rate of the associated video content before encoding and display is 60 fps.

[0018] Consider, without limitation, an embodiment in which original content is shot (or generated) at an original frame rate (e.g., 120 fps) with a 360 degree shutter angle. The receiving device can then render video output at various frame rates below the original frame rate by, for example, intelligently combining the original frames by averaging or other operations known in the art.

[0019] Although the combining process may be performed on nonlinearly encoded signals (e.g., using gamma, PQ, or HLG), the best image quality is obtained by combining the frames in the linear light domain by first converting the nonlinearly encoded signals to a linear light representation, then combining the converted frames, and finally re-encoding the output with a nonlinear transfer function. This process provides a more accurate simulation of physical camera exposure than combining in the nonlinear domain.

[0020] In general, the process of combining frames can be expressed in terms of the original frame rate, the target frame rate, the target shutter angle, and the number of frames to be combined as follows:

number

number

[0021] where n_frames is the number of combined frames, original_frame_rate is the frame rate of the original content, target_frame_rate is the rendered frame rate (where target_frame_rate≦original_frame_rate), and target_shutter_angle indicates the desired amount of motion blur. In this example, the maximum value of target_shutter_angle is 360 degrees, which corresponds to maximum motion blur. The minimum value of target_shutter_angle can be expressed as 360*(target_frame_rate / original_frame_rate), which corresponds to minimum motion blur. The maximum value of n_frames can be expressed as (original_frame_rate / target_frame_rate). The values ​​of target_frame_rate and target_shutter_angle should be chosen so that the value of n_frame is a non-zero integer.

[0022] In the special case where the original frame rate is 120 fps, equation (2) becomes:

number

number

[0023] 2 shows an exemplary process for combining consecutive original frames to render a target frame rate at a target shutter angle, according to one embodiment. Given an input sequence (205) at 120 fps and a 360-degree shutter angle, the process generates an output video sequence (210) at 24 fps and a 216-degree shutter angle by combining three of the input frames in a set of five consecutive frames (e.g., the first three consecutive frames) and dropping the other two. In some embodiments, output frame 01 of (210) can be generated by combining any of the input frames (205), such as frames 1, 3, and 5, or frames 2, 4, and 5, although it is noted that combining consecutive frames is believed to result in a better quality video output.

[0024] For example, it may be desirable to support original content at variable frame rates to manage artistic and stylistic effects. It may also be desirable for the variable input frame rates of the original content to be packaged into a "container" with a fixed frame rate to simplify content creation, exchange, and distribution. As an example, three embodiments of how to represent variable frame rate video data in a fixed frame rate container are presented. For clarity, and without limitation, the following description uses a fixed 120 fps container, but the approach can be easily extended to alternative frame rate containers.

[0025] First embodiment (Variable Frame Rate) The first embodiment is an explicit description of original content with a variable (non-constant) frame rate packaged into a container with a constant frame rate. For example, original content with different frame rates, e.g., 24, 30, 40, 60, or 120 fps, for different scenes may be packaged into a container with a constant frame rate of 120 fps. In this example, each input frame can be duplicated 5x, 4x, 3x, 2x, or 1x times and packaged into a common 120 fps container.

[0026] Figure 3 shows an example of an input video sequence A with a variable frame rate and variable shutter angle, represented by an encoded bitstream B with a fixed frame rate. At the decoder, the decoder then reconstructs an output video sequence C with a desired frame rate and shutter angle, which may vary from scene to scene. For example, as shown in Figure 3, to construct sequence B, some of the input frames are duplicated, some are coded as is (without duplication), and some are copied four times. Then, to construct sequence C, an arbitrary frame is selected from each duplicated frame to generate an output frame with a matching original frame rate and shutter angle.

[0027] In this embodiment, metadata is inserted into the bitstream to indicate the original (base) frame rate and shutter angle. The metadata can be signaled using high-level syntax such as the Sequence Parameter Set (SPS), Picture Parameter Set (PPS), slice or tile group headers. The presence of metadata allows encoders and decoders to perform useful functions such as: a) The encoder can ignore the duplicated frame, thereby increasing the encoding speed and simplifying the process. For example, all coding tree units (CTUs) in the duplicated frame can be coded using SKIP mode and reference index 0 in LIST0 of reference frames, which refers to the decoded frame to which the duplicated frame is copied. b) The decoder can bypass decoding of duplicate frames, thereby simplifying the process. For example, metadata in the bitstream can indicate that a frame is a duplicate of a previously decoded frame, which the decoder can recreate by copying and without decoding a new frame. c) By indicating the base frame rate, playback devices can optimize downstream processes, for example by adjusting frame rate conversion or noise reduction algorithms.

[0028] This embodiment allows end users to view content rendered at the frame rate intended by the content creator. This embodiment does not provide backward compatibility with devices that do not support the container's frame rate, e.g., 120 fps.

[0029] Tables 2 and 3 show example syntax for the Raw Byte Sequence Payload (RBSP) of the Sequence Parameter Set and Tile Group Header, where the proposed new syntax elements are shown in highlighted font. The remaining syntax follows that in the proposed specification for the Generic Video Codec (VVC) (Reference [2]).

[0030] For example, in SPS (see Table 2), you can add a flag to enable variable frame rates. sps_vfr_enabled_flag equal to 1 specifies that the Coded Video Sequence (CVS) can contain variable frame rate content. sps_vfr_enabled_flag equal to 0 specifies that the CVS can contain constant frame rate content. In tile_group header() (see Table 3), tile_group_vrf_info_present_flag equal to 1 specifies that the syntax elements tile_group_true_fr and tile_group_shutterangle are present in the syntax. tile_group_vrf_info_present_flag equal to 0 specifies that the syntax elements tile_group_true_fr and tile_group_shutterangle are not present in the syntax. If tile_group_vrf_info_present_flag is not present, it is implied to be 0. tile_group_true_fr indicates the actual frame rate of the video data transmitted in this bitstream. tile_group_shutterangle indicates the shutter angle corresponding to the actual frame rate of the video data transmitted in this bitstream. tile_group_skip_flag equal to 1 specifies that the current tile group is copied from another tile group. tile_group_skip_flag equal to 0 specifies that the current tile group is not copied from another tile group. tile_group_copy_pic_order_cnt_lsb specifies the picture order count modulo MaxPicOrderCntLsb of previously decoded pictures that the current picture copies from when tile_group_skip_flag is set to 1. [Table 2] [Table 3-1] [Table 3-2]

[0031] Second embodiment - Fixed Frame Rate Container The second embodiment enables a use case in which original content with a fixed frame rate and shutter angle can be rendered by a decoder at alternative frame rates and variable simulated shutter angles, as illustrated in FIG. 2. For example, if the original content has a frame rate of 120 fps and a shutter angle of 360 degrees (meaning the shutter is open for 1 / 120 seconds), the decoder can render multiple frame rates below 120 fps. For example, to decode 24 fps with a simulated shutter angle of 216 degrees, as described in Table 1, the decoder can combine three decoded frames and display them at 24 fps. Table 4 expands on Table 1 and shows how to combine different numbers of encoded frames to render at an output target frame rate and desired target shutter angle. Frame combining can be performed by simple pixel averaging, by weighted pixel averaging, where pixels from one frame are weighted more than pixels from other frames and all weights sum to 1, or by other filter interpolation schemes known in the art. In Table 4, the function Ce(a, b) represents a combination of the encoding frames a to b, and this combination can be performed by averaging, weighted averaging, filtering, or the like. [Table 4-1] [Table 4-2]

[0032] If the target shutter angle value is less than 360 degrees, the decoder can combine different sets of decoded frames. For example, from Table 1, given an original stream at 120 fps and 360 degrees, to generate a stream at 40 fps and a shutter angle of 240 degrees, the decoder needs to combine two of three possible frames. Therefore, it can combine either the first and second frames, or the second and third frames. The selection of which frames to combine can be described in terms of a "decoding phase," expressed as follows:

number

number

[0033] Generally, decode_phase_idx is in the range of [0, n_frames_max-n_frames]. For example, for an original sequence at 120 fps and a 360-degree shutter angle, for a target frame rate of 40 fps at a 240-degree shutter angle, n_frames_max=120 / 40=3. From equation (2), n_frames=2, so decode_phase_idx is in the range of [0,1]. Therefore, decode_phase_idx=0 indicates selecting frames with indexes 0 and 1, and decode_phase_idx=1 indicates selecting frames with indexes 1 and 2.

[0034] In this embodiment, the rendered variable frame rate intended by the content creator may be signaled as metadata, such as a supplemental enhancement information (SEI) message or as video usability information (VUI). Optionally, the rendered frame rate may be controlled by the receiver or the user. An example of frame rate conversion SEI messaging specifying the content creator's preferred frame rate and shutter angle is shown in Table 5. The SEI message may also indicate whether the combined frames are performed in the coded signal domain (e.g., gamma, PQ, etc.) or the linear light domain. Note that post-processing requires a frame buffer in addition to the decoder picture buffer (DPB). The SEI message may indicate the number of additional frame buffers required or some alternative methods for combining frames. For example, to reduce complexity, frames may be recombined at a reduced spatial resolution.

[0035] As shown in Table 4, for certain combinations of frame rate and shutter angle (e.g., 30 fps and 360 degrees, or 24 fps and 288 or 360 degrees), the decoder needs to combine three or more decoded frames, which increases the amount of buffer space required by the decoder. To reduce the burden of extra buffer space in the decoder, in some embodiments, certain combinations of frame rate and shutter angle may be limited to a set of allowed decoding parameters (e.g., by setting an appropriate coding profile and level).

[0036] Again, as an example, considering playback at 24 fps, the decoder may decide to display the same frame five times that would be displayed at an output frame rate of 120 fps. This is exactly the same as showing the frame once at an output frame rate of 24 fps. The advantage of maintaining a constant output frame rate is that the display can operate at a constant clock speed, which makes all the hardware much simpler. If the display can dynamically change its clock speed, it makes more sense to display the frame only once (1 / 24 of a second) instead of repeating the same frame five times (1 / 120 of a second each). The former approach may result in slightly higher image quality, slightly better optical efficiency, or slightly better power efficiency. Similar considerations are applicable to other frame rates.

[0037] Table 5 shows an example of frame rate conversion SEI messaging syntax according to one embodiment. [Table 5]

[0038] framerate_conversion_cancel_flag=1 indicates that the SEI message cancels the persistence of any previous frame rate conversion SEI messages in output order. framerate_conversion_cancel_flag=0 indicates that frame rate conversion information follows. base_frame_rate specifies the desired frame rate. base_shutter_angle specifies the desired shutter angle. decode_phase_idx_present_flag=1 specifies that decoding phase information is present, and decode_phase_idx_present_flag=0 specifies that decoding phase information is not present. decode_phase_idx indicates the offset index within a series of sequential frames with index values ​​0..(n_frames_max-1), where n_frames_max=120 / base_frame_rate. The value of decode_phase_idx must be in the range 0..(n_frames_max-n_frames), where n_frames=base_shutter_angle / (3*base_frame_rate). If decode_phase_idx is not present, it is inferred to be 0. conversion_domain_idc=0 specifies that frame combining is performed in the linear domain. conversion_domain_idc=1 specifies that frame combining is performed in the nonlinear domain. num_frame_buffers specifies the number of additional frame buffers (not counting DPBs). framerate_conversion_persistence_flag specifies the persistence of the frame rate conversion SEI message of the current layer. When framerate_conversion_persistence_flag=0, it specifies that the frame rate conversion SEI message applies only to the current decoded picture. Let picA be the current picture. When framerate_conversion_persistence_flag=1, it specifies that the frame rate conversion SEI message of the current layer persists in output order until one or more of the following conditions become true: - A new coded layer-wise video sequence (CLVS) for the current layer begins. - The bitstream ends. - Picture picB in the current layer in the access unit containing the framer rate conversion SEI message applicable to the current layer is output (PicOrderCnt(picB) is greater than PicOrderCnt(picA)), where PicOrderCnt(picB) and PicOrderCnt(picA) are the PicOrderCntVal values ​​of picB and picA, respectively, immediately after the invocation of the decoding process for the picture order count of picB.

[0039] Third embodiment - Multiple shutter angle coded input The third embodiment is a coding scheme that allows for the extraction of sub-frame rates from the bitstream, thus supporting backward compatibility. In HEVC, this is achieved through temporal scalability. Temporal layer scalability is enabled by assigning different values ​​to the temporal_id syntax element of decoded frames. This allows for simple extraction of bitstreams based on the temporal_id value. However, the HEVC-style approach to temporal scalability does not allow for rendering output frame rates with different shutter angles. For example, a base frame rate of 60 fps extracted from a 120 fps original always has a shutter angle of 180 degrees.

[0040] Another method described in ATSC 3.0 is that a 60 fps frame with a 360-degree shutter angle is emulated as a weighted average of two 120 fps frames. The emulated 60 fps frames are assigned a temporal_id value of 0 and are interleaved with original 120 fps frames assigned a temporal_id value of 1. When 60 fps is required, the decoder only needs to decode frames with a temporal_id of 0. When 120 fps is required, the decoder can recover the corresponding original 120 fps frames that were not explicitly transmitted by subtracting each temporal_id=1 frame (i.e., a 120 fps frame) from a scaled version of each corresponding temporal_id=0 frame (i.e., an emulated 60 fps frame), thereby reconstructing all of the original 120 fps frames.

[0041] In an embodiment of the present invention, a new algorithm is described that supports multiple target frame rates and target shutter angles in a backward compatible (BC) manner. It is proposed to preprocess the original 120 fps content at a base frame rate with several shutter angles. Then, at the decoder, other frame rates at various other shutter angles can be easily derived. The ATSC 3.0 approach can be considered as a special case of the proposed scheme, where frames with temporal_id=0 carry frames at 60 fps@360 shutter angle, and frames with temporal_id=1 carry frames at 60 fps@180 shutter angle.

[0042] As a first example, consider an input sequence at 120 fps and a 360-degree shutter angle, which is used to encode sequences with a base layer frame rate of 40 fps and shutter angles of 120, 240, and 360 degrees, as depicted in FIG. 4. In this scheme, the encoder calculates new frames by combining up to three of the original input frames. For example, encoded frame 2 (En-2), representing the input at 40 fps and 240 degrees, is generated by combining input frames 1 and 2, and encoded frame 3 (En-3), representing the input at 40 fps and 360 degrees, is generated by combining frame En-2 with input frame 3. At the decoder, to reconstruct the input sequence, decoded frame 2 (Dec-2) is generated by subtracting frame En-1 from frame En-2, and decoded frame 3 (Dec-3) is generated by subtracting frame En-2 from frame En-3. The three decoded frames represent the output at the base frame rate of 120 fps and a 360-degree shutter angle. Additional frame rates and shutter angles can be extrapolated using the decoded frames as shown in Table 6. In Table 6, the function Cs(a,b) indicates the combination of input frames a to b, which can be performed by averaging, weighted averaging, filtering, etc. [Table 6-1] [Table 6-2]

[0043] An advantage of this approach is that all 40 fps versions can be decoded without further processing, as shown in Table 6. Another advantage is that other frame rates can be derived with various shutter angles. For example, consider a decoder decoding at 30 fps and a shutter angle of 360. From Table 4, the output corresponds to a sequence of frames generated by Ce(1,4) = Cs(1,4), Cs(5,8), Cs(9,12), etc., which also matches the decoded sequence shown in Table 6, except in Table 6, Cs(5,8) = e6 - e4 + e8. In one embodiment, a look-up table (LUT) can be used to define how the decoded frames need to be combined to generate an output sequence at a specified output frame rate and emulated shutter angle.

[0044] In another example, it is proposed to combine up to five frames in the encoder to simplify the extraction of a 24 fps base layer at shutter angles of 72, 144, 216, 288, and 360 degrees, as shown below. This is desirable for movie content, which is best presented at 24 fps on legacy televisions. [Table 7-1] [Table 7-2]

[0045] As shown in Table 7, if the decoded frame rate matches the baseline frame rate (24 fps), in each group of five frames (e.g., e1-e5), the decoder can simply select one frame with the desired shutter angle (e.g., e2 for a 144-degree shutter angle). To decode at a different frame rate and a specific shutter angle, the decoder needs to determine how to appropriately combine the decoded frames (e.g., by addition or subtraction). For example, to decode at 30 fps and a 180-degree shutter angle, the following steps can be followed: a) A decoder can consider a hypothetical encoder transmitting at 120 fps and 360 degrees without any consideration for backward compatibility, then from Table 1, the decoder needs to combine 2 out of 4 frames to generate an output sequence at the desired frame rate and shutter angle. For example, as shown in Table 4, the sequence may include Ce(1,2) = Avg(s1,s2), Ce(5,6) = Avg(s5,s6), etc., where Avg(s1,s2) may indicate the average of frames s1 and s2. b) Assuming that by definition the coded frames can be expressed as e1 = s1, e2 = Avg(s1,s2), e3 = Avg(s1,s3), etc., it is easy to derive that the sequence of frames in step a) can also be expressed as follows: -Ce(1,2)=Avg(s1,s2)=e2 -Ce(5,6)=Avg(s5,s6)=Avg(s1,s5)-Avg(s1,s4)+s6=e5-e4+e6 -etc. As before, the appropriate combinations for the decoded frames can be pre-computed and made available as LUTs.

[0046] The advantage of the proposed method is that it provides options for both content creators and users, i.e., it allows for both editorial and user choices. For example, preprocessing within the encoder allows for the creation of a base frame rate with various shutter angles. Each shutter angle can be assigned a temporal_id value in the range [0,(n_frames-1)], where n_frames has a value equal to 120 divided by the base frame rate. (For example, if the base frame rate is 24 fps, the temporal_id is in the range [0,4].) This choice can be made to optimize compression efficiency or for aesthetic reasons. In some use cases, such as top-down streaming, multiple bitstreams with different base layers can be encoded and stored, and then offered to the user for selection.

[0047] In a second example of the disclosed method, multiple backward-compatible frame rates can be supported. Ideally, it would be possible to decode at 24 frames per second to obtain a 24 fps base layer, at 30 frames per second to obtain a 30 fps sequence, at 60 frames per second to obtain a 60 fps sequence, and so on. If a target shutter angle is not specified, a default target shutter angle as close to 180 degrees as possible within the shutter angles allowed for the source and target frame rates is recommended. For example, given the values ​​shown in Table 7, the preferred target shutter angles for 120, 60, 40, 30, and 24 fps are 360 ​​degrees, 180 degrees, 120 degrees, 180 degrees, and 216 degrees.

[0048] From the above examples, it can be seen that the choice of how to encode content can affect the complexity of decoding a particular base layer frame rate. One embodiment of the present invention is to adaptively select an encoding scheme based on the desired base layer frame rate. For movie content, this may be, for example, 24 fps, while for sports, it may be 60 fps.

[0049] An exemplary syntax for the BC embodiment of the present invention is shown in Tables 8 and 9 below. In SPS (Table 8), two syntax elements are added: SPS_hfr_BC_enabled_flag and SPS_base_framerate (if SPS_hfr_BC_enabled_flag is set to 1). sps_hfr_BC_enabled_flag=1 specifies that backward compatible high frame rate is enabled for Coded Video Sequences (CVS). sps_hfr_BC_enabled_flag=0 specifies that backward compatible high frame rate is not enabled for CVS. sps_base_framerate specifies the base frame rate of the current CVS. In the tile group header, if sps_hfr_BC_enabled_flag is set to 1, the syntax number_avg_frames is transmitted in the bitstream. number_avg_frames specifies the number of frames at the highest frame rate (e.g., 120 fps) that are combined to produce the current picture at the base frame rate. [Table 8] [Table 9]

[0050] Modification of the second embodiment (fixed frame rate) The HEVC (H.265) coding standard (Reference [1]) and the developing Versatile Video Coding Standard (commonly referred to as VVC, see Reference [2]) define a syntax element, pic_struct, that indicates whether a picture is to be represented as a frame or as one or more fields, and whether the decoded picture is to be repeated. A copy of Table D.2, "Interpretation of pic_struct", from HEVC is provided in the Appendix for reference.

[0051] As recognized by the inventors, it is important to note that the existing pic_struct syntax element can only support a certain subset of content frame rates when using a fixed frame rate coding container. For example, when using a 60 fps fixed frame rate container, the existing pic_struct syntax can support 30 fps by using frame doubling and 24 fps by using frame doubling and frame triples, alternating frame-by-frame with each other when fixed_pic_rate_within_cvs_flag=1. However, when using a 120 fps fixed frame rate container, the current pic_struct syntax cannot support frame rates of 24 fps or 30 fps. To mitigate this issue, two new methods have been proposed: one is an extension of the HEVC version, and the other is not.

[0052] Method 1: Backwards incompatible pic_struct VVC is still under development, thus allowing maximum freedom in designing the syntax. In one embodiment, it is proposed to remove the options for frame doubling and frame tripling in pic_struct and add a new syntax element num_frame_repetition_minus2, which indicates any frame repetition using a specific value in pic_struct and specifies the number of frames to repeat. An example of the proposed syntax is given in the table below, where Table 10 shows the changes to Table D.2.3 in HEVC and Table 11 shows the changes to Table D.2 shown in the Appendix. [Table 10] num_frame_repetition_minus2+2 indicates that if fixed_pic_rate_within_cvs_flag=1, frames should be presented to the display num_frame_repetition_minus2+2 times in succession, with a frame update interval equal to DpbOutputElementalInterval[n] as given by equation E-73. [Table 11]

[0053] Method 2: Extended version of HEVC version of pic_struct Since AVC and HEVC decoders have already been adopted, it may be desirable to simply extend the existing pic_struct syntax without removing the old options. In an embodiment, a new pic_struct=13, "frame repetition extension" value, and a new syntax element num_frame_repetition_minus4 are added. An example of the proposed syntax is given in Tables 12 and 13. For pic_struct values ​​0 through 12, the proposed syntax is the same as that in Table D.2 (as shown in the Appendix), so these values ​​have been omitted for simplicity. [Table 12] num_frame_repetition_minus4+4 indicates that if fixed_pic_rate_within_cvs_flag=1, the frames should be presented to the display num_frame_repetition_minus4+4 times in succession, with a frame update interval equal to DpbOutputElementalInterval[n] as given by equation E-73. [Table 13]

[0054] In HEVC, the parameter Frame_field_info_present_flag exists in the Video Usability Information (VUI), while the syntax elements pic_struct, source_scan_type, and duplicate_flag are in the pic_timing() SEI message. In an embodiment, it is proposed to move all related syntax elements, along with frame_field_info_present_flag, to the VUI. An example of the proposed syntax is shown in Table 14. [Table 14]

[0055] Alternative signaling of shutter angle information When dealing with variable frame rates, it is desirable to identify both the desired frame rate and the desired shutter angle. In conventional video coding standards, "video usability information" (VUI) provides essential information for proper display of video content, such as aspect ratio, color primaries, chroma subsampling, etc. The VUI may also provide frame rate information when a fixed picture rate is set to 1, but there is no support for shutter angle information. Embodiments allow for the use of different shutter angles for different temporal layers, and the decoder can use the shutter angle information to improve the final appearance on the display.

[0056] For example, HEVC supports temporal sublayers, which essentially use frame-dropping techniques to transition from higher to lower frame rates. The main problem with this is that each frame drop reduces the effective shutter angle. For example, 60 fps can be derived from 120 fps video by dropping every other frame. 30 fps can be derived by dropping three out of four frames, and 24 fps can be derived by dropping four out of five frames. Assuming a full 360-degree shutter for 120 Hz, simple frame dropping results in shutter angles of 180, 90, and 72 degrees for 60 fps, 30 fps, and 24 fps, respectively. [3] Experience has shown that shutter angles less than 180 degrees are generally unacceptable, especially for frame rates below 50 Hz. By providing shutter angle information, for example, if it is desired for a display to generate a cinematic effect from 120 Hz video with reduced shutter angles for each temporal layer, smart techniques can be applied to improve the final appearance.

[0057] In another example, we may want to support different temporal layers with the same shutter angle (e.g., a 60 fps sub-bitstream within a 120 fps bitstream). The main problem is that when a 120 fps video is displayed at 120 Hz, even / odd frames have different effective shutter angles. If the display has the relevant information, smart techniques can be applied to improve the final appearance. An example of the proposed syntax is shown in Table 15, where the E.2.1 VUI parameter syntax table in HEVC (reference [1]) is modified to support shutter angle information as described above. Note that in another embodiment, instead of expressing the shutter angle syntax in absolute degrees, it can also be expressed as a ratio of frame rate to shutter speed (see Equation (1)). [Table 15] vui_shutter_angle_info_present_flag=1 specifies that shutter angle information is present in the vui_parameters() syntax structure. vui_shutter_angle_info_present_flag=0 specifies that shutter angle information is not present in the vui_parameters() syntax structure. fixed_shutter_angle_within_cvs_flag=1 specifies that the shutter angle information is the same for all temporal sub-layers in the CVS. fixed_shutter_angle_within_cvs_flag=0 specifies that the shutter angle information is not the same for all temporal sub-layers in the CVS. fixed_shutter_angle specifies the shutter angle in degrees within the CVS. The value of fixed_shutter_angle ranges from 0 to 360. sub_layer_shutter_angle[i] specifies the shutter angle in degrees when HighestTid is equal to i. The value of sub_layer_shutter_angle[i] ranges from 0 to 360.

[0058] Gradual Frame Rate Update in Coded Video Sequences (CVS) Experiments show that for HDR content displayed on an HDR display, the frame rate must be increased based on the brightness of the content to achieve the same perceived motion juddering as standard dynamic range (SDR) playback at 100 nits. In most standards (e.g., AVC, HEVC, VVC), the video frame rate can be indicated in the VUI (contained in the SPS) using the vui_time_scale, vui_num_units_in_tick, and elemental_duration_in_tc_minus1[temporal_id_max] syntax elements. For example, see Table 16 below (see Section E.2.1 of Reference [1]). [Table 16] As discussed in reference [1], the variable ClockTick is derived as follows and is called ClockTick: ClockTick=vui_num_units_in_tick÷vui_time_scale picture_duration=ClockTick*(elemental_duration_in_tc_minus1[i]+1) frame_rate=1 / pic_duration

[0059] However, the frame rate can only be changed at certain times, for example, in HEVC only at intra random access point (IRAP) frames or only at the start of a new CVS. In HDR playback, for example, in the case of fade-in or fade-out, the picture brightness changes from frame to frame, so it may be necessary to change the frame rate or picture duration for each picture. To enable the frame rate or picture duration to be refreshed at any time (even every frame), in one embodiment, a new SEI message for "progressive refresh rate" is proposed, as shown in Table 17. [Table 17]

[0060] The new syntax num_units_in_tick has the same definition as vui_num_units_in_tick, and time_scale has the same definition as vui_time_scale. num_units_in_tick is the number of time units of a clock running at a frequency of time_scale Hz that corresponds to one increment (called a clock tick) of the clock tick counter. num_units_in_tick shall be greater than 0. The number of clock ticks in seconds is equal to num_units_in_tick divided by time_scale. For example, if the picture rate of a video signal is 25 Hz, time_scale may be equal to 27 000 000, num_units_in_tick equal to 1 080 000, and therefore a clock tick may be equal to 0.04 seconds. time_scale is the number of time units that elapse in one second. For example, a time coordinate system that measures time using a 27MHz clock has a time_scale of 27 000 000. The value of time_scale shall be greater than 0. The picture duration for pictures using the gradual_refresh_rate SEI message is defined as follows: picture_duration=num_units_in_tick÷time_scale.

[0061] Signaling shutter angle information via SEI message As mentioned above, Table 15 provides an example of VUI parameter syntax with shutter angle support. By way of example, and not limitation, Table 18 lists the same syntax elements, but now as part of an SEI message for shutter angle information. Note that the SEI messaging is used only as an example, and similar messaging may be constructed at other layers of high-level syntax, such as sequence parameter sets (SPS), picture parameter sets (PPS), slice or tile group headers, etc. [Table 18]

[0062] The shutter angle is usually expressed as 0 to 360 degrees. For example, a shutter angle of 180 degrees indicates that the exposure duration is 1 / 2 of the frame duration. The shutter angle can be expressed as shutter_angle = frame_rate * 360 * shutter_speed, where shutter_speed is the exposure duration and frame_rate is the reciprocal of the frame duration. The frame_rate for a given temporal sublayer Tid can be indicated by num_units_in_tick, time_scale, elemental_duration_in_tc_minus1[Tid]. For example, if fixed_pic_rate_within_cvs_flag[Tid] = 1, frame_rate= time_scale / (num_units_in_tick*(elemental_duration_in_tc_minus1[Tid]+1)).

[0063] In some embodiments, the shutter angle value (e.g., fixed_shutter_angle) may not be an integer, for example, 135.75 degrees. To allow for greater precision, in Table 21, u(9) (unsigned 9-bit) may be replaced with u(16) or some other suitable bit depth (e.g., 12-bit, 14-bit, or greater than 16-bit).

[0064] In some embodiments, it may be useful to express the shutter angle information in terms of "clock ticks." In VVC, the variable ClockTick is derived as follows:

number

number

number

number

[0065] Table 19 shows an example of SEI messaging as shown by equation (11). In this example, the shutter angle must be greater than 0 for a real-world camera. [Table 19] As mentioned above, the use of u(16) (unsigned 16 bits) for shutter angle precision is shown as an example, and is 360 / 2 16 =0.0055 precision. The precision can be adjusted based on the actual application. For example, if you use u(8), the precision is 360 / 2 8 =1.4063. NOTE: Shutter angles are expressed as greater than 0 degrees but less than or equal to 360 degrees. For example, a shutter angle of 180 degrees indicates that the exposure duration is 1 / 2 the frame duration. fixed_shutter_angle_within_cvs_flag=1 specifies that the shutter angle value is the same for all temporal sub-layers in the CVS. fixed_shutter_angle_in_cvs_flag=0 specifies that the shutter angle value is not the same for all temporal sub-layers in the CVS. fixed_shutter_angle_numer_minus1+1 specifies the numerator used to derive the shutter angle value. The value of fixed_shutter_angle_numer_minus1 must be in the range 0 to 65535. fixed_shutter_angle_demom_minus1+1 specifies the denominator used to derive the shutter angle value. The value of fixed_shutter_angle_demom_minus1 must be in the range 0 to 65535. The value of fixed_shutter_angle_numer_minus1 must be less than or equal to the value of fixed_shutter_angle_demom_minus1. The variable shutter angle, in degrees, is derived as follows: shutterAngle=360*(fixed_shutter_angle_numer_minus1+1)÷(fixed_shutter_angle_demom_minus1+1)) sub_layer_shutter_angle_numer_minus1[i]+1 specifies the numerator used to derive the shutter angle value when HighestTid is equal to i. The value of sub_layer_shutter_angle_numer_minus1[i] must be in the range 0 to 65535, inclusive. sub_layer_shutter_angle_demom_minus1[i]+1 specifies the denominator used to derive the shutter angle value when HighestTid is equal to i. The value of sub_layer_shutter_angle_demom_minus1[i] must be between 0 and 65535, inclusive. The value of sub_layer_shutter_angle_numer_minus1[i] must be less than or equal to the value of sub_layer_shutter_angle_denom_minus1[i]. The variable sublayer shutter angle[i], in degrees, is derived as follows: subLayerShutterAngle[i]=360* (sub_layer_shutter_angle_numer_minus1[i]+1)÷(sub_layer_shutter_angle_demom_minus1[i]+1)

[0066] In another embodiment, the frame duration (e.g., frame_duration) may be specified by some other means. For example, in DVB / ATSC, when fixed_pic_rate_within_cvs_flag[Tid]=1, frame_rate=time_scale / (num_units_in_tick*(elemental_duration_in_tc_minus1[Tid]+1)),frame_duration=1 / frame_rate.

[0067] Although the syntax in Table 19 and some subsequent tables assume that the shutter angle is always greater than zero, shutter angle=0 can be used to signal a creative intent that the content should be displayed without any motion blur. This is the case for moving graphics, animations, CGI textures, matte screens, etc. Thus, for example, signaling shutter angle=0 can be useful for mode decisions in transcoders (e.g., to select an edge-preserving transcoding mode) as well as in displays that receive shutter angle metadata via the CTA interface or 3GPP interface. For example, shutter angle=0 can be used to instruct a display that should not perform motion processing such as noise reduction, frame interpolation, etc. In such an embodiment, the syntax elements fixed_shutter_angle_nume_minus1 and sub_layer_shutter_angle_numer_minus1[i] may be replaced by the syntax elements fixed_shutter_angle_numer and sub_layer_shutter_angle_nume[i], where fixed_shutter_angle_numer specifies the numerator used to derive the shutter angle value. The value of fixed_shutter_angle_numer must be in the range from 0 to 65535, inclusive. sub_layer_shutter_angle_numer[i] specifies the numerator used to derive the shutter angle value when HighestTid is equal to i. The value of sub_layer_shutter_angle_numer[i] must be in the range 0 to 65535, inclusive.

[0068] In another embodiment, fixed_shutter_angle_denom_minus1 and sub_layer_shutter_angle_denom_minus1[i] are also replaced by the syntax elements fixed_shutter_angle_denom and sub_layer_shutter_angle_denom[i].

[0069] In one embodiment, the num_units_In_tick and time_scale syntax defined in SPS can be reused by setting general_hrd_parameters_present_flag=1 in VVC, as shown in Table 20. Under this scenario, the SEI message can be renamed as exposure duration SEI message. [Table 20] fixed_exposure_duration_within_cvs_flag=1 specifies that the effective exposure duration value is the same for all temporal sub-layers in the CVS. fixed_exposure_duration_numer_minus1+1 specifies the numerator used to derive the exposure duration value. The value of fixed_exposure_duration_numer_minus1 must be in the range 0 to 65535, inclusive. fixed_exposure_duration_demom_minus1+1 specifies the denominator used to derive the exposure duration value. fixed_exposure_duration_demom_minus1 must be in the range 0 to 65535, inclusive. The value of fixed_exposure_during_numer_minus1 must be less than or equal to the value of fixed_exposure_duration_demom_minus1. The variable fixedExposureDuration is derived as follows: fixedExposureDuration=(fixed_exposure_duration_numer_minus1+1)÷(fixed_exposure_duration_demom_minus1+1)*ClockTicks. sub_layer_exposure_duration_duration_numer_minus1[i]+1 specifies the numerator used to derive the exposure duration value when HighestTid is equal to i. The value of sub_layer_exposure_duration_numer_minus1[i] must be in the range 0 to 65535, inclusive. sub_layer_exposure_duration_demom_minus1[i]+1 specifies the denominator used to derive the exposure duration value when HighestTid is equal to i. The value of sub_layer_exposure_duration_demom_minus1[i] must be in the range 0 to 65535, inclusive. The value of sub_layer_exposure_duration_numer_minus1[i] must be less than or equal to the value of sub_layer_exposure_duration_demom_minus1[i]. The variable subLayerExposureDuration[i] for HighestTid equal to i is derived as follows: subLayerExposureDuration[i]=(sub_layer_exposure_duration_numer_minus1[i]+1)÷(sub_layer_exposure_duration_demom_minus1[i]+1)*ClockTick.

[0070] In another embodiment, clockTick can be explicitly defined by the syntax elements expo_num_units_in_tick and expo_time_scale, as shown in Table 21. The advantage here is that, like the previous embodiment, it does not depend on whether the general_hrd_parameters_present_flag is set to 1 in VVC.

number

[0071] As mentioned above, the syntax parameters sub_layer_exposure_duration_numer_minus1[i] and sub_layer_exposure_duration_denom_minus1[i] are also replaced by sub_layer_exposure_duration_numer[i] and sub_layer_exposure_duration_denom[i].

[0072] In another embodiment, the parameter ShutterInterval (i.e., exposure duration) can be defined by the syntax elements sii_num_units_in_shutter_interval and sii_time_scale, as shown in Table 22, where:

number

[0073] In an alternative embodiment, instead of using a numerator and denominator to signal the sub-layer shutter interval, a single value is used. An example of such syntax is shown in Table 23. [Table 23] Shutter interval information SEI message semantics The shutter interval information SEI message indicates the shutter interval for the associated video content before encoding and display, e.g., for camera-captured content, the amount of time the image sensor is exposed to generate a picture. sii_num_units_in_shutter specifies the number of time units of a clock running at a frequency of sii_time_scale Hz that correspond to one increment of the shutter clock tick counter. The shutter interval, defined by the variable ShutterInterval, is equal to sii_num_units_in_shutter_interval divided by sii_time_scale, in seconds. For example, when ShutterInterval is equal to 0.04 seconds, sii_time_scale may be equal to 27 000 000 and sii_num_units_in_shutter_interval may be equal to 1 080 000. sii_time_scale specifies the number of time units that pass in one second. For example, in a time coordinate system that measures time using a 27MHz clock, sii_time_scale is 27 000 000. If the value of sii_time_scale is greater than 0, the value of ShutterInterval is specified by: ShutterInterval=sii_num_units_in_shutter_interval÷sii_time_scale. Otherwise (value of sii_time_scale equals 0), ShutterInterval should be interpreted as unknown or unspecified. Note 1: A value of ShutterInterval equal to 0 may indicate that the associated video content includes screen-captured content, computer-generated content, or other non-camera-captured content. NOTE 2: A value of ShutterInterval greater than the reciprocal of the coded picture rate, the coded picture interval, may indicate that the coded picture rate is greater than the picture rate at which the associated video content was generated, e.g., if the coded picture rate is 120 Hz and the picture rate of the associated video content before encoding and display is 60 Hz. The coded picture interval for a given temporal sublayer Tid may be indicated by ClockTick and elemental_duration_in_tc_minus1[Tid]. For example, when fixed_pic_rate_within_cvs_flag[Tid]=1, the picture interval for a given temporal sublayer Tid defined by the variable PictureInterval[Tid] is specified by: PictureInterval[Tid]=ClockTick*(elemental_duration_in_tc_minus1[Tid]+1) fixed_shutter_interval_into_cvs_flag=1 specifies that the value of ShutterInterval is the same for all temporal sub-layers in the CVS. fixed_shutter_interval_within_cvs_flag=0 specifies that the value of ShutterInterval is not the same for all temporal sub-layers in the CVS. sub_layer_num_units_in_shutter_interval[i] specifies the number of time units of a clock running at a frequency of sii_time_scale Hz that corresponds to one increment of the shutter clock tick counter. The sublayer shutter interval defined by the variable subLayerShutterInterval[i] is equal to the quotient of sub_layer_num_units_in_shutter_interval[i] divided by sii_time_scale when HighestTid is equal to i, in seconds. If the value of fixed_shutter_interval_within_cvs_flag is equal to 0 and the value of sii_time_scale is greater than 0, the value of subLayerShutterInterval[i] is specified by: subLayerShutterInterval[i]=sub_layer_num_units_in_shutter_interval[i]÷sii_time_scale Otherwise (value of sii_time_scale equals 0), subLayerShutterInterval[i] should be interpreted as unknown or unspecified. If the value of fixed_shutter_interval_within_cvs_flag is not equal to 0, then subLayerShutterInterval[i] = ShutterInterval.

[0074] Table 24 provides an overview of the six approaches discussed in Tables 18-23 for providing SEI messaging related to shutter angle or exposure duration. [Table 24]

[0075] Variable Frame Rate Signaling As described in U.S. Provisional Application No. 62 / 883,195, filed August 6, 2019, many applications desire decoders to support playback at variable frame rates. Frame rate adaptation is typically part of the operation of a hypothetical reference decoder (HRD), as described, for example, in Appendix C of reference [2]. In one embodiment, it is proposed to signal via SEI messaging or other means a syntax element defining the picture presentation time (PPT) as a function of a 90 kHz clock. This is a kind of iteration of the nominal decoder picture buffer (DPB) output time specified in the HRD, but now using the 90 kHz Clock Ticks precision specified in the MPEG-2 system. The advantages of this SEI message are that a) if the HRD is not enabled, the PPT SEI message can still be used to indicate the timing of each frame, and b) it can facilitate conversion between bitstream timing and system timing.

[0076] Table 25 describes an example of the proposed PPT timing message syntax, which matches the syntax of the presentation timestamp (PTS) variable used in MPEG-2 Transport (H.222) (Reference [4]). [Table 25] PPT (Picture Presentation Time) Presentation time is related to decoding time as follows: The PPT is a 33-bit number coded in three separate fields. It represents the presentation time tp at the system target decoder of presentation unit k of elementary stream n. n The PPT value is specified in units of the system clock frequency divided by 300 (which produces 90 kHz). The Picture Presentation Time is derived from the PPT according to the following formula: PPT(k)=((system_clock_frequency x tp n (k)) / 300)%2 33 . where tp n (k) is the presentation unit P n (k) Presentation time.

[0077] References Each of the references cited herein is incorporated by reference in its entirety. [1] High Efficiency Video Coding (HEVC), H.265, Series H, Video Coding, ITU, (02 / 2018) [2] B. Bross, J. Chen, and S. Liu, "Variable Video Coding (VVC) (Draft 5)," JVET Output Document, JVET-N1001, v5, (updated May 14, 2019) [3] C. Carbonara, J. DeFilippis, M. Korpi, “High Frame Rate Capture and Generation,” SMPTE 2015 Technical Conference and Exhibition, October 26-29, 2015. [4] Infrastructure for Audiovisual Services - Transmission Multiplexing and Synchronization, H.222.0, Series H, Generic Coding of Moving Pictures and Associated Audio Information: Systems, ITU, 08 / 2018

[0078] Computer system implementation example Embodiments of the present invention may be implemented in a computer system, a system comprised of electronic circuits and components, an integrated circuit (IC) device such as a microcontroller, a field programmable gate array (FPGA) or another configurable or programmable logic device (PLD), a discrete-time or digital signal processor (DSP), an application-specific IC (ASIC), and / or an apparatus including one or more of such systems, devices, or components. The computer and / or IC may execute, control, or perform instructions related to frame rate scalability as described herein. The computer and / or IC may calculate any of the various parameters or values ​​related to frame rate scalability described herein. Image and video embodiments may be implemented in hardware, software, firmware, and various combinations thereof.

[0079] Certain embodiments of the present invention include a computer processor executing software instructions that cause the processor to perform the methods of the present invention. For example, one or more processors in a display, encoder, set-top box, transcoder, etc. can implement the frame rate scalability-related methods described above by executing software instructions in a program memory accessible to the processor. Embodiments of the present invention may also be provided in the form of a program product. A program product may include any non-transitory, tangible medium that carries a set of computer-readable signals that include instructions that, when executed by a data processor, cause the data processor to perform the methods of the present invention. A program product according to the present invention may be in any of a variety of non-transitory, tangible forms. A program product may include, for example, physical media such as magnetic data storage media including floppy diskettes and hard disk drives, optical data storage media including CD-ROMs and DVDs, ROMs, flash RAM, and the like. The computer-readable signals on the program product may optionally be compressed or encrypted. Where a component (e.g., a software module, processor, assembly, device, circuit, etc.) is referred to above, unless otherwise indicated, the reference to that component (including references to "means") should be interpreted as including any component equivalent that performs the function of the described component (e.g., functionally equivalent), including components that are not structurally equivalent to the disclosed structures that perform the function in exemplary embodiments of the invention.

[0080] Equivalents, Extensions, Substitutes and Others Thus, exemplary embodiments relating to frame rate scalability are described. In the foregoing specification, embodiments of the invention have been described with reference to numerous specific details that may vary from implementation to implementation. Accordingly, the sole and exclusive indication of what the invention is and is intended by the applicant to be the invention is the set of claims issued from this application, in the particular form in which such claims are issued, including any subsequent amendments. Definitions expressly set forth herein for terms contained in such claims prevail over the meaning of the terms used in the claims. Accordingly, no limitation, element, property, feature, advantage, or attribute not expressly recited in a claim should in any way limit the scope of such claim. Accordingly, the specification and drawings are to be regarded in an illustrative, and not restrictive, sense.

[0081] appendix This appendix provides a copy of Table D.2 and the associated pic_struct related information from the H.265 specification (Reference [1]). [Table D.2-1] [Table D.2-2] Semantics of the pic_struct syntax element pic_struct indicates whether a picture is to be displayed as a frame or as one or more fields, and, for display of frames when fixed_pic_rate_within_cvs_flag=1, may indicate the frame doubling or tripling repetition period for displays using a fixed frame update interval equal to DpbOutputElementalInterval[n] given by equation E-73. The interpretation of pic_struct is specified in Table D.2. Values ​​of pic_struct not listed in Table D.2 are reserved for future use by ITU-T|ISO / IEC and will not be present in bitstreams conforming to this version of this specification. Decoders shall ignore reserved values ​​of pic_struct. If present, it is a requirement of bitstream conformance that the value of pic_struct be constrained so that exactly one of the following conditions is true: The value of -pic_struct will be equal to 0, 7, or 8 for all pictures in the CVS. The value of -pic_struct will be equal to 1, 2, 9, 10, 11, or 12 for all pictures in the CVS. The value of -pic_struct will be equal to 3, 4, 5, or 6 for all pictures in the CVS. When fixed_pic_rate_within_cvs_flag=1, frame doubling is indicated by pic_struct equal to 7, which indicates that the frame should be displayed twice consecutively on the display with a frame refresh interval equal to DpbOutputElementalInterval[n], as given by equation E-73, and frame tripling is indicated by pic_struct equal to 8, which indicates that the frame should be displayed three times consecutively on the display with a frame refresh interval equal to DpbOutputElementalInterval[n], as given by equation E-73. Note 3: Frame doubling can be used to facilitate the display, for example, of 25 Hz progressive-scan video on a 50 Hz progressive-scan display, and 30 Hz progressive-scan video on a 60 Hz progressive-scan display. A combination of frame doubling and frame tripling, alternating on each other frame, can be used to facilitate the display of 24 Hz progressive-scan video on a 60 Hz progressive-scan display. The nominal vertical and horizontal sampling positions of the samples in the top and bottom fields for 4:2:0, 4:2:2, and 4:4:4 chroma formats are shown in Figures D.1, D.2, and D.3, respectively. The field association indicator (pic_struct equals 9 to 12) provides a hint for associating fields of complementary parity as frames. Field parity can be top or bottom; if one field's parity is top and the other field's parity is bottom, the parities of the two fields are considered complementary. When frame_field_info_present_flag=1, it is a requirement of bitstream conformance that the constraints specified in the third column of Table D.2 apply. NOTE 4: When frame_field_info_present_flag=0, a default value can often be inferred or indicated by other means. In the absence of other indications of the picture's intended display type, decoders SHOULD infer the value of pic_struct to be 0 when frame_field_info_present_flag=0.

Claims

1. 1. A non-transitory processor-readable medium storing instructions for generating an encoded video stream using a processor, the processor comprising: receiving one or more video pictures and associated shutter interval information; generating an encoded video bitstream comprising a picture data section including an encoding of the one or more video pictures and metadata including a shutter interval parameter for the one or more video pictures, the shutter interval parameter being determined by: a first flag indicating whether shutter angle information is present in the metadata, and if the first flag is 1, the metadata a shutter interval time scale parameter indicating the number of time units that elapse in one second; a shutter interval clock tick parameter indicating the number of time units of a clock operating at the frequency of said shutter interval time scale parameter; a shutter interval duration flag indicating whether exposure duration information for all temporal sub-layers of the picture data section is fixed; if the shutter interval duration flag indicates that the exposure duration information is fixed, then the decoded versions of the one or more video pictures for all the temporal sub-layers of the picture data section are decoded by calculating exposure duration values ​​based on the shutter interval time scale parameter and the shutter interval clock tick parameter; Otherwise, the metadata includes one or more arrays of sub-layer parameters, values ​​in the one or more arrays of sub-layer parameters combined with the shutter interval time scale parameter being used to calculate a corresponding sub-layer exposure duration value for each sub-layer to display a decoded version of the temporal sub-layers of the one or more video pictures.

2. 2. The non-transitory processor-readable medium of claim 1, wherein the one or more arrays of sub-layer parameters include an array of sub-layer shutter interval clock tick values, each indicating a number of time units of a clock operating at a frequency of the shutter interval time scale parameter.

3. The non-transitory processor-readable medium of claim 1 , wherein the exposure duration value is calculated as a quotient of the shutter interval clock tick parameter divided by the shutter interval time scale parameter.

Citation Information

Patent Citations

  • Transmission apparatus, transmission method, reception apparatus and reception method

    WO2015076277A1

  • Image processing device, image processing method, reception device and transmission device

    WO2016185947A1