Frame-rate scalable video coding
Frame rate scalability in video encoding addresses the challenge of delivering HDR content at varying frame rates and shutter angles, ensuring compatibility with existing devices by using metadata and signaling mechanisms to manage frame replication and extraction, enhancing encoding and decoding efficiency.
Patent Information
- Application Number
- JP2025053295
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-09-24
- Filing Date
- 2025-03-27
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2040-03-11
AI Technical Summary
Current video delivery systems are limited in their ability to efficiently deliver high dynamic range (HDR) content at varying frame rates and shutter angles, particularly for playback devices with different capabilities, necessitating the development of scalable video coding techniques to ensure compatibility with existing devices.
The implementation of frame rate scalability in video encoding through a processor that receives and decodes encoded video frames at varying frame rates and shutter angles, using metadata and signaling mechanisms to manage frame replication, combination, and extraction, enabling backward compatibility with existing playback devices.
Enables the delivery of HDR content at variable frame rates and shutter angles, ensuring compatibility with existing playback devices and allowing for artistic and stylistic control, while optimizing encoding and decoding processes.
Smart Images

Figure 2025094251000001_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims the benefit of priority from U.S. Provisional Application No. 62 / 816,521, filed Mar. 11, 2019; U.S. Provisional Application No. 62 / 850,985, filed May 21, 2019; U.S. Provisional Application No. 62 / 883,195, filed Aug. 6, 2019; and U.S. Provisional Application No. 62 / 904,744, filed Sep. 24, 2019, each of which is hereby incorporated by reference in its entirety.
[0002] This document generally relates to images. More particularly, one embodiment of the present invention relates to frame - rate scalable video coding.
Background Art
[0003] As used herein, the term "dynamic range" (DR) may relate, for example, to the ability of the human visual system (HVS) to perceive the range of intensities (e.g., luminance, luma) within an image from the darkest gray (black) to the brightest white (highlight). In this sense, DR relates to "scene - referred" intensities. DR may also relate to the ability of a display device to properly or generally render a particular width of intensity range. In this sense, DR relates to "display - referred" intensities. Unless explicitly specified to have a particular meaning at any point in the description herein, the terms should be presumed to be used interchangeably in either sense, for example.
[0004] As used herein, the term "high dynamic range" (HDR) relates to a DR width of 14 - 15 orders of magnitude of the human visual system (HVS). In practice, the DR that a human can simultaneously perceive over a wide range of intensity can be somewhat truncated with respect to HDR.
[0005] In practice, an image contains one or more color components (e.g., luma Y and chroma Cb and Cr), and each color component is represented with an accuracy of n bits per pixel (e.g., n = 8). Using linear luminance encoding, an image with n ≤ 8 (e.g., a 24-bit color JPEG image) is considered a standard dynamic range (SDR) image, and an image with n > 8 is considered a high dynamic range image. Also, HDR images can be stored and distributed using a high-precision (e.g., 16-bit) floating-point format such as the OpenEXR file format developed by Industrial Light and Magic.
[0006] Currently, the delivery of high dynamic range video content such as Dolby Vision from Dolby Laboratories or HDR10 in Blu-Ray is limited to 4K resolution (e.g., 4096×2160 or 3840×2160, etc.) and 60 frames per second (fps) by the capabilities of many playback devices. In future versions, content up to 8K resolution (e.g., 7680×4320) and 120 fps may be available for delivery and playback. To simplify HDR playback content ecosystems such as Dolby Vision, it is desirable that future content types be compatible with existing playback devices. Ideally, content producers should be able to adopt and deliver future HDR technologies without the need to derive and distribute special versions of content that are compatible with existing HDR devices (such as HDR10 or Dolby Vision). As understood by the inventors, improved techniques for scalable delivery of video content, particularly HDR content, are desired.
[0007] The approach described in this section is an approach that can be pursued, but it is not necessarily an approach that has been previously conceived or pursued. Therefore, unless otherwise specified, none of the approaches described in this section should be assumed to be eligible as prior art merely because they are included in this section. Similarly, unless otherwise specified, the problems identified with respect to one or more approaches should not be presumed to be recognized in any prior art based on this section.
Brief Description of the Drawings
[0008] Embodiments of the present invention are shown by way of example and not limitation in the figures of the accompanying drawings, and like reference numerals refer to like elements.
Figure 1
Figure 2
Figure 3
Figure 4
Modes for Carrying Out the Invention
[0009] Exemplary embodiments related to frame rate scalability for video encoding are described herein. In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present invention. It will be apparent, however, that the present invention may be practiced without these specific details. In other instances, well-known structures and devices are not described in exhaustive detail to avoid obscuring, obfuscating, or rendering the present invention unreadable.
[0010] Summary The exemplary embodiments described herein relate to frame rate scalability in video encoding. In one embodiment, a system having a processor receives an encoded bitstream including encoded video frames, where one or more of the encoded frames are encoded at a first frame rate and a first shutter angle. The processor receives a first flag indicating the presence of a group of encoded frames to be decoded at a second frame rate and a second shutter angle, accesses the encoded bitstream values for the second frame rate and the second shutter angle for the group of encoded frames, and generates decoded frames at the second frame rate and the second shutter angle based on the group of encoded frames, the first frame rate, the first shutter angle, the second frame rate, and the second shutter angle.
[0011] In a second embodiment, a decoder having a processor, receiving an encoded bitstream including a group of encoded video frames, wherein all of the encoded video frames in the encoded bitstream are encoded at a first frame rate, receiving the encoded bitstream, receiving the number of combined N frames, receiving a value for a baseline frame rate, Accessing a group of N consecutive encoded frames, wherein the i-th encoded frame within the group of N consecutive encoded frames represents the average of the input video frames up to the i-th frame encoded in the encoder at a baseline frame rate and the i-th shutter angle, based on a first shutter angle and a first frame rate, where i = 1, 2,... N, and accessing a group of N consecutive encoded frames, Accessing from an encoded bitstream or from user input values for a second frame rate and a second shutter angle to decode a group of N consecutive encoded frames at the second frame rate and the second shutter angle, Generating decoded frames at the second frame rate and the second shutter angle based on the group of N consecutive encoded frames, the first frame rate, the first shutter angle, the second frame rate, and the second shutter angle.
[0012] In a third embodiment, the encoded video stream structure includes An encoded picture section including encoding of a sequence of video pictures, A signaling section, A shutter interval time scale parameter indicating the number of time units elapsed in one second, A shutter interval clock tick parameter indicating the number of time units of a clock operating at the frequency of the shutter interval time scale parameter, where the shutter interval divided by the shutter interval time scale parameter represents an exposure duration value, and a shutter interval clock tick parameter, A signaling section including encoding of a shutter interval duration flag indicating whether exposure duration information is fixed for all temporal sublayers within the encoded picture section. When the shutter interval duration flag indicates that the exposure duration information is fixed, for all temporal sub-layers within the coded picture section, the coded version of the video picture sequence is decoded by calculating the exposure duration value based on the shutter interval time scale parameter and the shutter interval clock tick parameter; otherwise, The signaling section includes one or more arrays of sub-layer parameters, and the values in one or more arrays of sub-layer parameters combined with the shutter interval time scale parameter are used to calculate the corresponding sub-layer exposure duration value for each sub-layer to display the decoded version of the temporal sub-layer of the video picture sequence.
[0013] Example of a video delivery processing pipeline FIG. 1 shows an exemplary process of a conventional video delivery pipeline (100) showing various stages from video capture to video content display. A series of video frames (102) are captured or generated using an image generation block (105). The video frames (102) may be digitally captured (e.g., by a digital camera) or computer-generated (e.g., using computer animation) to provide video data (107). Alternatively, the video frames (102) may be captured on film by a film camera. The film is converted to a digital format to provide video data (107). In the production phase (110), the video data (107) is edited to provide a video production stream (112).
[0014] Next, the video data of the production stream (112) is supplied to the processor in block (115) for post-production editing. The post-production editing block (115) can include adjusting or modifying the color or brightness in specific regions of the image to improve the image quality or achieve a specific appearance of the image according to the creative intention of the video creator. This is what is called "color timing" or "color grading". Other edits (such as scene selection and sequencing, image cropping, addition of computer-generated visual special effects, judder or blur control, frame rate control, etc.) are performed in block (115), and during the final version of the production for distribution, the video image is displayed on the reference display (125). Following post-production (115), the video data of the final production (117) may be distributed to the encoding block (120) for downstream distribution to decoding and playback devices such as television sets, set-top boxes, movie theaters, etc. In some embodiments, the encoding block (120) can include an audio encoder and a video encoder, such as those defined by ATSC, DVB, DVD, Blu-Ray, and other distribution formats, to generate the encoded bitstream (122). At the receiver, the encoded bitstream (122) is decoded by the decoding unit (130) to generate a decoded signal (132) representing the same or an approximation of the signal (117). The receiver can be attached to a target display (140) that can have characteristics quite different from those of the reference display (125). In that case, a display management block (135) can be used to map the dynamic range of the decoded signal (132) to the characteristics of the target display (140) by generating a display map signal (137).
[0015] Scalable Encoding Scalable coding is already part of many video coding standards such as MPEG-2, AVC, and HEVC. In embodiments of the present invention, scalable coding is extended to improve performance and flexibility, particularly with respect to very high resolution HDR content.
[0016] As used herein, the term "shutter angle" refers to an adjustable shutter setting that controls the proportion of time during each frame interval that the film is exposed to light. For example, in one embodiment,
Number
[0017] This term has its origin in traditional mechanical rotating shutters, but modern digital cameras can also electronically adjust the shutter. Cinematographers can use the shutter angle to control the amount of motion blur and judder recorded in each frame. Note that alternative terms such as "exposure duration", "shutter interval", "shutter speed" may be used instead of "exposure time". Similarly, the term "frame duration" may be used instead of "frame interval". Alternatively, "frame interval" may be replaced with "1 / frame rate". The value of the exposure time is usually less than or equal to the duration of the frame. For example, a shutter angle of 180 degrees indicates that the exposure time is half of the frame duration. In some situations, the exposure time may be longer than the frame duration of the encoded video, for example, when the encoding frame rate is 120 fps and the frame rate of the relevant video content before encoding and display is 60 fps.
[0018] Consider embodiments where, rather than being limited, the original content is captured (or generated) at an original frame rate (e.g., 120 fps) with a 360 - degree shutter angle. Next, at the receiving device, the video output can be rendered at various frame rates below the original frame rate, for example, by averaging or other operations known in the art, by a judicious combination of the original frames.
[0019] The combining process may be performed on non - linear encoded signals (e.g., using gamma, PQ, or HLG), but the best image quality is obtained by first converting the non - linear encoded signal to a linear light representation, then combining the converted frames, and finally re - encoding the output with a non - linear transfer function, by combining the frames within the linear light region. This process provides a more accurate simulation of the physical camera exposure than combining in the non - linear region.
[0020] Generally, the process of combining frames can be expressed as follows in terms of the original frame rate, the target frame rate, the target shutter angle, and the number of frames to be combined.
Number
Number
[0021] Here, n_frames is the number of combined frames, original_frame_rate is the frame rate of the original content, target_frame_rate is the frame rate to be rendered (where target_frame_rate ≤ original_frame_rate), and target_shutter_angle indicates the amount of desired motion blur. In this example, the maximum value of target_shutter_angle is 360 degrees, corresponding to the maximum motion blur. The minimum value of target_shutter_angle can be expressed as 360*(target_frame_rate / original_frame_rate), corresponding to the minimum motion blur. The maximum value of n_frames can be expressed as (original_frame_rate / target_frame_rate). The values of target_frame_rate and target_shutter_angle should be selected such that the value of n_frame is a non-zero integer.
[0022] In the special case where the original frame rate is 120 fps, Equation (2) can be rewritten as
Equation
Equation
Table 1
[0023] Figure 2 shows an exemplary process for combining consecutive original frames to render a target frame rate at a target shutter angle, according to one embodiment. Given an input sequence (205) at 120 fps and a 360-degree shutter angle, the process combines three of the input frames (e.g., the first three consecutive frames) within a set of five consecutive frames and drops the other two to produce an output video sequence (210) at 24 fps and a 216-degree shutter angle. In some embodiments, output frame 01 of (210) can be produced by combining any of the input frames (205) such as frames 1, 3, and 5, or frames 2, 4, and 5, although it should be noted that combining consecutive frames is thought to result in better quality video output.
[0024] For example, it is desirable to support original content at variable frame rates in order to manage artistic and stylistic effects. Also, the variable input frame rate of the original content is desirably packaged into a "container" with a fixed frame rate in order to simplify the generation, exchange, and delivery of the content. As an example, three embodiments are presented regarding how to represent variable frame rate video data in a fixed frame rate container. For clarity and without limitation, the following description uses a fixed 120 fps container, although the approach can be easily extended to alternative frame rate containers.
[0025] First Embodiment (Variable frame rate) The first embodiment is an explicit description of original content having a variable (not constant) frame rate packaged in a container having a constant frame rate. For example, original content having different frame rates, e.g., 24, 30, 40, 60, or 120 fps for different scenes, may be packaged in a container having a constant frame rate of 120 fps. In this example, each input frame can be replicated 5x, 4x, 3x, 2x, or 1x and packaged into a common 120 fps container.
[0026] FIG. 3 shows an example of an input video sequence A having a variable frame rate and a variable shutter angle represented by an encoded bitstream B having a fixed frame rate. And in the decoder, the decoder reconstructs an output video sequence C at a desired frame rate and shutter angle that can vary for each scene. For example, as shown in FIG. 3, to construct sequence B, some of the input frames are replicated, some are encoded as is (without replication), and some are copied 4 times. Then, to construct sequence C, any one frame is selected from each replicated frame to generate an output frame with the same original frame rate and shutter angle.
[0027] In this embodiment, metadata is inserted into the bitstream to indicate the original (base) frame rate and shutter angle. The metadata can be signaled using high-level syntax such as a sequence parameter set (SPS), a picture parameter set (PPS), a slice or tile group header. The presence of the metadata enables the encoder and decoder to perform useful functions as follows. a) The encoder can ignore the duplicated frames, thereby improving the encoding speed and simplifying the process. For example, all coding tree units (CTUs) within the duplicated frame can be encoded using the SKIP mode and reference index 0 in LIST0 of the reference frames. This refers to the decoded frame to which the duplicated frame is copied. b) The decoder can bypass the decoding of the duplicated frames, thereby simplifying the process. For example, the metadata within the bitstream can indicate that the frame is a copy of a previously decoded frame that the decoder can reproduce by copying and without decoding a new frame. c) The playback device can optimize the downstream process by indicating the base frame rate, for example, by adjusting frame rate conversion or noise reduction algorithms.
[0028] This embodiment enables the end - user to view the content rendered at the frame rate intended by the content creator. This embodiment does not provide backward compatibility with devices that do not support the frame rate of the container, e.g., 120fps.
[0029] Tables 2 and 3 show examples of the syntax of the raw byte sequence payload (RBSP) of the sequence parameter set and the tile group header, where the proposed new syntax elements are shown in bold - faced font. The remaining syntax follows the syntax in the proposed specification of the Versatile Video Coding (VVC) (Reference [2]).
[0030] As an example, in the SPS (see Table 2), a flag can be added to enable variable frame rate. That the sps_vfr_enabled_flag is equal to 1 indicates that the coded video sequence (CVS) can include variable frame rate content. That the sps_vfr_enabled_flag is equal to 0 indicates that the CVS can include fixed frame rate content. In the tile_group header() (see Table 3), That the tile_group_vrf_info_present_flag is equal to 1 indicates that the syntax elements tile_group_true_fr and tile_group_shutterangle are present in the syntax. That the tile_group_vrf_info_present_flag is equal to 0 indicates that the syntax elements tile_group_true_fr and tile_group_shutterangle are not present in the syntax. If the tile_group_vrf_info_present_flag does not exist, it is implied to be 0. tile_group_true_fr indicates the actual frame rate of the video data transmitted in this bitstream. tile_group_shutterangle indicates the shutter angle corresponding to the actual frame rate of the video data transmitted in this bitstream. That the tile_group_skip_flag is equal to 1 indicates that the current tile group is copied from another tile group. That the tile_group_skip_flag is equal to 0 indicates that the current tile group is not copied from another tile group. tile_group_copy_pic_order_cnt_lsb indicates the picture order count modulo MaxPicOrderCntLsb of the previously decoded picture that the current picture copies when the tile_group_skip_flag is set to 1.
Table 2
Table 3-1
Table 3-2
[0031] Second Embodiment - Fixed Frame Rate Container The second embodiment enables a use case where original content having a fixed frame rate and shutter angle can be rendered by a decoder at an alternative frame rate and a simulated variable shutter angle, as illustrated in FIG. 2. For example, if the original content has a frame rate of 120 fps and a shutter angle of 360 degrees (which means the shutter is open for 1 / 120 second), the decoder can render multiple frame rates below 120 fps. For example, as described in Table 1, to decode 24 fps at a simulated shutter angle of 216 degrees, the decoder can combine three decoded frames and display them at 24 fps. Table 4 is an expansion of Table 1 and shows how to combine different numbers of encoded frames to render at the output target frame rate and the desired target shutter angle. Frame combination can be performed by simple pixel averaging, weighted pixel averaging where pixels from one frame are weighted more than pixels from other frames and the sum of all weights is 1, or by other filter interpolation schemes known in the art. In Table 4, the function Ce(a,b) indicates a combination of encoded frames a to b, and this combination can be performed by averaging, weighted averaging, filtering, etc.
Table 4-1
Table 4-2
[0032] When the value of the target shutter angle is less than 360 degrees, the decoder can combine different sets of decoded frames. For example, from Table 1, given an original stream at 120 fps and 360 degrees, to generate a stream at 40 fps and a shutter angle of 240 degrees, the decoder needs to combine two of the three possible frames. Thus, either the first and second frames, or the second and third frames can be combined. The choice of which frames to combine can be explained in terms of the "decoding phase" expressed as follows.
Number
Number
[0033] Generally, decode_phase_idx ranges from [0, n_frames_max - n_frames]. For example, for an original sequence at 120 fps and a shutter angle of 360 degrees, for a target frame rate of 40 fps at a shutter angle of 240 degrees, n_frames_max = 120 / 40 = 3. From Equation (2), since n_frames = 2, decode_phase_idx ranges from [0, 1]. Thus, decode_phase_idx = 0 indicates selecting the frames at indices 0 and 1, and decode_phase_idx = 1 indicates selecting the frames at indices 1 and 2.
[0034] In this embodiment, the intended rendered variable frame rate by the content creator may be signaled as metadata such as a supplemental enhancement information (SEI) message or video usability information (VUI). Optionally, the rendered frame rate may be controlled by the receiver or the user. An example of frame rate conversion SEI messaging that specifies the preferred frame rate and shutter angle of the content creator is shown in Table 5. The SEI message can also indicate whether the combined frame is to be performed in the coded signal domain (e.g., gamma, PQ, etc.) or in the linear light domain. Note that for post-processing, a frame buffer is required in addition to the decoder picture buffer (DPB). The SEI message can indicate the number of additional frame buffers required or some alternative ways to combine frames. For example, to reduce complexity, frames may be recombined at a reduced spatial resolution.
[0035] As shown in Table 4, for certain combinations of frame rate and shutter angle (e.g., 30 fps and 360 degrees, or 24 fps and 288 or 360 degrees), the decoder needs to combine three or more decoded frames, which increases the amount of buffer space required by the decoder. To reduce the burden of the extra buffer space in the decoder, in some embodiments, certain combinations of frame rate and shutter angle may be limited to a set of acceptable decoding parameters (e.g., by setting an appropriate coding profile and level).
[0036] Again, as an example, considering the case of playback at 24fps, the decoder can determine to display the same frame to be displayed at an output frame rate of 120fps five times. This is exactly the same as showing the frame once at an output frame rate of 24fps. The advantage of maintaining a constant output frame rate is that the display can operate at a constant clock speed, which simplifies all hardware significantly. If the display can dynamically change the clock speed, it would make sense to display the frame only once (1 / 24 seconds) instead of repeating the same frame five times (each 1 / 120 seconds). The former approach may result in slightly higher image quality, slightly better optical efficiency, or slightly better power efficiency. Similar considerations apply to other frame rates.
[0037] Table 5 shows an example of the frame rate conversion SEI message syntax according to one embodiment. [Table 5]
[0038] framerate_conversion_cancel_flag = 1 indicates that the SEI message cancels the persistence of a previous frame rate conversion SEI message in output order. framerate_conversion_cancel_flag = 0 indicates that frame rate conversion information follows. base_frame_rate specifies the desired frame rate. base_shutter_angle specifies the desired shutter angle. decode_phase_idx_present_flag = 1 specifies that decoding phase information exists. decode_phase_idx_present_flag = 0 specifies that decoding phase information does not exist. decode_phase_idx indicates the offset index within a sequence of sequential frames having index values 0..(n_frames_max-1), where n_frames_max = 120 / base_frame_rate. The value of decode_phase_idx must be in the range 0..(n_frames_max - n_frames), where n_frames = base_shutter_angle / (3*base_frame_rate). If decode_phase_idx does not exist, it is assumed to be 0. conversion_domain_idc = 0 specifies that frame combination is performed in the linear domain. conversion_domain_idc = 1 specifies that frame combination is performed in the non-linear domain. num_frame_buffers specifies the number of additional frame buffers (DPB is not counted). framerate_conversion_persistence_flag specifies the persistence of the frame rate conversion SEI message for the current layer. When framerate_conversion_persistence_flag = 0, it specifies that the frame rate conversion SEI message is applied only to the current decoded picture. Let picA be the current picture. framerate_conversion_persistence_flag = 1 specifies that the frame rate conversion SEI message for the current layer persists in output order until one or more of the following conditions become true. - A new coded layer-wise video sequence (CLVS) for the current layer starts. - The bitstream ends. - A picture picB within the current layer in an access unit containing a frame rate conversion SEI message applicable to the current layer is output (PicOrderCnt(picB) is greater than PicOrderCnt(picA)), where PicOrderCnt(picB) and PicOrderCnt(picA) are the PicOrderCntVal values of picB and picA respectively, and it is immediately after the call of the decoding process for the picture order count of picB.
[0039] Third Embodiment - Input Encoded at Multiple Shutter Angles The third embodiment is an encoding method that enables extraction of the sub-frame rate from the bitstream and thus supports backward compatibility. In HEVC, this is achieved by temporal scalability. The scalability of the temporal layer becomes effective by assigning different values to the temporal_id syntax element of the decoded frame. Thereby, the bitstream can be simply extracted based on the temporal_id value. However, the HEVC-style approach to temporal scalability does not enable rendering the output frame rate at different shutter angles. For example, a base frame rate of 60fps extracted from an original of 120fps always has a shutter angle of 180 degrees.
[0040] In ATSC 3.0, other methods are described where 60fps frames with a 360-degree shutter angle are emulated as a weighted average of two 120fps frames. The emulated 60fps frames are assigned a temporal_id value of 0 and are interleaved with the original 120fps frames assigned a temporal_id value of 1. When 60fps is required, the decoder need only decode frames with a temporal_id of 0. When 120fps is required, the decoder subtracts each temporal_id=1 frame (i.e., the 120fps frame) from a scaled version of the corresponding temporal_id=0 frame (i.e., the emulated 60fps frame) to recover the corresponding original 120fps frame that was not explicitly transmitted, thereby being able to reconstruct all of the original 120fps frames.
[0041] In embodiments of the present invention, a new algorithm is described that supports multiple target frame rates and target shutter angles in a backward compatibility (BC) manner. It is proposed to preprocess the original 120fps content at the base frame rate at several shutter angles. Then, at the decoder, other frame rates at various other shutter angles can be easily derived. The ATSC 3.0 approach can be considered a special case of the proposed scheme, where frames with temporal_id=0 carry frames at 60fps@360 shutter angle and frames with temporal_id=1 carry frames at 60fps@180 shutter angle.
[0042] As a first example, consider an input sequence at 120 fps and a 360 shutter angle, used to encode a sequence having a base layer frame rate of 40 fps and shutter angles of 120, 240, and 360 degrees, as depicted in FIG. 4. In this scheme, the encoder calculates new frames by combining up to three of the original input frames. For example, encoded frame 2 (En-2), representing an input at 40 fps and 240 degrees, is generated by combining input frames 1 and 2, and encoded frame 3 (En-3), representing an input at 40 fps and 360 degrees, is generated by combining frame En-2 with input frame 3. At the decoder, to reconstruct the input sequence, decoded frame 2 (Dec-2) is generated by subtracting frame En-1 from frame En-2, and decoded frame 3 (Dec-3) is generated by subtracting frame En-2 from frame En-3. The three decoded frames represent the output at a base frame rate of 120 fps and a shutter angle of 360 degrees. Additional frame rates and shutter angles can be extrapolated using the decoded frames, as shown in Table 6. In Table 6, the function Cs(a,b) indicates the combination of input frames a to b, which can be performed by means such as averaging, weighted averaging, filtering, etc.
Table 6-1
Table 6-2
[0043] The advantages of this approach are, as shown in Table 6, that all 40fps versions can be decoded without further processing. Another advantage is that other frame rates can be derived at various shutter angles. For example, consider a decoder that decodes at 30fps and a shutter angle of 360. From Table 4, the output corresponds to a sequence of frames generated by Ce(1,4)=Cs(1,4), Cs(5,8), Cs(9,12), etc., which also match the decoding sequences shown in Table 6, where Cs(5,8)=e6 - e4 + e8. In one embodiment, a look-up table (LUT) can be used to define how the decoded frames need to be combined to generate the output sequence at a specified output frame rate and an emulated shutter angle.
[0044] In another example, as shown below, it is proposed to combine up to five frames within the encoder to simplify the extraction of the 24fps base layer at shutter angles of 72, 144, 216, 288, and 360 degrees. This is desirable for movie content that is best presented at 24fps on legacy TVs.
Table 7-1
Table 7-2
[0045] As shown in Table 7, when the decoded frame rate matches the baseline frame rate (24fps), in each group of five frames (e.g., e1~e5), the decoder can simply select one frame at the desired shutter angle (e.g., e2 for a shutter angle of 144 degrees). To decode at different frame rates and specific shutter angles, the decoder needs to determine how to appropriately combine the decoded frames (e.g., by addition or subtraction). For example, to decode at 30fps and a shutter angle of 180 degrees, the following steps can be followed. a) The decoder can consider a virtual encoder that transmits at 120 fps and 360 degrees without any backward compatibility considerations. Then, from Table 1, the decoder needs to combine 2 out of 4 frames to generate the output sequence at the desired frame rate and shutter angle. For example, as shown in Table 4, the sequence includes Ce(1,2)=Avg(s1,s2), Ce(5,6)=Avg(s5,s6), etc., where Avg(s1,s2) can represent the average of frames s1 and s2. b) Based on the premise that by definition, the encoded frames can be expressed as e1=s1, e2=Avg(s1,s2), e3=Avg(s1,s3), etc., it can be easily derived that the frame sequence in step a) can also be expressed as follows. -Ce(1,2)=Avg(s1,s2)=e2 -Ce(5,6)=Avg(s5,s6)=Avg(s1,s5)-Avg(s1,s4)+s6=e5-e4+e6 -etc. As before, the appropriate combination of decoded frames can be pre-calculated and made available as a LUT.
[0046] The advantage of the proposed method is to provide options for both content creators and users, that is, to enable directory / editing choices and user selections. For example, the pre-processing content within the encoder can enable creating a base frame rate at various shutter angles. A temporal_id value in the range of [0, (n_frames - 1)] can be assigned to each shutter angle, where n_frames has a value equal to 120 divided by the base frame rate. (For example, when the base frame rate is 24 fps, the temporal_id is in the range of [0, 4].) It can be selected to optimize compression efficiency or for aesthetic reasons. In some use cases, for example, in the case of high - level streaming, multiple bitstreams with different base layers can be encoded and saved for the user to provide and select.
[0047] In a second example of the disclosed method, multiple backward compatibility frame rates may be supported. Ideally, it is desired that it be possible to decode at 24 frames per second to obtain a base layer at 24 fps, at 30 frames per second to obtain a 30 fps sequence, at 60 frames per second to obtain a 60 fps sequence, and so on. If the target shutter angle is not specified, a default target shutter angle that is as close as possible to 180 degrees among the shutter angles allowed for the source frame rate and the target frame rate is recommended. For example, in the values shown in Table 7, the preferred target shutter angles for fps at 120, 60, 40, 30, and 24 are 360 degrees, 180 degrees, 120 degrees, 180 degrees, and 216 degrees, respectively.
[0048] From the above examples, it is recognized that the choice of how to encode the content can affect the complexity of decoding a particular base layer frame rate. One embodiment of the present invention is to adaptively select an encoding method based on the desired base layer frame rate. In the case of movie content, this can be, for example, 24 fps, while in the case of sports, it can be 60 fps.
[0049] Exemplary syntax for the BC embodiment of the present invention is shown in Tables 8 and 9 below. In the SPS (Table 8), two syntax elements, SPS_hfr_BC_enabled_flag and SPS_base_framerate, are added (when SPS_hfr_BC_enabled_flag is set to 1). sps_hfr_BC_enabled_flag = 1 specifies that backward compatible high frame rates are enabled in the coded video sequence (CVS). sps_hfr_BC_enabled_flag = 0 specifies that backward compatible high frame rates are not enabled in the CVS. sps_base_framerate specifies the base frame rate of the current CVS. In the tile group header, when sps_hfr_BC_enabled_flag is set to 1, the syntax number_avg_frames is transmitted in the bitstream. number_avg_frames specifies the number of frames at the highest frame rate (e.g., 120 fps) that are combined to generate the current picture at the base frame rate.
Table 8
Table 9
[0050] Variant of the second embodiment (fixed frame rate) In the HEVC (H.265) coding standard (see Reference [1]) and the developing Versatile Video Coding Standard (commonly referred to as VVC; see Reference [2]), a syntax element pic_struct is defined that indicates whether a picture is to be displayed as a frame or as one or more fields, and whether the decoded picture is to be repeated. A copy of Table D.2, "Interpretation of pic_struct," from HEVC is provided in the appendix for reference.
[0051] As recognized by the inventors, it is important to note that existing pic_struct syntax elements can only support a specific subset of content frame rates when using a fixed frame rate coding container. For example, when using a fixed frame rate container of 60fps, the existing pic_struct syntax can support 30fps by using frame doubling when fixed_pic_rate_within_cvs_flag = 1, and 24fps by using frame doubling and frame tripling, alternating every other frame. However, when using a fixed frame rate container of 120fps, the current pic_struct syntax cannot support frame rates of 24fps or 30fps. To mitigate this problem, two new methods have been proposed. One is an extension of the HEVC version, and the other is not.
[0052] Method 1: Non-backward compatible pic_struct VVC is still under development, and thus the syntax can be designed with maximum freedom. In one embodiment, it is proposed to remove the options for frame doubling and frame tripling in pic_struct, and add a new syntax element num_frame_repetition_minus2 that uses a specific value of pic_struct to indicate any frame repetition and specifies the number of frames to repeat. An example of the proposed syntax is described in the following table, where Table 10 shows the changes to Table D.2.3 in HEVC, and Table 11 shows the changes to Table D.2 shown in the appendix.
Table 10
[0053] Method 2: Extended version of pic_struct for HEVC Since AVC and HEVC decoders are already adopted, it may be desirable to simply extend the existing pic_struct syntax without removing the old options. In an embodiment, a new pic_struct = 13, "frame repetition extension" value, and a new syntax element num_frame_repetition_minus4 are added. An example of the proposed syntax is described in Tables 12 and 13. For pic_struct values 0 to 12, the proposed syntax is the same as that in Table D.2 (as shown in the appendix), so these values are omitted for simplicity. [Table 12] num_frame_repetition_minus4 + 4 indicates that when fixed_pic_rate_within_cvs_flag = 1, the frame needs to be continuously displayed on the display num_frame_repetition_minus4 + 4 times at a frame update interval equal to DpbOutputElementalInterval[n] given by Equation E-73. [Table 13]
[0054] In HEVC, the parameter Frame_field_info_present_flag is present in the Video User Information (VUI), while the syntax elements pic_struct, source_scan_type, and duplicate_flag are in the pic_timing() SEI message. In an embodiment, it is proposed to move all related syntax elements to the VUI together with the frame_field_info_present_flag. An example of the proposed syntax is shown in Table 14. [Table 14]
[0055] Alternative Signaling of Shutter Angle Information When dealing with variable frame rates, it is desirable to identify both the desired frame rate and the desired shutter angle. In conventional video coding standards, "Video User Information" (VUI) provides essential information for the proper display of video content such as aspect ratio, color primaries, chroma subsampling, etc. The VUI may provide frame rate information when the fixed pixel rate is set to 1, but there is no support for shutter angle information. Embodiments enable the use of different shutter angles for different temporal layers, and the decoder can use the shutter angle information to improve the final appearance on the display.
[0056] For example, HEVC supports a temporal sublayer that essentially uses a frame dropping technique to transition from a higher frame rate to a lower frame rate. The main problem with this is that the effective shutter angle decreases with each frame drop. For example, 60fps can be derived from a 120fps video by dropping every other frame. 30fps can be derived by dropping 3 out of 4 frames, and 24fps can be derived by dropping 4 out of 5 frames. Assuming a full 360-degree shutter for 120Hz, with simple frame dropping, the shutter angles for 60fps, 30fps, and 24fps would be 180, 90, and 72 degrees respectively [3]. Experience has shown that shutter angles less than 180 degrees are generally unacceptable, especially at frame rates below 50Hz. By providing shutter angle information, smart technology can be applied to improve the final appearance, for example, if a display desires to generate a cinematic effect from a 120Hz video with reduced shutter angles for each temporal layer.
[0057] In another example, there may be a desire to support different temporal layers (e.g., a 60fps sub-bitstream within a 120fps bitstream) with the same shutter angle. And the main problem is that when a 120fps video is displayed at 120Hz, even / odd frames have different effective shutter angles. If the display has the relevant information, smart technology can be applied to improve the final appearance. An example of the proposed syntax is shown in Table 15, where the E.2.1 VUI parameter syntax table in HEVC (reference [1]) is modified to support shutter angle information as described above. Note that in another embodiment, instead of expressing the shutter angle syntax in absolute degrees, it can also be expressed as the ratio of the frame rate to the shutter speed (see Equation (1)).
Table 15
[0058] Stepwise Frame Rate Update within the Coded Video Sequence (CVS) Experiments show that for HDR content displayed on an HDR display, in order to perceive the same motion judder as in standard dynamic range (SDR) playback at 100 nits (nits) display, it is necessary to increase the frame rate based on the brightness of the content. In most standards (such as AVC, HEVC, VVC, etc.), the video frame rate can be indicated in the VUI (included in the SPS) using the vui_time_scale, vui_num_units_in_tick, elemental_duration_in_tc_minus1[temporal_id_max] syntax elements. For example, as shown in Table 16 below (see Section E.2.1 of Reference [1]).
Table 16
[0059] However, the frame rate can only be changed at specific times, for example, in HEVC, only in an Intra Random Access Point (IRAP) frame or only at the start of a new CVS. In HDR playback, in the case of fade-in or fade-out, since the brightness of the picture changes frame by frame, it may be necessary to change the frame rate or the duration of the picture for each picture. In order to make the frame rate or picture duration refreshable at any time (even frame by frame), in one embodiment, as shown in Table 17, a new SEI message for "progressive refresh rate" is proposed.
Table 17
[0060] The definition of the new syntax num_units_in_tick is the same as vui_num_units_in_tick, and the definition of time_scale is the same as vui_time_scale. The number of units in a tick, num_units_in_tick, is the number of time units of a clock that operates at a frequency of time_scale Hz corresponding to one increment of the clock tick counter (referred to as a clock tick). It is assumed that num_units_in_tick is greater than 0. The clock tick in seconds is equal to the quotient of dividing num_units_in_tick by time_scale. For example, if the picture rate of a video signal is 25 Hz, time_scale is equal to 27,000,000, num_units_in_tick is equal to 1,080,000, and thus the clock tick may be equal to 0.04 seconds. time_scale is the number of time units that elapse in one second. For example, a time coordinate system that measures time using a 27 MHz clock has a time_scale of 27,000,000. It is assumed that the value of time_scale is greater than 0. The picture duration of a picture that uses the gradual_refresh_rate SEI message is defined as follows. picture_duration = num_units_in_tick ÷ time_scale.
[0061] Signaling of shutter angle information by SEI messages As described above, Table 15 provides an example of the VUI parameter syntax having shutter angle support. As an example, but not by way of limitation, Table 18 enumerates the same syntax elements, but here they are listed as part of an SEI message for shutter angle information. It should be noted that SEI messaging is used only as an example, and similar messaging may be constructed in other layers of the high-level syntax, such as the sequence parameter set (SPS), picture parameter set (PPS), slice, or tile group header.
Table 18
[0062] The shutter angle is usually represented in the range of 0 to 360 degrees. For example, a shutter angle of 180 degrees indicates that the exposure duration is 1 / 2 of the frame duration. The shutter angle can be expressed as shutter_angle = frame_rate * 360 * shutter_speed, where shutter_speed is the exposure duration and frame_rate is the reciprocal of the frame duration. The frame_rate for a given temporal sublayer Tid can be indicated by num_units_in_tick, time_scale, and elemental_duration_in_tc_minus1[Tid]. For example, when fixed_pic_rate_within_cvs_flag[Tid] = 1, frame_rate = time_scale / (num_units_in_tick * (elemental_duration_in_tc_minus1[Tid] + 1)).
[0063] In some embodiments, the value of the shutter angle (e.g., fixed_shutter_angle) may not be an integer and may be, for example, 135.75 degrees. To allow for higher precision, in Table 21, u(9) (unsigned 9-bit) can be replaced with u(16) or some other suitable bit depth (e.g., 12-bit, 14-bit, or more than 16-bit).
[0064] In some embodiments, it may be beneficial to represent the shutter angle information in terms of "clock ticks". In VVC, the variable ClockTick is derived as follows.
Number
Number
Number
Number
[0065] Table 19 shows an example of SEI messaging as represented by Equation (11). In this example, the shutter angle must be greater than 0 for a real-world camera.
Table 19
[0066] In another embodiment, the frame duration (e.g., frame_duration) may be specified by some other means. For example, in DVB / ATSC, when fixed_pic_rate_within_cvs_flag[Tid] = 1, frame_rate = time_scale / (num_units_in_tick * (elemental_duration_in_tc_minus1[Tid] + 1)), frame_duration = 1 / frame_rate.
[0067] Some of the syntax in Table 19 and subsequent tables assumes that the shutter angle is always greater than zero, but a shutter angle = 0 can be used to signal a creative intent that the content should be displayed without any motion blur. This is the case for moving graphics, animations, CGI textures, matte screens, etc. Thus, for example, signaling a shutter angle = 0 can be useful for mode determination in a transcoder (e.g., to select a transcoding mode that preserves edges) as well as in a display that receives shutter angle metadata via a CTA interface or a 3GPP® interface. For example, the shutter angle = 0 can be used to instruct a display that should not perform motion processing such as noise reduction, frame interpolation, etc. In such embodiments, the syntax elements fixed_shutter_angle_nume_minus1 and sub_layer_shutter_angle_numer_minus1[i] can be replaced by the syntax elements fixed_shutter_angle_numer and sub_layer_shutter_angle_nume[i], where fixed_shutter_angle_numer specifies the numerator used to derive the shutter angle value. The value of fixed_shutter_angle_numer must be in the range from 0 to 65535 (including both ends of the interval). sub_layer_shutter_angle_numer[i] specifies the numerator used to derive the shutter angle value when HighestTid is equal to i. The value of sub_layer_shutter_angle_numer[i] must be in the range from 0 to 65535 (including both ends of the interval).
[0068] In another embodiment, fixed_shutter_angle_denom_minus1 and sub_layer_shutter_angle_denom_minus1[i] are also replaced by the syntax elements fixed_shutter_angle_denom and sub_layer_shutter_angle_denom[i].
[0069] In one embodiment, as shown in Table 20, by setting general_hrd_parameters_present_flag = 1 in VVC, the num_units_In_tick and time_scale syntax defined in the SPS can be reused. Under this scenario, the SEI message can be renamed as the exposure duration SEI message.
Table 20
[0070] In another embodiment, as shown in Table 21, clockTick can be explicitly defined by the syntax elements expo_num_units_in_tick and expo_time_scale. The advantage here is, similar to the previous embodiment, that in VVC, it does not depend on whether the general_hrd_parameters_present_flag set is set to 1.
Number
Table 21
[0071] As described above, the syntax parameters, sub_layer_exposure_duration_numer_minus1[i] and sub_layer_exposure_duration_denom_minus1[i], are also replaced by sub_layer_exposure_duration_numer[i] and sub_layer_exposure_duration_denom[i].
[0072] In another embodiment, as shown in Table 22, the parameters ShutterInterval (i.e., exposure duration) can be defined by the syntax elements sii_num_units_in_shutter_interval and sii_time_scale, where
Number
Table 22
[0073] In an alternative embodiment, instead of using a numerator and a denominator to signal the sublayer shutter interval, a single value is used. An example of such syntax is shown in Table 23.
Table 23
[0074] Table 24 provides an overview of the six approaches discussed in Tables 18 - 23 for providing SEI messaging related to shutter angle or exposure duration.
Table 24
[0075] Variable Frame Rate Signaling As described in U.S. Provisional Application No. 62 / 883,195, filed on August 06, 2019, in many applications, it is desirable for a decoder to support playback at variable frame rates. Frame rate adaptation is typically part of the operation in a Hypothetical Reference Decoder (HRD), as described, for example, in Annex C of Reference [2]. In one embodiment, it is proposed to signal, via SEI messaging or other means, a syntax element that defines the Picture Presentation Time (PPT) as a function of a 90 kHz clock. This is a kind of repetition of the nominal decoder picture buffer (DPB) output time specified in the HRD, but currently uses the 90 kHz ClockTicks accuracy specified in the MPEG-2 system. The advantages of this SEI message are: a) when the HRD is not enabled, the PPT SEI message can still be used to indicate the timing of each frame; b) it can facilitate the conversion of bitstream timing and system timing.
[0076] Table 25 illustrates an example of the syntax of the proposed PPT timing message that matches the syntax of the Presentation Time Stamp (PTS) variable used in MPEG-2 Transport (H.222) (Reference [4]). [Table 25] PPT (Picture Presentation Time) The presentation time is related to the decoding time as follows. The PPT is a 33-bit number encoded in three separate fields. This represents the presentation time tp(k) of the presentation unit k of the elementary stream n in the system target decoder. The value of the PPT is specified in units obtained by dividing the system clock frequency by 300 (90 kHz is generated). The picture presentation time is derived from the PPT according to the following formula. n (k). The picture presentation time is derived from the PPT according to the following formula. PPT(k) = ((system_clock_frequency x tp n (k)) / 300) % 2 33 。 Here, tp n (k) is the presentation time of the presentation unit P n (k).
[0077] References Each of the references listed in this specification is hereby incorporated by reference in its entirety. [1] High Efficiency Video Coding (HEVC), H.265, Series H, Video Coding, ITU, (02 / 2018) [2] B. Bross, J. Chen, and S. Liu, "Versatile Video Coding (VVC) (Draft 5)", JVET Output Document, JVET-N1001, v5, (uploaded on May 14, 2019) [3] C. Carbonara, J. DeFilippis, M. Korpi, "High Frame Rate Capture and Generation", SMPTE, 2015 Technical Conference and Exhibition, October 26 - 29, 2015 [4] Infrastructure for Audiovisual Services - Transmission Multiplexing and Synchronization, H.222.0, Series H, General Coding of Moving Pictures and Associated Audio Information: Systems, ITU, 08 / 2018
[0078] Computer System Implementation Example Embodiments of the present invention may be implemented in a computer system, a system composed of electronic circuits and components, an integrated circuit (IC) device such as a microcontroller, a field programmable gate array (FPGA), or another configurable or programmable logic device (PLD), a discrete-time or digital signal processor (DSP), an application specific integrated circuit (ASIC), and / or an apparatus including one or more of such systems, devices, or components. The computer and / or IC can execute, control, or perform instructions regarding frame rate scalability as described herein. The computer and / or IC may calculate any of the various parameters or values related to the frame rate scalability described herein. Embodiments of images and videos can be implemented in hardware, software, firmware, and various combinations thereof.
[0079] Certain embodiments of the present invention include a computer processor that executes software instructions that cause the processor to perform the method of the present invention. For example, one or more processors, encoders, set-top boxes, transcoders, etc. within a display can implement the method related to the above-described frame rate scalability by executing software instructions in a program memory accessible to the processor. Embodiments of the present invention may also be provided in the form of a program product. The program product can include any non-transitory and tangible medium that carries a set of computer-readable signals that, when executed by a data processor, cause the data processor to perform the method of the present invention. The program product according to the present invention may be in any of a wide variety of non-transitory and tangible forms. The program product can include, for example, physical media such as a floppy disk, a magnetic data storage medium including a hard disk drive, an optical data storage medium including a CD ROM, a DVD, an electronic data storage medium including a ROM, a flash RAM, etc. The computer-readable signals on the program product can optionally be compressed or encrypted. When an element (e.g., a software module, a processor, an assembly, a device, a circuit, etc.) is referred to above, unless otherwise indicated, the reference to that element (including the reference to "means") shall be construed to include equivalents of the recited element (e.g., functionally equivalent) that perform the function of the recited element, including elements that are not structurally equivalent to the disclosed structure that performs the function in the exemplary embodiments of the present invention.
[0080] Equivalents, extensions, alternatives, and others Accordingly, exemplary embodiments related to frame rate scalability are described. In the foregoing specification, embodiments of the present invention have been described with reference to numerous specific details that may vary from implementation to implementation. Accordingly, the only exclusive indicator of what the present invention is and what is intended by the applicant to be the present invention is the set of claims issued from this application in the specific form in which such claims are issued, including subsequent amendments. The definitions expressly set forth herein for the terms contained in the claims shall take precedence over the meanings of the terms used in the claims. Accordingly, no limitation, element, characteristic, feature, advantage, or attribute not expressly recited in the claims should limit the scope of such claims in any way. Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense.
[0081] Appendix This appendix provides a copy of Table D.2 and related pic_struct related information from the H.265 specification (Reference [1]).
Table D.2-1
Table D.2-2
Claims
1. 1. A non-transitory processor-readable medium having stored thereon instructions for generating an encoded video stream using a processor, the processor comprising: Receiving one or more video pictures and associated shutter interval information; generating an encoded video bitstream comprising a picture data section including an encoding of the one or more video pictures and metadata including a shutter interval parameter for the one or more video pictures, the shutter interval parameter being: a shutter interval time scale parameter indicating the number of time units that elapse in one second; a shutter interval clock tick parameter indicating the number of time units of a clock operating at the frequency of said shutter interval time scale parameter; a shutter interval duration flag indicating whether exposure duration information for all temporal sub-layers of the picture data section is fixed; if the shutter interval duration flag indicates that the exposure duration information is fixed, then the decoded version of the one or more video pictures for all the temporal sub-layers of the picture data section is decoded by calculating an exposure duration value based on the shutter interval time scale parameter and the shutter interval clock tick parameter; Otherwise, the metadata includes one or more arrays of sub-layer parameters, values in the one or more arrays of sub-layer parameters in combination with the shutter interval time scale parameter are used to calculate a corresponding sub-layer exposure duration value for each sub-layer to display a decoded version of the temporal sub-layers of the one or more video pictures.
2. 2. The non-transitory processor-readable medium of claim 1 , wherein the one or more arrays of sub-layer parameters include an array of sub-layer shutter interval clock tick values, each of which indicates a number of time units of a clock operating at a frequency of the shutter interval time scale parameter.
3. The non-transitory processor-readable medium of claim 1 , wherein the exposure duration value is calculated as a quotient of the shutter interval clock tick parameter divided by the shutter interval time scale parameter.
Citation Information
Patent Citations
Frame Rate Scalable Video Coding
JP7411727B2
Transmission apparatus, transmission method, reception apparatus and reception method
WO2015076277A1
Image processing device, image processing method, reception device and transmission device
WO2016185947A1