Frame Rate Scalable Video Coding
Frame rate scalability in video encoding through adjustable shutter angles and metadata signaling addresses the challenge of delivering HDR content at higher resolutions and frame rates, ensuring compatibility with existing devices.
Patent Information
- Application Number
- JP2024188070
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-03-25
- Filing Date
- 2024-10-25
- Publication Date
- 2025-07-03
- Estimated Expiration
- 2041-06-15
AI Technical Summary
Current video delivery systems are limited in their ability to support high dynamic range (HDR) content at resolutions higher than 4K and frame rates beyond 60 frames per second, lacking compatibility with existing playback devices, necessitating separate content versions for different formats.
Implement frame rate scalability in video encoding by combining original frames at adjustable shutter angles to render content at various frame rates and shutter angles, using metadata to signal intended frame rates and shutter angles, and employing scalable coding techniques to ensure compatibility with diverse playback devices.
Enables the delivery of HDR content at higher resolutions and frame rates while maintaining compatibility with existing devices, allowing content creators to distribute a single version compatible with multiple playback formats.
Smart Images

Figure 0007702558000047 
Figure 0007702558000048 
Figure 0007702558000049
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims the benefit of priority from U.S. Application No. 16 / 901,911, filed Jun. 15, 2020, and U.S. Application No. 17 / 212,701, filed Mar. 25, 2021, each of which is incorporated herein by reference in its entirety.
[0002] This document generally relates to images. More particularly, one embodiment of the present invention relates to frame - rate - scalable video coding.
Background Art
[0003] As used herein, the term "dynamic range" (DR) may relate, for example, to the ability of the human visual system (HVS) to perceive the range of intensities (e.g., luminance, luma) within an image from the darkest gray (black) to the brightest white (highlight). In this sense, DR relates to "scene - referred" intensities. DR may also relate to the ability of a display device to properly or approximately render a specific width of intensity range. In this sense, DR relates to "display - referred" intensities. Unless explicitly specified to have a particular meaning at any point in the description herein, the term should be presumed to be used interchangeably in either sense.
Summary of the Invention
[0004] As used herein, the term "high - dynamic range" (HDR) relates to a DR width spanning 14 - 15 digits of the human visual system (HVS). In practice, the DR that a human can simultaneously perceive over a wide range of intensity can be somewhat truncated with respect to HDR.
[0005] In practice, an image contains one or more color components (e.g., luma Y and chroma Cb and Cr), and each color component is represented with a precision of n bits per pixel (e.g., n = 8). Using linear luminance encoding, an image with n ≤ 8 (e.g., a 24-bit color JPEG image) is considered a standard dynamic range (SDR) image, and an image with n > 8 is considered a high dynamic range image. Also, HDR images can be stored and distributed using a high-precision (e.g., 16-bit) floating-point format such as the OpenEXR file format developed by Industrial Light and Magic.
[0006] Currently, the delivery of high dynamic range video content such as Dolby Vision from Dolby Laboratories or HDR10 in Blu-Ray is limited to 4K resolution (e.g., 4096×2160 or 3840×2160, etc.) and 60 frames per second (fps) by the capabilities of many playback devices. In future versions, it is expected that content up to 8K resolution (e.g., 7680×4320) and 120 fps will be available for delivery and playback. To simplify HDR playback content ecosystems such as Dolby Vision, it is desirable that future content types be compatible with existing playback devices. Ideally, content producers should be able to adopt and deliver future HDR technologies without having to derive and distribute special versions of content that are compatible with existing HDR devices (such as HDR10 or Dolby Vision). As understood by the inventors, improved techniques for scalable delivery of video content, particularly HDR content, are desired.
[0007] The approach described in this section is an approach that can be pursued, but it is not necessarily an approach that has been previously considered or pursued. Therefore, unless otherwise stated, none of the approaches described in this section should be assumed to be eligible as prior art merely because they are included in this section. Similarly, unless otherwise stated, the problems identified with respect to one or more approaches should not be assumed to be recognized in any prior art based on this section.
Brief Description of the Drawings
[0008] Embodiments of the present invention are shown by way of example and not limitation in the figures of the accompanying drawings, and like reference numerals refer to like elements.
Figure 1
Figure 2
Figure 3
Figure 4
Modes for Carrying Out the Invention
[0009] Exemplary embodiments related to frame rate scalability for video encoding are described herein. In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present invention. It will be apparent, however, that the present invention may be practiced without these specific details. In other instances, well-known structures and devices are not described in exhaustive detail to avoid obscuring, obfuscating, or rendering the present invention unreadable.
[0010] Summary The exemplary embodiments described herein relate to frame rate scalability in video encoding. In one embodiment, a system having a processor receives an encoded bitstream including encoded video frames, and one or more of the encoded frames are encoded at a first frame rate and a first shutter angle. The processor receives a first flag indicating the presence of a group of encoded frames to be decoded at a second frame rate and a second shutter angle, accesses the encoded bitstream values for the second frame rate and the second shutter angle for the group of encoded frames, and generates decoded frames at the second frame rate and the second shutter angle based on the group of encoded frames, the first frame rate, the first shutter angle, the second frame rate, and the second shutter angle.
[0011] In a second embodiment, a decoder having a processor, receiving an encoded bitstream including a group of encoded video frames, wherein all of the encoded video frames in the encoded bitstream are encoded at a first frame rate, receiving the number of N combined frames, receiving a value for a baseline frame rate, Accessing a group of N consecutive encoded frames, wherein the i-th encoded frame within the group of N consecutive encoded frames represents the average of the input video frames up to the i-th frame encoded at the encoder at the baseline frame rate and the i-th shutter angle, based on the first shutter angle and the first frame rate, where i = 1, 2,... N, accessing the group of N consecutive encoded frames, and Accessing from the encoded bitstream or from user input values for the second frame rate and the second shutter angle to decode a group of N consecutive encoded frames at the second frame rate and the second shutter angle, and Generating decoded frames at the second frame rate and the second shutter angle based on the group of N consecutive encoded frames, the first frame rate, the first shutter angle, the second frame rate, and the second shutter angle.
[0012] In a third embodiment, the encoded video stream structure includes An encoded picture section including encoding of a sequence of video pictures, and A signaling section, A shutter interval time scale parameter indicating the number of time units elapsed per second, and A shutter interval clock tick parameter indicating the number of time units of a clock operating at the frequency of the shutter interval time scale parameter, where the shutter interval divided by the shutter interval time scale parameter represents an exposure duration value, the shutter interval clock tick parameter, and A signaling section including encoding of a shutter interval duration flag indicating whether exposure duration information is fixed for all temporal sublayers within the encoded picture section. When the shutter interval duration flag indicates that the exposure duration information is fixed, for all temporal sublayers within the coded picture section, the coded version of the video picture sequence is decoded by calculating the exposure duration value based on the shutter interval time scale parameter and the shutter interval clock tick parameter; otherwise, The signaling section contains one or more arrays of sublayer parameters, and the values in one or more arrays of sublayer parameters combined with the shutter interval time scale parameter are used to calculate the corresponding sublayer exposure duration value for each sublayer in order to display the decoded version of the temporal sublayer of the video picture sequence.
[0013] Example of a video delivery processing pipeline Figure 1 shows an exemplary process of a conventional video delivery pipeline (100) showing various stages from video capture to video content display. A series of video frames (102) are captured or generated using an image generation block (105). The video frames (102) may be digitally captured (e.g., by a digital camera) or computer-generated (e.g., using computer animation) to provide video data (107). Alternatively, the video frames (102) may be captured on film by a film camera. The film is converted to a digital format to provide video data (107). In the production phase (110), the video data (107) is edited to provide a video production stream (112).
[0014] Next, the video data of the production stream (112) is supplied to the processor in block (115) for post-production editing. The post-production editing block (115) can include adjusting or modifying the color or brightness in specific regions of the image to improve the image quality or achieve a particular look of the image according to the creative intent of the video creator. This is what is referred to as "color timing" or "color grading". Other edits (e.g., scene selection and sequencing, image cropping, addition of computer-generated visual special effects, judder or blur control, frame rate control, etc.) are performed in block (115), and during the final version of the production for distribution, the video image is displayed on the reference display (125). Following post-production (115), the video data of the final production (117) may be distributed to the encoding block (120) for downstream distribution to a decoding and playback device such as a television set, set-top box, movie theater, etc. In some embodiments, the encoding block (120) can include an audio encoder and a video encoder such as those defined by ATSC, DVB, DVD, Blu-Ray, and other distribution formats to generate an encoded bitstream (122). At the receiver, the encoded bitstream (122) is decoded by the decoding unit (130) to generate a decoded signal (132) that represents the same or an approximation of the signal (117). The receiver can be attached to a target display (140) that can have characteristics quite different from those of the reference display (125). In that case, the display management block (135) can be used to map the dynamic range of the decoded signal (132) to the characteristics of the target display (140) by generating a display map signal (137).
[0015] Scalable Encoding Scalable coding is already part of many video coding standards such as MPEG-2, AVC, and HEVC. In embodiments of the present invention, scalable coding is extended to improve performance and flexibility, particularly with respect to very high resolution HDR content.
[0016] As used herein, the term "shutter angle" refers to an adjustable shutter setting that controls the proportion of time that film is exposed to light during each frame interval. For example, in one embodiment,
Number
[0017] This term has its origin in traditional mechanical rotating shutters, but modern digital cameras can also electronically adjust the shutter. Cinematographers can use the shutter angle to control the amount of motion blur and judder recorded in each frame. Note that alternative terms such as "exposure duration", "shutter interval", "shutter speed" may be used instead of "exposure time". Similarly, the term "frame duration" may be used instead of "frame interval". Alternatively, "frame interval" may be replaced with "1 / frame rate". The value of the exposure time is usually less than or equal to the duration of the frame. For example, a shutter angle of 180 degrees indicates that the exposure time is half of the frame duration. In some situations, the exposure time may be longer than the frame duration of the encoded video, for example, when the encoded frame rate is 120 fps and the frame rate of the relevant video content before encoding and display is 60 fps.
[0018] Consider embodiments where, rather than being limited, the original content is captured (or generated) at a 360-degree shutter angle at the original frame rate (e.g., 120 fps). Next, at the receiving device, the video output can be rendered at various frame rates below the original frame rate, for example, by averaging or by other operations known in the art, by a judicious combination of the original frames.
[0019] The combining process may be performed on the non-linearly encoded signal (e.g., using gamma, PQ, or HLG), but the best image quality is obtained by first converting the non-linearly encoded signal to linear light representation, then combining the converted frames, and finally re-encoding the output with the non-linear transfer function, by combining the frames within the linear light region. This process provides a more accurate simulation of the physical camera exposure than combining in the non-linear region.
[0020] Generally, the process of combining frames can be expressed as follows in terms of the original frame rate, the target frame rate, the target shutter angle, and the number of frames to be combined.
Number
Number
[0021] Here, n_frames is the number of combined frames, original_frame_rate is the frame rate of the original content, target_frame_rate is the frame rate to be rendered (where target_frame_rate ≤ original_frame_rate), and target_shutter_angle indicates the amount of desired motion blur. In this example, the maximum value of target_shutter_angle is 360 degrees, corresponding to the maximum motion blur. The minimum value of target_shutter_angle can be expressed as 360*(target_frame_rate / original_frame_rate), corresponding to the minimum motion blur. The maximum value of n_frames can be expressed as (original_frame_rate / target_frame_rate). The values of target_frame_rate and target_shutter_angle should be selected such that the value of n_frame is a non-zero integer.
[0022] In the special case where the original frame rate is 120 fps, Equation (2) becomes [Equation] which can be rewritten as [Equation] and is equivalent to The relationship between the values of target_frame_rate, n_frames, and target_shutter_angle for the case of original_frame_rate = 120 fps is shown in Table 1. In Table 1, "NA" indicates that the corresponding combination of target frame rate and number of combined frames is not allowed. [Table 1]
[0023] Figure 2 shows an exemplary process for combining consecutive original frames to render a target frame rate at a target shutter angle. Given an input sequence (205) at 120 fps and a 360-degree shutter angle, the process combines three of the input frames (e.g., the first three consecutive frames) within a set of five consecutive frames and drops the other two to produce an output video sequence (210) at 24 fps and a 216-degree shutter angle. In some embodiments, output frame 01 of (210) can be produced by combining any of the input frames (205) such as frames 1, 3, and 5, or frames 2, 4, and 5, although it should be noted that combining consecutive frames is thought to result in better quality video output.
[0024] For example, it is desirable to support original content at variable frame rates in order to manage artistic and stylistic effects. Also, the variable input frame rate of the original content is desirably packaged into a "container" with a fixed frame rate to simplify the generation, exchange, and delivery of the content. As an example, three embodiments are presented regarding how to represent variable frame rate video data in a fixed frame rate container. For clarity and without limitation, the following description uses a fixed 120 fps container, although the approach can be easily extended to alternative frame rate containers.
[0025] First Embodiment (Variable Frame Rate) The first embodiment is an explicit description of original content having a variable (not constant) frame rate packaged in a container having a constant frame rate. For example, original content having different frame rates, e.g., 24, 30, 40, 60, or 120 fps for different scenes, may be packaged in a container having a constant frame rate of 120 fps. In this example, each input frame can be replicated 5x, 4x, 3x, 2x, or 1x and packaged into a common 120 fps container.
[0026] Figure 3 shows an example of an input video sequence A having a variable frame rate and a variable shutter angle represented by an encoded bitstream B having a fixed frame rate. And in the decoder, the decoder reconstructs an output video sequence C at a desired frame rate and shutter angle that can vary for each scene. For example, as shown in Figure 3, to construct sequence B, some of the input frames are replicated, some are encoded as is (without replication), and some are copied 4 times. Then, to construct sequence C, any one frame is selected from each replicated frame to generate an output frame with the same original frame rate and shutter angle.
[0027] In this embodiment, metadata is inserted into the bitstream to indicate the original (base) frame rate and shutter angle. The metadata can be signaled using high-level syntax such as a sequence parameter set (SPS), a picture parameter set (PPS), a slice or tile group header. The presence of the metadata enables the encoder and decoder to perform useful functions as follows. a) The encoder can ignore the duplicated frames, thereby improving the encoding speed and simplifying the process. For example, all coding tree units (CTUs) within the duplicated frame can be encoded using the SKIP mode and the reference index 0 in LIST0 of the reference frames. This refers to the decoded frame to which the duplicated frame is copied. b) The decoder can bypass the decoding of the duplicated frames, thereby simplifying the process. For example, the metadata within the bitstream can indicate that the frame is a copy of a previously decoded frame that the decoder can reproduce by copying and without decoding a new frame. c) The playback device can optimize the downstream process by indicating the base frame rate, for example, by adjusting frame rate conversion or noise reduction algorithms.
[0028] This embodiment enables the end user to view the content rendered at the frame rate intended by the content creator. This embodiment does not provide backward compatibility with devices that do not support the frame rate of the container, for example, 120 fps.
[0029] Tables 2 and 3 show examples of the syntax of the raw byte sequence payload (RBSP) of the sequence parameter set and the tile group header, where the proposed new syntax elements are shown in italic font. The remaining syntax follows the syntax in the proposed specification of the Versatile Video Coding (VVC) (Reference [2]).
[0030] As an example, in the SPS (see Table 2), a flag can be added to enable variable frame rate. That the sps_vfr_enabled_flag is equal to 1 indicates that variable frame rate content can be included in the coded video sequence (CVS). That the sps_vfr_enabled_flag is equal to 0 indicates that fixed frame rate content can be included in the CVS. In the tile_group header() (see Table 3), That the tile_group_vrf_info_present_flag is equal to 1 indicates that the syntax elements tile_group_true_fr and tile_group_shutterangle are present in the syntax. That the tile_group_vrf_info_present_flag is equal to 0 indicates that the syntax elements tile_group_true_fr and tile_group_shutterangle are not present in the syntax. If the tile_group_vrf_info_present_flag does not exist, it is implied to be 0. tile_group_true_fr indicates the actual frame rate of the video data transmitted in this bitstream. tile_group_shutterangle indicates the shutter angle corresponding to the actual frame rate of the video data transmitted in this bitstream. That the tile_group_skip_flag is equal to 1 indicates that the current tile group is copied from another tile group. That the tile_group_skip_flag is equal to 0 indicates that the current tile group is not copied from another tile group. tile_group_copy_pic_order_cnt_lsb indicates the picture order count modulo MaxPicOrderCntLsb of the previously decoded picture that the current picture copies when the tile_group_skip_flag is set to 1.
Table 2
Table 3
[0031] Second Embodiment - Fixed Frame Rate Container The second embodiment enables a use case where original content having a fixed frame rate and shutter angle can be rendered by a decoder at an alternative frame rate and a simulated variable shutter angle, as illustrated in FIG. 2. For example, if the original content has a frame rate of 120 fps and a shutter angle of 360 degrees (meaning the shutter is open for 1 / 120 second), the decoder can render multiple frame rates below 120 fps. For example, as described in Table 1, to decode 24 fps at a simulated shutter angle of 216 degrees, the decoder can combine three decoded frames and display them at 24 fps. Table 4 is an expansion of Table 1 and shows how to combine different numbers of encoded frames to render at the output target frame rate and the desired target shutter angle. Frame combination can be performed by simple pixel averaging, weighted pixel averaging where pixels from one frame are weighted more than pixels from other frames and the sum of all weights is 1, or by other filter interpolation schemes known in the art. In Table 4, the function Ce(a,b) indicates a combination of encoded frames a - b, and this combination can be performed by averaging, weighted averaging, filtering, etc.
Table 4 - 1
Table 4 - 2
[0032] When the value of the target shutter angle is less than 360 degrees, the decoder can combine different sets of decoded frames. For example, from Table 1, when an original stream at 120 fps and 360 degrees is given, to generate a stream at 40 fps and a shutter angle of 240 degrees, the decoder needs to combine two of the three possible frames. Therefore, either the first and second frames or the second and third frames can be combined. The choice of which frames to combine can be explained in terms of the "decoding phase" expressed as follows.
Number
Number
[0033] Generally, decode_phase_idx ranges from [0, n_frames_max - n_frames]. For example, for the original sequence at 120 fps and a shutter angle of 360 degrees, for a target frame rate of 40 fps at a shutter angle of 240 degrees, n_frames_max = 120 / 40 = 3. From Equation (2), n_frames = 2, so decode_phase_idx ranges from [0, 1]. Therefore, decode_phase_idx = 0 indicates selecting the frames at indices 0 and 1, and decode_phase_idx = 1 indicates selecting the frames at indices 1 and 2.
[0034] In this embodiment, the intended rendered variable frame rate by the content creator may be signaled as metadata such as a supplemental enhancement information (SEI) message or video usability information (VUI). Optionally, the rendered frame rate may be controlled by the receiver or the user. An example of frame rate conversion SEI messaging that specifies the preferred frame rate and shutter angle of the content creator is shown in Table 5. The SEI message can also indicate whether the combined frame is to be performed in the coded signal domain (e.g., gamma, PQ, etc.) or the linear light domain. Note that for post-processing, a frame buffer is required in addition to the decoder picture buffer (DPB). The SEI message may indicate the number of additional frame buffers required or some alternative ways to combine frames. For example, to reduce complexity, the frames may be recombined at a reduced spatial resolution.
[0035] As shown in Table 4, for certain combinations of frame rate and shutter angle (e.g., 30 fps and 360 degrees, or 24 fps and 288 or 360 degrees), the decoder needs to combine three or more decoded frames, which increases the amount of buffer space required by the decoder. To reduce the burden of the extra buffer space in the decoder, in some embodiments, certain combinations of frame rate and shutter angle may be limited to a set of acceptable decoding parameters (e.g., by setting an appropriate coding profile and level).
[0036] Again, as an example, considering the case of playback at 24fps, the decoder can determine to display the same frame to be shown at an output frame rate of 120fps five times. This is exactly the same as showing the frame once at an output frame rate of 24fps. The advantage of maintaining a constant output frame rate is that the display can operate at a constant clock speed, which simplifies all hardware significantly. If the display can dynamically change the clock speed, it would make sense to display the frame only once (1 / 24 seconds) instead of repeating the same frame five times (each 1 / 120 seconds). The former approach may result in slightly higher image quality, slightly better optical efficiency, or slightly better power efficiency. Similar considerations apply to other frame rates.
[0037] Table 5 shows an example of the frame rate conversion SEI message syntax according to an embodiment. [Table 5] framerate_conversion_cancel_flag = 1 indicates that the SEI message cancels the persistence of a previous frame rate conversion SEI message in output order. framerate_conversion_cancel_flag = 0 indicates that frame rate conversion information follows. base_frame_rate specifies the desired frame rate. base_shutter_angle specifies the desired shutter angle. decode_phase_idx_present_flag = 1 specifies that decoding phase information exists. decode_phase_idx_present_flag = 0 specifies that decoding phase information does not exist. decode_phase_idx indicates the offset index within a series of sequential frames having index values 0..(n_frames_max-1), where n_frames_max = 120 / base_frame_rate. The value of decode_phase_idx must be in the range 0..(n_frames_max - n_frames), where n_frames = base_shutter_angle / (3*base_frame_rate). If decode_phase_idx does not exist, it is assumed to be 0. conversion_domain_idc = 0 specifies that frame combination is performed in the linear domain. conversion_domain_idc = 1 specifies that frame combination is performed in the non-linear domain. num_frame_buffers specifies the number of additional frame buffers (DPB is not counted). framerate_conversion_persistence_flag specifies the persistence of the frame rate conversion SEI message for the current layer. When framerate_conversion_persistence_flag = 0, it specifies that the frame rate conversion SEI message is applied only to the current decoded picture. Let picA be the current picture. framerate_conversion_persistence_flag = 1 specifies that the frame rate conversion SEI message for the current layer persists in output order until one or more of the following conditions become true. - A new coded layer-wise video sequence (CLVS) for the current layer starts. - The bitstream ends. - A picture picB within the current layer in an access unit that includes a frame rate conversion SEI message applicable to the current layer is output (PicOrderCnt(picB) is greater than PicOrderCnt(picA)), where PicOrderCnt(picB) and PicOrderCnt(picA) are the PicOrderCntVal values of picB and picA, respectively, and it is immediately after the call to the decoding process for the picture order count of picB.
[0038] Third Embodiment - Input Encoded at Multiple Shutter Angles The third embodiment is an encoding method that enables extraction of the sub-frame rate from the bitstream and thus supports backward compatibility. In HEVC, this is achieved by temporal scalability. The scalability of the temporal layer becomes effective by assigning different values to the temporal_id syntax element of the decoded frame. Thereby, the bitstream can be simply extracted based on the temporal_id value. However, the HEVC-style approach to temporal scalability does not enable rendering the output frame rate at different shutter angles. For example, a base frame rate of 60fps extracted from an original of 120fps always has a shutter angle of 180 degrees.
[0039] In ATSC 3.0, other methods are described where a 60fps frame with a 360-degree shutter angle is emulated as a weighted average of two 120fps frames. The emulated 60fps frame is assigned a temporal_id value of 0 and is interleaved with the original 120fps frames assigned a temporal_id value of 1. When 60fps is required, the decoder needs to decode only the frames with a temporal_id of 0. When 120fps is required, the decoder subtracts each temporal_id=1 frame (i.e., the 120fps frame) from a scaled version of the corresponding temporal_id=0 frame (i.e., the emulated 60fps frame) to recover the corresponding original 120fps frame that was not explicitly transmitted, thereby being able to reconstruct all of the original 120fps frames.
[0040] In an embodiment of the present invention, a new algorithm is described that supports multiple target frame rates and target shutter angles in a backward compatibility (BC) manner. It is proposed to preprocess the original 120fps content at a base frame rate at several shutter angles. Then, at the decoder, other frame rates at various other shutter angles can be easily derived. The ATSC 3.0 approach can be considered a special case of the proposed scheme, where frames with temporal_id=0 carry frames at 60fps@360 shutter angle and frames with temporal_id=1 carry frames at 60fps@180 shutter angle.
[0041] As a first example, consider an input sequence at 120 fps and a 360 shutter angle, which is used to encode a sequence having a base layer frame rate of 40 fps and shutter angles of 120, 240, and 360 degrees, as depicted in FIG. 4. In this scheme, the encoder calculates new frames by combining up to three of the original input frames. For example, encoded frame 2 (En-2) representing an input at 40 fps and 240 degrees is generated by combining input frames 1 and 2, and encoded frame 3 (En-3) representing an input at 40 fps and 360 degrees is generated by combining frame En-2 with input frame 3. At the decoder, to reconstruct the input sequence, decoded frame 2 (Dec-2) is generated by subtracting frame En-1 from frame En-2, and decoded frame 3 (Dec-3) is generated by subtracting frame En-2 from frame En-3. The three decoded frames represent the output at a base frame rate of 120 fps and a shutter angle of 360 degrees. Additional frame rates and shutter angles can be extrapolated using the decoded frames, as shown in Table 6. In Table 6, the function Cs(a,b) indicates the combination of input frames a - b, which can be performed by means such as averaging, weighted averaging, filtering, etc.
Table 6-1
Table 6-2
[0042] The advantage of this approach is that, as shown in Table 6, all 40fps versions can be decoded without further processing. Another advantage is that other frame rates can be derived at various shutter angles. For example, consider a decoder that decodes at 30fps and a shutter angle of 360. From Table 4, the output corresponds to a sequence of frames generated by Ce(1,4)=Cs(1,4), Cs(5,8), Cs(9,12), etc., which also match the decoding sequence shown in Table 6, where Cs(5,8)=e6 - e4 + e8. In one embodiment, a look-up table (LUT) can be used to define how the decoded frames need to be combined to generate the output sequence at the specified output frame rate and the emulated shutter angle.
[0043] In another example, as shown below, it is proposed to combine up to five frames within the encoder to simplify the extraction of the 24fps base layer at shutter angles of 72, 144, 216, 288, and 360 degrees. This is desirable for movie content that is best presented at 24fps on legacy TVs.
Table 7-1
Table 7-2
[0044] As shown in Table 7, when the decoded frame rate matches the baseline frame rate (24fps), in each group of five frames (e.g., e1~e5), the decoder can simply select one frame at the desired shutter angle (e.g., e2 for a shutter angle of 144 degrees). To decode at different frame rates and specific shutter angles, the decoder needs to determine how to appropriately combine the decoded frames (e.g., by addition or subtraction). For example, to decode at 30fps and a shutter angle of 180 degrees, the following steps can be followed. a) The decoder can consider a virtual encoder that transmits at 120 fps and 360 degrees without any backward compatibility considerations. Then, from Table 1, the decoder needs to combine 2 out of 4 frames to generate the output sequence at the desired frame rate and shutter angle. For example, as shown in Table 4, the sequence includes Ce(1,2)=Avg(s1,s2), Ce(5,6)=Avg(s5,s6), etc., where Avg(s1,s2) can represent the average of frames s1 and s2. b) Based on the premise that by definition, the encoded frames can be expressed as e1=s1, e2=Avg(s1,s2), e3=Avg(s1,s3), etc., it can be easily derived that the frame sequence in step a) can also be expressed as follows. -Ce(1,2)=Avg(s1,s2)=e2 -Ce(5,6)=Avg(s5,s6)=Avg(s1,s5)-Avg(s1,s4)+s6=e5-e4+e6 -etc. As before, the appropriate combination of decoded frames can be pre-calculated and made available as a LUT.
[0045] The advantage of the proposed method is to provide options for both content creators and users, that is, to enable directory / editing selections and user selections. For example, the pre-processing content within the encoder can enable creating the base frame rate at various shutter angles. Each shutter angle can be assigned a temporal_id value in the range of [0,(n_frames-1)], where n_frames has a value equal to 120 divided by the base frame rate. (For example, when the base frame rate is 24 fps, the temporal_id is in the range of [0,4].) It can be selected to optimize the compression efficiency or for aesthetic reasons. In some use cases, for example, in the case of high-bitrate streaming, multiple bitstreams with different base layers can be encoded and saved for the user to provide and select.
[0046] In a second example of the disclosed method, multiple backward compatibility frame rates may be supported. Ideally, it is desired that it be possible to decode at 24 frames per second to obtain a base layer of 24 fps, at 30 frames per second to obtain a 30 fps sequence, at 60 frames per second to obtain a 60 fps sequence, and so on. If the target shutter angle is not specified, a default target shutter angle that is as close as possible to 180 degrees is recommended among the shutter angles allowed for the source frame rate and the target frame rate. For example, in the values shown in Table 7, the preferred target shutter angles for fps at 120, 60, 40, 30, and 24 are 360 degrees, 180 degrees, 120 degrees, 180 degrees, and 216 degrees, respectively.
[0047] From the above examples, it is recognized that the choice of how to encode the content can affect the complexity of decoding a particular base layer frame rate. One embodiment of the present invention is to adaptively select an encoding method based on the desired base layer frame rate. In the case of movie content, this can be, for example, 24 fps, while in the case of sports, it can be 60 fps.
[0048] Exemplary syntax for the BC embodiment of the present invention is shown in Tables 8 and 9 below. In the SPS (Table 8), two syntax elements, SPS_hfr_BC_enabled_flag and SPS_base_framerate, are added (when SPS_hfr_BC_enabled_flag is set to 1). sps_hfr_BC_enabled_flag = 1 specifies that backward-compatible high frame rates are enabled in the coded video sequence (CVS). sps_hfr_BC_enabled_flag = 0 specifies that backward-compatible high frame rates are not enabled in the CVS. sps_base_framerate specifies the base frame rate of the current CVS. In the tile group header, when sps_hfr_BC_enabled_flag is set to 1, the syntax number_avg_frames is transmitted in the bitstream. number_avg_frames specifies the number of frames at the highest frame rate (e.g., 120 fps) that are combined to generate the current picture at the base frame rate.
Table 8
Table 9
[0049] Variation of the second embodiment (fixed frame rate) In the HEVC (H.265) coding standard (see Reference [1]) and the developing Versatile Video Coding Standard (commonly referred to as VVC. See Reference [2]), a syntax element pic_struct is defined to indicate whether a picture is to be displayed as a frame or as one or more fields, and whether to repeat the decoded picture. A copy of Table D.2 "Interpretation of pic_struct" from HEVC is provided in the appendix for reference.
[0050] As recognized by the inventors, it is important to note that existing pic_struct syntax elements can only support a specific subset of content frame rates when using a fixed frame rate coding container. For example, when using a fixed frame rate container of 60 fps, the existing pic_struct syntax can support 30 fps by using frame doubling when fixed_pic_rate_within_cvs_flag = 1, and 24 fps by using frame doubling and frame tripling, alternately for every other frame. However, when using a fixed frame rate container of 120 fps, the current pic_struct syntax cannot support frame rates of 24 fps or 30 fps. To mitigate this problem, two new methods have been proposed. One is an extension of the HEVC version, and the other is not.
[0051] Method 1: Non-backward compatible pic_struct VVC is still under development, and thus the syntax can be designed with maximum freedom. In one embodiment, it is proposed to remove the options for frame doubling and frame tripling in pic_struct, and add a new syntax element num_frame_repetition_minus2 that uses a specific value of pic_struct to indicate any frame repetition and specifies the number of frames to repeat. An example of the proposed syntax is described in the following table, where Table 10 shows the changes to Table D.2.3 in HEVC, and Table 11 shows the changes to Table D.2 shown in the appendix.
Table 10
[0052] Method 2: Extended version of pic_struct for HEVC Since AVC and HEVC decoders are already adopted, it may be desirable to simply extend the existing pic_struct syntax without removing the old options. In an embodiment, a new pic_struct = 13, "frame repetition extension" value, and a new syntax element num_frame_repetition_minus4 are added. An example of the proposed syntax is described in Tables 12 and 13. For pic_struct values 0 to 12, the proposed syntax is the same as that in Table D.2 (as shown in the appendix), so these values are omitted for simplicity. [Table 12] num_frame_repetition_minus4 + 4 indicates that, when fixed_pic_rate_within_cvs_flag = 1, the frame needs to be continuously displayed on the display for num_frame_repetition_minus4 + 4 times at a frame update interval equal to DpbOutputElementalInterval[n] given by Equation E-73. [Table 13]
[0053] In HEVC, the parameter Frame_field_info_present_flag is present in the Video User Information (VUI), while the syntax elements pic_struct, source_scan_type, and duplicate_flag are in the pic_timing() SEI message. In an embodiment, it is proposed to move all related syntax elements to the VUI together with the frame_field_info_present_flag. An example of the proposed syntax is shown in Table 14.
Table 14
[0054] Alternative Signaling of Shutter Angle Information When dealing with variable frame rates, it is desirable to identify both the desired frame rate and the desired shutter angle. In conventional video coding standards, "Video User Information" (VUI) provides essential information for the proper display of video content such as aspect ratio, color primaries, chroma subsampling, etc. The VUI may provide frame rate information when the fixed pixel rate is set to 1, but there is no support for shutter angle information. Embodiments enable the use of different shutter angles for different temporal layers, and the decoder can use the shutter angle information to improve the final appearance on the display.
[0055] For example, HEVC supports a temporal sublayer that essentially uses a frame dropping technique to transition from a higher frame rate to a lower frame rate. The main problem with this is that the effective shutter angle decreases with each frame drop. For example, 60fps can be derived from a 120fps video by dropping every other frame. 30fps can be derived by dropping 3 out of 4 frames, and 24fps can be derived by dropping 4 out of 5 frames. Assuming a full 360-degree shutter for 120Hz, with simple frame dropping, the shutter angles for 60fps, 30fps, and 24fps would be 180, 90, and 72 degrees respectively [3]. Experience has shown that shutter angles less than 180 degrees are generally unacceptable, especially at frame rates below 50Hz. By providing shutter angle information, for example, when it is desirable for a display to generate a cinematic effect from a 120Hz video with reduced shutter angles for each temporal layer, smart technology can be applied to improve the final appearance.
[0056] In another example, there may be a desire to support different temporal layers (e.g., a 60fps sub-bitstream within a 120fps bitstream) at the same shutter angle. And the main problem is that when a 120fps video is displayed at 120Hz, the even / odd frames have different effective shutter angles. If the display has the relevant information, smart technology can be applied to improve the final appearance. An example of the proposed syntax is shown in Table 15, where the E.2.1 VUI parameter syntax table in HEVC (reference [1]) is modified to support shutter angle information as described above. Note that in another embodiment, instead of expressing the shutter angle syntax in absolute degrees, it can also be expressed as the ratio of the frame rate to the shutter speed (see Equation (1)).
Table 15
[0057] Stepwise Frame Rate Update within the Coded Video Sequence (CVS) Experiments show that for HDR content displayed on an HDR display, in order to perceive the same motion judder as in standard dynamic range (SDR) playback at 100 nits display, it is necessary to increase the frame rate based on the brightness of the content. In most standards (such as AVC, HEVC, VVC, etc.), the video frame rate can be indicated in the VUI (included in the SPS) using the vui_time_scale, vui_num_units_in_tick, elemental_duration_in_tc_minus1[temporal_id_max] syntax elements. For example, as shown in Table 16 below (see Section E.2.1 of Reference [1]).
Table 16
[0058] However, the frame rate can be changed only at specific points in time, for example, in the HEVC intra-random access point (IRAP) frame only, or only at the start of a new CVS. In HDR playback, in the case of fade-in or fade-out, since the brightness of the picture changes frame by frame, it may be necessary to change the frame rate or the duration of the picture for each picture. In order to make the frame rate or the picture duration refreshable at any point in time (even frame by frame), in one embodiment, as shown in Table 17, a new SEI message for "progressive refresh rate" is proposed.
Table 17
[0059] The definition of the new syntax num_units_in_tick is the same as vui_num_units_in_tick, and the definition of time_scale is the same as vui_time_scale. The num_units_in_tick is the number of time units of a clock operating at a frequency of time_scale Hz corresponding to one increment of the clock tick counter (referred to as a clock tick). It is assumed that num_units_in_tick is greater than 0. The clock tick in seconds is equal to the quotient of dividing num_units_in_tick by time_scale. For example, when the picture rate of a video signal is 25 Hz, time_scale is equal to 27,000,000, num_units_in_tick is equal to 1,080,000, and thus the clock tick may be equal to 0.04 seconds. The time_scale is the number of time units elapsed in one second. For example, a time coordinate system for measuring time using a 27 MHz clock has a time_scale of 27,000,000. It is assumed that the value of time_scale is greater than 0. The picture duration of a picture using the gradual_refresh_rate SEI message is defined as follows. picture_duration = num_units_in_tick ÷ time_scale.
[0060] Signaling of shutter angle information by SEI messages As described above, Table 15 provides an example of the VUI parameter syntax having shutter angle support. As an example, but not limited to, Table 18 enumerates the same syntax elements, but here it is listed as part of an SEI message for shutter angle information. Note that SEI messaging is used only as an example, and similar messaging may be constructed in other layers of the high-level syntax such as the sequence parameter set (SPS), picture parameter set (PPS), slice, or tile group header.
Table 18
[0061] The shutter angle is usually represented in the range of 0 to 360 degrees. For example, a shutter angle of 180 degrees indicates that the exposure duration is 1 / 2 of the frame duration. The shutter angle can be expressed as shutter_angle = frame_rate * 360 * shutter_speed, where shutter_speed is the exposure duration and frame_rate is the reciprocal of the frame duration. The frame_rate for a given temporal sublayer Tid can be indicated by num_units_in_tick, time_scale, and elemental_duration_in_tc_minus1[Tid]. For example, when fixed_pic_rate_within_cvs_flag[Tid] = 1, frame_rate = time_scale / (num_units_in_tick * (elemental_duration_in_tc_minus1[Tid] + 1)).
[0062] In some embodiments, the value of the shutter angle (e.g., fixed_shutter_angle) may not be an integer and may be, for example, 135.75 degrees. To enable higher precision, in Table 21, u(9) (unsigned 9-bit) can be replaced with u(16) or some other appropriate bit depth (e.g., 12-bit, 14-bit, or more than 16-bit).
[0063] In some embodiments, it may be beneficial to represent the shutter angle information from the perspective of "clock ticks". In VVC, the variable ClockTick is derived as follows.
Number
Number
Number
Number
[0064] Table 19 shows an example of SEI messaging as shown by Equation (11). In this example, the shutter angle must be greater than 0 for a real-world camera.
Table 19
[0065] In another embodiment, the frame duration (e.g., frame_duration) may be specified by some other means. For example, in DVB / ATSC, when fixed_pic_rate_within_cvs_flag[Tid] = 1, frame_rate = time_scale / (num_units_in_tick * (elemental_duration_in_tc_minus1[Tid] + 1)), frame_duration = 1 / frame_rate.
[0066] Some of the syntax in Table 19 and subsequent tables assumes that the shutter angle is always greater than zero, but shutter angle = 0 can be used to signal a creative intent that the content should be displayed without any motion blur. This is the case for moving graphics, animations, CGI textures, matte screens, etc. Thus, for example, signaling shutter angle = 0 can be useful for mode determination in a transcoder (e.g., to select a transcoding mode that preserves edges) and in a display that receives shutter angle metadata via a CTA interface or a 3GPP (registered trademark) interface. For example, using shutter angle = 0, a display that should not perform motion processing such as noise removal, frame interpolation, etc. can be instructed. In such embodiments, the syntax elements fixed_shutter_angle_nume_minus1 and sub_layer_shutter_angle_numer_minus1[i] can be replaced by the syntax elements fixed_shutter_angle_numer and sub_layer_shutter_angle_nume[i], where fixed_shutter_angle_numer specifies the numerator used to derive the shutter angle value. The value of fixed_shutter_angle_numer must be in the range from 0 to 65535 (including both ends of the interval). sub_layer_shutter_angle_numer[i] specifies the numerator used to derive the shutter angle value when HighestTid is equal to i. The value of sub_layer_shutter_angle_numer[i] must be in the range from 0 to 65535 (including both ends of the interval).
[0067] In another embodiment, fixed_shutter_angle_denom_minus1 and sub_layer_shutter_angle_denom_minus1[i] may also be replaced by the syntax elements fixed_shutter_angle_denom and sub_layer_shutter_angle_denom[i].
[0068] In one embodiment, as shown in Table 20, by setting general_hrd_parameters_present_flag = 1 in VVC, the num_units_In_tick and time_scale syntax defined in the SPS can be reused. Under this scenario, the SEI message can be renamed as the exposure duration SEI message.
Table 20
[0069] In another embodiment, as shown in Table 21, clockTick can be explicitly defined by the syntax elements expo_num_units_in_tick and expo_time_scale. The advantage here is, as in the previous embodiment, that it does not depend on whether the general_hrd_parameters_present_flag set is set to 1 in VVC. [Number] [Table 21] expo_num_units_in_tick is the number of time units of a clock operating at a frequency time scale Hz corresponding to one increment of the clock tick counter (referred to as a clock tick). expo_num_units_in_tick must be greater than 0. The clock tick defined by the variable clockTic is, in seconds, equal to the quotient of expo_num_units_in_tick divided by expo_time_scale. expo_time_scale is the number of time units elapsed in one second. clockTick = expo_num_units_in_tick ÷ expo_time_scale. Note: The two syntax elements (expo_num_units_in_tick and expo_time_scale) are defined to measure the exposure time. If num_units_in_tick and time_scale exist, clockTick must be less than or equal to ClockTick. fixed_exposure_duration_within_cvs_flag = 1 specifies that the effective exposure duration value is the same for all temporal sub-layers within the CVS. fixed_exposure_duration_in_CVS_flag = 0 specifies that the effective exposure duration values are not the same for all temporal sub-layers within the CVS. When fixed_exposure_duration_within_cvs_flag = 1, the variable fixedExposureDuration is set to clockTick. sub_layer_exposure_duration_numer_minus1[i] + 1 specifies the numerator used to derive the exposure duration value when HighestTid is equal to i. The value of sub_layer_exposure_duration_numer_minus1[i] must be in the range from 0 to 65535 (including both ends of the interval). sub_layer_exposure_duration_demom_minus1[i] + 1 specifies the denominator used to derive the exposure duration value when HighestTid is equal to i. The value of sub_layer_exposure_duration_demom_minus1[i] must be in the range from 0 to 65535 (including both ends of the interval). The value of sub_layer_exposure_duration_numer_minus1[i] must be less than or equal to the value of sub_layer_exposure_duration_demom_minus1[i]. The variable subLayerExposureDuration[i] for HigestTid equal to i is derived as follows. subLayerExposureDuration[i] = (sub_layer_exposure_duration_numer_minus1[i] + 1) ÷ (sub_layer_exposure_duration_denom_minus1[i] + 1) * clockTick.
[0070] As described above, the syntax parameters, sub_layer_exposure_duration_numer_minus1[i] and sub_layer_exposure_duration_denom_minus1[i], are also replaced by sub_layer_exposure_duration_numer[i] and sub_layer_exposure_duration_denom[i].
[0071] In another embodiment, as shown in Table 22, the parameters ShutterInterval (i.e., exposure duration) can be defined by the syntax elements sii_num_units_in_shutter_interval and sii_time_scale, where [Number] [Table 22] Shutter Interval Information SEI Message Semantics The Shutter Interval Information SEI message indicates the shutter interval for the associated video content before encoding and display, e.g., camera capture content, the amount of time the image sensor was exposed to generate a picture. sii_num_units_in_shutter_interval specifies the number of time units of a clock operating at a frequency of sii_time_scale Hz corresponding to one increment of the shutter clock tick counter. The shutter interval defined by the variable ShutterInterval is, in seconds, equal to the quotient of sii_num_units_in_shutter_interval divided by sii_time_scale. For example, when ShutterInterval is equal to 0.04 seconds, sii_time_scale may be equal to 27 000 000 and sii_num_units_in_shutter_interval may be equal to 1 080 000. The sii_time_scale specifies the number of time units that elapse in one second. For example, in a time coordinate system that measures time using a 27 MHz clock, sii_time_scale will be 27,000,000. If the value of sii_time_scale is greater than 0, the value of ShutterInterval is specified as follows. ShutterInterval = sii_num_units_in_shutter_interval ÷ sii_time_scale. Otherwise (if the value of sii_time_scale is equal to 0), ShutterInterval must be interpreted as unknown or not specified. Note 1: A value of ShutterInterval equal to 0 specifies that the associated video content includes screen capture content, computer-generated content, or other non-camera capture content. Note 2: A value of ShutterInterval greater than the reciprocal of the coded picture rate, the coded picture interval, may indicate that the coded picture rate is greater than the picture rate at which the associated video content was created, for example, when the coded picture rate is 120 Hz and the picture rate of the associated video content before coding and display is 60 Hz. The coding interval for a given temporal sublayer Tid may be indicated by ClockTick and elemental_duration_in_tc_minus1[Tid]. For example, when fixed_pic_rate_within_cvs_flag[Tid] = 1, the picture interval for a given temporal sublayer Tid defined by the variable PictureInterval[Tid] may be specified as follows. PictureInterval[Tid] = ClockTick * (elemental_duration_in_tc_minus1[Tid] + 1). fixed_shutter_interval_into_cvs_flag = 1 specifies that the value of ShutterInterval is the same for all temporal sub - layers within the CVS. fixed_shutter_interval_within_cvs_flag = 0 specifies that the value of ShutterInterval is not the same for all temporal sub - layers within the CVS. sub_layer_shutter_interval_numer[i] specifies, in seconds, the numerator used to derive the sub - layer shutter interval defined by the variable subLayerShutterInterval[i] when HighestTid is equal to i. sub_layer_shutter_interval_denom[i] specifies, in seconds, the denominator used to derive the sub - layer shutter interval defined by the variable subLayerShutterInterval[i] when HighestTid is equal to i. The value of subLayerShutterInterval[i] for HighestTid equal to i is derived as follows. When the value of fixed_shutter_interval_within_cvs_flag is equal to 0 and the value of sub_layer_shutter_interval_denom[i] is greater than 0, subLayerShutterInterval[i]=ShutterInterval * sub_layer_shutter_interval_numer[i]÷sub_layer_shutter_interval_denom[i]. Otherwise (when the value of sub_layer_shutter_interval_denom[i] is equal to 0), subLayerShutterInterval[i] should be interpreted as unknown or not specified. When the value of fixed_shutter_interval_within_cvs_flag is not equal to 0, subLayerShutterInterval[i]=ShutterInterval.
[0072] In an alternative embodiment, instead of using a numerator and a denominator to signal the sublayer shutter interval, a single value is used. An example of such syntax is shown in Table 23.
Table 23
[0073] Table 24 provides an overview of the six approaches discussed in Tables 18 - 23 for providing SEI messaging related to shutter angle or exposure duration.
Table 24
[0074] Variable Frame Rate Signaling As described in U.S. Provisional Application No. 62 / 883,195, filed on August 06, 2019, in many applications, it is desirable for a decoder to support playback at variable frame rates. Frame rate adaptation is typically part of the operation in a Hypothetical Reference Decoder (HRD), as described in Annex C of Ref. [2], for example. In one embodiment, it is proposed to signal syntax elements that define Picture Presentation Time (PPT) as a function of a 90 kHz clock via SEI messaging or other means. This is a kind of repetition of the nominal decoder picture buffer (DPB) output time specified in the HRD, but currently uses the 90 kHz ClockTicks accuracy specified in the MPEG-2 system. The advantages of this SEI message are: a) when the HRD is not enabled, the PPT SEI message can still be used to indicate the timing of each frame; b) it can facilitate the conversion of bitstream timing and system timing.
[0075] Table 25 illustrates an example of the syntax of the proposed PPT timing message that matches the syntax of the Presentation Time Stamp (PTS) variable used in MPEG-2 Transport (H.222) (Ref. [4]).
Table 25
[0076] Shutter Interval Messaging in AVC In one embodiment, it is proposed that if a Shutter Interval Information (SII) SEI message exists for any picture in the coded video sequence (CVS), it must be present in the first access unit of the CVS. Unlike HEVC, the temporal index (used to identify the sublayer index) does not exist in the AVC single-layer bitstream. To address this problem when the shutter interval is not fixed within the CVS, it is proposed that a SEI message with shutter interval information exists for each picture to identify the sublayer index of the current picture and to assign a value of sii_sub_layer_idx to each picture. Other shutter interval-related information is presented only for the first access unit of the CVS and persists until a new CVS starts or the bitstream ends.
[0077] In AVC, an access unit is defined as a set of NAL units whose decoding order is consecutive and which contains exactly one primary coded picture. In addition to the primary coded picture, the access unit may also contain one or more redundant coded pictures, one auxiliary coded picture, or other NAL units that do not contain slices or slice data partitions of the coded pictures. Decoding of an access unit always results in decoded pictures.
[0078] Table 26 shows exemplary syntax element values when the shutter interval is fixed with respect to the CVS. Table 27 shows exemplary syntax element values of the first and subsequent shutter interval information SEI messages when the shutter interval is different for different sub-layers. In Tables 26 and 27, cells with "(none)" indicate that the value is not signaled in the shutter interval information SEI message for the corresponding syntax element. [Table 26] [Table 27]
[0079] Table 28 shows an exemplary syntax structure for SII SEI messaging in AVC. [Table 28]
[0080] The shutter interval information SEI message indicates the shutter interval of the associated video source picture before encoding and display. For example, in the case of camera-captured content, the shutter interval is the amount of time the image sensor is exposed to generate each source picture. sii_sub_layer_idx specifies the shutter interval time sub-layer index of the current picture. When the current access unit is the first access unit of the CVS, the value of sii_sub_layer_idx is equal to 0. When fixed_shutter_interval_within_cvs_flag is equal to 1, the value of sii_sub_layer_idx must be equal to 0. Otherwise, fixed_shutter_interval_within_cvs_flag is equal to 0, and the value of sii_sub_layer_idx is less than or equal to the value of sii_max_sub_layers_minus1. The shutter_interval_info_present_flag being equal to 1 indicates the presence of the syntax elements sii_time_scale, fixed_shutter_interval_within_cvs_flag, and sii_num_units_in_shutter_interval or sii_max_sub_layers_minus1 and sub_layer_num_units_in_shutter_interval[i]. The value of shutter_interval_info_present_flag shall be equal to 1 when the current access unit is the first access unit of the CVS. Otherwise, the current access unit is not the first access unit of the CVS, and the value of shutter_interval_info_present_flag shall be equal to 0. sii_time_scale specifies the number of time units that elapse in one second. The value of sii_time_scale is greater than 0. For example, a time coordinate system that measures time using a 27 MHz clock has an sii_time_scale of 27 000 000. A fixed_shutter_interval_within_cvs_flag equal to 1 specifies that the indicated shutter interval is the same for all pictures in the CVS. A fixed_shutter_interval_within_cvs_flag equal to 0 specifies that the indicated shutter interval may not be the same for all pictures in the CVS. sii_num_units_in_shutter, when fixed_shutter_interval_within_cvs_flag is equal to 1, specifies the number of time units of a clock operating at a frequency of sii_time_scale Hz corresponding to the indicated shutter interval for each picture within the CVS. The value 0 may be used to indicate that the associated video content includes screen capture content, computer-generated content, or other non-camera capture content. The indicated shutter interval, as indicated by the variable shutter interval, is equal to the quotient of sii_num_units_in_shutter_interval divided by sii_time_scale, in seconds. For example, to represent a shutter interval equal to 0.04 seconds, sii_time_scale can be equal to 27,000,000 and sii_num_units_in_shutter_interval can be equal to 1,080,000. sii_max_sub_layers_minus1 + 1 specifies the maximum number of shutter interval time sub-layer indices that can exist in the CVS. sub_layer_num_units_in_shutter_interval[i], if it exists, specifies the number of time units of a clock operating at a frequency of sii_time_scale Hz corresponding to the shutter interval of each picture in the CVS where the value of sii_sub_layer_idx is equal to i. The sub-layer shutter interval of each picture, represented by the variable subLayerShutterInterval[i], where the value of sii_sub_layer_idx is equal to i, is equal to the quotient of sub_layer_num_units_in_shutter_interval[i] divided by sii_time_scale, in seconds. Therefore, the variable subLayerShutterInterval[i], corresponding to the indicated shutter interval of each picture in the sub-layer representation having a TemporalId equal to i in the CVS, is derived as follows: if(fixed_shutter_interval_within_cvs_flag) subLayerShutterInterval[i] = sii_num_units_in_shutter_interval ÷ sii_time_scale Otherwise, subLayerShutterInterval[i] = sub_layer_num_units_in_shutter_interval[i] ÷ sii_time_scale When a shutter interval information SEI message exists for any access unit within the CVS, the shutter interval information SEI message exists for the IDR access unit which is the first access unit of the CVS. All shutter interval information SEI messages applied to the same access unit shall have the same content. sii_time_scale and fixed_shutter_interval_within_cvs_flag survive from the first access unit of the CVS until a new CVS starts or the bitstream ends. When the value of fixed_shutter_interval_within_cvs_flag is equal to 0, the shutter interval information SEI message exists for each picture within the CVS. If it exists, sii_num_units_in_shutter_interval, sii_max_sub_layers_minus1, and sub_layer_num_units_in_shutter_interval[i] survive from the first access unit of the CVS until a new CVS is started or the bitstream ends. References Each of the references listed in this specification is hereby incorporated by reference in its entirety. [1] High Efficiency Video Coding (HEVC), H.265, Series H, Coding of moving pictures, ITU, (02 / 2018) [2] B. Bross, J. Chen, and S. Liu, "Versatile Video Coding (VVC) (Draft 5)", JVET Output Document, JVET-N1001, v5, (uploaded on May 14, 2019) [3] C. Carbonara, J. DeFilippis, M. Korpi, "High Frame Rate Capture and Generation", SMPTE, 2015 Technical Conference and Exhibition, October 26 - 29, 2015 [4]Infrastructure for Audio-Visual Services - Transmission Multiplexing and Synchronization, H.222.0, Series H, Generic Coding of Moving Pictures and Associated Audio Information: Systems, ITU, 08 / 2018
[0081] Example of computer system implementation Embodiments of the present invention can be implemented in a computer system, a system composed of electronic circuits and components, an integrated circuit (IC) device such as a microcontroller, a field programmable gate array (FPGA), or another configurable or programmable logic device (PLD), a discrete-time or digital signal processor (DSP), an application-specific IC (ASIC), and / or a device including one or more of such systems, devices, or components. The computer and / or IC can execute, control, or perform instructions regarding frame rate scalability as described herein. The computer and / or IC may calculate any of the various parameters or values related to the frame rate scalability described herein. Embodiments of images and videos can be implemented in hardware, software, firmware, and various combinations thereof.
[0082] Certain embodiments of the present invention include a computer processor that executes software instructions that cause the processor to perform the methods of the present invention. For example, one or more processors within a display, an encoder, a set-top box, a transcoder, etc. can implement the methods related to the above-mentioned frame rate scalability by executing software instructions in a program memory accessible to the processor. Embodiments of the present invention may also be provided in the form of a program product. The program product can include any non-transitory and tangible medium that carries a set of computer-readable signals that, when executed by a data processor, cause the data processor to perform the methods of the present invention. The program product according to the present invention can be in any of a variety of non-transitory and tangible forms. The program product can include, for example, physical media such as floppy disks, magnetic data storage media including hard disk drives, optical data storage media including CD ROMs, DVDs, electronic data storage media including ROMs, flash RAMs, etc. The computer-readable signals on the program product can optionally be compressed or encrypted. When components (e.g., software modules, processors, assemblies, devices, circuits, etc.) are referred to above, unless otherwise indicated, the reference to such components (including references to "means") shall be construed to include equivalents of the described components (e.g., functionally equivalent) that perform the functions of the described components, including components that are not structurally equivalent to the disclosed structures that perform the functions in exemplary embodiments of the present invention.
[0083] Equivalents, extensions, alternatives, and others Accordingly, exemplary embodiments related to frame rate scalability are described. In the foregoing specification, embodiments of the present invention have been described with reference to numerous specific details that may vary from implementation to implementation. Accordingly, the only exclusive indicator of what the present invention is and what is intended by the applicant to be the present invention is the set of claims issued from this application in the specific form in which such claims are issued, including subsequent amendments. The definitions explicitly set forth herein for the terms contained in the claims shall take precedence over the meanings of the terms used in the claims. Accordingly, no limitation, element, characteristic, feature, advantage or attribute not explicitly recited in the claims should limit the scope of the claims in any way. Accordingly, the specification and drawings should be regarded in an illustrative rather than a restrictive sense.
[0084] Appendix This appendix provides a copy of Table D.2 and related pic_struct related information from the H.265 specification (Reference [1]). [Table D.2-1] [Table D.2-2] Semantics of pic_struct Syntax Elements The pic_struct indicates whether a picture is to be displayed as a frame or as one or more fields, and for frame display when fixed_pic_rate_within_cvs_flag = 1, it can indicate the frame doubling or tripling repetition period of a display that uses a fixed frame update interval equal to DpbOutputElementalInterval[n] given by Equation E-73. The interpretation of pic_struct is defined in Table D.2. Values of pic_struct not listed in Table D.2 are reserved for future use by ITU-T|ISO / IEC and do not exist in a bitstream conforming to this version of this specification. The decoder shall ignore reserved values of pic_struct. When present, it is a requirement for bitstream conformance that the value of pic_struct be constrained such that exactly one of the following conditions is true. - The value of pic_struct is equal to 0, 7, or 8 for all pictures within the CVS. - The value of pic_struct is equal to 1, 2, 9, 10, 11, or 12 for all pictures within the CVS. - The value of pic_struct is equal to 3, 4, 5, or 6 for all pictures within the CVS. When fixed_pic_rate_within_cvs_flag = 1, frame doubling is indicated by a pic_struct equal to 7, which indicates that it should be displayed twice consecutively on the display with a frame refresh interval equal to DpbOutputElementalInterval[n] as given by Equation E-73, and frame tripling is indicated by a pic_struct equal to 8, which indicates that it should be displayed three times consecutively on the display with a frame refresh interval equal to DpbOutputElementalInterval[n] as given by Equation E-73. Note 3: Using frame doubling can ease display. For example, in a 50 Hz progressive scan display, 25 Hz progressive scan video can be used, and in a 60 Hz progressive scan display, 30 Hz progressive scan video can be used. By alternately combining frame doubling and frame tripling on every other frame, the display of 24 Hz progressive scan video on a 60 Hz progressive scan display can be eased. The nominal vertical and horizontal sampling positions of samples in the top and bottom fields for 4:2:0, 4:2:2, and 4:4:4 chroma formats are shown in Figures D.1, D.2, and D.3, respectively. The field association indicator (where pic_struct equals 9 to 12) provides a hint for associating complementary parity fields as a frame. The parity of a field can be top or bottom, and when the parity of one field is top and the parity of the other field is bottom, the parities of the two fields are considered complementary. When frame_field_info_present_flag = 1, it is a bitstream conformity requirement that the constraints specified in the third column of Table D.2 apply. Note 4: When frame_field_info_present_flag = 0, in many cases, default values can be inferred or indicated by other means. In the absence of other indications of the intended display type of a picture, the decoder should infer the value of pic_struct to be 0 when frame_field_info_present_flag = 0.
Claims
Claim 1 When executed by a processor, cause the processor to encode an input video picture to generate an encoded bitstream, generate metadata parameters indicating shutter interval information for the encoded bitstream, generate an output video stream including the encoded bitstream and the metadata parameters, where the metadata parameters include a shutter interval time sublayer index of the current picture in the encoded bitstream, when the shutter interval time sublayer index is equal to zero, the metadata parameters include a shutter interval information presence flag, and when the shutter interval information presence flag is equal to 1, the metadata parameters include a shutter interval time scale parameter indicating the number of time units elapsed per second, and a fixed shutter interval duration flag indicating whether the shutter interval duration information is fixed for all pictures in the encoded bitstream, when the fixed shutter interval duration flag indicates that the shutter interval duration information is fixed, the metadata parameters include a shutter interval clock tick parameter indicating the number of time units of a clock operating at the frequency of the shutter interval time scale parameter, otherwise, the metadata parameters include an array of one or more sublayer shutter interval clock tick parameters indicating the number of time units of a clock at the frequency of the shutter interval time scale parameter for one or more sublayers in the encoded bitstream, and for a first sublayer in the encoded bitstream, the corresponding sublayer shutter interval clock tick parameter divided by the shutter interval time scale parameter indicates the exposure duration value for all video pictures in the first sublayer of the encoded bitstream, A non - transitory processor - readable medium storing the instructions. Claim 2 When the fixed shutter interval duration flag indicates that the shutter interval duration information is not fixed, the metadata parameter includes a maximum sublayer parameter that calculates the maximum number of one or more sublayers in the encoded bitstream, the non-transitory processor-readable medium according to claim 1. **Claim 3** When the fixed shutter interval duration flag is equal to 1, the shutter interval time sublayer index is equal to zero; otherwise, when the fixed shutter interval duration flag is equal to zero, the value of the shutter interval time sublayer index is less than or equal to the maximum sublayer parameter, the non-transitory processor-readable medium according to claim 2. **Claim 4** The shutter interval clock tick parameter divided by the shutter interval time scale parameter indicates the exposure duration value for all video pictures in the encoded bitstream, the non-transitory processor-readable medium according to claim 1. **Claim 5** A non-transitory processor-readable medium that stores instructions that, when executed by a processor, cause the processor to transmit an encoded bitstream, wherein generating the encoded bitstream comprises encoding an input video picture to generate the encoded bitstream; and generating metadata parameters indicating shutter interval information for the encoded bitstream; and generating an output video stream including the encoded bitstream and the metadata parameters, wherein the metadata parameters include a shutter interval time sublayer index of a current picture in the encoded bitstream, when the shutter interval time sublayer index is equal to zero, the metadata parameter includes a shutter interval information presence flag, and when the shutter interval information presence flag is equal to 1, the metadata parameter includes a shutter interval time scale parameter indicating the number of time units elapsed per second; and a fixed shutter interval duration flag indicating whether the shutter interval duration information is fixed for all pictures in the encoded bitstream. When the fixed shutter interval duration flag indicates that the shutter interval duration information is fixed, the metadata parameter includes a shutter interval clock tick parameter indicating the number of time units of a clock operating at the frequency of the shutter interval time scale parameter, otherwise, the metadata parameter includes an array of one or more sublayer shutter interval clock tick parameters indicating the number of time units of a clock at the frequency of the shutter interval time scale parameter for one or more sublayers in the encoded bitstream, and for a first sublayer in the encoded bitstream, the corresponding sublayer shutter interval clock tick parameter divided by the shutter interval time scale parameter indicates the exposure duration value for all video pictures in the first sublayer of the encoded bitstream. A non-transitory processor-readable medium storing the instructions. **Claim 6** When the fixed shutter interval duration flag indicates that the shutter interval duration information is not fixed, the metadata parameter includes a maximum sublayer parameter for calculating the maximum number of one or more sublayers in encoding the input video picture. The non-transitory processor-readable medium according to claim 5. **Claim 7** When the fixed shutter interval duration flag is equal to 1, the shutter interval time sublayer index is equal to zero; otherwise, when the fixed shutter interval duration flag is equal to zero, the value of the shutter interval time sublayer index is less than or equal to the maximum sublayer parameter. The non-transitory processor-readable medium according to claim 6. **Claim 8** The shutter interval clock tick parameter divided by the shutter interval time scale parameter indicates the exposure duration value for all video pictures in the encoded bitstream. The non-transitory processor-readable medium according to claim 5.
Citation Information
Patent Citations
Transmission apparatus, transmission method, reception apparatus and reception method
WO2015076277A1
Image processing device, image processing method, reception device and transmission device
WO2016185947A1
Frame-rate-conversion metadata
WO2019067762A1
Frame-rate scalable video coding
WO2020185853A2