Frame-rate scalable video coding
Patent Information
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- DOLBY LABORATORIES LICENSING CORP
- Filing Date
- 2020-03-11
- Publication Date
- 2026-08-05
Smart Images

Figure 112025001522045-PAT00005_ABST
Abstract
Description
Technology Field
[0001] Cross-reference for related applications
[0002] The present application claims the benefit of priority of U.S. Provisional Application No. 62 / 816,521 filed March 11, 2019; U.S. Provisional Application No. 62 / 850,985 filed May 21, 2019; U.S. Provisional Application No. 62 / 883,195 filed August 6, 2019; and U.S. Provisional Application No. 62 / 904,744 filed September 24, 2019, each incorporated by reference in its entirety.
[0003] Technology field
[0004] This document generally relates to images. Specifically, embodiments of the present invention relate to frame rate scalable video coding. Background Technology
[0005] As used herein, the term ‘dynamic range’ (DR) may relate to the ability of the human visual system (HVS) to perceive a range of intensity (e.g., luminance, luma) in an image, e.g., from the darkest gray (black) to the brightest white (highlight). In this sense, DR relates to ‘scene-referred’ intensity. DR may also relate to the ability of a display device to render a range of intensity of a specific width appropriately or approximately. In this sense, DR relates to ‘display-referred’ intensity. Unless a specific meaning is explicitly designated as having particular significance at any point in the description herein, it should be inferred that the term may be used in any sense, e.g., interchangeably.
[0006] As used herein, the term HDR (high dynamic range) relates to a DR range spanning 14 to 15 orders of magnitude of the human visual system (HVS). In reality, the DR, which allows humans to perceive a wide range of intensity simultaneously, may be somewhat truncated in relation to HDR.
[0007] In reality, an image contains one or more color components (e.g., luminance Y and saturation Cb and Cr), where each color component is per pixel n - Bit precision (e.g., n It is expressed as = 8). If linear luminance coding is used, n Images with ≤ 8 (e.g., 24-bit color JPEG images) are considered Standard Dynamic Range (SDR) images, and n Images can be considered as images with enhanced dynamic range. HDR images can also be stored and distributed using high-precision (e.g., 16-bit) floating-point formats, such as the OpenEXR file format developed by Industrial Light and Magic.
[0008] Currently, the distribution of HDR video content, such as Dolby Laboratories' Dolby Vision or Blu-ray's HDR10, is limited to 4K resolution (e.g., 4096 x 2160 or 3840 x 2160, etc.) and 60 frames per second (fps) due to the performance limitations of many playback devices. In future versions, it is expected that content at up to 8K resolution (e.g., 7680 x 4320) and 120fps will be available for distribution and playback. To simplify the ecosystem of HDR playback content such as Dolby Vision, it is desirable for future content types to be compatible with existing playback devices. Ideally, content creators should be able to adopt and distribute future HDR technologies without having to induce and distribute special versions of content compatible with existing HDR devices (e.g., HDR10 or Dolby Vision). As recognized by the inventors herein, improved technology is required for the scalable distribution of video content, particularly HDR content.
[0009] The approaches described in this section are approaches that may be pursued, but are not necessarily approaches that have been previously conceived or pursued. Therefore, unless otherwise indicated, the approaches described in this section should not be assumed to qualify as prior art merely because they are included in this section. Similarly, problems identified in relation to one or more of the approaches should not be assumed to have been recognized in the prior art based on this section, unless otherwise stated. Brief explanation of the drawing
[0010] Embodiments of the present invention are illustrated in the accompanying drawings as examples rather than limitations, and similar reference numerals refer to similar elements. Figure 1 illustrates an exemplary process for a video delivery pipeline. FIG. 2 illustrates an exemplary process of combining consecutive original frames to render a target frame rate at a target shutter angle according to an embodiment of the present invention. FIG. 3 illustrates an exemplary representation of an input sequence having a variable input frame rate and a variable shutter angle in a container having a fixed frame rate according to an embodiment of the present invention. FIG. 4 illustrates an exemplary representation of temporal scalability at various frame rates and shutter angles having backward compatibility according to an embodiment of the present invention. Specific details for implementing the invention
[0011] Exemplary embodiments related to frame rate scalability for video coding are described herein. In the following description, numerous specific details are presented for illustrative purposes to provide a complete understanding of various embodiments of the invention. However, it will be apparent that various embodiments of the invention may be practiced without these specific details. In other examples, known structures and devices are not described in full detail to avoid unnecessarily obscuring, making unclear, or ambiguous embodiments of the invention.
[0012] summation
[0013] The exemplary embodiments described herein relate to frame rate scalability of video coding. In one embodiment, a system having a processor receives a coded bitstream containing coded video frames, and one or more coded frames are encoded at a first frame rate and a first shutter angle. The processor receives a first flag indicating the existence of a group of coded frames to be decoded at a second frame rate and a second shutter angle, accesses a second frame rate and a second shutter angle value for the group of coded frames from the coded bitstream, and generates a frame decoded at a second frame rate and a second shutter angle based on the group of coded frames, the first frame rate, the first shutter angle, the second frame rate, and the second shutter angle.
[0014] In the second embodiment, a decoder having a processor:
[0015] Receive a coded bitstream containing a group of coded video frames—all coded video frames within the coded bitstream are encoded at a first frame rate;
[0016] Receive the number of combined frames N;
[0017] Receive a value for the baseline frame rate;
[0018] Access N consecutive coded frame groups and - within N consecutive coded frame groups i The nth coded frame (here i = 1, 2, ..., N) is the baseline frame rate and based on the first shutter angle and the first frame rate. i Represents the average of up to i input video frames encoded within the encoder at the i-th shutter angle;
[0019] To decode a group of N consecutive coded frames at a second frame rate and a second shutter angle, access the values of the second frame rate and the second shutter angle from the coded bitstream or from user input; and
[0020] Based on N consecutive coded frame groups, a first frame rate, a first shutter angle, a second frame rate, and a second shutter angle, a frame decoded at a second frame rate and a second shutter angle is generated.
[0021] In the third embodiment, the encoded video stream structure is:
[0022] An encoded image section including the encoding of a video image sequence; and
[0023] A shutter interval time scale parameter indicating the number of time units passing per second;
[0024] A shutter interval clock tick parameter representing the number of time units of a clock operating at the frequency of the above shutter interval time scale parameter—the above shutter interval clock tick parameter divided by the above shutter interval time scale parameter represents an exposure duration value;
[0025] It includes a signaling section comprising encoding of a shutter interval duration flag indicating whether exposure duration information is fixed for all time sublayers within the above-mentioned encoded image section, and
[0026] If the shutter interval duration flag indicates that the exposure duration information is fixed, a decoded version of the video image sequence for all the time sublayers within the encoded image section is decoded by calculating the exposure duration value based on the shutter interval time scale parameter and the shutter interval clock tick parameter, and otherwise
[0027] The above signaling section includes one or more arrays of sublayer parameters and calculates a corresponding sublayer exposure duration for displaying a decoded version of the time sublayer of the video image sequence for each sublayer using the values of one or more arrays of sublayer parameters combined with the shutter interval time scale parameter.
[0028] Exemplary video delivery processing pipeline
[0029] FIG. 1 illustrates an exemplary process of a conventional video delivery pipeline (100) showing various stages from video capture to video content display. A sequence of video frames (102) is captured or generated using an image generation block (105). The video frames (102) may be digitally captured (e.g. by a digital camera) or generated by a computer (e.g. using computer animation) to provide video data (107). Alternatively, the video frames (102) may be captured on film by a film camera. The film is converted into a digital format to provide video data (107). In the production stage (110), the video data (107) is edited to provide a video production stream (112).
[0030] Video data of the production stream (112) is provided to the processor in the block (115) for post-production editing. Post-production editing in the block (115) may include adjusting or modifying color or brightness in specific areas of the image to improve image quality or to achieve a specific appearance of the image according to the creative intent of the video producer. This may also be referred to as "color timing" or "color grading." Other editing (e.g., scene selection and ordering, image cropping, addition of computer-generated visual effects, judder or blur control, frame rate control, etc.) may be performed in the block (115) to produce a final version (117) of the production for distribution. During post-production editing (115), the video image may be viewed on a reference display (125). After post-production (115), the video data of the final production (117) may be transferred to an encoding block (120) to be delivered downstream to a decoding and playback device such as a television set, set-top box, or movie theater. In some embodiments, the coding block (120) may include audio and video encoders, such as those defined by ATSC, DVB, DVD, Blu-ray, and other delivery formats, to generate a coded bitstream (122). In the receiver, the coded bitstream (122) is decoded by a decoding unit (130) to generate a decoded signal (132) that represents an approximation identical to or close to the signal (117). The receiver may be attached to a target display (140) that may have characteristics entirely different from those of a reference display (125). In this case, a display management block (135) may be used to map the dynamic range of the decoded signal (132) to the characteristics of the target display (140) by generating a display-mapped signal (137).
[0031] Scalable Coding
[0032] Scalable coding is already part of many video coding standards, such as MPEG-2, AVC, and HEVC. In embodiments of the present invention, scalable coding is extended to improve performance and flexibility, particularly because it is relevant to ultra-high resolution HDR content.
[0033] As used herein, the term "shutter angle" refers to an adjustable shutter setting that controls the ratio of the time the film is exposed to light during each frame interval. For example, in one embodiment, it is as follows.
[0034] (1)
[0035] This term originates from traditional mechanical rotary shutters, but modern digital cameras can also electronically control the shutter. Cinematographers can use the shutter angle to control the amount of motion blur or judder recorded in each frame. Instead of using "exposure time," alternative terms such as "exposure duration," "shutter interval," and "shutter speed" may be used. Similarly, the term "frame duration" may be used instead of "frame interval." Alternatively, "frame interval" can be replaced with "1 / frame rate." The exposure time value is generally less than or equal to the frame duration. For example, a shutter angle of 180 degrees indicates that the exposure time is half the frame duration. In some situations, the exposure time may be greater than the frame duration of the encoded video, for example, if the encoded frame rate is 120fps and the frame rate of the associated video content prior to encoding and display is 60fps.
[0036] Consider, without limitation, an embodiment in which original content is captured (or generated) at an original frame rate (e.g., 120fps) with a shutter angle of 360 degrees. Then, at a receiving device, video output can be rendered at various frame rates equal to or lower than the original frame rate by a fair combination of original frames, for example, by averaging or other operations known in the art.
[0037] The combining process can be performed with non-linearly encoded signals (e.g., using Gamma, PQ, or HLG), but the best image quality can be achieved by combining frames in the linear light domain by first converting the non-linearly encoded signal into a linear light representation, combining the converted frames, and finally re-encoding the output with a non-linear transfer function. This process provides a more accurate simulation of actual camera exposure than combining in the non-linear domain.
[0038] Generally, the frame combining process can be expressed as follows in terms of the original frame rate, target frame rate, target shutter angle, and the number of frames to be combined, and
[0039] n_frames=(target_shutter_angle / 360)*(original_frame_rate / target_frame_rate) (2)
[0040] This is equivalent to the following equation and
[0041] target_shutter_angle=360*n_frames*(target_frame_rate / original_frame_rate) (3)
[0042] Here, n_frames is the number of frames to be combined, original_frame_rate is the frame rate of the original content, target_frame_rate is the frame rate to be rendered (where target_frame_rate ≤ original_frame_rate), and target_shutter_angle represents the desired amount of motion blur. In this example, the maximum value of target_shutter_angle is 360 degrees and corresponds to maximum motion blur. The minimum value of target_shutter_angle can be expressed as 360*(target_frame_rate / original_frame_rate) and corresponds to minimum motion blur. The maximum value of n_frames can be expressed as (original_frame_rate / target_frame_rate). The values of target_frame_rate and target_shutter_angle must be selected so that the value of n_frames is a non-zero integer.
[0043] In the special case where the original frame rate is 120fps, Equation (2) can be rewritten as follows, and
[0044] n_frames = target_shutter_angle / (3*target_frame_rate) (4)
[0045] This is equivalent to the following equation.
[0046] target_shutter_angle = 3*n_frames*target_frame_rate (5)
[0047] The relationships between the target_frame_rate, n_frames, and target_shutter_angle values are shown in Table 1 for the case of original_frame_rate = 120fps. In Table 1, "NA" indicates that the corresponding combination of the number of frames to be combined with the target frame rate is not allowed.
[0048] Table 1: Relationship between Target Frame Rate, Combined Frames, and Target Shutter Angle for Original Frame Rate of 120fps
[0049] Target frame rate (fps) Number of frames to be combined 5 4 3 2 1 Target shutter angle (degrees) 24 360 288 216 144 72 30 NA 360 270 180 90 40 NA NA 360 240 120 60 NA NA NA 360 180
[0050] FIG. 2 illustrates an exemplary process for combining consecutive original frames to render a target frame rate at a target shutter angle according to one embodiment. Given an input sequence (205) of 120fps and a shutter angle of 360 degrees, the process combines three input frames (e.g., the first three consecutive frames) from a set of five consecutive frames and drops the remaining two to generate an output video sequence (210) of 24fps and a shutter angle of 216 degrees. Note that in some embodiments, the output frame-01 of (210) may be generated by combining alternative input frames (205), such as frames 1, 3, and 5, or frames 2, 4, and 5, etc. However, it can be expected that combining consecutive frames will yield a video output of better quality.
[0051] For example, it is desirable to support original content with a variable frame rate to manage artistic and stylistic effects. Additionally, the variable input frame rate of the original content is desirable to be packaged into a "container" having a fixed frame rate to simplify content creation, exchange, and distribution. As an example, three embodiments of a method for representing variable frame rate video data in a fixed frame rate container are presented. For clarity and without limitation, the following description uses a fixed 120fps container, but this approach can be easily extended to alternative frame rate containers.
[0052] First embodiment (Variable Frame Rate)
[0053] The first embodiment is an explicit description of source content having a variable (non-constant) frame rate packaged in a container having a constant frame rate. For example, source content having different frame rates of 24, 30, 40, 60, or 120fps for different scenes may be packaged in a container having a constant frame rate of 120fps. In this example, each input frame may be duplicated 5x, 4x, 3x, 2x, or 1x and packaged into a common 120fps container.
[0054] FIG. 3 illustrates an example of an input video sequence A having a variable frame rate and a variable shutter angle, represented by a bitstream B coded at a fixed frame rate. Subsequently, in a decoder, the decoder reconstructs an output video sequence C at a desired frame rate and shutter angle that can be changed from scene to scene. For example, as shown in FIG. 3, to construct sequence B, some input frames are duplicated, some are coded as is (without duplication), and some are copied four times. Then, to construct sequence C, one frame is selected from each set of duplicated frames to generate an output frame that matches the original frame rate and shutter angle.
[0055] In this embodiment, metadata is inserted into the bitstream to indicate the source (base) frame rate and shutter angle. The metadata may be signaled using high-level syntax such as Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Slice, or Tile Group headers. With the metadata present, the encoder and decoder can perform the following useful functions.
[0056] a) The encoder can ignore duplicate frames, thereby increasing encoding speed and simplifying processing. For example, all coding tree units (CTUs) of a duplicate frame can be encoded using SKIP mode and reference index 0 in LIST 0 of reference frames that reference the decoded frame to which the duplicate frame is copied.
[0057] b) The decoder can simplify processing by bypassing the decoding of duplicate frames. For example, bitstream metadata may indicate that a frame is a duplicate of a previously decoded frame, which the decoder can copy and play back without decoding the new frame.
[0058] c) The playback device can optimize downstream processing by displaying the base frame rate, for example, by adjusting the frame rate conversion or noise reduction algorithm.
[0059] This embodiment enables the end user to view content rendered at the frame rate intended by the content creator. This embodiment does not provide backward compatibility with devices that do not support the container's frame rate, for example, 120fps.
[0060] Tables 2 and 3 illustrate exemplary syntax of raw byte sequence payloads (RBSB) for sequence parameter sets and Tile Group headers, where the new syntax elements proposed herein are depicted in italics. The remaining syntax follows the syntax of the standard proposed in the Versatile Video Codec (VVC) (Reference [2]).
[0061] For example, in SPS (see Table 2), a flag can be added to enable variable frame rates.
[0062] sps_vfr_enabled_flagA value of 1 specifies that the coded video sequence (CVS) can contain variable frame rate content. A value of sps_vfr_enabled_flag of 0 specifies that the CVS contains fixed frame rate content.
[0063] In tile_group header() (see Table 3),
[0064] tile_group_vrf_info_present_flag A value of 1 specifies that the syntax elements tile_group_true_fr and tile_group_shutterangle are in the syntax. A value of tile_group_vrf_info_present_flag of 0 specifies that the syntax elements tile_group_true_fr and tile_group_shutterangle are not in the syntax. If tile_group_vrf_info_present_flag is not present, it is inferred to be 0.
[0065] tile_group_true_fr represents the actual frame rate of the video data transmitted in this bitstream. tile_group_shutterangle represents the shutter angle corresponding to the actual frame rate of the video data transmitted in this bitstream.
[0066] tile_group_skip_flag A value of 1 specifies that the current tile group has been copied from another tile group. A value of tile_group_skip_flag of 0 specifies that the current tile group has not been copied from another tile group.
[0067] tile_group_copy_pic_order_cnt_lsb MaxPicOrderCntLsb is set as the image order count module for the previously decoded image from which the current image was copied when tile_group_skip_flag is set to 1.
[0068] Table 2: Exemplary RBSP Syntax with Parameter Sets for Variable Frame Rate Content
[0069] seq_parameter_set_rbsp( ) { explainer (Descriptor) sps_max_sub_layers_minus1 u(3) sps_reserved_zero_5bits u(5) profile_tier_level( sps_max_sub_layers_minus1 ) sps_seq_parameter_set_id ue(v) ... sps_vfr_enabled_flag u(1) ... sps_extension_flag u(1) if( sps_extension_flag ) while( more_rbsp_data( ) ) sps_extension_data_flag u(1) rbsp_trailing_bits( ) }
[0070] Table 3: Exemplary Tile Group header syntax supporting variable frame rate content
[0071] tile_group_header( ) { explainer tile_group_pic_parameter_set_id ue(v) if( NumTilesInPic > 1 ) { tile_group_address u(v) num_tiles_in_tile_group_minus1 ue(v) } tile_group_type ue(v) tile_group_pic_order_cnt_lsb u(v) if( sps_vfr_enabled_flag ) { tile_group_vfr_info_present_flag u(1) if(tile_group_vfr_info_present_flag) { tile_group_true_fr u(9) tile_group_shutterangle u(9) } tile_group_skip_flag u(1) } if( tile_group_skip_flag ) tile_group_copy_pic_order_cnt_lsb u(v) else{ ALL OTHER TILE_GROUP_SYNTAX } if( num_tiles_in_tile_group_minus1 > 0 ) { offset_len_minus1 ue(v) for( i = 0; i < num_tiles_in_tile_group_minus1; i++ ) entry_point_offset_minus1 [ i ] u(v) } byte_alignment( ) }
[0072] Second Embodiment - Fixed Frame Rate Container
[0073] The second embodiment enables a use case in which original content having a fixed frame rate and shutter angle can be rendered by a decoder at an alternative frame rate and a variable simulated shutter angle, as illustrated in FIG. 2. For example, if the frame rate of the original content is 120fps and the shutter angle is 360 degrees (shutter open 1 / 120 second), the decoder can render at various frame rates of 120fps or less. For example, to decode 24fps at a simulated shutter angle of 216 degrees, as described in Table 1, the decoder can combine three decoded frames and display at 24fps. Table 4 extends Table 1 to illustrate a method of combining different numbers of encoded frames to render at an output target frame rate and a desired target shutter angle. Combining frames can be performed by simple pixel averaging, weighted pixel averaging where pixels in a specific frame may have higher weights than pixels in another frame and the sum of all weights is 1, or other filter interpolation methods known in the field. In Table 4, the function Ce(a, b) represents a combination of encoded frames a to b, where the combination can be performed by averaging, weighted averaging, filtering, etc.
[0074] Table 4: Example of combining input frames at 120fps to generate output frames at target fps and shutter angle values
[0075] input s1 s2 s3 s4 s5 s6 s7 s8 s9 s10 Enc. e1 e2 e3 e4 e5 e6 e7 e8 e9 e10 Dec. 120fps @360 e1 e2 e3 e4 e5 e6 e7 e8 e9 e10 Dec. 60fps @360 Ce(1,2) Ce(3,4) Ce(5,6) Ce(7,8) Ce(9,10) @180 e1 e3 e5 e7 e9 Dec. 40fps @360 Ce(1,3) Ce(4,6) Ce(7,9) Ce(10,12) @240 Ce(1,2) Ce(4,5) Ce(7,8) Ce(10,11) @120 e1 e4 e7 e10 Dec. 30fps @360 Ce(1,4) Ce(5,8) Ce(9,12) @270 C(1,3) Ce(5,7) Ce(9,11) @180 Ce(1,2) Ce(5,6) Ce(9,10) @90 e1 e5 e9 Dec. 24fps @360 Ce(1,5) Ce(6,10) @288 Ce(1,4) Ce(6,9) @216 Ce(1,3) Ce(6,8) @144 Ce(1,2) Ce(6,7) @72 e1 e6
[0076] When the value of the target shutter angle is less than 360 degrees, the decoder can combine different sets of decoded frames. For example, in Table 1, given an original stream of 120fps at 360 degrees, to generate a stream at 40fps and a shutter angle of 240 degrees, the decoder must combine two of the three possible frames. Thus, the first and second frames or the second and third frames can be combined. The selection of frames to be combined can be described as a "decoding step" expressed as follows.
[0077] decode_phase = decode_phase_idx*(360 / n_frames) (6)
[0078] Here, decode_phase_idx represents an offset index within a series of frame sets having index values within [0, n_frames_max-1], where n_frames is given by Equation (2), and
[0079] n_frames_max = orig_frame_rate / target_frame_rate (7)
[0080] am.
[0081] Generally, the range of decode_phase_idx is [0, n_frames_max-n_frames]. For example, for an original sequence of 120fps and a 360-degree shutter angle, n_frames_max = 120 / 40 = 3 for a target frame rate of 40fps at a 240-degree shutter angle. From Equation (2), n_frames = 2, so the range of decode_phase_idx is [0, 1]. Thus, decode_phase_idx = 0 indicates that frames with indices 0 and 1 are selected, and decode_phase_idx = 1 indicates that frames with indices 1 and 2 are selected.
[0082] In this embodiment, the rendered variable frame rate intended by the content creator may be signaled as metadata, such as Additional Enhancement Information (SEI) messages or Video Usability Information (VUI). Optionally, the rendered frame rate may be controlled by the receiver or user. An example of a frame rate conversion SEI message specifying the content creator's preferred frame rate and shutter angle is shown in Table 5. The SEI message may also indicate whether frame combining is performed in a coded signal domain (e.g., gamma, PQ, etc.) or a linear light domain. Note that post-processing requires a frame buffer along with a decoder picture buffer (DPB). The SEI message may indicate the number of additional frame buffers required or other methods for combining frames. For example, to reduce complexity, frames may be recombined at reduced spatial resolution.
[0083] As shown in Table 4, for specific combinations of frame rate and shutter angle (e.g., 30fps and 360 degrees or 24fps and 288 or 360 degrees), the decoder may need to combine three or more decoded frames, which increases the number of buffer spaces required by the decoder. To reduce the burden of additional buffer space on the decoder, in some embodiments, specific combinations of frame rate and shutter angle may deviate from the limitations on the allowed set of decoding parameters (e.g., by setting appropriate coding profiles and levels).
[0084] To give another example, consider the case of playback at 24fps. The decoder might decide to display the same frame five times to achieve a 120fps output frame rate. This is exactly the same as displaying a frame only once at a 24fps output frame rate. The advantage of maintaining a constant output frame rate is that the display can run at a constant clock speed, which makes all the hardware much simpler. If the display can dynamically change the clock speed, displaying a frame only once (for 1 / 24 seconds) might be more reasonable than repeating the same frame five times (each for 1 / 120 seconds). The former approach can result in slightly higher image quality, better optical efficiency, or better power efficiency. Similar considerations apply to other frame rates as well.
[0085] Table 5 shows an example of a frame rate conversion SEI message syntax according to an embodiment.
[0086] Table 5: Examples of SEI message syntax that allows frame rate conversion
[0087] framerate_conversion( payloadSize ) { explainer framerate_conversion_cancel_flag u(1) if(!frame_conversion_cancel_flag) { base_frame_rate u(9) base_shutter_angle u(9) decode_phase_idx_present_flag u(1) if(decode_phase_idx_present_flag) { decode_phase_idx u(3) } conversion_domain_idc u(1) num_frame_buffer u(3) framerate_conversion_persistence_flag u(1) } }
[0088] framerate_conversion_cancel_flag A value of 1 indicates that the SEI message cancels the persistence of the previous frame rate conversion SEI message in the output order. A framerate_conversion_cancel_flag value of 0 indicates that the frame rate conversion information is followed.
[0089] base_frame_rate Specifies the desired frame rate.
[0090] base_shutter_angle Specifies the desired shutter angle.
[0091] decode_phase_idx_present_flag A value of 1 indicates that decoding phase information exists. A value of 0 indicates that decoding phase information does not exist.
[0092] decode_phase_idx represents the offset index within a series of frame sets with index values of 0..(n_frames_max-1), where n_frames_max = 120 / base_frame_rate. The value of decode_phase_idx must be in the range 0..(n_frames_max-n_frames), where n_frames = base_shutter_angle / (3*base_frame_rate). If decode_phase_idx is missing, it is inferred to be 0.
[0093] conversion_domain_idc A value of 0 indicates that frame combination is performed in the linear domain. A value of 1 indicates that frame combination is performed in the non-linear domain.
[0094] num_frame_buffers Specifies the number of additional frame buffers (DPB is not calculated).
[0095] framerate_conversion_persistence_flag specifies the persistence of framerate conversion SEI messages for the current layer. A framerate_conversion_persistence_flag of 0 specifies that framerate conversion SEI messages apply only to the currently decoded image. Let picA be the current image. A framerate_conversion_persistence_flag of 1 specifies that framerate conversion SEI messages persist for the current layer in output order until one or more of the following conditions are true:
[0096] - The new coded layer-by-layer video sequence (CLVS) of the current layer starts.
[0097] - The bitstream ends.
[0098] - An image picB in the current layer of an access unit containing a frame rate conversion SEI message that can be applied to the current layer where PicOrderCnt(picB) is greater than PicOrderCnt(picA) is output, where PicOrderCnt(picB) and PicOrderCnt(picA) are the PicOrderCntVal values of picB and picA, respectively, immediately after the call of the decoding process for the image order count for picB.
[0099] Third Embodiment—Input encoded with multiple shutter angles
[0100] The third embodiment is a coding method that supports backward compatibility by allowing the extraction of subframe rates from a bitstream. In HEVC, this is achieved by temporal scalability. Temporal layer scalability is enabled by assigning different values to the temporal_id syntax element for the decoded frame. Thus, the bitstream can be simply extracted based on the temporal_id value. However, the HEVC-style approach to temporal scalability cannot render output frame rates with different shutter angles. For example, a 60fps base frame rate extracted from a 120fps source always has a shutter angle of 180 degrees.
[0101] In ATSC 3.0, an alternative method is described in which a 60fps frame with a 360-degree shutter angle is emulated as a weighted average of two 120fps frames. The emulated 60fps frames are assigned a temporal_id value of 0 and alternately combined with original 120fps frames assigned a temporal_id value of 1. If 60fps is required, the decoder only needs to decode the frames with temporal_id 0. If 120fps is required, the decoder can reconstruct all original 120fps frames by subtracting each temporal_id = 1 frame (i.e., 120fps frame) from the scaled version of each corresponding temporal_id = 0 frame (i.e., emulated 60fps frame) to recover the corresponding original 120fps frames that were not explicitly transmitted.
[0102] In an embodiment of the present invention, a new algorithm is described that supports multiple target frame rates and target shutter angles in a backward compatible (BC) manner. The proposal is to preprocess original 120fps content at a base frame rate at various shutter angles. Then, in the decoder, different frame rates at various different shutter angles can be simply derived. The ATSC 3.0 approach can be considered a special case of the proposed method, where frames with temporal_id = 0 are delivered at 60fps @ 360 shutter angles and frames with temporal_id = 1 are delivered at 60fps @ 180 shutter angles.
[0103] As a first example, consider an input sequence of 120fps and 360 shutter angles used to encode sequences with a base layer frame rate of 40fps and shutter angles of 120, 240, and 360 degrees, as shown in FIG. 4. In this method, the encoder calculates a new frame by combining up to three original input frames. For example, input frames 1 and 2 are combined to generate an encoded frame 2 (En-2) representing the input at 40fps and 240 degrees, and frame En-2 is combined with input frame 3 to generate an encoded frame 3 (En-3) representing the input at 40fps and 360 degrees. In the decoder, to reconstruct the input sequence, frame En-1 is subtracted from frame En-2 to generate a decoded frame 2 (Dec-2), and frame En-2 is subtracted from frame En-3 to generate a decoded frame 3 (Dec-3). The three decoded frames represent the output at a base frame rate of 120fps and a shutter angle of 360 degrees. Additional frame rates and shutter angles can be extrapolated using the decoded frames as shown in Table 6. In Table 6, the function Cs(a, b) represents a combination of input frames a and b, where the combination can be performed by averaging, weighted averaging, filtering, etc.
[0104] Table 6: Examples of frame combinations with a 40fps baseline
[0105] Input frames 120fps@360 s1 s2 s3 s4 s5 s6 s7 s8 s9 Encoded frames 120fps e1=s1 e2=Cs(1,2) e3=Cs(1,3) e4=s4 e5=Cs(4,5) e6=Cs(4,6) e7=s7 e8=Cs(7.8) e9=Cs(7,9) Decoding 120fps @360 e1=s1 e2-e1=s2 e3-e2=s3 e4=s4 e5-e4=s5 e6-e4=s6 e7=s7 e8-e7=s8 e9-e8=s9 Decoding 60fps @360 e2 e3-e2+e4=Cs(3,4) e6-e4=Cs(5,6) e8=Cs(7,8) e9-e8+e10 @180 e1 e3-e2=s3 e5-e4=s5 e7 e9-e8 Decoding 40fps @360 e3=Cs(1,3) e6 e9 @240 e2=Cs(1,2) e5 e8 @120 e1=s1 e4 e7 Decoding 30fps @360 e3+e4=Cs(1,4) e6-e4+e8=Cs(5,8) e9-e8+e12 @270 e3=Cs(1,3) e6-e5+e7=Cs(5,7) e9-e8+e11 @180 e2=Cs(1,2) e6-e4=Cs(5,6) e9-e8+e10 @90 e1 e5-e4=s5 e9-e8 Decoding 24fps @360 e3+e5=Cs(1,5) e6-e5+e9+e10=Cs(6,10) @288 e3+e4=Cs(1,4) e6-e5+e9=Cs(6,9) @216 e3=Cs(1,3) e6-e5+e8=Cs(6,8) @144 e2=Cs(1,2) e6-e5+e7=Cs(6,7) @72 e1=s1 e6-e5=s6
[0106] The advantage of this approach is that, as shown in Table 6, all 40fps versions can be decoded without additional processing. Another advantage is that different frame rates can be derived at various shutter angles. For example, assume a decoder decoding at 30fps and a shutter angle of 360. From Table 4, the output corresponds to a frame sequence generated by Ce(1,4) = Cs(1,4), Cs(5,8), Cs(9,12), etc., which also matches the decoding sequence shown in Table 6. However, in Table 6, Cs(5,8) = e6-e4 + e8. In the embodiment, a method can be defined to combine decoded frames to generate an output sequence at a specified output frame rate and an emulated shutter angle using a lookup table (LUT).
[0107] In another example, as shown below, it is proposed to combine up to five frames in the encoder to simplify the extraction of the 24fps base layer at shutter angles of 72, 144, 216, 288, and 360 degrees. This is suitable for movie content that is best represented at 24fps on conventional TVs.
[0108] Table 7: Examples of frame combinations with a 24fps baseline
[0109] Input frames 120fps@360 s1 s2 s3 s4 s5 s6 s7 s8 s9 Encoded frame e1=s1 e2 = Cs(1,2) e3=Cs(1,3) e4=Cs(1,4) e5=Cs(1,5) e6=s6 e7 = Cs(6,7) e8 = Cs(6,8) e9 = Cs(6,9) Decoding 120fps @360 e1 e2-e1 e3-e2 e4-e3 e5-e4 e6 e7-e6 e8-e7 e9-e8 60fps decoding @360 e2 e4-e2 e5-e4+e6 e8-e6 e10-e8 @180 e1 e3-e2 e5-e4 e7-e6 e9-e8 Decoding 40fps @360 e3 e5-e3+e6 e9-e6 @240 e2 e5-e3 e8-e6 @120 e1 e4-e3 e7-e6 30fps decoding @360 e4 e5-e4+e8 e10-e8+e12 @270 e3 e5-e4+e7 e10-e8+e11 @180 e2 e5-e4+e6 e10-e8 @90 e1 e5-e4 e9-e8 24fps decoding @360 e5 e10 @288 e4 e9 @216 e3 e8 @144 e2 e7 @72 e1 e6
[0110] As shown in Table 7, when the decoding frame rate matches the baseline frame rate (24fps), the decoder can simply select one frame at a desired shutter angle within each group of five frames (e.g., e1 to e5) (e.g., e2 for a 144-degree shutter angle). To decode at different frame rates and specific shutter angles, the decoder must determine how to appropriately combine the decoded frames (e.g., addition or subtraction). For example, to decode at 30fps and a 180-degree shutter angle, the following steps can be followed.
[0111] a) The decoder may consider a virtual encoder transmitting at 120fps and 360 degrees without considering backward compatibility, and if so, from Table 1, the decoder must combine two of four frames to generate an output sequence at the desired frame rate and shutter angle. For example, as shown in Table 4, the sequence includes Ce(1,2) = Avg(s1, s2), Ce(5,6) = Avg(s5, s6), etc., where Avg(s1, s2) can represent the average of frames s1 and s2.
[0112] b) Given that, by definition, encoded frames can be expressed as e1 = s1, e2 = Avg(s1, s2), e3 = Avg(s1, s3), etc., it can be easily derived that the frame sequence of step a) can also be expressed as follows.
[0113] Ce(1,2) = Avg(s1,s2) = e2
[0114] ㆍCe(5,6) = Avg(s5,s6) = Avg(s1,s5) - Avg(s1,s4) + s6 = e5-e4+e6
[0115] ㆍetc.
[0116] As before, an appropriate combination of decoded frames can be pre-calculated and used as an LUT.
[0117] The advantage of the proposed method is that it provides options to both content creators and users. In other words, it enables director / editor choices and user choices. For example, preprocessing content in an encoder can create base frame rates with various shutter angles. Each shutter angle can be assigned a temporal_id value in the range [0, (n_frames - 1)], where n_frames is 120 divided by the base frame rate (e.g., if the base frame rate is 24fps, the temporal_id is in the range [0, 4]). The choice may be made to optimize compression efficiency or for aesthetic reasons. In some use cases, such as over-the-top streaming, multiple bitstreams with different base layers can be encoded and stored, and provided for the user to select.
[0118] In the second example of the disclosed method, multiple backward-compatible frame rates may be supported. Ideally, one may want to be able to decode 24 frames per second to obtain a 24fps base layer, 30 frames per second to obtain a 30fps sequence, 60 frames per second to obtain a 60fps sequence, and so on. If a target shutter angle is not specified, among the allowable shutter angles for the source and target frame rates, a default target shutter angle as close as possible to 180 degrees is recommended. For example, for the values shown in Table 7, the preferred target shutter angles for 120, 60, 40, 30, and 24 fps are 360, 180, 120, 180, and 216 degrees, respectively.
[0119] From the examples above, it can be seen that the choice of method for encoding content can affect the complexity of specific base layer frame rate decoding. One embodiment of the present invention is to adaptively select an encoding method based on a desired base layer frame rate. For example, it may be 24fps for movie content and 60fps for sports.
[0120] Example syntax for the BC embodiments of the present invention is shown in Tables 8 and 9 below.
[0121] In SPS (Table 8), two syntax elements are added: sps_hfr_BC_enabled_flag and sps_base_framerate (when sps_hfr_BC_enabled_flag is set to 1).
[0122] sps_hfr_BC_enabled_flag A value of 1 specifies that backward-compatible high frame rates can be used in coded video sequences (CVS). A value of sps_hfr_BC_enabled_flag of 0 specifies that backward-compatible high frame rates cannot be used in CVS.
[0123] sps_base_framerate Specifies the default frame rate for the current CVS.
[0124] If sps_hfr_BC_enabled_flag is set to 1 in the tile group header, the number_avg_frames syntax is transmitted as a bitstream.
[0125] number_avg_frames specifies the number of frames of the highest frame rate (e.g., 120fps) combined to create the current image as the base frame rate.
[0126] Table 8: Exemplary RBSP syntax for input at various shutter angles
[0127] seq_parameter_set_rbsp( ) { explainer sps_max_sub_layers_minus1 u(3) sps_reserved_zero_5bits u(5) profile_tier_level( sps_max_sub_layers_minus1 ) sps_seq_parameter_set_id ue(v) ... sps_hfr_BC_enabled_flag u(1) if(sps_hfr_BC_enabled_flag) { u(1) sps_base_frame_rate u(9) } ... sps_extension_flag u(1) if( sps_extension_flag ) while( more_rbsp_data( ) ) sps_extension_data_flag u(1) rbsp_trailing_bits( ) }
[0128] Table 9: Exemplary set of image parameters for input at various shutter angles RBSB syntax
[0129] pic_parameter_set_rbsp( ) { explainer pps_pic_parameter_set_id ue(v) pps_seq_parameter_set_id ue(v) ... if( sps_hfr_BC_enabled_flag ) number_avg_frames ... se(v) rbsp_trailing_bits( ) }
[0130] Variation of the second embodiment (fixed frame rate)
[0131] The HEVC (H.265) coding standard (Reference [1]) and the Versatile Video Coding Standard under development (commonly referred to as VVC, see Reference [2]) define pic_struct, a syntactic element that indicates whether an image should be displayed as a frame or in one or more fields, and whether the decoded image should be repeated. For easy reference, a copy of Table D.2, “Interpretation of” of HEVC is provided in the appendix.
[0132] As understood by the inventors, it is important to note that existing pic_struct syntax elements can only support a specific subset of content frame rates when using a fixed frame rate coding container. For example, when using a fixed frame rate container of 60fps, the existing pic_struct syntax can support 30fps using frame doubling when fixed_pic_rate_within_cvs_flag is 1, and can support 24fps using an alternating combination of frame doubling and frame tripping. However, when using a fixed frame rate container of 120fps, the current pic_struct syntax cannot support frame rates of 24fps or 30fps. To mitigate this problem, two new methods have been proposed. One is an extension of the HEVC version, and the other is not.
[0133] Method 1: Non-backward compatible pic_struct
[0134] Since VVC is still under development, the syntax can be designed with maximum freedom. In the examples, it is proposed to remove options for frame doubling and frame tripping from pic_struct, use specific values of pic_struct to indicate arbitrary frame repetition, and add a new syntax element num_frame_repetition_minus2 to specify the number of frames to repeat. Examples of the proposed syntax are described in the following table, where Table 10 shows the changes to Table D.2.3 in HEVC and Table 11 shows the changes to Table D.2 shown in the appendix.
[0135] Table 10: Exemplary Image Timing SEI Message Syntax, Method 1
[0136] pic_timing( payloadSize ) { explainer if(frame_field_info_present_flag) { pic_struct u(4) if( pic_struct == 7) u(4) num_frame_repetition_minus2 u(4) source_scan_type u(2) duplicate_flag u(1) } ...(as the original)
[0137] num_frame_repetition_minus2 + 2 indicates that when fixed_pic_rate_within_cvs_flag is 1, frames should be displayed num_frame_repetition_minus2 + 2 times consecutively at a frame refresh interval equal to DpbOutputElementalInterval[n] given by E-73.
[0138] Table 11: Example of pic_struct modified according to Method 1
[0139] value Instructed image display limits 0 (Progressive) Frame field_seq_flag must be 0 1 Top field field_seq_flag must be 1 2 Bottom field field_seq_flag must be 1 3 The top field, the bottom field, in that order. field_seq_flag must be 0 4 Bottom field, top field, in that order field_seq_flag must be 0 5 Top field, bottom field, repeat top field, in that order field_seq_flag must be 0 6 Bottom field, top field, repeat bottom field, in that order field_seq_flag must be 0 7 Frame repetition field_seq_flag must be 0 fixed_pic_rate_within_cvs_flag must be 1 8 In the output order, the upper field is paired with the previous lower field. field_seq_flag must be 1 9 In the output order, the lower field is paired with the previous upper field. field_seq_flag must be 1 10 In the output order, the top field is paired with the next bottom field. field_seq_flag must be 1 11 In the output order, the bottom field is paired with the next top field. field_seq_flag must be 1
[0140] Method 2: Extended HEVC version of pic_struct
[0141] Since AVC and HEVC decoders are already deployed, it may be desirable to simply extend the existing pic_struct syntax without removing existing options. In one embodiment, a new pic_struct = 13, a "frame repetition extension" value, and a new syntax element, num_frame_repetition_minus4, are added. Examples of the proposed syntax are described in Tables 12 and 13. When the pic_struct value is 0-12, the proposed syntax is identical to the syntax in Table D.2 (see Appendix), so this value is omitted for simplicity.
[0142] Table 12: Exemplary image timing SEI message syntax, Method 2
[0143] pic_timing( payloadSize ) { explainer if(frame_field_info_present_flag) { pic_struct u(4) if( pic_struct == 13) u(4) num_frame_repetition_minus4 u(4) source_scan_type u(2) duplicate_flag u(1) } ...(as the original)
[0144] num_frame_repetition_minus4 + 4 indicates that when fixed_pic_rate_within_cvs_flag is 1, frames should be displayed consecutively num_frame_repetition_minus4 + 4 times at a frame refresh interval equal to DpbOutputElementalInterval[n] given by E-73.
[0145] Table 13: Example of modified pic_struct, Method 2
[0146] value Instructed image display limits 0-12 Same as Table D.2 Same as Table D.2 13 Frame repetition expansion field_seq_flag must be 0 fixed_pic_rate_within_cvs_flag 1 month ago
[0147] In HEVC, the parameter frame_field_info_present_flag exists in the video usability information (VUI), but the syntax elements pic_struct, source_scan_type, and duplicate_flag are in the pic_timing() SEI message. In one embodiment, it is proposed to move all related syntax elements, along with frame_field_info_present_flag, to the VUI. An example of the proposed syntax is shown in Table 14.
[0148] Figure 14: Use the pic_struct user interface VUI user interface
[0149] vui_parameters( ) { 설명자 ... u(1) field_seq_flag u(1) frame_field_info_present_flag u(1) if( frame_field_info_present_flag ) { pic_struct u(4) source_scan_type u(2) duplicate_flag u(1) } ... }
[0150] Alternative signaling of shutter angle information
[0151] When dealing with variable frame rates, it is desirable to identify both the desired frame rate and the desired shutter angle. In previous video coding standards, "Video Usability Information (VUI) provides essential information for the proper display of video content, such as aspect ratio, primary colors, and chroma subsampling. VUI may also provide frame rate information when the fixed pic rate is set to 1, but shutter angle information is not supported. An embodiment allows different shutter angles to be used for different time layers, and the decoder may use shutter angle information to improve the final appearance on the display.
[0152] For example, HEVC supports a temporal sublayer that moves from a higher frame rate to a lower frame rate using frame drop technology by default. The biggest problem with this is that the effective shutter angle decreases with each frame dropped. For example, 60fps can be achieved by dropping every other frame in 120fps video. 30fps can be achieved by dropping 3 out of 4 frames. 24fps can be achieved by dropping 4 out of 5 frames. Assuming a full 360-degree shutter at 120Hz, the shutter angles for 60fps, 30fps, and 24fps with simple frame drop are 180, 90, and 72 degrees, respectively[3]. Experience shows that shutter angles of less than 180 degrees are generally unacceptable, especially when the frame rate is less than 50Hz. By providing shutter angle information, for example, if it is desirable for a display to generate a cinematic effect from a 120Hz video with a reduced shutter angle for each time layer, the final appearance can be improved by applying smart technology.
[0153] In another example, you may want to support different time layers with the same shutter angle (i.e., a 60fps sub-bitstream within a 120fps bitstream). The biggest problem then is that when displaying 120fps video at 120Hz, even and odd frames have different effective shutter angles. If relevant information is available in the display, smart technology can be applied to improve the final appearance. An example of the proposed syntax is shown in Table 15, and the HEVC E.2.1 VUI parameter syntax table (Reference [1]) has been modified to support shutter angle information as mentioned. Note that in another embodiment, the shutter_angle syntax can be expressed alternatively as a ratio of the frame rate to the shutter speed instead of as an absolute angle (see Equation (1)).
[0154] Page 15: Use the VUI graphical application
[0155] vui_parameters( ) { 설명자 ... vui_timing_info_present_flag u(1) if( vui_timing_info_present_flag ) { vui_num_units_in_tick u(32) vui_time_scale u(32) vui_poc_proportional_to_timing_flag u(1) if( vui_poc_proportional_to_timing_flag ) vui_num_ticks_poc_diff_one_minus1 ue(v) vui_hrd_parameters_present_flag u(1) if( vui_hrd_parameters_present_flag ) hrd_parameters( 1, sps_max_sub_layers_minus1 ) } vui_shutter_angle_info_present_flag u(1) if( vui_shutter_angle_info_present_flag ) { fixed_shutter_angle_within_cvs_flag u(1) if(fixed_shutter_angle_with_cvs_flag ) fixed_shutter_angle u(9) else { for( i = 0; i <= sps_max_sub_layers_minus1; i++ ) { sub_layer_shutter_angle[ i ] u(9) } } ... }
[0156] vui_shutter_angle_info_present_flag A value of 1 specifies that shutter angle information exists in the vui_parameters() syntax structure. A value of 0 for vui_shutter_angle_info_present_flag specifies that shutter angle information does not exist in the vui_parameters() syntax structure.
[0157] fixed_shutter_angle_within_cvs_flag A value of 1 specifies that the shutter angle information is the same for all time sublayers within CVS. fixed_shutter_angle_within_cvs_flag A value of 0 specifies that shutter angle information may not be the same for all time sublayers within CVS.
[0158] fixed_shutter_angle Specifies the shutter angle in degrees within the CVS. The value of fixed_shutter_angle is in the range of 0 to 360.
[0159] sub_layer_shutter_angle[ i ] Specifies the shutter angle in degrees when HighestTid is i. The value of sub_layer_shutter_angle[ i ] is in the range of 0 to 360.
[0160] Gradual frame rate update within a coded video sequence (CVS)
[0161] Experiments show that for HDR content displayed on an HDR display, the frame rate must be increased based on the brightness of the content to detect the same motion judder as standard dynamic range (SDR) playback on a 100-nit display. In most standards (AVC, HEVC, VVC, etc.), the video frame rate, for example, as shown in Table 16 below (see Section E.2.1 of Reference [1]), vui_time_scale, vui_num_units_in_tick and elemental_duration_in_tc_minus1 [temporal_id_max] It can be displayed within the VUI (included in SPS) using syntax elements.
[0162] Table 16: VUI syntax elements for representing frame rates in HEVC
[0163] vui_parameters( ) { explainer ... vui_timing_info_present_flag u(1) if( vui_timing_info_present_flag ) { vui_num_units_in_tick u(32) vui_time_scale u(32) vui_poc_proportional_to_timing_flag u(1) if( vui_poc_proportional_to_timing_flag ) vui_num_ticks_poc_diff_one_minus1 ue(v) vui_hrd_parameters_present_flag u(1) if( vui_hrd_parameters_present_flag ) hrd_parameters( 1, sps_max_sub_layers_minus1 ) } ...
[0164] As discussed in reference [1],
[0165] The variable ClockTick is derived as follows and is referred to as the clock tick.
[0166] ClockTick = vui_num_units_in_tick ÷ vui_time_scale
[0167] picture_duration= ClockTick*(elemental_duration_in_tc_minus1[ i ] + 1 )
[0168] frame_rate = 1 / pic_duration
[0169] However, the frame rate may be changed at specific points in time, for example, in HEVC, only at intra-random access point (IRAP) frames or when a new CVS begins. In the case of HDR playback, since the brightness of the image changes frame by frame when there is a fade-in or fade-out case, the frame rate or image duration may need to be changed for all images. To allow frame rate or image duration refresh at any point in time (even on a frame-by-frame basis), in one embodiment, a new SEI message for a "gradual refresh rate" is proposed as shown in Table 17.
[0170] Table 17: Exemplary syntax for supporting progressive refresh frame rates within SEI messages
[0171] gradual_refresh_rate(payloadSize) { explainer num_units_in_tick u(32) time_scale u(32) }
[0172] New syntax num_units_in_tick The definition of is the same as vui_num_units_in_tick, and time_scale The definition of is the same as vui_time_scale.
[0173] num_units_in_ticknum_units_in_tick is the number of time units of the clock operating at the frequency time_scale Hz corresponding to one increment (referred to as a clock tick) of the clock tick counter. num_units_in_tick must be greater than 0. A clock tick in seconds is equal to the quotient of num_units_in_tick divided by time_scale. For example, when the image rate of a video signal is 25Hz, time_scale can be 27,000,000 and num_units_in_tick can be 1,080,000, and consequently, the clock tick can be 0.04 seconds.
[0174] time_scale is the number of time units passing in one second. For example, a time coordinate system measuring time using a 27 MHz clock has a time_scale of 27,000,000. The value of time_scale must be greater than 0.
[0175] The duration of an image using the gradual_refresh_rate SEI message is defined as follows.
[0176] picture_duration = num_units_in_tick ÷ time_scale
[0177] Signaling shutter angle information via SEI message
[0178] As previously discussed, Table 15 provides an exemplary VUI parameter syntax that supports shutter angles. For example, without limitation, Table 18 lists the same syntax elements, but now they are part of an SEI message for shutter angle information. The SEI message is used for example only, and similar messages can be constructed at other layers of high-level syntax, such as Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Slice or Tile Group headers, etc.
[0179] Table 18: Exemplary SEI message syntax for shutter angle information
[0180] shutter_angle_information(payloadSize) { explainer fixed_shutter_angle_within_cvs_flag u(1) if(fixed_shutter_angle_within_cvs_flag) fixed_shutter_angle u(9) else { for( i = 0; i <= sps_max_sub_layers_minus1; i++ ) sub_layer_shutter_angle [ i ] u(9) } }
[0181] The shutter angle is generally expressed as an angle ranging from 0 to 360 degrees. For example, a shutter angle of 180 degrees indicates that the exposure duration is half the frame duration. The shutter angle can be expressed as shutter_angle = frame_rate * 360 * shutter_speed, where shutter_speed is the exposure duration and frame_rate is the reciprocal of the frame duration. The frame_rate for a given temporal sublayer Tid can be represented as num_units_in_tick, time_scale, elemental_duration_in_tc_minus1 [Tid]. For example, when fixed_pic_rate_within_cvs_flag [Tid] is equal to 1:
[0182] frame_rate=time_scale / (num_units_in_tick *(elemental_duration_in_tc_minus1[Tid]+1))
[0183] am.
[0184] In some embodiments, the value of the shutter angle (e.g., fixed_shutter_angle) may not be an integer and may be, for example, 135.75 degrees. To increase precision, in Table 21, u(9) (unsigned 9 bits) may be replaced with u(16) or another appropriate bit depth (e.g., 12 bits, 14 bits, or 16 bits or more).
[0185] In some embodiments, it may be beneficial to express shutter angle information in terms of "clock ticks." In VVC, the variable ClockTick is derived as follows:
[0186] ClockTick = num_units_in_tick ÷ time_scale (8)
[0187] Then, the frame duration and exposure duration can be expressed as multiples or fractions of clock ticks.
[0188] exposure_duration = fN*ClockTick , (9)
[0189] frame_duration = fM*ClockTick (10)
[0190] Here, fN and fM are floating-point values, and fN It is fM.
[0191] then
[0192] shutter_angle frame_rate * 360*shutter_speed =
[0193] =(1 / frame_duration)*360*exposure_duration = (11)
[0194] =(exposure_duration*360) / frame_duration =
[0195] =(fN*ClockTick*360) / (fM*ClockTick) =
[0196] =(fN / fM ) * 360 =(Numerator / Denominator) * 360, and
[0197] Here, the numerator and denominator are integers close to the fN / fM ratio.
[0198] Table 19 is an exemplary SEI message represented by Equation (11). In this example, for a real camera, the shutter angle must be greater than 0.
[0199] Table 19: Exemplary SEI message for shutter angle information based on clock ticks
[0200] shutter_angle_information( payloadSize ) { 설명자 fixed_shutter_angle_within_cvs_flag u(1) if(fixed_shutter_angle_within_cvs_flag) { fixed_shutter_angle_numer_minus1 u(16) fixed_shutter_angle_denom_minus1 u(16) } else { for( i = 0; i <= sps_max_sub_layers_minus1; i++ ) { sub_layer_shutter_angle_numer_minus1 [ i ] u(16) sub_layer_shutter_angle_denom_minus1 [ i ] u(16) } }
[0201] As previously discussed, using u(16) (unsigned 16-bit) for shutter angle precision is illustrated as an example, which is 360 / 2 16= corresponds to a precision of 0.0055. The precision can be adjusted depending on the actual application. For example, if u(8) is used, the precision is 360 / 2 8 = 1.4063.
[0202] Note - The shutter angle is expressed as an angle greater than 0 but less than or equal to 360 degrees. For example, a shutter angle of 180 degrees indicates that the exposure duration is half the frame duration.
[0203] fixed_shutter_angle_within_cvs_flag A value of 1 specifies that the shutter angle information is the same for all time sublayers within the CVS. A value of 0 for fixed_shutter_angle_within_cvs_flag specifies that the shutter angle information may not be the same for all time sublayers within the CVS.
[0204] fixed_shutter_angle_numer_minus1 + 1 specifies the numerator used to derive the shutter angle value. The fixed_shutter_angle_numer_minus1 value must be in the range of 0 to 65535.
[0205] fixed_shutter_angle_demom_minus1 + 1 specifies the denominator used to derive the shutter angle value. The fixed_shutter_angle_demom_minus1 value must be in the range of 0 to 65535.
[0206] The value of fixed_shutter_angle_numer_minus1 must be less than or equal to the value of fixed_shutter_angle_demom_minus1.
[0207] The angle variable shutterAngle is derived as follows:
[0208] shutterAngle = 360*(fixed_shutter_angle_numer_minus1+1) ÷ (fixed_shutter_angle_demom_minus1 + 1))
[0209] sub_layer_shutter_angle_numer_minus1[ i ] + 1 specifies the numerator used to derive the shutter angle value when HighestTid is i. The value of sub_layer_shutter_angle_numer_minus1[ i ] must be in the range of 0 to 65535.
[0210] sub_layer_shutter_angle_demom_minus1[ i ] + 1 specifies the denominator used to derive the shutter angle value when HighestTid is i. The value of sub_layer_shutter_angle_demom_minus1[ i ] must be in the range of 0 to 65535.
[0211] The value of sub_layer_shutter_angle_numer_minus1[ i ] must be less than or equal to the value of sub_layer_shutter_angle_denom_minus1[ i ].
[0212] The angle variable subLayerShutterAngle[ i ] is derived as follows:
[0213] subLayerShutterAngle[ i ] = 360 *
[0214] (sub_layer_shutter_angle_numer_minus1[ i ] + 1) ÷ sub_layer_shutter_angle_demom_minus1[ i ] + 1)
[0215] In another embodiment, the frame duration (e.g., frame_duration) is specified by some other means. For example, in DVB / ATSC, when fixed_pic_rate_within_cvs_flag[ Tid ] is 1:
[0216] frame_rate=time_scale / (num_units_in_tick*(elemental_duration_in_tc_minus1[Tid]+1)),
[0217] frame_duration = 1 / frame_rate
[0218] am.
[0219] The syntax of Table 19 and some subsequent tables assumes that the shutter angle will always be greater than 0. However, shutter angle = 0 may be used to indicate a creative intent where content should be displayed without motion blur. This may be the case for moving graphics, animations, CGI textures, and matte screens. For example, signaling shutter angle = 0 can be useful for mode determination in transcoders (e.g., selecting an edge-preserving transcoding mode) and in displays receiving shutter angle metadata via a CTA interface or 3GPP interface. For example, shutter angle = 0 may be used to indicate a display that should not perform any motion processing, such as denoising or frame interpolation. In these embodiments, the syntax element fixed_shutter_angle_numer_minus1 and sub_layer_shutter_angle_numer_minus1[ i ] is a syntax element fixed_shutter_angle_numer and sub_layer_shutter_angle_numer[ i ] It can be replaced with, and here
[0220] fixed_shutter_angle_numer Specifies the numerator used to derive the shutter angle value. The fixed_shutter_angle_numer value must be in the range of 0 to 65535.
[0221] sub_layer_shutter_angle_numer[ i ] Specifies the numerator used to derive the shutter angle value when HighestTid is i. The value of sub_layer_shutter_angle_numer[ i ] must be in the range of 0 to 65535.
[0222] In another embodiment, fixed_shutter_angle_denom_minus1 and sub_layer_shutter_angle_denom_minus1[ i ] may also be replaced with the syntax elements fixed_shutter_angle_denom and sub_layer_shutter_angle_denom[ i ].
[0223] In one embodiment, as shown in Table 20, the num_units_in_tick and time_scale syntax defined in SPS can be reused by setting general_hrd_parameters_present_flag in VVC to 1. In this scenario, the SEI message can be renamed as the exposure duration SEI message.
[0224] 표 20: 노출 지속시간 시그널링을 위한 예시적인 SEI 메시지
[0225] exposure_duration_information( payloadSize ) { 설명자 fixed_exposure_duration_within_cvs_flag u(1) if(fixed_shutter_angle_within_cvs_flag) { fixed_exposure_duration_numer_minus1 u(16) fixed_exposure_duration_denom_minus1 u(16) } else { for( i = 0; i <= sps_max_sub_layers_minus1; i++ ) { sub_layer_exposure_duration_numer_minus1 [ i ] u(16) sub_layer_exposure_duration_denom_minus1 [ i ] u(16) } }
[0226] fixed_exposure_duration_within_cvs_flag A value of 1 specifies that the effective exposure duration value is the same for all time sublayers within the CVS. A value of fixed_exposure_duration_within_cvs_flag of 0 specifies that the effective exposure duration value may not be the same for all time sublayers within the CVS.
[0227] fixed_exposure_duration_numer_minus1 + 1 specifies the numerator used to derive the exposure duration value. The fixed_exposure_duration_numer_minus1 value must be in the range of 0 to 65535.
[0228] fixed_exposure_duration_demom_minus1 + 1 specifies the denominator used to derive the exposure duration value. The fixed_exposure_duration_demom_minus1 value must be in the range of 0 to 65535.
[0229] The value of fixed_exposure_during_numer_minus1 must be less than or equal to the value of fixed_exposure_duration_demom_minus1.
[0230] The variable fixedExposureDuration is derived as follows:
[0231] fixedExposureDuration =( fixed_exposure_duration_numer_minus1 + 1 ) ÷ ( fixed_exposure_duration_demom_minus1 + 1 ) * ClockTicks
[0232] sub_layer_exposure_duration_numer_minus1[ i ] + 1 specifies the numerator used to derive the exposure duration value when HighestTid is i. The value of sub_layer_exposure_duration_numer_minus1[ i ] must be in the range of 0 to 65535.
[0233] sub_layer_exposure_duration_demom_minus1[ i ] +1 specifies the denominator used to derive the exposure duration value when HighestTid is i. The value of sub_layer_exposure_duration_demom_minus1[ i ] must be in the range of 0 to 65535.
[0234] The value of sub_layer_exposure_duration_numer_minus1[ i ] must be less than or equal to the value of sub_layer_exposure_duration_demom_minus1[ i ].
[0235] The variable subLayerExposureDuration[ i ] for a HigestTid like i is derived as follows:
[0236] subLayerExposureDuration[i]=(sub_layer_exposure_duration_numer_minus1[i]+1)÷(sub_layer_exposure_duration_demom_minus1[i]+1)*ClockTicks
[0237] In another embodiment, as shown in Table 21, clockTick can be explicitly defined by the syntax elements expo_num_units_in_tick and expo_time_scale. The advantage here is that, as in the previous embodiment, it does not depend on whether general_hrd_parameters_present_flag is set to 1 in VVC, and then
[0238] clockTick = expo_num_units_in_tick ÷ expo_time_scale (12)
[0239] am.
[0240] 표 21: 노출 시간 시그널링을 위한 예시적인 SEI 메시지
[0241] exposure_duration_information( payloadSize ) { 설명자 expo_num_units_in_tick u(32) expo_time_scale u(32) fixed_exposure_duration_within_cvs_flag u(1) if(!fixed_exposure_duration_within_cvs_flag) for( i = 0; i <= sps_max_sub_layers_minus1; i++ ) { sub_layer_exposure_duration_numer_minus1 [ i ] u(16) sub_layer_exposure_duration_denom_minus1 [ i ] u(16) } }
[0242] expo_num_units_in_tick is the number of time units of a clock operating at a frequency time_scale Hz corresponding to one increment (referred to as a clock tick) of the clock tick counter. expo_num_units_in_tick must be greater than 0. A clock tick in seconds defined by the variable clockTick is equal to the quotient of expo_num_units_in_tick divided by expo_time_scale.
[0243] expo_time_scale is the number of time units that pass in one second.
[0244] clockTick = expo_num_units_in_tick ÷ expo_time_scale
[0245] Note: The two syntax elements expo_num_units_in_tick and expo_time_scale are defined to measure exposure duration.
[0246] When num_units_in_tick and time_scale exist, clockTick must be less than or equal to ClockTick, which is a requirement for bitstream conformity.
[0247] fixed_exposure_duration_within_cvs_flag A value of 1 specifies that the effective exposure duration value is the same for all time sublayers within the CVS. A value of fixed_exposure_duration_within_cvs_flag of 0 specifies that the effective exposure duration value may not be the same for all time sublayers within the CVS. When fixed_exposure_duration_within_cvs_flag is 1, the variable fixedExposureDuration is set to be the same as clockTick.
[0248] sub_layer_exposure_duration_numer_minus1[ i ] + 1 specifies the numerator used to derive the exposure duration value when HighestTid is i. The value of sub_layer_exposure_duration_numer_minus1[ i ] must be in the range of 0 to 65535.
[0249] sub_layer_exposure_duration_demom_minus1[ i ] + 1 specifies the denominator used to derive the exposure duration value when HighestTid is i. The value of sub_layer_exposure_duration_demom_minus1[ i ] must be in the range of 0 to 65535.
[0250] The value of sub_layer_exposure_duration_numer_minus1[ i ] must be less than or equal to the value of sub_layer_exposure_duration_demom_minus1[ i ].
[0251] The variable subLayerExposureDuration[ i ] for a HigestTid like i is derived as follows:
[0252] subLayerExposureDuration[i]=(sub_layer_exposure_duration_numer_minus1[i]+1)÷(sub_layer_exposure_duration_denom_minus1[i]+1)*clockTick.
[0253] As previously discussed, the syntax parameters sub_layer_exposure_duration_numer_minus1[ i ] and sub_layer_exposure_duration_denom_minus1[ i ] can be replaced with sub_layer_exposure_duration_numer[ i ] and sub_layer_exposure_duration_denom[ i ].
[0254] In another embodiment, as shown in Table 22, the parameter ShutterInterval (i.e., exposure duration) can be defined by the syntax elements sii_num_units_in_shutter_interval and sii_time_scale, where
[0255] ShutterInterval=sii_num_units_in_shutter_interval÷sii_time_scale (13)
[0256] am.
[0257] 표 22: 노출 지속시간(셔터 간격 정보) 시그널링을 위한 예시적인 SEI 메시지
[0258] shutter_interval_information( payloadSize ) { 설명자 sii_num_units_in_shutter_interval u(32) sii_time_scale u(32) fixed_shutter_interval_within_cvs_flag u(1) if(!fixed_shutter_interval_within_cvs_flag) for( i = 0; i <= sps_max_sub_layers_minus1; i++ ) { sub_layer_shutter_interval_numer [ i ] u(16) sub_layer_shutter_interval_denom [ i ] u(16) }
[0259] Shutter Interval Information SEI Message Meaning
[0260] The shutter interval information SEI message indicates the shutter interval for the associated video content prior to encoding and display. For example, in the case of camera capture content, it is the time the image sensor was exposed to generate an image.
[0261] sii_num_units_in_shutter_intervalSpecifies the number of time units of the clock operating at a frequency sii_time_scale Hz corresponding to one increment of the shutter clock tick counter. The shutter interval in seconds defined by the variable ShutterInterval is equal to the quotient of sii_num_units_in_shutter_interval divided by sii_time_scale. For example, when ShutterInterval is 0.04 seconds, sii_time_scale can be 27,000,000 and sii_num_units_in_shutter_interval can be 1,080,000.
[0262] sii_time_scale specifies the number of time units passing in one second. For example, the sii_time_scale of a time coordinate system that measures time using a 27 MHz clock is 27,000,000.
[0263] When the value of sii_time_scale is greater than 0, the value of ShutterInterval is set as follows:
[0264] ShutterInterval = sii_num_units_in_shutter_interval ÷ sii_time_scale
[0265] Otherwise (if the value of sii_time_scale is 0), ShutterInterval should be interpreted as unknown or unspecified.
[0266] Note 1 - A ShutterInterval value of 0 may indicate that the associated video content includes screen capture content, computer-generated content, or other non-camera capture content.
[0267] Note 2 - The inverse of the coded image rate, a ShutterInterval value greater than the coded image interval, may indicate that the coded image rate is greater than the image rate at which the associated video content was generated—for example, when the coded image rate is 120Hz and the image rate of the associated video content prior to encoding and display is 60Hz. The coded interval for a given time sublayer Tid can be represented by ClockTick and elemental_duration_in_tc_minus1[ Tid ]. For example, when fixed_pic_rate_within_cvs_flag[ Tid ] is 1, the image interval for a given time sublayer Tid, defined by the variable PictureInterval[ Tid ], can be specified as follows:
[0268] PictureInterval[Tid]=ClockTick*(elemental_duration_in_tc_minus1[Tid]+1)
[0269] fixed_shutter_interval_within_cvs_flag A value of 1 specifies that the ShutterInterval value is the same for all time sublayers within the CVS. A value of fixed_shutter_interval_within_cvs_flag of 0 specifies that the ShutterInterval value may not be the same for all time sublayers within the CVS.
[0270] sub_layer_shutter_interval_numer[i] specifies the numerator used to derive the sublayer shutter interval in seconds defined by the variable subLayerShutterInterval[i] when HighestTid is equal to i.
[0271] sub_layer_shutter_interval_denom[i] specifies the denominator used to derive the sublayer shutter interval in seconds defined by the variable subLayerShutterInterval[i] when HighestTid is equal to i.
[0272] The value of subLayerShutterInterval[i] for a HigestTid equal to i is derived as follows. When the value of fixed_shutter_interval_within_cvs_flag is 0 and the value of sub_layer_shutter_interval_denom[i] is greater than 0:
[0273] subLayerShutterInterval[i] = ShutterInterval * sub_layer_shutter_interval_numer[i] ÷ sub_layer_shutter_interval_denom[i]
[0274] Otherwise (if sub_layer_shutter_interval_denom[i] is 0), subLayerShutterInterval[i] should be interpreted as unknown or unspecified. When fixed_shutter_interval_within_cvs_flag is not 0,
[0275] subLayerShutterInterval[ i ] = ShutterInterval
[0276] In an alternative embodiment, instead of using a numerator and a denominator to signal the sublayer shutter interval, a single value is used. An example of this syntax is shown in Table 23.
[0277] Table 23: Exemplary SEI messages for shutter interval signaling
[0278] shutter_interval_information( payloadSize ) { explainer sii_num_units_in_shutter_interval u(32) sii_time_scale u(32) fixed_shutter_interval_within_cvs_flag u(1) if(!fixed_shutter_interval_within_cvs_flag) for( i = 0; i <= sps_max_sub_layers_minus1; i++ ) { sub_layer_num_units_in_shutter_interval [ i ] u(32) } }
[0279] Shutter Interval Information SEI Message Meaning
[0280] The shutter interval information SEI message indicates the shutter interval for the associated video content prior to encoding and display. For example, in the case of camera capture content, it is the time the image sensor was exposed to generate an image.
[0281] sii_num_units_in_shutterSpecifies the number of time units of the clock operating at a frequency sii_time_scale Hz corresponding to one increment of the shutter clock tick counter. The shutter interval in seconds defined by the variable ShutterInterval is equal to the quotient of sii_num_units_in_shutter_interval divided by sii_time_scale. For example, when ShutterInterval is 0.04 seconds, sii_time_scale can be 27,000,000 and sii_num_units_in_shutter_interval can be 1,080,000.
[0282] sii_time_scale specifies the number of time units passing in one second. For example, the sii_time_scale of a time coordinate system that measures time using a 27 MHz clock is 27,000,000.
[0283] When the value of sii_time_scale is greater than 0, the value of ShutterInterval is set as follows:
[0284] ShutterInterval = sii_num_units_in_shutter_interval ÷ sii_time_scale
[0285] Otherwise (if the value of sii_time_scale is 0), ShutterInterval should be interpreted as unknown or unspecified.
[0286] Note 1 - A ShutterInterval value of 0 may indicate that the associated video content includes screen capture content, computer-generated content, or other non-camera capture content.
[0287] Note 2 - The inverse of the coded image rate, a ShutterInterval value greater than the coded image interval, may indicate that the coded image rate is greater than the image rate at which the associated video content was generated—for example, when the coded image rate is 120Hz and the image rate of the associated video content prior to encoding and display is 60Hz. The coded image interval for a given time sublayer Tid can be represented by ClockTick and elemental_duration_in_tc_minus1[ Tid ]. For example, when fixed_pic_rate_within_cvs_flag[ Tid ] is 1, the image interval for a given time sublayer Tid, defined by the variable PictureInterval[ Tid ], can be specified as follows:
[0288] PictureInterval[Tid]=ClockTick*(elemental_duration_in_tc_minus1[Tid]+1)
[0289] fixed_shutter_interval_within_cvs_flag A value of 1 specifies that the ShutterInterval value is the same for all time sublayers within the CVS. A value of fixed_shutter_interval_within_cvs_flag of 0 specifies that the ShutterInterval value may not be the same for all time sublayers within the CVS.
[0290] sub_layer_num_units_in_shutter_interval[ i ] is the number of time units of the clock operating at a frequency sii_time_scale Hz corresponding to one increment of the shutter clock tick counter. The sublayer shutter interval in seconds defined by the variable subLayerShutterInterval[ i ] is equal to the quotient of sub_layer_num_units_in_shutter_interval[ i ] divided by sii_time_scale when HighestTid is i.
[0291] When the value of fixed_shutter_interval_within_cvs_flag is 0 and the value of sii_time_scale is greater than 0, the value of subLayerShutterInterval[ i ] is set as follows:
[0292] subLayerShutterInterval[i] = sub_layer_num_units_in_shutter_interval[i] ÷ sii_time_scale
[0293] Otherwise (if sii_time_scale is 0), subLayerShutterInterval[ i ] should be interpreted as unknown or unspecified. If fixed_shutter_interval_within_cvs_flag is not 0,
[0294] subLayerShutterInterval[ i ] = ShutterInterval
[0295] am.
[0296] Table 24 provides a summary of the branch approaches discussed in Tables 18 through 23 to provide SEI messages regarding shutter angle or exposure duration.
[0297] Table 24: Summary of SEI Message Approaches for Signal Shutter Angle Information Signaling
[0298] Table number Key signal elements and dependencies 18 The shutter angle (0 to 360) is explicitly signaled. 19 The shutter angle is expressed as the ratio of the numerator and denominator values and is adjusted to 360 (the clock tick value is implied). 20 The exposure duration is signaled as the ratio of the numerator and denominator values (the clock tick value is implied). 21 Exposure duration is signaled as the ratio of the numerator and denominator values; clock tick values are explicitly signaled as the ratio of the two values. 22 Shutter interval information is signaled as the ratio of two values—the number of clock tick units of exposure and the exposure time scale—; sublayer-related exposure time is signaled as the ratio of two values. 23 Shutter interval information or exposure duration is signaled as the ratio of two values—the number of clock ticks of exposure and the exposure time scale; sublayer-related exposure time is signaled as the number of clock ticks of exposure within each sublayer.
[0299] Variable frame rate signaling
[0300] As discussed in U.S. Provisional Application No. 62 / 883,195 filed on August 6, 2019, it is desirable for decoders to support playback at variable frame rates in many applications. Frame rate adaptation is generally part of the operation in a virtual reference decoder (HRD), as described, for example, in Annex C of Reference [2]. In one embodiment, it is proposed to signal a syntactic element defining the Picture Presentation Time (PPT) as a function of a 90 kHz clock via SEI messaging or other means. This is a kind of repetition of the nominal decoder picture buffer (DPB) output time as specified in the HRD, but now uses the 90 kHz clock tick precision as specified in the MPEG-2 system. The advantages of this SEI message are as follows: a) timing for each frame can be indicated using the PPT SEI message even when the HRD is unavailable; b) conversion between bitstream timing and system timing can be easily performed.
[0301] Table 25 describes an example of a proposed PPT timing message syntax that matches the syntax of the presentation time stamp (PTS) variable used in MPEG-2 transmission (H.222) (Reference [4]).
[0302] Table 25: Exemplary syntax of a video presentation time message
[0303] picture_presentation_time(payloadSize) { explainer PPT u(33) }
[0304] PPT( Video presentation time )
[0305] - The presentation time must be related to the decoding time as follows. PPT is a 33-bit number coded in three separate fields. In a System Goal decoder of presentation unit k for base stream n, it represents the presentation time tpn(k). The value of PPT is specified in units of a period (90 kHz) in which the system clock frequency is divided by 300. The image presentation time is derived from PPT according to the equation below, and
[0306] PPT(k) = ((system_clock_frequency x tp n (k)) / 300)%2 33
[0307] tp here n (k) is the presentation time of the presentation unit Pn(k).
[0308] References
[0309] Each of the references listed herein is incorporated by reference in its entirety.
[0310] [1] High-efficiency video coding , H.265, Series H, Video Coding, ITU,(02 / 2018).
[0311] [2] B. Bross, J. Chen, and S. Liu, " Multipurpose video coding (draft 5),"JVET Output Document, JVET-N1001, v5, uploaded on May 14, 2019.
[0312] [3] C. Carbonara, J. DeFilippis, M. Korpi, " High Frame Rate Capture and Production ,"SMPTE 2015 Annual Technical Conference and Exhibition, October 26-29, 2015.
[0313] [4] Audiovisual service infrastructure - Transmission multiplexing and synchronization , H.222.0, Series H, General Coding of Video and Related Audio Information: System, ITU, 08 / 2018.
[0314] Exemplary computer system implementation
[0315] Embodiments of the present invention may be implemented as computer systems, systems composed of electronic circuits and components, integrated circuit (IC) devices such as microcontrollers, field programmable gate arrays (FPGAs), or other configurable or programmable logic devices (PLDs), discrete-time or digital signal processors (DSPs), application-specific ICs (ASICs), and / or devices comprising one or more of these systems, devices, or components. Computers and / or ICs may perform, control, or execute instructions related to frame rate scalability as described herein. Computers and / or ICs may calculate any various parameters or values related to frame rate scalability as described herein. Image and video embodiments may be implemented as hardware, software, firmware, and various combinations thereof.
[0316] A specific embodiment of the present invention includes a computer processor that executes software instructions that cause the processor to perform the method of the present invention. For example, one or more processors, such as a display, encoder, set-top box, or transcoder, may implement the method related to the frame rate scalability described above by executing software instructions in a program memory accessible to the processor. Embodiments of the present invention may also be provided in the form of a program product. The program product may include any non-transient and tangible medium having a set of computer-readable signals that, when executed by a data processor, cause the data processor to perform the method of the present invention. The program product according to the present invention may exist in various non-transient and tangible forms. The program product may include physical media such as, for example, a magnetic data storage medium including a floppy diskette, a hard disk drive, an optical data storage medium including a CD-ROM or DVD, a ROM, or an electronic data storage medium including a flash RAM. The computer-readable signals of the program product may optionally be compressed or encrypted.
[0317] Where a component (e.g., software module, processor, assembly, device, circuit, etc.) is mentioned above, unless otherwise specified, a reference to such component (including a reference to "means") shall be interpreted to include an equivalent of such component, which is any component that performs the function of the described component (e.g., functionally equivalent), including a component that is not structurally equivalent to the disclosed structure performing the function of the exemplary embodiment described in the present invention.
[0318] Equivalent, extension, alternative, and others
[0319] Exemplary embodiments related to frame rate scalability are described. In the foregoing specification, embodiments of the invention have been described with reference to numerous specific details that may vary depending on the implementation. Accordingly, the sole and exclusive indicator of what the invention is and what the applicant intended to be is the set of claims issued from this application, as in the specific form in which such claims are issued, including subsequent corrections. All definitions expressly provided herein for terms included in such claims govern the meaning of such terms as used in the claims. Accordingly, any limitation, element, characteristic, feature, benefit, or attribute not expressly mentioned in the claims should not limit such claims in any way. Accordingly, the specification and drawings should be considered in an exemplary sense rather than a restrictive sense.
[0320] supplement
[0321] This appendix provides a copy of Table D.2 of the H.265 standard and the associated pic_struct related information (Reference [1]).
[0322] Table D.2 - Analysis of pic_struct
[0323] value Instructed image display limits 0 (Progressive) Frame field_seq_flag must be 0 1 Top field field_seq_flag must be 1 2 Bottom field field_seq_flag must be 1 3 The top field, the bottom field, in that order. field_seq_flag must be 0 4 Bottom field, top field, in that order field_seq_flag must be 0 5 Top field, bottom field, repeat top field, in that order field_seq_flag must be 0 6 Bottom field, top field, repeat bottom field, in that order field_seq_flag must be 0 7 Frame doubling field_seq_flag must be 0 and fixed_pic_rate_within_cvs_flag must be 1. 8 Frame tripleling field_seq_flag must be 0 and fixed_pic_rate_within_cvs_flag must be 1. 9 In the output order, the upper field is paired with the previous lower field. field_seq_flag must be 1 10 In the output order, the lower field is paired with the previous upper field. field_seq_flag must be 1 11 In the output order, the top field is paired with the next bottom field. field_seq_flag must be 1 12 In the output order, the bottom field is paired with the next top field. field_seq_flag must be 1
[0324] Meaning of pic_struct syntax elements
[0325] pic_struct indicates whether the image should be displayed as a frame or as one or more fields, and when fixed_pic_rate_within_cvs_flag is 1, it may indicate the frame doubling or tripleting repetition cycle for displays using a fixed frame refresh interval, such as DpbOutputElementalInterval [n] given by E-73 for frame display. The interpretation of pic_struct is specified in Table D.2. Values of pic_struct not listed in Table D.2 are reserved for future use by ITU-T | ISO / IEC and must not exist in bitstreams compliant with this version of the standard. Decoders must ignore reserved values of pic_struct.
[0326] If present, the bitstream conformance requirement is that the value of pic_struct must be restricted to exactly one of the following conditions:
[0327] - The value of pic_struct is 0, 7, or 8 for all images in CVS.
[0328] - The value of pic_struct is 1, 2, 9, 10, 11, or 12 for all images in CVS.
[0329] - The value of pic_struct is 3, 4, 5, or 6 for all images in CVS.
[0330] When fixed_pic_rate_within_cvs_flag is 1, frame doubling is indicated by pic_struct being 7, which means that the frame must be displayed twice in succession at a frame refresh interval equal to DpbOutputElementalInterval [n] given by E-73, and frame tripleting is indicated by pic_struct being 8, which means that the frame must be displayed three times in succession at a frame refresh interval equal to DpbOutputElementalInterval [n] given by E-73.
[0331] Note 3 - Frame doubling can be used, for example, to display 25Hz progressive scan video on a 50Hz progressive scan display, or 30Hz progressive scan video on a 60Hz progressive scan display. Frame doubling and frame tripleting can be used in an alternating combination to display 24Hz progressive scan video on a 60Hz progressive scan display.
[0332] The nominal vertical and horizontal sampling positions of samples in the upper and lower fields for 4:2:0, 4:2:2, and 4:4:4 chroma formats are shown in Figs. D.1, D.2, and D.3, respectively.
[0333] The association indicator for a field (pic_struct is 9 to 12) provides a hint to associate complementary parity fields together in a frame. The parity of a field can be up or down, and if the parity of one field is up and the parity of the other field is down, the parities of the two fields are considered complementary.
[0334] When frame_field_info_present_flag is 1, the constraint specified in the third column of Table D.2 is a requirement for bitstream conformity.
[0335] Note 4 - When frame_field_info_present_flag is 0, in many cases the default value may be inferred or displayed in a different way. If there is no other indication of the image's intended display type, the decoder should infer that the value of pic_struct is 0 when frame_field_info_present_flag is 0.
Claims
Claim 1 A non-transient processor-readable medium storing instructions for decoding a video bitstream by a processor, wherein the processor receives: a coded video bitstream comprising an image data section containing one or more encodings of video images and a supplemental enhancement information (SEI) message comprising shutter interval parameters, wherein the shutter interval parameters include: a shutter interval time scale parameter indicating the number of time units passing in one second; a shutter interval clock tick parameter indicating the number of time units of a clock operating at the frequency of the shutter interval time scale parameter; and a shutter interval duration flag indicating whether exposure duration information is fixed for all time sublayers within the image data section; And if the shutter interval duration flag indicates that the exposure duration information is fixed, one or more decoded versions of the video images for all the time sublayers within the image data section are decoded by calculating an exposure duration value based on the shutter interval time scale parameter and the shutter interval clock tick parameter; otherwise, the SEI message includes one or more arrays of sublayer parameters, and the values of the one or more arrays of sublayer parameters combined with the shutter interval time scale parameter are used to calculate a corresponding sublayer exposure duration value for displaying a decoded version of the time sublayer of one or more video images for each sublayer; and a non-transient processor-readable medium configured to decode the one or more video images based on the shutter interval parameter. Claim 2 A non-transient processor-readable medium according to claim 1, wherein one or more arrays of the sublayer parameters each comprise an array of sublayer shutter interval clock tick values representing the number of time units of the clock operating at the frequency of the shutter interval time scale parameter. Claim 3 A non-transient processor-readable medium, wherein the exposure duration value is calculated as the quotient obtained by dividing the shutter interval clock tick parameter by the shutter interval time scale parameter in claim 1.
Citation Information
Patent Citations
Transmission device, transmission method, reception device, and reception method
US20160234500A1