Scalable drawing surface size Video coding
Drawing surface size scalability in video coding addresses the challenge of distributing HDR content across different resolutions by dividing images into overlapping regions and using adaptive slice addressing, enhancing decoding efficiency and compatibility with existing devices.
Patent Information
- Application Number
- JP2024228155
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-12-10
- Filing Date
- 2024-12-25
- Publication Date
- 2025-09-08
- Estimated Expiration
- 2040-08-05
AI Technical Summary
Current video coding standards struggle with efficiently distributing and decoding high dynamic range (HDR) content at various resolutions, particularly beyond 4K, due to limitations in existing playback devices and the need for special content versions for different display sizes.
Implementing drawing surface size scalability in video coding by dividing images into overlapping regions and using layer-adaptive slice addressing, with enhanced in-loop filtering and reference picture resampling to support decoding of different resolutions independently, while maintaining coding efficiency and visual quality.
Enables flexible and efficient decoding of HDR content at various resolutions using a single-loop decoder, reducing coding inefficiencies and boundary artifacts, and ensuring compatibility with existing playback devices.
Smart Images

Figure 0007735524000020 
Figure 0007735524000021 
Figure 0007735524000022
Abstract
Description
[Technical Field]
[0001] [Related Applications] This application claims the benefit of priority to U.S. Provisional Application Nos. 62 / 883,195, filed August 6, 2019, 62 / 902,818, filed September 19, 2019, and 62 / 945,931, filed December 10, 2019.
[0002] [Technical field] FIELD OF THE INVENTION This disclosure relates generally to images. More particularly, embodiments of the present invention relate to surface size scalable video coding. [Background technology]
[0003] As used herein, the term "dynamic range (DR)" may relate to the ability of the human visual system (HVS) to perceive a range of intensities (e.g., luminance, luma) within an image, e.g., from darkest gray (black) to brightest white (highlight). In this scene, DR relates to "scene-referred" intensity. DR may also relate to the ability of a display device to properly or approximately render an intensity range of a particular width. In this scene, DR relates to "display-referred" intensity. At any point in the description herein, unless it is explicitly specified that a particular scene has particular importance, it should be presumed that the terms may be used interchangeably, e.g., for either scene.
[0004] As used herein, the term "high dynamic range (HDR)" refers to a DR width that spans 14-15 times or more the magnitude of the human visual system (HVS). In fact, DR, where humans can simultaneously perceive a wide range of intensities, may be omitted in some way in connection with HDR.
[0005] In practice, an image comprises one or more color components (e.g., luma Y and chroma Cb and Cr), each represented with n bits of precision per pixel (e.g., n=8). Using nonlinear luminance coding, images with n≦8 (e.g., color 24-bit JPEG images) are considered to be standard dynamic range (SDR) images, while images with n>8 may be considered to be extended dynamic range images. HDR images may be stored and distributed using high-definition (e.g., 16-bit) floating-point formats such as the OpenEXR file format developed by Industrial Light and Magic.
[0006] Currently, the distribution of video high dynamic range content, such as Dolby Laboratories' Dolby Vision or HDR10 in Blue-Ray®, is limited to 4K resolution (e.g., 4096 x 2160, 3840 x 2160, etc.) and 60 frames per second (fps) due to the capabilities of many playback devices. In future versions, content with up to 8K resolution (e.g., 7680 x 4320) and 120 fps is expected to be available for distribution and playback. To simplify the HDR playback content ecosystem, such as Dolby Vision, future content types should be compatible with existing playback devices. Ideally, content creators should be able to adopt and distribute future HDR technologies (such as HDR10 or Dolby Vision) without having to derive and distribute special versions of their content compatible with existing HDR devices. As recognized by the inventors herein, improved techniques for scalable distribution of video content, particularly HDR content, are desirable.
[0007] The approaches described in this section are approaches that could be pursued, but not necessarily approaches that have been previously conceived or pursued. Thus, unless otherwise indicated, it should not be assumed that any of the approaches described in this section qualify as prior art merely by virtue of their inclusion in this section. Similarly, problems identified with one or more approaches should not be assumed to have been recognized in any prior art under this section unless otherwise indicated. [Brief explanation of the drawings]
[0008] Embodiments of the present invention are illustrated by way of example, and not by way of limitation, in the accompanying figures, in which like reference symbols represent similar elements and in which:
[0009] [Figure 1] 1 illustrates an exemplary process for a video delivery pipeline.
[0010] [Figure 2A] 10 shows an example of a picture sub-region for defining the display area of the input content according to the resolution of the target display.
[0011] [Figure 2B] 2B illustrates an example of boundary-spanning restrictions in tile representations, according to an embodiment, for the picture region of FIG. 2A.
[0012] [Figure 2C] 1 illustrates an example of layer-adaptive slice addressing, according to an embodiment.
[0013] [Figure 3] 1 shows an example of spatial scalability according to the prior art.
[0014] [Figure 4] 10 illustrates an example of drawing surface scalability, according to an embodiment.
[0015] [Figure 5] 1 illustrates examples of base layer and enhancement layer pictures and corresponding adaptive windows, according to an embodiment.
[0016] [Figure 6A] 1 illustrates an exemplary process flow for supporting drawing surface size scalability, according to an embodiment of the present invention. [Figure 6B] 1 illustrates an exemplary process flow for supporting drawing surface size scalability, according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0017] Exemplary embodiments relating to drawing surface size scalability for video coding are described herein. In the following detailed description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of various embodiments of the present invention. It will be apparent, however, that various embodiments of the present invention may be practiced without some of these specific details. In other instances, well-known structures and devices are not described in exhaustive detail to avoid unnecessarily occluding, obscuring, or obscuring embodiments of the present invention.
[0018] <Summary> SUMMARY OF THE INVENTION Exemplary embodiments described herein relate to drawing surface size scalability in video coding. In embodiments, a processor includes: receiving an offset parameter for an adaptation window in a first layer; accessing a reference picture width and a reference picture height of a coding region in a reference layer; receiving offset parameters for a first region of interest (ROI) in the first layer; receiving an offset parameter for a second ROI in the reference layer; calculating a first picture width and a first picture height for a coding region in the first layer based on the offset parameter for the adaptive window; calculating a second picture width and a second picture height for a current ROI in the first layer based on the first picture width, the first picture height, and the offset parameter for the first ROI; calculating a third picture width and a third picture height for a reference ROI in the reference layer based on the reference picture width, the reference picture height, and the offset parameter for the second ROI in the reference layer; calculating a horizontal magnification ratio based on the second picture width and the third picture width; calculating a vertical scaling factor based on the second picture height and the third picture height; Scaling the reference ROI based on the horizontal and vertical scale factors to generate a scaled reference ROI; An output picture is generated based on the current ROI and the scaled reference ROI.
[0019] In a second embodiment, the decoder comprises: receiving an offset parameter for an adaptation window in a first layer; accessing a reference picture width and a reference picture height of a coding region in a reference layer; receiving adjusted offset parameters for a first region of interest (ROI) in the first layer, the adjusted offset parameters combining offset parameters for the first ROI with the offset parameters for the fitting window in the first layer; receiving adjusted offset parameters for a second ROI in the reference layer, the adjusted offset parameters combining offset parameters for the second ROI with offset parameters for a fitting window in the reference layer; calculating a first picture width and a first picture height for a current ROI in the first layer based on the adjusted offset parameters for the first ROI in the first layer; calculating a second picture width and a second picture height for a reference ROI in the reference layer based on the adjusted offset parameter for the second ROI in the reference layer; Calculating a horizontal magnification based on the first picture width and the second picture width; Calculating a vertical scaling factor based on the first picture height and the second picture height; Scaling the reference ROI based on the horizontal and vertical scale factors to generate a scaled reference ROI; An output picture is generated based on the current ROI and the scaled reference ROI.
[0020] <Exemplary Video Streaming Processing Pipeline> 1 illustrates an exemplary process for a conventional video distribution pipeline 100, showing various stages from video capture to video content display. A sequence of video frames 102 is captured or generated using an image generation block 105. The video frames 102 may be captured digitally (e.g., by a digital camera) or generated computer-generated (e.g., using computer animation) to provide video data 107. Alternatively, the video frames 102 may be captured on film using a film camera. The film is converted to a digital format to provide video data 107. In a production stage 110, the video data 107 is edited to provide a video production stream 112.
[0021] The video data in the production stream 112 is then provided to a processor for post-production editing in block 115. The post-production editing in block 115 may include adjusting or changing the color or brightness of specific areas of the image to improve image quality or achieve a particular look according to the video producer's creative intent. This is sometimes referred to as "color timing" or "color grading." Other editing (e.g., scene selection and sequencing, image blackening, addition of computer-generated visual special effects, judder or blur control, frame rate control, etc.) may be performed in block 115 to generate a final version of the production 117 for distribution. During post-production editing 115, the video image is displayed on a reference display 125. Following post-production 115, the final production video data 117 may be delivered to a coding block 120 for downstream delivery to decoding and playback devices such as television sets, set-top boxes, movie theaters, etc. In some embodiments, coding block 120 may include audio and video encoders, such as those defined by ATSC, DVB, DVD, Blu-Ray, and other distribution formats, to generate a coded bit stream 122. At the receiver, coded bit stream 122 is decoded by a decoding unit 130 to generate a decoded signal 132 that represents the same as, or a very close approximation of, signal 117. The receiver may be attached to a target display 140, which may have characteristics quite different from those of the reference display 125. In that case, a display management block 135 may be used to map the dynamic range of decoded signal 132 to the characteristics of target display 140 by generating a display-mapped signal 137.
[0022] <Extensible coding> Scalable coding is already part of many video coding standards, such as MPEG-2, AVC, and HEVC. In embodiments of the present invention, scalable coding is extended to improve performance and flexibility, especially when related to ultra-high resolution HDR content.
[0023] <Extensibility of drawing surface size> As is known in the art, spatial scalability is primarily used to allow decoders to generate content at various resolutions. In embodiments of the present invention, spatial or drawing surface scalability is designed to allow extraction of different regions of an image. For example, a content creator may choose to frame content (i.e., specify a display area) differently for large displays than for small displays. For example, the area framed for display may depend on the size of the screen or the distance from the screen to the viewer. Embodiments of the present invention divide an image into overlapping regions (typically rectangular) and encode them such that a selected number of sub-regions can be decoded for presentation independently of the other sub-regions.
[0024] An example is shown in FIG. 2A , where various regions contain and / or are contained by other regions. As an example, smallest region 215 has 2K resolution and largest region 205 has 8K resolution. The base layer bitstream corresponds to the smallest spatial region, while additional layers in the bitstream correspond to increasingly larger image regions. Thus, a 2K display displays content only within 2K region 215. A 4K display displays content in both 2K and 4K regions (within 210), and an 8K display displays everything within the larger 205. In another example, a 2K display may display a downsampled version of 4K content, and a 4K display may display a downsampled version of 8K content. Ideally, the base layer region can be decoded by legacy devices, while the other regions can be used by future devices to expand the drawing surface size.
[0025] Existing coding standards, such as HEVC, can enable drawing surface scalability using tiles. In a tiled representation, a frame is divided into a set of rectangular, non-overlapping regions. A receiver can decide to decode and display only the set of tiles required for display. In HEVC, coding dependencies between tiles are disabled. Specifically, entropy coding and reconstruction dependencies are not allowed to cross tile boundaries. This includes motion vector prediction, intra-prediction, and context selection. (In-loop filtering is the only exception; it is allowed across boundaries but can be disabled by a flag in the bitstream.) Furthermore, to ensure that the base layer is independently decodable, encoder-side constraints on motion constrained tiles (MCTS) are required, requiring the temporal motion constrained tile set supplemental enhancement information (SEI) message. For bitstream extraction and adaptation purposes, the motion constrained tile set extraction information set SEI message is necessary. The deficiency of tile definition in HEVC, especially with the ability to independently decode, is a loss of coding efficiency.
[0026] In an alternative implementation, HEVC enables surface scalability by using pan-scan rectangle SEI messages to extract a region of interest (ROI). SEI messaging specifies a rectangular extent, but provides information or constraints that allow the ROI to be decoded independently of other regions. Typically, a decoder needs to decode the entire image to obtain the ROI.
[0027] In an embodiment, a novel solution is proposed by improving the HEVC tile concept. For example, given the region shown in FIG. 2A, in an embodiment, independent decoding is required only for region 2K. As shown in FIG. 2B, for tiles within the 2K range, the proposed method allows boundary-spanning prediction (intra / inter) and entropy coding. For 4K, it allows boundary-spanning prediction (intra / inter) and entropy coding from 2K and within 4K. For 8K, it allows boundary-spanning prediction (intra / inter) and entropy coding from 2K and 4K and within 8K. Here, it is proposed to assign layer_id 0 to 2K, layer_id 1 to 4K, and layer_id 2 to 8K. Given a currently decoded layer_id=N, boundary-spanning tile prediction and entropy coding are only allowed from layer_ids less than or equal to N. In this case, the loss of coding efficiency is reduced compared to HEVC-type tiles. An example syntax is shown in Tables 1 and 2 below, where new syntax elements proposed for the proposed VVC (Versatile Video Codec) draft specification of Ref [2] are shown highlighted in grey. Table 1: Example sequence parameter set RBSP syntax that allows for drawing surface resizing [Table 1] Table 2: Example Picture Parameter RBSP Syntax for Drawing Surface Resize [Table 2]
[0028] In SPS (Table 1), the flag sps_canvas_tile_enabled_flag has been added. sps_canvas_tile_enabled_flag equal to 1 specifies that drawing surface tiling is currently enabled in CVS. sps_canvas_tile_enabled_flag equal to 0 specifies that drawing surface tiling is currently not enabled in CVS.
[0029] In PPS (Table 2), a new layer_id information parameter, tile_layer_id[i], specifies the layer id of the ith drawing surface tile. If tile_layer_id values are constrained to be consecutive starting from 0, in an embodiment, the maximum possible value of tile_layer_id may be NumTilesInPic-1, in accordance with the proposed VVC working draft (Ref. [2]).
[0030] Tiles are used as an illustration, but "bricks," slices, and sub-pictures, as defined in VVC and known in the art, can also be constructed similarly.
[0031] <Layer-adaptive slice addressing> As the inventors have recognized, in certain streaming applications the following features may be desirable: (1) When using network abstraction layer (NAL) units in a video coding layer (VCL), a 2K resolution bitstream must be self-contained, and all of its NAL units must have the same value of nuh_layer_id (e.g., layer 0). An additional bitstream enabling 4K resolution must also be self-contained, and its NAL units must have the same value of nuh_layer_id (e.g., layer 1) but different from the nuh_layer_id of the 2K layer. Finally, any additional bitstream enabling 8K resolution must also be self-contained, and its NAL units must have the same value of nuh_layer_id (e.g., layer 2) but different from the nuh_layer_id values of the 2K and 4K layers. Therefore, by analyzing the NAL unit header, it must be possible to extract a bitstream having a target resolution or region of interest (e.g., 2K, 4K, or 8K) using the nuh_layer_id. (2) In non-VCL NAL units, the stream and picture parameter set headers (e.g., SPS, PPS, etc.) must be self-contained for each resolution. (3) At the target resolution, the bitstream extraction process must be able to discard NAL units that are not needed for the target resolution. After bitstream extraction for the target resolution, the bitstream conforms to the single-layer profile, so the decoder can only decode the single-resolution bitstream. Note that 2K, 4K, and 8K resolutions are provided by way of example only, not limitation, and the same method can be applied to any number of different spatial resolutions or regions of interest. For example, one may start with a picture of the highest possible resolution (e.g., res_layer[0]=8K) and define sub-layers or regions of interest at the lowest resolution, where res_layer[i]>res_layer[i+1], for i=1, 2, ..., N-1, and N denotes the total number of layers. Next, one may want to decode a particular sub-layer without first decoding the entire picture. This helps the decoder reduce complexity, save power, etc.
[0032] To meet the above requirements, the following method is proposed in the embodiment: A high-level syntax that reuses the video parameter set (VPS) syntax to specify layer information, including the number of layers, inter-layer dependency relationships, layer representation formats, DPB size, and other information relevant to defining bitstream conformance, including layer sets, output layer sets, profile tier levels, and timing-related parameters. The signal parameter sets (SPS) associated with each different layer, picture resolution, adaptation window, sub-picture, etc., should conform to a distinct resolution (e.g., 2K, 4K, or 8K). In the picture parameter sets (PPS) associated with each different layer, tile, brick, slice, etc., the information should conform to the distinct resolution (e.g., 2K, 4K, or 8K). If the distinct regions are set to be the same in the VCS, the tile / brick / slice information may also be set in the SPS. In the slice header, slice_address should be set to the lowest target resolution containing the slice. As mentioned above, in independent layer decoding, during prediction a layer can only use tile / brick / slice neighboring information from lower layers and / or the same layer.
[0033] VVC (Ref. [2]) defines a slice as an integer number of bricks of a picture, contained exclusively in a single NAL unit. A brick is defined as a row of CTUs in a rectangular area within a particular tile in a picture. A CTU (coding tree unit) is a block of samples with luma and chroma information.
[0034] In our 2K / 4K / 8K example, in an embodiment, the value of slice_address (which indicates the slice address of the slice) may need to have a different slice_address value in the 2K bitstream than in the 4K or 8K bitstream. Therefore, a conversion of slice_address from low resolution to high resolution may be required. Therefore, in an embodiment, such information is provided in the VPS layer.
[0035] FIG. 2C shows an example of a 4K picture with one sub-layer (e.g., for 2K and 4K). Consider picture 220 with nine tiles and three slices. Suppose the gray tiles specify a 2K resolution area. In a 2K bitstream, the slice_address of the gray area must be 0, but in a 4K bitstream, the slice_address of the gray area must be 1. The proposed new syntax allows specifying the slice_address according to the resolution layer. For example, in a VPS, slice_address conversion information may be added to specify that for null_layer_id=1, the slice_address is changed to 1 for 4K. To simplify implementation, embodiments may wish to constrain that the slice information for each resolution must be kept the same in the coded video stream (CVS). An example syntax for a VPS, based on the HEVC video parameter set RBSP syntax, is shown in Table 3 (Section 7.3.2.1 of Ref. [1]). Information can also be conveyed through other layers of the high-level syntax (HLS), such as SPS, PPS, slice headers, and SEI messages. Table 3: Example syntax for a VPS supporting layer-adaptive slice addressing [Table 3] vps_layer_slice_info_present_flag equal to 1 specifies that slice information is present in the VPS() syntax structure. vps_layer_slice_info_present_flag equal to 0 specifies that slice information is not present in the VPS() syntax structure. Specifying num_slices_in_layer_minus1[i] as plus 1 specifies the number of slices in the i-th layer. The value of num_slices_in_layer_minus1[i] is equal to num_slices_in_pic_minus1 in the i-th layer. layer_slice_address[i][j][k] specifies the i-th layer slice address of the target of the k-th slice in the j-th layer.
[0036] By way of example, returning to the example of FIG. 2C, picture 220 includes two layers: Layer 0 (e.g. 2K) has one slice 230 (gray) with slice address 0. Layer 1 (e.g. 4K) has three slices (225, 230, and 235) with slice addresses 0, 1, and 2. When decoding layer 1 (i=1), within layer 0 (j=0), slice 0 (k=0) 230 has slice address 1, and therefore, according to the notation in Table 3, layer_slice_address[1][0][0]=1.
[0037] <SEI message after filtering> When implementing drawing surface scalability using bricks / tiles / slices / subpictures, a potential problem is the implementation of in-loop filtering across boundaries (e.g., deblocking, SAO, ALF). For example, Ref. [4] describes a problem when component windows are coded using independent regions (or subpictures). When coding an entire picture using independent coding regions (which can be implemented, for example, by bricks / tiles / slices / subpictures, etc.), in-loop filtering across independent coding regions can result in drift and boundary artifacts. For drawing surface size applications, it is important to have good visual quality for both high-resolution and low-resolution video. For high-resolution video, boundary artifacts must be reduced. Therefore, in-loop filtering across independent coding regions (e.g., deblocking filters) must be enabled. For low-resolution video, drift and boundary artifacts must also be minimized.
[0038] [4] proposes a solution that stores sub-picture border padding for inter prediction. This approach can be implemented with encoder-only constraints to prohibit motion vectors that use those pixels affected by in-loop filtering. Alternatively, embodiments propose to solve this problem using post-filtering that is communicated to the decoder via SEI messaging.
[0039] First, it is proposed that in-loop filtering across independent coding regions (e.g., slice boundaries in regions 225 and 230) be disabled. Filtering across independent coding regions for the entire picture may be performed in post-filtering processing. Post-filtering may include one or more of deblocking, SAO, ALF, or other filters. Deblocking may be the most important filter to remove ROI boundary artifacts. Typically, the decoder or display / user may have their own choice of what filter should be used. Table 4 shows an example syntax for SEI messaging after ROI-related filtering. Table 4: Example syntax after ROI-related filtering [Table 4] As an example, the syntax parameters may be defined as follows:
[0040] deblocking_enabled_flag equal to 1 specifies that deblocking processing may be applied to independent ROI boundaries of the reconstructed picture for display. deblocking_enabled_flag equal to 0 specifies that deblocking processing shall not be applied to independent ROI boundaries of the reconstructed picture for display.
[0041] sao_enabled_flag equal to 1 specifies that sample adaptive offset (SAO) processing may be applied to independent ROI boundaries of the reconstructed picture for display purposes. sao_enabled_flag equal to 0 specifies that sample adaptive offset (SAO) processing shall not be applied to independent ROI boundaries of the reconstructed picture for display purposes.
[0042] alf_enabled_flag equal to 1 specifies that adaptive loop filter process (ALF) may be applied to the independent ROI boundary of the reconstructed picture for display. alf_enabled_flag equal to 0 specifies that adaptive loop filter process shall not be applied to the independent ROI boundary of the reconstructed picture for display.
[0043] user_defined_filter_enabled_flag equal to 1 specifies that user-defined filtering may be applied to independent ROI boundaries of the reconstructed picture for display purposes. user_defined_filter_enabled_flag equal to 0 specifies that user-defined filtering shall not be applied to independent ROI boundaries of the reconstructed picture for display purposes.
[0044] In an embodiment, the SEI messaging in Table 4 can be simplified by removing one or more of the proposed flags. If all flags are removed, then the presence of the SEI message independent_ROI_across_boundary_filter(payloadSize){} alone indicates to the decoder that a post-filter should be used to mitigate ROI-related boundary artifacts.
[0045] <Region of Interest (ROI) expandability> The latest specification of VVC (Ref. [2]), as discussed in more detail in Ref. [3], describes spatial, quality, and view scalability using a combination of reference picture resampling (RPR) and reference picture selection (RPS). It is based on single-loop decoding and block-based on-the-fly resampling. RPS is used to define the predictive relationship between the base layer and one or more enhancement layers, or more specifically, between the coded pictures assigned to either the base layer or one or more enhancement layers. RPR is used to code a subset of pictures, i.e., spatial enhancement layer pictures, at a higher / lower resolution than the base layer while predicting from a smaller / higher base layer picture. Figure 3 shows an example of spatial scalability under the RPS / RPR framework.
[0046] As shown in FIG. 3, the bitstream includes two streams: a low-resolution (LR) stream 305 (e.g., standard definition (HD), 2K, etc.) and a higher-resolution (HR) stream 310 (e.g., HD, 2K, 4K, 8K, etc.). The arrows indicate possible inter-predictive coding dependencies. For example, HR frame 310-P1 depends on LR frame 305-I. To predict blocks in 310-P1, the decoder needs to upscale 305-I. Similarly, HR frame 310-P2 may depend on HR frame 310-P1 and LR frame 305-P1. Any prediction from LR frame 305-P1 requires spatial upscaling from LR to HR. In other embodiments, the order of the LR and HR frames may be reversed, such that the base layer may be the HR stream and the enhancement layer may be the LR stream. Note that scaling of base layer pictures is not performed explicitly as in SHVC. Instead, it is incorporated in inter-layer motion compensation and calculated on the fly. In Ref. [2], the scalability ratio is implicitly derived using a cropping window.
[0047] ROI scalability is supported in HEVC (Ref. [1]) as part of Annex H "Scalable high efficiency video coding," commonly referred to as SHVC. For example, Section F.7.3.2.3.4 defines syntax elements related to scaled_ref_layer_offset_present_flag[i] and ref_region_offset_present_flag[i]. The related parameters are derived in equations (H-2) through (H-21) and (H-67) through (H-68). VVC does not yet support region of interest (ROI) scalability. As the inventors have recognized, supporting ROI scalability may enable drawing surface size scalability using the same single-loop VVC decoder without requiring scalability extensions as in SHVC.
[0048] As an example, given the three data layers (e.g., 2K, 4K, and 8K) shown in FIG. 2B, FIG. 4 shows an exemplary embodiment of a bitstream that supports drawing surface size scalability using the existing RPS / RPR framework.
[0049] As shown in Figure 4, the bitstream allocates its pictures to three layers or streams: a 2K stream 402, a 4K stream 405, and an 8K stream 410. The arrows indicate examples of possible inter-predictive coding dependencies. For example, a pixel block in an 8K frame 410-P2 may depend on blocks in the 8K frame 410-P1, the 4K frame 405-P2, and the 2K frame 402-P1. Compared with conventional scalability schemes using multiple loop decoders, the proposed ROI scalability scheme has the following advantages and disadvantages: -Advantages: Requires a single loop decoder and does not require any other tools. The decoder does not need to care about how to handle brick / tile / slice / subpicture boundary issues. Disadvantages: To decode the enhancement layer, decoded pictures of both the base layer and the enhancement layer are needed in the decoded picture buffer (DPB), thus requiring a larger DPB size than a non-scalable solution. Since both the base layer and the enhancement layer need to be decoded, it may also require a faster decoder speed.
[0050] The main difference in enabling ROI scalability support between the embodiments proposed for SHVC and VVC is that in SHVC, picture resolution must be the same for all pictures in the same layer. However, in VVC, with RPR support, pictures in the same layer may have different resolutions. For example, in FIG. 3, in SHVC, 305-I, 305-P1, and 305-P2 must have the same spatial resolution. However, in VVC, with RPR support, 305-I, 305-P1, and 305-P2 can have different resolutions. For example, 305-I and 305-P1 can have a first lower resolution (e.g., 720p), and 305-P2 can have a second lower resolution (e.g., 480p). Embodiments of the present invention aim to support both ROI scalability between different layers and RPR for pictures in the same layer. Another main difference is that in SHVC, motion vectors from inter-layer prediction are constrained to be 0. However, in VVC, there is no such constraint and a motion vector can be either 0 or non-0. This reduces the constraints for identifying inter-layer correspondence.
[0051] The VVC coding tree only allows coding of whole coding units (CUs). Most standard formats code picture regions in multiples of 4 or 8 pixels, but non-standard formats may need to be padded at the encoder to fit the minimum CTU size. The same problem existed in HEVC. It was solved by creating a "conformance window" that specifies the picture range that is considered for conformance to the picture output. The conformance window was also added to VVC (Ref. [2]) and is specified by four variables: conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset. For ease of reference, the following section is copied from Ref. [2].
[0052] conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset specify the sample of the picture in the CVS output from the decoding process in terms of the rectangular area specified in the picture coordinates for output. When conformance_window_flag is equal to 0, the values of conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset are inferred to be equal to 0. The adaptive cropping window constrains the luma samples by a horizontal picture coordinate from SubWidthC*conf_win_left_offsettopic_width_in_luma_samples-(SubWidthC*conf_win_right_offset+1), inclusive, and a vertical coordinate from SubHeightC*conf_win_top_offsettopic_height_in_luma_samples-(SubHeightC*conf_win_bottom_offset+1). The value of SubWidthC*(conf_win_left_offset+conf_win_right_offset) must be less than pic_width_in_luma_samples, and the value of SubHeightC*(conf_win_top_offset+conf_win_bottom_offset) must be less than pic_height_in_luma_samples. The variables PicOutputWidthL and PicOutputHeightL are derived as follows:
number
[0053] In a first embodiment, the newly defined ROI offset is combined with the existing offset of the adaptation window to derive a scaling factor. An exemplary embodiment of the proposed syntax elements is shown in Figure 5, which shows a base layer picture 520 and an enhancement layer picture 502 with their corresponding adaptation windows. The following ROI syntax elements are defined: Base Layer (BL)
number
number
[0054] The width 512 and height 514 of the EL picture 502 can be calculated using the adaptive window parameters of the enhancement layer using equations (7-43) and (7-44) above. (For example, pic_width_in_luma_samples may correspond to width 512, and PicOutputWidth may correspond to the width of the dotted window 518.)
[0055] As an example, Table 5 shows how pic_parameter_set_rbsp(), defined in Section 7.3.2.4 of Ref. [2], is modified to support new syntax elements (edits are highlighted in gray). Table 5: Example syntax for supporting ROI extensibility in VVC [Table 5] num_ref_loc_offsets specifies the number of reference layer location offsets that exist in the PPS. The value of num_ref_loc_offsets should be in the range 0 to vps_max_layers_minus1, inclusive. ref_loc_offset_layer_id[i] specifies the nuh_layer_id value for which the i-th reference layer position offset parameter is specified. NOTE: ref_loc_offset_layer_id[i] does not have to be directly between the reference layers, for example when the spatial correspondence of an auxiliary picture to its associated primary picture is specified. The i-th reference layer position offset parameter includes the i-th scaled reference layer offset parameter and the i-th reference region offset parameter. scaled_ref_layer_offset_present_flag[i] equal to 1 specifies that the i-th scaled reference layer offset parameter is present in the PPS. scaled_ref_layer_offset_present_flag[i] equal to 0 specifies that the i-th scaled reference layer offset parameter is not present in the PPS. When not present, the value of scaled_ref_layer_offset_present_flag[i] is inferred to be equal to 0. The i-th scaled reference layer offset parameter specifies the spatial correspondence of the picture referencing this PPS relative to the reference region in the composite picture with nuh_layer_id equal to ref_loc_offset_layer_id[i]. The sum of scaled_ref_layer_left_offset[ref_loc_offset_layer_id[i]] and conf_win_left_offset specifies the horizontal offset between the sample in the current picture that is co-located with the top-left luma sample of the reference region in the decoded picture with nuh_layer_id equal to ref_loc_offset_layer_id[i] and the top-left luma sample of the current picture in units of subWC luma samples, where subWC is equal to SubWidthC of the picture that references this PPS. The value of the sum of scaled_ref_layer_left_offset[ref_loc_offset_layer_id[i]] and conf_win_left_offset shall be inclusively -2. 14 ~2 14Should be in the range -1. When not present, the value of scaled_ref_layer_left_offset[ref_loc_offset_layer_id[i]] is inferred to be equal to 0. The sum of scaled_ref_layer_top_offset[ref_loc_offset_layer_id[i]] and conf_win_top_offset specifies the vertical offset between the sample in the current picture that is co-located with the top-left luma sample of the reference region in the decoded picture with nuh_layer_id equal to ref_loc_offset_layer_id[i] and the top-left luma sample of the current picture in units of subHC luma samples, where subHC equals the SubHeightC of the picture that references this PPS. The value of the sum of scaled_ref_layer_top_offset[ref_loc_offset_layer_id[i]] and conf_win_top_offset shall be inclusively limited to -2. 14 ~2 14 Should be in the range -1. When not present, the value of scaled_ref_layer_top_offset[ref_loc_offset_layer_id[i]] is inferred to be equal to 0. The sum of scaled_ref_layer_right_offset[ref_loc_offset_layer_id[i]] and conf_win_right_offset specifies the horizontal offset between the sample in the current picture that is co-located with the bottom-right luma sample of the reference region in the decoded picture with nuh_layer_id equal to ref_loc_offset_layer_id[i] and the bottom-right luma sample of the current picture in units of subWC luma samples, where subWC is equal to SubWidthC of the picture that references this PPS. The value of the sum of scaled_ref_layer_right_offset[ref_loc_offset_layer_id[i]] and conf_win_right_offset shall be inclusively -2. 14 ~2 14Should be in the range -1. When not present, the value of scaled_ref_layer_right_offset[ref_loc_offset_layer_id[i]] is inferred to be equal to 0. The sum of scaled_ref_layer_bottom_offset[ref_loc_offset_layer_id[i]] and conf_win_bottom_offset specifies the vertical offset between the sample in the current picture that is co-located with the bottom-right luma sample of the reference region in the decoded picture with nuh_layer_id equal to ref_loc_offset_layer_id[i] and the bottom-right luma sample of the current picture in units of subHC luma samples, where subHC equals the SubHeightC of the picture that references this PPS. The value of the sum of scaled_ref_layer_bottom_offset[ref_loc_offset_layer_id[i]] and conf_win_bottom_offset shall be -2 inclusive. 14 ~2 14 Should be in the range -1. When not present, the value of scaled_ref_layer_bottom_offset[ref_loc_offset_layer_id[i]] is inferred to be equal to 0. Let currTopLeftSample, currBotRightSample, colRefRegionTopLeftSample, and colRefRegionBotRightSample be the top-left luma sample of the current picture, the bottom-right luma sample of the current picture, the sample in the current picture that is co-located with the top-left luma sample of the reference region in the decoded picture with nuh_layer_id equal to ref_loc_offset_layer_id[i], and the sample in the current picture that is co-located with the bottom-right luma sample of the reference region in the decoded picture with nuh_layer_id equal to ref_loc_offset_layer_id[i], respectively. When the value of (scaled_ref_layer_left_offset[ref_loc_offset_layer_id[i]]+conf_win_left_offset) is greater than 0, colRefRegionTopLeftSample is located to the right of currTopLeftSample. When the value of (scaled_ref_layer_left_offset[ref_loc_offset_layer_id[i]]+conf_win_left_offset) is less than 0, colRefRegionTopLeftSample is located to the left of currTopLeftSample. When the value of (scaled_ref_layer_top_offset[ref_loc_offset_layer_id[i]]+conf_win_top_offset) is greater than 0, colRefRegionTopLeftSample is located below currTopLeftSample. When the value of (scaled_ref_layer_top_offset[ref_loc_offset_layer_id[i]]+conf_win_top_offset) is less than 0, colRefRegionTopLeftSample is located above currTopLeftSample. When the value of (scaled_ref_layer_right_offset[ref_loc_offset_layer_id[i]]+conf_win_right_offset) is greater than 0, colRefRegionBotRightSample is located to the left of currBotRightSample. When the value of (scaled_ref_layer_right_offset[ref_loc_offset_layer_id[i]]+conf_win_right_offset) is less than 0, colRefRegionTopLeftSample is located to the right of currBotRightSample. When the value of (scaled_ref_layer_bottom_offset[ref_loc_offset_layer_id[i]]+conf_win_bottom_offset) is greater than 0, colRefRegionBotRightSample is located above currBotRightSample. When the value of (scaled_ref_layer_bottom_offset[ref_loc_offset_layer_id[i]]+conf_win_bottom_offset) is less than 0, colRefRegionTopLeftSample is located below currBotRightSample. ref_region_offset_present_flag[i] equal to 1 specifies that the offset parameters of the i-th reference region are present in the PPS. ref_region_offset_present_flag[i] equal to 0 specifies that the offset parameters of the i-th reference region are not present in the PPS. When not present, the value of ref_region_offset_present_flag[i] is inferred to be equal to 0. The offset parameter of the ith reference area specifies the spatial correspondence of the reference area in the decoded picture with nuh_layer_id equal to ref_loc_offset_layer_id[i] relative to the same decoded picture. Let refConfLeftOffset[ref_loc_offset_layer_id[i]], refConfTopOffset[ref_loc_offset_layer_id[i]], refConfRightOffset[ref_loc_offset_layer_id[i]], and refConfBottomOffset[ref_loc_offset_layer_id[i]] be the values of conf_win_left_offset, conf_win_top_offset, conf_win_right_offset, and conf_win_bottom_offset, respectively, of the decoded picture having nuh_layer_id equal to ref_loc_offset_layer_id[i]. The sum of ref_region_left_offset[ref_loc_offset_layer_id[i]] and refConfLeftOffset[ref_loc_offset_layer_id[i]] specifies the horizontal offset between the top-left luma sample of the reference region in the decoded picture with nuh_layer_id equal to ref_loc_offset_layer_id[i] and the top-left luma sample of the same decoded picture in units of subWC luma samples, where subWC is equal to SubWidthC of the layer with nuh_layer_id equal to ref_loc_offset_layer_id[i]. The value of the sum of ref_region_left_offset[ref_loc_offset_layer_id[i]] and refConfLeftOffset[ref_loc_offset_layer_id[i]] is -2 inclusive. 14 ~2 14 Should be in the range -1. When not present, the value of ref_region_left_offset[ref_loc_offset_layer_id[i]] is inferred to be equal to 0. The sum of ref_region_top_offset[ref_loc_offset_layer_id[i]] and refConfTopOffset[ref_loc_offset_layer_id[i]] specifies the vertical offset between the top-left luma sample of the reference region in the decoded picture with nuh_layer_id equal to ref_loc_offset_layer_id[i] and the top-left luma sample of the same decoded picture in units of subHC luma samples, where subHC is equal to SubHeightC of the layer with nuh_layer_id equal to ref_loc_offset_layer_id[i]. The value of the sum of ref_region_top_offset[ref_loc_offset_layer_id[i]] and refConfTopOffset[ref_loc_offset_layer_id[i]] is -2 inclusive. 14 ~2 14Should be in the range -1. When not present, the value of ref_region_top_offset[ref_loc_offset_layer_id[i]] is inferred to be equal to 0. The sum of ref_region_right_offset[ref_loc_offset_layer_id[i]] and refConfRightOffset[ref_loc_offset_layer_id[i]] specifies the horizontal offset between the bottom right luma sample of the reference region in the decoded picture with nuh_layer_id equal to ref_loc_offset_layer_id[i] and the bottom right luma sample of the same decoded picture in units of subWC luma samples, where subWC is equal to SubWidthC of the layer with nuh_layer_id equal to ref_loc_offset_layer_id[i]. The value of the sum of ref_layer_right_offset[ref_loc_offset_layer_id[i]] and refConfRightOffset[ref_loc_offset_layer_id[i]] is -2 inclusive. 14 ~2 14 Should be in the range -1. When not present, the value of ref_region_right_offset[ref_loc_offset_layer_id[i]] is inferred to be equal to 0. The sum of ref_region_bottom_offset[ref_loc_offset_layer_id[i]] and refConfBottomOffset[ref_loc_offset_layer_id[i]] specifies the vertical offset between the bottom right luma sample of the reference region in the decoded picture with nuh_layer_id equal to ref_loc_offset_layer_id[i] and the bottom right luma sample of the same decoded picture in units of subHC luma samples, where subHC is equal to SubHeightC of the layer with nuh_layer_id equal to ref_loc_offset_layer_id[i]. The value of the sum of ref_layer_bottom_offset[ref_loc_offset_layer_id[i]] and refConfBottomOffset[ref_loc_offset_layer_id[i]] is -2 inclusive. 14 ~2 14 Should be in the range -1. When not present, the value of ref_region_bottom_offset[ref_loc_offset_layer_id[i]] is inferred to be equal to 0. Let refPicTopLeftSample, refPicBotRightSample, refRegionTopLeftSample, and refRegionBotRightSample be the top-left luma sample of the decoded picture with nuh_layer_id equal to ref_loc_offset_layer_id[i], the bottom-right luma sample of the decoded picture with nuh_layer_id equal to ref_loc_offset_layer_id[i], the top-left luma sample of the reference area in the decoded picture with nuh_layer_id equal to ref_loc_offset_layer_id[i], and the bottom-right luma sample of the reference area in the decoded picture with nuh_layer_id equal to ref_loc_offset_layer_id[i], respectively. When the value of (ref_region_left_offset[ref_loc_offset_layer_id[i]]+refConfLeftOffset[ref_loc_offset_layer_id[i]]) is greater than 0, refRegionTopLeftSample is located to the right of refPicTopLeftSample. When the value of (ref_region_left_offset[ref_loc_offset_layer_id[i]]+refConfLeftOffset[ref_loc_offset_layer_id[i]]) is less than 0, refRegionTopLeftSample is located to the left of refPicTopLeftSample. When the value of (ref_region_top_offset[ref_loc_offset_layer_id[i]]+refConfTopOffset[ref_loc_offset_layer_id[i]]) is greater than 0, refRegionTopLeftSample is located below refPicTopLeftSample. When the value of (ref_region_top_offset[ref_loc_offset_layer_id[i]]+refConfTopOffset[ref_loc_offset_layer_id[i]]) is less than 0, refRegionTopLeftSample is located above refPicTopLeftSample. When the value of (ref_region_right_offset[ref_loc_offset_layer_id[i]]+refConfRightOffset[ref_loc_offset_layer_id[i]]) is greater than 0, refRegionBotRightSample is located to the left of refPicBotRightSample. When the value of (ref_region_right_offset[ref_loc_offset_layer_id[i]]+refConfRightOffset[ref_loc_offset_layer_id[i]]) is less than 0, refRegionBotRightSample is located to the right of refPicBotRightSample. When the value of (ref_region_bottom_offset[ref_loc_offset_layer_id[i]]+refConfBottomOffset[ref_loc_offset_layer_id[i]]) is greater than 0, refRegionBotRightSample is located above refPicBotRightSample. When the value of (ref_region_bottom_offset[ref_loc_offset_layer_id[i]]+refConfBottomOffset[ref_loc_offset_layer_id[i]]) is less than 0, refRegionBotRightSample is located below refPicBotRightSample.
[0056] Given the proposed syntax elements, in an embodiment, but not by way of limitation, the corresponding VVC section may be modified as follows: Formulas marked (7-xx) and (8-xx) indicate new formulas that need to be inserted into the VVC specification and are renumbered as necessary.
number
number
number
[0057] To support ROI scalability, the specification must be modified as follows:
number
number
number
number
[0058]
number
number
[0059] In another embodiment, unlike SHVC, VVC has no constraints on the size of motion vectors during inter-layer coding, so there is no need to consider the top-left position of the reference layer ROI and the scaled reference layer (current picture) ROI when finding pixel correspondence between ROI regions. Therefore, in all the above equations, references to fRefLeftOffset, fRefTopOffset, fCurLeftOffset and fCurTopOffset can be removed.
[0060] 6A provides an exemplary summary of the above-described process flow. As shown in FIG. 6A, in step 605, the decoder may receive syntax parameters related to an adaptation window (e.g., conf_win_xxx_offset, where xxx is left, top, right, or bottom), a scaled reference layer offset (e.g., scaled_ref_layer_xxx_offset[]) for the current picture, and a reference layer region offset (e.g., ref_region_xxx_offset[]). If not inter-coding (step 610), decoding proceeds as in single-layer decoding; otherwise, in step 615, the decoder calculates adaptation windows for both the reference and current pictures (e.g., using equations (7-43) and (7-44)). If it is not inter-layer coding (step 620), it is still necessary to calculate the RPR scaling factor for inter-prediction for pictures with different resolutions in the same layer, and then decoding proceeds as in single-layer decoding; otherwise (with inter-layer coding), the decoder calculates the scaling factors for the current and reference pictures based on the received offsets (e.g., by calculating hori_scale_fp and vert_scale_fp in equations (8-753) and (8-754)).
[0061] As described above (see, for example, equations (8-x1) to (8-x2)), the decoder needs to calculate the width and height (e.g., fRefWidth and fRefHeight) of the reference ROI by subtracting the left and right offset values (e.g., RefLayerRegionLeftOffset and RefLayerRegionRightOffset) and top and bottom offsets (e.g., RefLayerRegionTopOffset and RefLayerRegionBottomOffset) of the reference layer from the PicOutputWidthL and PicOutputHeightL of the reference layer picture.
[0062] Similarly (see, e.g., equations (8-x3) to (8-x4)), the decoder needs to calculate the width and height (e.g., fCurWidth and fCurHeight) of the current ROI by subtracting the left and right offset values (e.g., ScaledRefLayerLeftOffset and ScaledRefLayerRightOffset) and the top and bottom offsets (e.g., ScaledRefLayerTopOffset and ScaledRefLayerBottomOffset) from the PicOutputWidthL and PicOutputHeightL of the current layer picture. Given these adjusted sizes of the current and reference ROIs, only minimal additional modifications are required (shown in the gray highlighted portion above), and the decoder determines the horizontal and vertical scaling factors (see, e.g., equations (8-753) and (8-754)) as in the existing VVC RPR block (e.g., processing equations (8-755) to (8-766)).
[0063] In equations (8-x5) to (8-x8), adjusted left and top offsets are also calculated to determine the correct positions of the reference and current ROIs relative to the upper left corner of the fitting window for proper pixel correspondence.
[0064] In a second embodiment, the definitions of the ref_region_xxx_offset[] and scaled_ref_region_xxx_offset[] offsets may be redefined to combine both the adaptive window offset and the ROI offset (e.g., by adding them together). For example, in Table 5, scaled_ref_layer_xxx_offset may be replaced by scaled_ref_layer_xxx_offset_sum, which is defined as follows:
number
[0065] Similar definitions can be generated for ref_region_xxx_offset_sum, where xxx = bottom, top, left, and right. As will be explained, these parameters allow the decoder to skip step 615, since the processing in step 615 may be combined with the processing in step 625.
[0066] For example, in Figure 6A, a) In step 615, PicOutputWidthL may be calculated by subtracting the adaptive window left and right offsets from the picture width (see, for example, equation (7-43)). b) Set fCurWidth=PicOutputWidthL. c) Next, in step 625, we adjust fCurWidth (e.g., see (8-x3)) by subtracting the ScaledRefLayer left and right offsets, which from equation (7-xx) are based on the scaled_ref_layer left and right offsets. For example, if we are considering adjusting only the width of the current ROI, in simplified notation (i.e., by ignoring the SubWidthC scaling parameter), we can calculate it as follows: Picture Output Width = Picture Width -(Fit window left offset + Fit window right offset) (2) ROI current width = picture output width - (ROI current left offset + ROI current right offset) (3) By combining equations (2) and (3) together, ROI current width = Picture Width -((Fit window left offset + ROI current left offset) + (Fit window right offset + ROI current right offset)) (4)
[0067] If it is as follows, ROI current left total offset = fitting window left offset + ROI current left offset ROI current right total offset = fitting window right offset + ROI current right offset Equation (4) can be simplified as follows: ROI current width = picture width - ((ROI current left total offset) + (ROI current right total offset)) (5) The definition of the new "sum" offset (eg, ROI current left sum offset) corresponds to the definition of ref_region_left_offset_sum defined in equation (1).
[0068] Therefore, if we redefine the scaled_ref_layer left and right offsets to include the sum of the layer's conf_win_left_offset, as described above, we proceed to blocks 615 and 625 and calculate the width and height of the current and reference ROIs (e.g., equations (2) and (3) can be combined into one (e.g., equation (5))) (e.g., step 630).
[0069] As shown in Figure 6B, steps 615 and 625 can be combined into a single step 630. Compared to Figure 6A, this approach saves some additions, but the revised offsets (e.g., scaled_ref_layer_left_offset_sum[]) are now larger quantities and therefore they require more bits to be coded in the bitstream. Note that the conf_win_xxx_offset values can be different for each layer and can be extracted by the PPS information in each layer.
[0070] In a third embodiment, horizontal and vertical scale factors (e.g., hori_scale_fp and vert_scale_fp) may be explicitly signaled between inter-layer pictures. In such a scenario, the horizontal and vertical scale factors, as well as the top and left offsets, need to be communicated for each layer.
[0071] A similar approach is applicable to embodiments with pictures incorporating multiple ROIs, each with any upsampling and downsampling filters. <References> Each of the references listed herein is incorporated herein by reference in its entirety. [1] High efficiency video coding, H.265, Series H, Coding of moving video, ITU, (02 / 2018) [2] B. Bross, J. Chen, and S. Liu, “Versatile Video Coding (Draft 6),” JVET output document, JVET-O2001, vE, uploaded July 31, 2019 [3] S. Wenger, et al., “AHG8: Spatial scalability using reference picture resampling,” JVET-O0045, JVET Meeting, Gothenburg, SE, July 2019 [4] R. Skupin et al., AHG12: “On filtering of independently coded region,” JVET-O0494 (v3), JVET Meeting, Gothenburg, SE, July 2019 Exemplary Computer System Implementation Embodiments of the present invention may be implemented by a computer system, a system configured in electronic circuits and components, an integrated circuit (IC) device such as a microcontroller, a field programmable gate array (FPGA) or other configurable or programmable logic device (PLD), a discrete time or digital signal processor (DSP), an application specific IC (ASIC), and / or an apparatus including one or more of such systems, devices, or components. The computer and / or IC may execute, control, or perform instructions related to drawing surface size scalability as described herein. The computer and / or IC may calculate any of the various parameters or values related to drawing surface size scalability as described herein. Image and video embodiments may be implemented in hardware, software, firmware, and various combinations thereof.
[0072] A specific implementation of the present invention includes a computer processor executing software instructions that cause the processor to perform the methods of the present invention. For example, one or more processors in a display, encoder, set-top box, transcoder, etc. may implement the methods related to the scalability of drawing surface size described above by executing software instructions in a program memory accessible to the processor. An embodiment of the present invention may be provided in the form of a program product. The program product may include any non-transitory tangible medium that carries a set of computer-readable signals containing instructions that, when executed by a data processor, cause the data processor to perform the methods of the present invention. A program product according to the present invention may be in any of a variety of non-transitory tangible forms. The program product may include physical media such as magnetic data storage media including floppy disks, hard disk drives, optical data storage media including CD-ROMs and DVDs, ROMs, electronic data storage media including flash RAM, etc. The computer-readable signals on the program product may be optically compressed or encrypted.
[0073] Although components (e.g., software modules, processors, components, devices, circuits, etc.) have been referred to above, unless otherwise indicated, references to those components (including references to "means") should be interpreted to include equivalents of those components, any components that perform the functions of the described components (e.g., are functionally equivalent), and components that are not structurally equivalent to the disclosed structures that perform the functions in the illustrated exemplary embodiments of the present invention.
[0074] <Equivalents, Extensions, Alternatives and Miscellaneous> Thus, exemplary embodiments relating to drawing surface size scalability are described. In the foregoing specification, embodiments of the invention have been described with reference to numerous specific details that may vary from implementation to implementation. Accordingly, the sole and exclusive indication of what the invention is, and what Applicant intends to be the invention, is set forth in the claims as issued in particular form hereby, including any subsequent amendments. Any definitions expressly set forth herein for terms contained in such claims shall control the meaning of such terms as used in the claims. Accordingly, any limitation, element, feature, advantage, or attribute not expressly recited in the claims should not in any way limit the scope of the claims. The specification and drawings are, therefore, to be considered illustrative, and not restrictive.
Claims
1. 1. A device for decoding a coding bitstream, the device comprising: an input for receiving a coding bitstream; a processor for processing coded pictures in the coding bitstream; and for a current picture in the coding bitstream, the processor: receiving a current picture width and a current picture height having unsigned integer values; receiving a first offset parameter that determines a rectangular area on the current picture, the first offset parameter including a signed integer value; calculating a current range width and a current range height of the rectangular range on the current picture based on the current picture width, the current picture height, and the first offset parameter; For a reference range, accessing a reference range width, a reference range height, a reference range left offset, and a reference range top offset, and accessing the reference range width and the reference range height means, for a reference picture, Access the reference picture width and reference picture height, receiving a second offset parameter that determines a rectangular area within the reference picture, the second offset parameter comprising a signed integer value; calculating the reference range width and the reference range height of the rectangular range within the reference picture based on the reference picture width, the reference picture height, and the second offset parameter; Calculating the reference range left offset and the reference range top offset based on the second offset parameter; This includes: Calculate the horizontal scale factor (hori_scale_fp) and vertical scale factor (vert_scale_fp) as follows: [Equation 1] fRefWidth represents the reference range width, fCurWidth represents the current range width, fRefHeight represents the reference range height, and fCurHeight represents the current range height; Calculating a left offset adjustment and a top offset adjustment for the current range based on the first offset parameter; The apparatus performs motion compensation based on the horizontal magnification, the vertical magnification, the left offset adjustment, the top offset adjustment, the reference range left offset, and the reference range top offset.
2. The apparatus of claim 1 , wherein the first offset parameters include a left offset, a top offset, a right offset, and a bottom offset.
3. One or more of the left offset, the top offset, the right offset, or the bottom offset is −2 14 From 2 14 3. The device of claim 2, having a value between
Citation Information
Patent Citations
Apparatus, a method and a computer program for video coding and decoding
US20150195573A1
Method, an apparatus and a computer program product for coding a 360-degree panoramic video
US20170085917A1
Image decoder, image encoder, and encoded data converter
WO2015053287A1