Method and apparatus for video coding with scalable canvas size

Through the canvas size scalability technology, the scaling factor is calculated using the consistency window and reference layer offset parameters, which solves the compatibility problem of high dynamic range content among devices with different resolutions, and realizes the effective decoding and display of a single coded stream on multiple devices.

CN114208202BActive Publication Date: 2025-05-13DOLBY LABORATORIES LICENSING CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080055974.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-12-10
Filing Date
2020-08-05
Publication Date
2025-05-13
Estimated Expiration
2040-08-05

AI Technical Summary

Technical Problem

Existing video encoding technologies are difficult to effectively support the compatibility and scalability of high dynamic range content among devices with different resolutions, resulting in content producers having to make multiple versions for different devices, increasing workload and complexity.

Method used

By introducing canvas size scalability technology, the scalability factor is calculated using the consistency window offset parameters and reference layer offset parameters to generate output pictures that are suitable for different resolutions, and support the decoding and display of a single coded stream on different devices.

Benefits of technology

The video content compatibility between devices with different resolutions is achieved, the workload of content producers is reduced, encoding efficiency and decoder flexibility is improved, and the display complexity of the device for high dynamic range content is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114208202B_ABST
    Figure CN114208202B_ABST
Patent Text Reader

Abstract

The present disclosure relates to methods and apparatus for video coding with scalable canvas size. Methods and systems for implementing canvas size scalability across the same or different bitstream layers of a video coding bitstream are described. A consistency window, a reference region of interest (ROI) in a reference layer, and offset parameters of a current ROI in a current layer are received. The width and height of the current ROI and the reference ROI are calculated based on the offset parameters, and the width and height are used to generate a width scaling factor and a height scaling factor, which are used by a reference picture resampling unit to generate an output picture based on the current ROI and the reference ROI.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the benefit of priority to U.S. Provisional Application Serial No. 62 / 883,195 filed on August 6, 2019, Serial No. 62 / 902,818 filed on September 19, 2019, and Serial No. 62 / 945,931 filed on December 10, 2019. Technical Field

[0003] This document relates generally to images. More specifically, embodiments of the present invention relate to canvas size scalable video coding. Background Art

[0004] As used herein, the term "dynamic range (DR)" may relate to the ability of the human visual system (HVS) to perceive a range of intensities (e.g., luminance, luma) in an image, such as from the darkest gray (black) to the brightest white (highlight). In this sense, DR is related to "scene-referred" intensities. DR may also relate to the ability of a display device to fully or approximately render a range of intensities of a particular breadth. In this sense, DR is related to "display-referred" intensities. Unless a particular sense is explicitly specified at any point in the description herein to have a particular meaning, it should be inferred that the terms may be used in either sense, such as interchangeably.

[0005] As used herein, the term "high dynamic range (HDR)" refers to a DR width that spans 14 to 15 orders of magnitude across the human visual system (HVS). In practice, DR, where humans can simultaneously perceive a wide range of intensities, may be slightly truncated relative to HDR.

[0006] In practice, an image includes one or more color components (e.g., luminance Y and chrominance Cb and Cr), where each color component is represented by n bits of precision per pixel (e.g., n=8). Using linear luminance encoding, images where n≤8 (e.g., color 24-bit JPEG images) are considered standard dynamic range (SDR) images, while images where n>8 can be considered enhanced dynamic range images. HDR images can also be stored and distributed using high-precision (e.g., 16-bit) floating-point formats, such as the OpenEXR file format developed by Industrial Light and Magic.

[0007] Currently, the distribution of video high dynamic range content (such as Dolby Vision from Dolby Laboratories or HDR10 in Blu-ray) is limited by the performance of many playback devices to 4K resolution (e.g., 4096×2160 or 3840×2160, etc.) and 60 frames per second (fps). In future versions, it is expected that content with up to 8K resolution (e.g., 7680×4320) and 120fps will be available for distribution and playback. It is expected that future content types will be compatible with existing playback devices to simplify the HDR playback content ecosystem, such as Dolby Vision. Ideally, content producers should be able to adopt and distribute future HDR technologies without having to also obtain and distribute special versions of content that are compatible with existing HDR devices (such as HDR10 or Dolby Vision). As understood by the inventors herein, improved techniques for scalable distribution of video content, especially HDR content, are desired.

[0008] The approaches described in this section are approaches that could be pursued, but not necessarily approaches that have been previously conceived or pursued. Therefore, unless otherwise indicated, it should not be assumed that any approach described in this section qualifies as prior art simply by virtue of its inclusion in this section. Similarly, unless otherwise indicated, issues identified with respect to one or more approaches should not be assumed to qualify as prior art based on this section. Summary of the invention

[0009] Example embodiments described herein relate to canvas size scalability in video coding.In an embodiment, a processor receives an offset parameter of a consistency window in a first layer;

[0010] Access to the reference picture width and reference picture height of the coding region in the reference layer;

[0011] receiving an offset parameter of a first region of interest (ROI) in a first layer;

[0012] receiving an offset parameter of a second ROI in the reference layer;

[0013] Calculating a first picture width and a first picture height of a coding region in a first layer based on an offset parameter of the consistency window;

[0014] Calculating a second picture width and a second picture height of a current ROI in the first layer based on the first picture width, the first picture height, and an offset parameter of the first ROI in the first layer;

[0015] calculating a third picture width and a third picture height of the reference ROI in the reference layer based on the reference picture width, the reference picture height, and an offset parameter of the second ROI in the reference layer;

[0016] Calculating a horizontal scaling factor based on the second image width and the third image width;

[0017] Calculating a vertical scaling factor based on the second image height and the third image height;

[0018] scaling the reference ROI based on the horizontal scaling factor and the vertical scaling factor to generate a scaled reference ROI; and

[0019] Generate an output image based on the current ROI and the scaled reference ROI.

[0020] In a second embodiment, the decoder:

[0021] receiving an offset parameter of a consistency window in a first layer;

[0022] Access to the reference picture width and reference picture height of the coding region in the reference layer;

[0023] receiving adjusted offset parameters for a first region of interest (ROI) in the first layer, wherein the adjusted offset parameters combine the offset parameters of the first ROI with the offset parameters of the consistency window in the first layer;

[0024] receiving adjusted offset parameters for a second ROI in the reference layer, wherein the adjusted offset parameters combine the offset parameters for the second ROI with the offset parameters for the consistency window in the reference layer;

[0025] calculating a first picture width and a first picture height of a current ROI in the first layer based on the adjusted offset parameter of the first ROI in the first layer;

[0026] calculating a second picture width and a second picture height of the reference ROI in the reference layer based on the adjusted offset parameter of the second ROI in the reference layer;

[0027] Calculating a horizontal scaling factor based on a first picture width and a second picture width;

[0028] Calculating a vertical scaling factor based on a first picture height and a second picture height;

[0029] scaling the reference ROI based on the horizontal scaling factor and the vertical scaling factor to generate a scaled reference ROI; and

[0030] Generate output image based on current ROI and scaled reference ROI BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Embodiments of the invention are illustrated by way of example and not limitation in the accompanying drawings and in which like reference numerals refer to similar elements and in which:

[0032] Figure 1 Depicts an example process of a video transmission pipeline;

[0033] Figure 2A depicts an example of a picture sub-region for defining an input content viewing area according to a resolution of a target display;

[0034] against Figure 2A The picture area, Figure 2B depicts an example of cross-boundary restrictions in a tile representation according to an embodiment;

[0035] Figure 2C depicts an example of layer adaptive stripe addressing according to an embodiment;

[0036] Figure 3 An example of spatial scalability according to the prior art is depicted;

[0037] Figure 4 depicts an example of canvas scalability according to an embodiment;

[0038] Figure 5 depicts examples of base layer pictures and enhancement layer pictures and corresponding consistency windows according to an embodiment; and

[0039] Fig. 6A and Figure 6B An example process flow for supporting canvas size scalability according to an embodiment of the present invention is depicted. DETAILED DESCRIPTION

[0040] Example embodiments related to canvas size scalability for video coding are described herein. In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of various embodiments of the present invention. However, it is apparent that various embodiments of the present invention may be practiced without these specific details. In other cases, well-known structures and devices are not described in detail in order to avoid unnecessarily obscuring, obscuring, or confusing embodiments of the present invention.

[0041] Example Video Transmission Processing Pipeline

[0042] Figure 1 An example process of a conventional video transmission pipeline (100) is depicted, showing various stages from video capture to video content display. An image generation block (105) is used to capture or generate a sequence of video frames (102). The video frames (102) may be digitally captured (e.g., by a digital camera) or generated by a computer (e.g., using computer animation) to provide video data (107). Alternatively, the video frames (102) may be captured on film by a film camera. The film is converted to a digital format to provide video data (107). In a production stage (110), the video data (107) is edited to provide a video production stream (112).

[0043] The video data (112) of the production stream is then provided to a processor at block (115) for post-production editing. Block (115) post-production editing may include adjusting or modifying the color or brightness in specific areas of the image to enhance the image quality or achieve a specific look of the image in accordance with the creative intent of the video creator. This is sometimes referred to as "color timing" or "color grading." Other edits (e.g., scene selection and sequencing, image cropping, adding computer-generated visual effects, jitter or blur control, frame rate control, etc.) may be performed at block (115) to produce a final version (117) of the work for distribution. During post-production editing (115), the video image is viewed on a reference display (125).

[0044] After post-production (115), the video data of the final product (117) can be transmitted to an encoding block (120) for downstream transmission to decoding and playback devices such as televisions, set-top boxes, movie theaters, etc. In some embodiments, the encoding block (120) may include an audio encoder and a video encoder such as those defined by ATSC, DVB, DVD, Blu-ray and other transmission formats to generate an encoded bitstream (122). In the receiver, the encoded bitstream (122) is decoded by a decoding unit (130) to generate a decoded signal (132) representing an identical or nearly similar version of the signal (117). The receiver can be attached to a target display (140), which can have completely different characteristics from the reference display (125). In this case, a display management block (135) can be used to map the dynamic range of the decoded signal (132) to the characteristics of the target display (140) by generating a display mapping signal (137).

[0045] Scalable Coding

[0046] Scalable coding is already part of many video coding standards such as MPEG-2, AVC and HEVC. In an embodiment of the present invention, scalable coding is extended to improve performance and flexibility, especially when scalable coding involves very high resolution HDR content.

[0047] Canvas size scalability

[0048] As is known in the art, spatial scalability is primarily used to allow decoders to create content at various resolutions. In an embodiment of the present invention, spatial scalability or canvas scalability is designed to allow different areas of an image to be extracted. For example, a content producer may choose to frame content differently (i.e., specify viewing areas) for large and small displays. For example, the framed area to be displayed may depend on the size of the screen or the distance from the screen to the viewer. An embodiment of the present invention divides an image into overlapping areas (usually rectangular) and encodes these areas in a manner that a selected number of sub-areas can be decoded for presentation independently of other sub-areas.

[0049] Figure 2A An example is shown in which various regions cover other regions and / or are covered by other regions. As an example, the smallest region (215) has a 2K resolution and the largest region (205) has an 8K resolution. The base layer bitstream corresponds to the smallest spatial region, while the multiple additional layers in the bitstream correspond to increasingly larger image regions. Therefore, a 2K display will only display content within the 2K region (215). A 4K display will display content in both the 2K region and the 4K region (the region within 210), while an 8K display will display everything within the 205 border. In another example, a 2K display can display a downsampled version of the 4K content, and a 4K display can display a downsampled version of the 8K content. Ideally, the base layer region can be decoded by legacy devices, while other regions can be used by future devices to expand the canvas size.

[0050] Existing coding standards such as HEVC can achieve canvas scalability using tiles. In tile representation, the frame is divided into a set of rectangular non-overlapping areas. The receiver can decide to decode and display only the set of tiles required for display. In HEVC, coding dependencies between tiles are disabled. Specifically, entropy coding and reconstruction dependencies are not allowed to cross tile boundaries. This includes motion vector prediction, intra-frame prediction, and context selection. (Loop filtering is the only exception that is allowed to cross boundaries, but it can be disabled by a flag in the bitstream.) In addition, in order to allow independent decoding of the base layer, encoder-side constraints on temporal motion constraint tiles (MCTS) are required, and message delivery of supplementary enhancement information (SEI) for temporal motion constraint tile sets is required. For the purpose of bitstream extraction and consistency, SEI messages for extraction information sets of motion constraint tile sets are required. The disadvantage of tile definition in HEVC, especially in terms of independent decoding capability, is the loss of coding efficiency.

[0051] In an alternative implementation, HEVC allows the use of full scan rectangular SEI messages to extract regions of interest (ROIs) for canvas scalability. SEI messages pass the specified rectangular region, but do not provide information or constraints that enable the ROI to be decoded independently of other regions. Typically, the decoder needs to decode the entire image to obtain the ROI.

[0052] In an embodiment, a novel solution is proposed by improving the HEVC tile concept. Figure 2A In the region depicted in FIG. 1 , in an embodiment, only region 2K (215) needs to be decoded independently. Figure 2B As illustrated, for tiles within 2K, the proposed method allows cross-border prediction (intra / inter) and entropy coding. 4K tiles allow cross-border prediction (intra / inter) and entropy coding from 2K and within 4K. 8K tiles allow cross-border prediction (intra / inter) and entropy coding from 2K and 4K and within 8K. It is recommended to assign layer_id0 to 2K, layer_id 1 to 4K, and layer_id 2 to 8K. Given the current decoding layer_id=N, tile cross-border prediction and entropy coding are only allowed from layer_id less than or equal to N. In this case, the loss in coding efficiency is reduced compared to HEVC-style tiles. Example syntax is shown in Tables 1 and 2 below, where new syntax elements proposed based on the draft Versatile Video Coding (VVC) specification in reference [2] are used. Font description.

[0053] Table 1: Example sequence parameter set RBSP syntax for enabling canvas resizing

[0054]

[0055] Table 2: Example picture parameter RBSP syntax for canvas resizing

[0056]

[0057] In SPS (Table 1), the flag sps_canvas_tile_enabled_flag is added.

[0058] sps_canvas_tile_enabled_flag equal to 1 specifies that canvas tiles are enabled in the current CVS. sps_canvas_tile_enabled_flag equal to 0 specifies that canvas tiles are not enabled in the current CVS.

[0059] In the PPS (Table 2), the new layer_id information parameter tile_layer_id[i] specifies the layer id of the i-th canvas tile. If tile_layer_id is restricted to be a continuous value starting from 0, then in an embodiment, according to the proposed VVC working draft (reference [2]), the maximum possible value of tile_layer_id will be NumTilesInPic-1.

[0060] Although tiles are used as an illustration, "bricks", slices, and sub-pictures may also be configured in a similar manner, as defined in VVC and known in the art.

[0061] Layer Adaptive Stripe Addressing

[0062] As appreciated by the inventors, in certain streaming applications, the following features may be desirable:

[0063] 1) When using Network Abstraction Layer (NAL) units in the Video Coding Layer (VCL), the bitstream for 2K resolution should be self-contained and all its NAL units must have the same nuh_layer_id value (e.g., layer 0). Additional bitstreams for enabling 4K resolution should also be self-contained and their NAL units must have the same nuh_layer_id value (e.g., layer 1), but different from the nuh_layer_id of the 2K layer. Finally, any additional bitstreams for enabling 8K resolution should also be self-contained and their NAL units must have the same nuh_layer_id value (e.g., layer 2), but different from the nuh_layer_id values ​​of the 2K and 4K layers. Therefore, by parsing the NAL unit header using the nuh_layer_id, it should be possible to extract a bitstream with the target resolution or (multiple) regions of interest (e.g., 2K, 4K, or 8K).

[0064] 2) For non-VCL NAL units, stream and picture parameter set headers (e.g., SPS, PPS, etc.) should be self-contained for each resolution.

[0065] 3) For the target resolution, the bitstream extraction process should be able to discard NAL units that are not needed for the target resolution. After the bitstream is extracted for the target resolution, the bitstream will conform to the single-layer profile, so the decoder can simply decode the single resolution bitstream.

[0066] Note that 2K, 4K, and 8K resolutions are provided by way of example only and not limitation, and the same approach should be applicable to any number of different spatial resolutions or regions of interest. For example, starting with a picture at the highest possible resolution (e.g., res_layer[0]=8K), sub-layers or regions of interest can be defined at lower resolutions, where res_layer[i]>res_layer[i+1] for i=1, 2, ..., N-1, where N represents the total number of layers. It is then desirable to decode a specific sub-layer without first decoding the entire picture. This can help the decoder reduce complexity, thereby saving power, etc.

[0067] In order to meet the above requirements, in an embodiment, the following method is proposed:

[0068] - In the high-level syntax, the video parameter set (VPS) syntax can be reused to specify layer information, including the number of layers, dependencies between layers, representation formats of layers, DPB sizes, and other information related to defining the consistency of the bitstream, including layer sets, output layer sets, profile layer levels, and timing-related parameters.

[0069] - For the signal parameter set (SPS) associated with each different layer, picture resolution, consistency window, sub-picture, etc. should conform to different resolutions (e.g., 2K, 4K or 8K).

[0070] - For the picture parameter set (PPS) associated with each different layer, tile, block, slice, etc., the information should be consistent with the different resolutions (e.g. 2K, 4K or 8K). The tile / block / slice information can also be set in the SPS if different regions are set to be the same in the CVS.

[0071] - For the slice header, slice_address should be set to the lowest target resolution that includes the slice.

[0072] - As mentioned before, for independent layer decoding, during prediction a layer can only use tile / block / slice neighbor information from lower layers and / or the same layer.

[0073] VVC (reference [2]) defines a slice as an integer number of "blocks in a picture" that are contained only in a single NAL unit. A block is defined as a rectangular area of ​​CTU rows within a specific tile in a picture. A CTU (Coding Tree Unit) is a block of samples with luma and chroma information.

[0074] In the 2K / 4K / 8K example, in an embodiment, the slice_address value (which represents the slice address of the slice) of the 2K bitstream may need to have a different slice_address value than the 4K bitstream or 8K bitstream. Therefore, the slice_address may need to be converted from a lower resolution to a higher resolution. Therefore, in an embodiment, this information is provided at the VPS layer.

[0075] Figure 2C Such an example of a 4K picture with one sublayer is depicted (e.g., 2K and 4K cases). Consider a picture (220) with nine tiles and three slices. Let the gray tiles specify an area of ​​2K resolution. For a 2K bitstream, the slice_address of the gray area should be 0; however, for a 4K bitstream, the slice_address of the gray area should be 1. The proposed new syntax allows slice_address to be specified based on the resolution layer. For example, in the VPS, for nul_layer_id=1, slice_address conversion information can be added to specify that the slice_address is modified to 1 in the 4K case. In order to simplify the implementation, in an embodiment, it may be desirable to restrict the slice information for each resolution to remain the same within the coded video stream (CVS). Table 3 shows an example syntax in a VPS based on the HEVC video parameter set RBSP syntax (Section 7.3.2.1 in reference [1]). Information may also be carried through other high-level syntax (HLS) layers such as SPS, PPS, slice headers and SEI messages.

[0076] Table 3: Example syntax for supporting layer adaptive stripe addressing in VPS

[0077]

[0078] vps_layer_slice_info_present_flag equal to 1 specifies that slice information is present in the VPS() syntax structure. vps_layer_slice_info_present_flag equal to 0 specifies that slice information is not present in the VPS() syntax structure.

[0079] num_slices_in_layer_minus1[i] specifies the number of slices in the i-th layer plus 1. The value of num_slices_in_layer_minus1[i] is equal to num_slices_in_pic_minus1 in the i-th layer.

[0080] layer_slice_address[i][j][k] specifies the target i-th layer slice address for the k-th slice in the j-th layer.

[0081] As an example, return to Figure 2C For example, picture 220 includes two layers:

[0082] In layer 0 (e.g. 2K), there is one (grey) stripe 230 with stripe address 0

[0083] In layer 1 (say 4K), there are three stripes (225, 230, and 235) with stripe addresses 0, 1, and 2

[0084] When decoding layer 1 (i=1), slice 0 (k=0) (230) in layer 0 (j=0) should have slice address 1, so following the notes in Table 3, layer_slice_address[1][0][0]=1.

[0085] Post-filtering SEI messaging

[0086] When canvas scalability is implemented using blocks / tiles / strips / sub-pictures, a potential problem is the cross-border implementation of loop filtering (e.g., deblocking, SAO, ALF). As an example, reference [4] describes the problem when the synthesis window is encoded using independent regions (or sub-pictures). When the full picture is encoded using independent coded regions (as an example, which can be implemented by blocks / tiles / strips / sub-pictures, etc.), loop filtering across independent coded regions can cause drift and boundary artifacts. For canvas size applications, it is very important that both high-resolution video and low-resolution video have good visual quality. For high-resolution video, boundary artifacts should be mitigated, so loop filtering (especially deblocking filter) across independent coded regions should be enabled. For low-resolution video, drift and boundary artifacts should also be minimized.

[0087] In reference [4], a solution is proposed to extend the sub-picture boundary padding for inter-frame prediction. The method can be implemented by the following encoder-only constraint, that is, prohibiting the use of motion vectors of those pixels affected by loop filtering. Alternatively, in an embodiment, it is proposed to use post-filtering to solve this problem, which is transmitted to the decoder through SEI messaging.

[0088] First, it is proposed to disable loop filtering across independent coding regions (e.g., slice boundaries in regions 225 and 230). Filtering of independent coding regions of the full picture can be done in a post-filtering process. Post-filtering can include one or more of deblocking, SAO, ALF, or other filters. Deblocking can be the most important filter for removing ROI boundary artifacts. Typically, the decoder or display / user can choose the filter to use. Table 4 depicts an example syntax for SEI messaging for ROI-related post-filtering.

[0089] Table 4: Example syntax for ROI-related post-filtering

[0090]

[0091] As an example, the syntax parameters may be defined as follows:

[0092] deblocking_enabled_flag equal to 1 specifies that a deblocking process may be applied to independent ROI boundaries of a reconstructed picture for display purposes. deblocking_enabled_flag equal to 0 specifies that a deblocking process may not be applied to independent ROI boundaries of a reconstructed picture for display purposes.

[0093] sao_enabled_flag equal to 1 specifies that the sample adaptive offset (SAO) process may be applied to reconstruct independent ROI boundaries of the picture for display purposes. sao_enabled_flag equal to 0 specifies that the sample adaptive process may not be applied to reconstruct independent ROI boundaries of the picture for display purposes.

[0094] alf_enabled_flag equal to 1 specifies that an adaptive loop filter process (ALF) may be applied to reconstruct independent ROI boundaries of a picture for display purposes. alf_enabled_flag equal to 0 specifies that an adaptive loop filter process may not be applied to reconstruct independent ROI boundaries of a picture for display purposes.

[0095] user_defined_filter_enabled_flag equal to 1 specifies that a user-defined filter process may be applied to independent ROI boundaries of reconstructed pictures for display purposes. user_defined_filter_enabled_flag equal to 0 specifies that a user-defined filter process may not be applied to independent ROI boundaries of reconstructed pictures for display purposes.

[0096] In an embodiment, the SEI messaging in Table 4 may be simplified by removing one or more of the proposed flags. If all flags are removed, the only SEI message present, independent_ROI_across_boundary_filter(payloadSize){}, will indicate to the decoder that a post filter should be used to mitigate boundary artifacts associated with the ROI.

[0097] Region of Interest (ROI) Scalability

[0098] The latest specification of VVC (reference [2]) describes spatial scalability, quality scalability and view scalability (as discussed in more detail in reference [3]) using a combination of reference picture resampling (RPR) and reference picture selection (RPS). It is based on single-loop decoding and block-based on-the-fly resampling. RPS is used to define the prediction relationship between a base layer and one or more enhancement layers, or more specifically, to define the prediction relationship between coded pictures assigned to a base layer or one or more enhancement layers. RPR is used to encode a subset of pictures (i.e., those pictures of (multiple) spatial enhancement layers) at a higher / smaller resolution than the base layer, while predicting from the smaller / higher base layer pictures. Figure 3 An example of spatial scalability according to the RPS / RPR framework is depicted.

[0099] like Figure 3 As depicted in , the bitstream includes two streams: a low resolution (LR) stream (305) (e.g., standard definition, HD, 2K, etc.) and a higher resolution (HR) stream 310 (e.g., HD, 2K, 4K, 8K, etc.). The arrows represent possible inter-frame coding dependencies. For example, HR frame 310-P1 depends on LR frame 305-I. In order to predict the blocks in 310-P1, the decoder needs to scale up 305-I. Similarly, HR frame 310-P2 can depend on HR frame 310-P1 and LR frame 305-P1. Any prediction from LR frame 305-P1 requires proportional spatial upscaling from LR to HR. In other embodiments, the order of LR frames and HR frames can also be reversed, so that the base layer can be the HR stream and the enhancement layer can be the LR stream. Note that the scaling of the base layer pictures is not performed explicitly as in SHVC. Instead, the scaling is absorbed and calculated on the fly in the inter-layer motion compensation. In reference [2], a cropping window is implicitly used to obtain the scalability ratio.

[0100] HEVC (reference [1]) supports ROI scalability as part of Annex H "Scalable High Efficiency Video Coding" (commonly referred to as SHVC). For example, in Section F.7.3.2.3.4, syntax elements related to scaled_ref_layer_offset_present_flag[i] and ref_region_offset_present_flag[i] are defined. The relevant parameters are derived from equations (H-2) to (H-21) and (H-67) to (H-68). VVC does not yet support region of interest (ROI) scalability. As the inventors appreciate, support for ROI scalability allows canvas size scalability to be achieved using the same single-loop VVC decoder without the need for scalability extensions as in SHVC.

[0101] As an example, given Figure 2B The three data tiers depicted in (e.g., 2K, 4K, and 8K), Figure 4 An example embodiment of a bitstream that uses the existing RPS / RPR framework to support canvas size scalability is depicted.

[0102] like Figure 4 As depicted, the bitstream distributes its pictures into three layers or streams, namely, a 2K stream (402), a 4K stream (405), and an 8K stream (410). The arrows represent examples of possible inter-frame coding dependencies. For example, a pixel block in an 8K frame 410-P2 may depend on blocks in an 8K frame 410-P1, a 4K frame 405-P2, and a 2K frame 402-P1. Compared to previous scalability schemes using a multi-loop decoder, the proposed ROI scalability scheme has the following advantages and disadvantages:

[0103] Advantages: Requires a single-loop decoder without any additional tools. The decoder does not need to worry about how to handle block / tile / slice / sub-picture boundaries.

[0104] Disadvantages: To decode the enhancement layer, both base layer decoded pictures and enhancement layer decoded pictures are needed in the Decoded Picture Buffer (DPB), thus requiring a larger DPB size than a non-scalable solution. A higher decoder speed may also be required since both base layer and enhancement layer need to be decoded.

[0105] The key difference between the implementation of ROI scalability support in SHVC and the proposed embodiment for VVC is that in SHVC, all pictures in the same layer are required to have the same picture resolution. However, in VVC, due to the support of RPR, pictures in the same layer can have different resolutions. For example, in Figure 3In SHVC, 305-I, 305-P1 and 305-P2 are required to have the same spatial resolution. But in VVC, due to the support of RPR, 305-I, 305-P1 and 305-P2 can have different resolutions. For example, 305-I and 305-P1 can have a first low resolution (e.g., 720p), while 305-P2 can have a second low resolution (e.g., 480p). Embodiments of the present invention are intended to support both ROI scalability across different layers and RPR for pictures of the same layer. Another major difference is that in SHVC, the motion vector from inter-layer prediction is constrained to be zero. But for VVC, such a constraint does not exist, and the motion vector can be either zero or non-zero. This reduces the constraints on identifying inter-layer correspondences.

[0106] VVC's coding tree only allows encoding of full coding units (CUs). While most standard formats encode picture areas in multiples of four or eight pixels, non-standard formats may require padding at the encoder to match the minimum CTU size. The same problem exists in HEVC. The problem is solved by creating a "consistency window" that specifies the picture area that is considered to be conforming to the picture output. The consistency window was also added to VVC (reference [2]) and is specified by four variables: conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset. For ease of reference, the following part is copied from reference [2].

[0107] "conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset" specify the picture samples in the CVS output from the decoding process according to the rectangular area for output specified in picture coordinates. When conformance_window_flag is equal to 0, the values ​​of conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset are inferred to be equal to 0.

[0108] The conforming cropping window contains luma samples with horizontal picture coordinates from SubWidthC*conf_win_left_offset to pic_width_in_luma_samples-(SubWidthC*conf_win_right_offset+1) and vertical picture coordinates from SubHeightC*conf_win_top_offset to pic_height_in_luma_samples-(SubHeightC*conf_win_bottom_offset+1) (inclusive).

[0109] The value of SubWidthC*(conf_win_left_offset+conf_win_right_offset) should be less than pic_width_in_luma_samples, and the value of SubHeightC*(conf_win_top_offset+conf_win_bottom_offset) should be less than pic_height_in_luma_samples.

[0110] The variables PicOutputWidthL and PicOutputHeightL are obtained as follows:

[0111] PicOutputWidthL=pic_width_in_luma_samples- (7-43)

[0112] SubWidthC*(conf_win_right_offset+conf_win_left_offset)

[0113] PicOutputHeightL=pic_height_in_pic-size_units- (7-44)

[0114] SubHeightC*(conf_win_bottom_offset+conf_win_top_offset)”

[0115] In a first embodiment, the newly defined ROI offset is combined with the existing offset of the consistency window to obtain the scaling factor. Figure 5 An example embodiment of the proposed syntax elements is depicted in FIG, which depicts a base layer picture (520) and an enhancement layer picture (502) and their corresponding consistency windows. The following ROI syntax elements are defined:

[0116] Base Layer (BL)

[0117] rer_region_top_offset(528)

[0118] ·ref_region_bottom_offset(530)

[0119] ref_region_left_offset(524)

[0120] ·ref_region_right_offset(526)

[0121] Note that the width (522) and height (532) of the BL picture (520) may be calculated using the consistency window parameters of the base layer using equations (7-43) and (7-44) above. (For example, pic_width_in_luma_samples may correspond to the width 522, and PicOutputWidth may correspond to the width of the dashed window 540).

[0122] Enhancement Layer

[0123] ·scaled_ref_region_top_offset(508)

[0124] ·scaled_ref_region_bottom_offset(510)

[0125] ·scaled_ref_region_left_offset(504)

[0126] ·scaled_ref_region_right_offset(506)

[0127] Note that the width (512) and height (514) of the EL picture (502) may be calculated using the consistency window parameters of the enhancement layer using equations (7-43) and (7-44) above. (For example, pic_width_in_luma_samples may correspond to width 512, and PicOutputWidth may correspond to the width of the dotted window 518).

[0128] As an example, Table 5 shows how pic_parameter_set_rbsp() defined in section 7.3.2.4 of reference [2] is modified (the edited content is highlighted in gray) to support the new syntax elements.

[0129] Table 5: Example syntax in VVC for supporting ROI scalability

[0130]

[0131] num_ref_loc_offsets specifies the number of reference layer location offsets present in the PPS. The value of num_ref_loc_offsets should be in the range of 0 to vps_max_layers_minus1, inclusive.

[0132] ref_loc_offset_layer_id[i] specifies the nuh_layer_id value for which the i-th reference layer location offset parameter is specified.

[0133] NOTE - ref_loc_offset_layer_id[i] does not need to be in the direct reference layer, e.g. when the spatial correspondence of an auxiliary picture to its associated primary picture is specified.

[0134] The i-th reference layer position offset parameter is composed of the i-th scaled reference layer offset parameter and the i-th reference region offset parameter.

[0135] scaled_ref_layer_offset_present_flag[i] equal to 1 specifies that the i-th scaled reference layer offset parameter is present in the PPS. scaled_ref_layer_offset_present_flag[i] equal to 0 specifies that the i-th scaled reference layer offset parameter is not present in the PPS. When not present, the value of scaled_ref_layer_offset_present_flag[i] is inferred to be equal to 0.

[0136] The i-th scaled reference layer offset parameter specifies the spatial correspondence of the picture that references this PPS relative to the reference region in the decoded picture whose nuh_layer_id is equal to ref_loc_offset_layer_id[i].

[0137] scaled_ref_layer_left_offset[ref_loc_offset_layer_id[i]] plus conf_win_left_offset specifies the horizontal offset in units of subWC luma samples between the sample in the current picture that is juxtaposed with the upper-left luma sample in the reference area of ​​the decoded picture with nuh_layer_id equal to ref_loc_offset_layer_id[i]] and the upper-left luma sample in the current picture, where subWC is equal to SubWidthC of the picture referencing this PPS. The value of scaled_ref_layer_left_offset[ref_loc_offset_layer_id[i]] plus conf_win_left_offset should be between -2 and 1. 14 to 2 14 In the range of -1 (inclusive), when not present, the value of scaled_ref_layer_left_offset[ref_loc_offset_layer_id[i]] is inferred to be equal to 0.

[0138] scaled_ref_layer_top_offset[ref_loc_offset_layer_id[i]] plus conf_win_top_offset specifies the vertical offset in units of subHC luma samples between the sample in the current picture that is juxtaposed with the top-left luma sample in the reference area of ​​the decoded picture with nuh_layer_id equal to ref_loc_offset_layer_id[i], and the top-left luma sample in the current picture, where subHC is equal to SubHeightC of the picture referencing this PPS. The value of scaled_ref_layer_top_offset[ref_loc_offset_layer_id[i]] plus conf_win_top_offset should be between -2 and 1. 14 to 2 14 In the range of -1 (inclusive), when not present, the value of scaled_ref_layer_top_offset[ref_loc_offset_layer_id[i]] is inferred to be equal to 0.

[0139] scaled_ref_layer_right_offset[ref_loc_offset_layer_id[i]] plus conf_win_right_offset specifies the horizontal offset in units of subWC luma samples between the sample in the current picture that is juxtaposed with the lower-right luma sample in the reference area of ​​the decoded picture with nuh_layer_id equal to ref_loc_offset_layer_id[i], where subWC is equal to SubWidthC of the picture referencing this PPS. The value of scaled_ref_layer_right_offset[ref_loc_offset_layer_id[i]] plus conf_win_right_offset should be between -2 and 1. 14 to 2 14 In the range of -1 (inclusive), when not present, the value of scaled_ref_layer_right_offset[ref_loc_offset_layer_id[i]] is inferred to be equal to 0.

[0140] scaled_ref_layer_bottom_offset[ref_loc_offset_layer_id[i]] plus conf_win_bottom_offset specifies the vertical offset in units of subHC luma samples between the sample in the current picture that is juxtaposed with the lower-right luma sample in the reference area of ​​the decoded picture with nuh_layer_id equal to ref_loc_offset_layer_id[i], and the lower-right luma sample in the current picture, where subHC is equal to SubHeightC of the picture referencing this PPS. The value of scaled_ref_layer_bottom_offset[ref_loc_offset_layer_id[i]] plus conf_win_bottom_offset should be between -2 and 1. 14 to 2 14 In the range of -1 (inclusive), when not present, the value of scaled_ref_layer_bottom_offset[ref_loc_offset_layer_id[i]] is inferred to be equal to 0.

[0141] Let currTopLeftSample, currBotRightSample, colRefRegionTopLeftSample, and colRefRegionBotRightSample be, respectively, the top-left luma sample of the current picture, the bottom-right luma sample of the current picture, the sample in the current picture juxtaposed with the top-left luma sample of the reference region in the decoded picture whose nuh_layer_id is equal to ref_loc_offset_layer_id[i], and the sample in the current picture juxtaposed with the bottom-right luma sample of the reference region in the decoded picture whose nuh_layer_id is equal to ref_loc_offset_layer_id[i].

[0142] When the value of (scaled_ref_layer_left_offset[ref_loc_offset_layer_id[i]]+conf_win_left_offset) is greater than 0, colRefRegionTopLeftSample is located to the right of currTopLeftSample. When the value of (scaled_ref_layer_left_offset[ref_loc_offset_layer_id[i]]+conf_win_left_offset) is less than 0, colRefRegionTopLeftSample is located to the left of currTopLeftSample.

[0143] When the value of (scaled_ref_layer_top_offset[ref_loc_offset_layer_id[i]]+conf_win_top_offset) is greater than 0, colRefRegionTopLeftSample is below currTopLeftSample. When the value of (scaled_ref_layer_top_offset[ref_loc_offset_layer_id[i]]+conf_win_top_offset) is less than 0, colRefRegionTopLeftSample is above currTopLeftSample.

[0144] When the value of (scaled_ref_layer_right_offset[ref_loc_offset_layer_id[i]]+conf_win_right_offset) is greater than 0, colRefRegionBotRightSample is located to the left of currBotRightSample. When the value of (scaled_ref_layer_right_offset[ref_loc_offset_layer_id[i]]+conf_win_right_offset) is less than 0, colRefRegionTopLeftSample is located to the right of currBotRightSample.

[0145] When the value of (scaled_ref_layer_bottom_offset[ref_loc_offset_layer_id[i]]+conf_win_bottom_offset) is greater than 0, colRefRegionBotRightSample is above currBotRightSample. When the value of (scaled_ref_layer_bottom_offset[ref_loc_offset_layer_id[i]]+conf_win_bottom_offset) is less than 0, colRefRegionTopLeftSample is below currBotRightSample.

[0146] ref_region_offset_present_flag[i] equal to 1 specifies that the i-th reference region offset parameter is present in the PPS. ref_region_offset_present_flag[i] equal to 0 specifies that the i-th reference region offset parameter is not present in the PPS. When not present, the value of ref_region_offset_present_flag[i] is inferred to be equal to 0.

[0147] The i-th reference region offset parameter specifies the spatial correspondence of the reference region in the decoded picture with nuh_layer_id equal to ref_loc_offset_layer_id[i] relative to the same decoded picture.

[0148] Let refConfLeftOffset[ref_loc_offset_layer_id[i]], refConfTopOffset[ref_loc_offset_layer_id[i]], refConfRightOffset[ref_loc_offset_layer_id[i]], and refConfBottomOffset[ref_loc_offset_layer_id[i]] be the values ​​of conf_win_left_offset, conf_win_top_offset, conf_win_right_offset, and conf_win_bottom_offset, respectively, of the decoded picture whose nuh_layer_id is equal to ref_loc_offset_layer_id[i].

[0149] ref_region_left_offset[ref_loc_offset_layer_id[i]] plus refConfLeftOffset[ref_loc_offset_layer_id[i]] specifies the horizontal offset in units of subWC luma samples between the top-left luma sample of the reference region in the decoded picture with nuh_layer_id equal to ref_loc_offset_layer_id[i] and the top-left luma sample of the same decoded picture, where subWC is equal to SubWidthC of the layer with nuh_layer_id equal to ref_loc_offset_layer_id[i]. The value of ref_region_left_offset[ref_loc_offset_layer_id[i]] plus refConfLeftOffset[ref_loc_offset_layer_id[i]] shall be between -2 and 14 to 2 14 In the range of -1 (inclusive), when not present, the value of ref_region_left_offset[ref_loc_offset_layer_id[i]] is inferred to be equal to 0.

[0150] ref_region_top_offset[ref_loc_offset_layer_id[i]] plus refConfTopOffset[ref_loc_offset_layer_id[i]] specifies the vertical offset in units of subHC luma samples between the top left luma sample of the reference region in the decoded picture with nuh_layer_id equal to ref_loc_offset_layer_id[i] and the top left luma sample of the same decoded picture, where subHC is equal to SubHeightC of the layer with nuh_layer_id equal to ref_loc_offset_layer_id[i]. The value of ref_region_top_offset[ref_loc_offset_layer_id[i]] plus refConfTopOffset[ref_loc_offset_layer_id[i]] shall be between -2 and 14 to 2 14 In the range of -1 (inclusive), when not present, the value of ref_region_top_offset[ref_loc_offset_layer_id[i]] is inferred to be equal to 0.

[0151] ref_region_right_offset[ref_loc_offset_layer_id[i]] plus refConfRightOffset[ref_loc_offset_layer_id[i]] specifies the horizontal offset in units of subWC luma samples between the lower right luma sample of the reference region in the decoded picture with nuh_layer_id equal to ref_loc_offset_layer_id[i]] and the lower right luma sample of the same decoded picture, where subWC is equal to SubWidthC of the layer with nuh_layer_id equal to ref_loc_offset_layer_id[i]. The value of ref_layer_right_offset[ref_loc_offset_layer_id[i]] plus refConfRightOffset[ref_loc_offset_layer_id[i]] shall be between -2 and 14 to 2 14 In the range of -1 (inclusive), when not present, the value of ref_region_right_offset[ref_loc_offset_layer_id[i]] is inferred to be equal to 0.

[0152] ref_region_bottom_offset[ref_loc_offset_layer_id[i]] plus refConfBottomOffset[ref_loc_offset_layer_id[i]] specifies the vertical offset in units of subHC luma samples between the bottom-right luma sample of the reference region in the decoded picture with nuh_layer_id equal to ref_loc_offset_layer_id[i] and the bottom-right luma sample of the same decoded picture, where subHC is equal to SubHeightC of the layer with nuh_layer_id equal to ref_loc_offset_layer_id[i]. The value of ref_layer_bottom_offset[ref_loc_offset_layer_id[i]] plus refConfBottomOffset[ref_loc_offset_layer_id[i]] shall be between -2 and 14 to 2 14 In the range of -1 (inclusive), when not present, the value of ref_region_bottom_offset[ref_loc_offset_layer_id[i]] is inferred to be equal to 0.

[0153] Let refPicTopLeftSample, refPicBotRightSample, refRegionTopLeftSample and refRegionBotRightSample be the top-left luminance sample of the decoded picture with nuh_layer_id equal to ref_loc_offset_layer_id[i], the bottom-right luminance sample of the decoded picture with nuh_layer_id equal to ref_loc_offset_layer_id[i], the top-left luminance sample of the reference region in the decoded picture with nuh_layer_id equal to ref_loc_offset_layer_id[i], and the bottom-right luminance sample of the reference region in the decoded picture with nuh_layer_id equal to ref_loc_offset_layer_id[i].

[0154] When the value of (ref_region_left_offset[ref_loc_offset_layer_id[i]]+refConfLeftOffset[ref_loc_offset_layer_id[i]]) is greater than 0, refRegionTopLeftSample is located to the right of refPicTopLeftSample. When the value of (ref_region_left_offset[ref_loc_offset_layer_id[i]]+refConfLeftOffset[ref_loc_offset_layer_id[i]]) is less than 0, refRegionTopLeftSample is located to the left of refPicTopLeftSample.

[0155] When the value of (ref_region_top_offset[ref_loc_offset_layer_id[i]]+refConfTopOffset[ref_loc_offset_layer_id[i]]) is greater than 0, refRegionTopLeftSample is located below refPicTopLeftSample. When the value of (ref_region_top_offset[ref_loc_offset_layer_id[i]]+refConfTopOffset[ref_loc_offset_layer_id[i]]) is less than 0, refRegionTopLeftSample is located above refPicTopLeftSample.

[0156] When the value of (ref_region_right_offset[ref_loc_offset_layer_id[i]]+refConfRightOffset[ref_loc_offset_layer_id[i]]) is greater than 0, refRegionBotRightSample is located to the left of refPicBotRightSample. When the value of (ref_region_right_offset[ref_loc_offset_layer_id[i]]+refConfRightOffset[ref_loc_offset_layer_id[i]]) is less than 0, refRegionBotRightSample is located to the right of refPicBotRightSample.

[0157] When the value of (ref_region_bottom_offset[ref_loc_offset_layer_id[i]]+refConfBottomOffset[ref_loc_offset_layer_id[i]]) is greater than 0, refRegionBotRightSample is above refPicBotRightSample. When the value of (ref_region_bottom_offset[ref_loc_offset_layer_id[i]]+refConfBottomOffset[ref_loc_offset_layer_id[i]]) is less than 0, refRegionBotRightSample is below refPicBotRightSample.

[0158] Given the proposed syntax elements, in an embodiment and without limitation, the corresponding VVC section may be modified as follows: The equations labeled (7-xx) and (8-xx) represent new equations that need to be inserted into the VVC specification and will be renumbered as necessary.

[0159]

[0160] The variables PicOutputWidthL and PicOutputHeightL are obtained as follows:

[0161] PicOutputWidthL=pic_width_in_luma_samples- (7-43)

[0162] SubWidthC*(conf_win_right_offset+conf_win_left_offset)

[0163] PicOutputHeightL=pic_height_in_pic_size_units- (7-44)

[0164] SubHeightC*(conf_win_bottom_offset+conf_win_top_offset)

[0165] The variable fRefWidth is set equal to PicOutputWidthL of the reference picture in luma samples.

[0166] The variable fRefHeight is set equal to PicOutputHeightL of the reference picture in luma samples.

[0167] The variable refConfWinLeftOffset is set equal to the ConfWinLeftOffset of the reference picture in luma samples.

[0168] The variable refConfWinTopOffset is set equal to the ConfWinTopOffset of the reference picture in luma samples.

[0169] …

[0170] - If cIdx is equal to 0, the following applies:

[0171] -The scaling factor and its fixed-point representation are defined as

[0172] hori_scale_fp=((fRefWidth<<14)+(PicOutputWidthL>>1)) / PicOutputWidthL

[0173] (8-753)

[0174] vert_scale_fp=((fRefHeight<<14)+(PicOutputHeightL>>1)) / PicOutputHeightL

[0175] (8-754)

[0176] -Let (refxSb L ,refySb L ) and (refx L , refy L ) is the motion vector given in units of 1 / 16 samples

[0177] The brightness position pointed to by (refMvLX[0], refMvLX[1]). Variable refxSb L ,refx L 、refySb L

[0178] and refy L As follows:

[0179] sbX L =((xSb<<4)+refMvLX[0])*hori_scale_fp (8-755)

[0180] refx L =((Sign(refxSb L )*((Abs(refxSb L)+128)>>8)+x L *((hori_scale_fp+8)>>4))+32)>>6 (8-756)

[0181] sb L =((ySb<<4)+refMvLX[1])*vert_scale_fp (8-757)

[0182] refyL=((Sign(refySb L )*((Abs(refySb L )+128)>>8)+yL*((vert_scale_fp+8)>>4))+32)>>6 (8-758)

[0183] - Otherwise (cIdx is not equal to 0), the following applies:

[0184] -Let (refxSb C ,refySb C ) and (refx C , refy C ) is the chroma position pointed to by the motion vector (mvLX[0], mvLX[1]) given in units of 1 / 32 samples. The variable refxSb C 、refySb C ,refx C and refy C As follows:

[0185] sbX C =((xSb / SubWidthC<<5)+mvLX[0])*hori_scale_fp

[0186] (8-763)

[0187]

[0188] sb C =((ySb / SubHeightC<<5)+mvLX[1])*vert_scale_fp

[0189] (8-765)

[0190] refy C =((Sign(refySb C )*((Abs(refySb C )+256)>>9)+yC*((vert_scale_fp+8)>>4))+16)>>5 (8-766)

[0191] To support ROI (canvas size) scalability, the specification should be modified as follows:

[0192]

[0193] The variables PicOutputWidthL and PicOutputHeightL are obtained as follows:

[0194] PicOutputWidthL=pic_width_in_luma_samples- (7-43)

[0195] SubWidthC*(conf_win_right_offset+conf_win_left_offset)

[0196] PicOutputHeightL=pic_height_in_pic_size_units- (7-44)

[0197] SubHeightC*(conf_win_bottom_offset+conf_win_top_offset)

[0198] The variable rLId specifies the value of nuh_layer_id of the direct reference layer picture.

[0199]

[0200] The variable fRefWidth is set equal to PicOutputWidthL of the reference picture in luma samples.

[0201] The variable fRefHeight is set equal to PicOutputHeightL of the reference picture in luma samples.

[0202]

[0203] - If cIdx is equal to 0, the following applies:

[0204] -The scaling factor and its fixed-point representation are defined as

[0205]

[0206] -Let (refxSb L ,refySb L ) and (refx L , refy L) is the luminance position pointed to by the motion vector (refMvLX[0], refMvLX[1]) given in units of 1 / 16 samples. L ,refx L 、refySb L and refy L As follows:

[0207]

[0208]

[0209] - Otherwise (cIdx is not equal to 0), the following applies:

[0210] -Let (refxSb C ,refySb C ) and (refx C , refy C ) is the chroma position pointed to by the motion vector (mvLX[0], mvLX[1]) given in units of 1 / 32 samples. The variable refxSb C 、refySb C ,refx C and refy C As follows:

[0211]

[0212] In another embodiment, because VVC has no restrictions on the size of motion vectors during inter-layer coding, unlike SHVC, it may not be necessary to consider the top / left position of the reference layer ROI and the scaled reference layer (current picture) ROI when finding the pixel correspondence between ROI regions. Therefore, in all the above equations, references to fRefLeftOffset, fRefTopOffset, fCurLeftOffset, and fCurTopOffset can be removed.

[0213] Fig. 6A An example summary of the above process flow is provided. Fig. 6AAs shown, in step 605, the decoder may receive syntax parameters related to the consistency window (e.g., conf_win_xxx_offset, where xxx is left, top, right, or bottom), the scaled reference layer offset of the current picture (e.g., scaled_ref_layer_xxx_offset[]), and the reference layer region offset (e.g., ref_region_xxx_offset[]). If inter-frame coding is not present (step 610), decoding is performed like single-layer decoding, otherwise, in step 615, the decoder calculates the consistency window for both the reference picture and the current picture (e.g., using equations (7-43) and (7-44)). If there is no inter-layer coding (step 620), then in step 622, the RPR scaling factor still needs to be calculated to perform inter-frame prediction for pictures with different resolutions in the same layer, and then decoded like single-layer decoding. Otherwise (with inter-layer coding), the decoder calculates the scaling factors of the current picture and the reference picture based on the received offsets (for example, by calculating hori_scale_fp and vert_scale_fp in equations (8-753) and (8-754)).

[0214] As described above (e.g., see equations (8-x1) to (8-x2)), the decoder needs to calculate the width and height of the reference ROI (e.g., fRefWidth and fRefHeight) by subtracting the left and right offset values ​​of the reference layer (e.g., RefLayerRegionLeftOffset and RefLayerRegionRightOffset) and the top and bottom offsets (e.g., RefLayerRegionTopOffset and RefLayerRegionBottomOffset) from the PicOutputWidthL and PicOutputHeightL of the reference layer picture.

[0215] Similarly, (e.g., see equations (8-x3) to (8-x4)), the decoder needs to calculate the width and height of the current ROI (e.g., fCurWidth and fCurHeight) by subtracting the left and right offset values ​​(e.g., ScaledRefLayerLeftOffset and ScaledRefLayerRightOffset) and the top and bottom offsets (e.g., ScaledRefLayerTopOffset and ScaledRefLayerBottomOffset) from the PicOutputWidthL and PicOutputHeightL of the current layer picture. Given these adjusted sizes of the current ROI and the reference ROI, the decoder determines the horizontal scaling factor and the vertical scaling factor (e.g., see equations (8-753) and (8-754)) in the same way as in the existing VVCRPR block (e.g., processing from equation (8-755) to equation (8-766)), with only minimal additional modifications (as highlighted in gray in the figure above).

[0216] In equations (8-x5) to (8-x8), the adjusted left offset and top offset are also calculated to determine the correct position of the reference ROI and the current ROI relative to the upper left corner of the consistency window to achieve proper pixel correspondence.

[0217] In a second embodiment, the definitions of ref_region_xxx_offset[] and scaled_ref_region_xxx_offset[] offsets may be redefined to combine both the consistency window offset and the ROI offset (e.g., by adding these offsets together). For example, in Table 5, scaled_ref_layer_xxx_offset may be replaced with scaled_ref_layer_xxx_offset_sum, defined as follows:

[0218] scaled_ref_layer_left_offset_sum[]=scaled_ref_layer_left_offset[]+conf_win_left_offset

[0219] scaled_ref_layer_top_offset_sum[]=scaled_ref_layer_top_offset[]+conf_win_top_offset

[0220] scaled_ref_layer_right_offset_sum[]=

[0221] scaled_ref_layer_right_offset[]+conf_win_right_offset (1)

[0222] scaled_ref_layer_bottom_offset_sum[]=

[0223] scaled_ref_layer_bottom_offset[]+conf_win_bottom_offset

[0224] Similar definitions can also be generated for ref_region_xxx_offset_sum, where xxx = bottom, top, left, and right. As explained below, these parameters allow the decoder to skip step 615, because the processing in step 615 can be combined with the processing in step 625.

[0225] As an example, in Fig. 6A middle:

[0226] a) In step 615, PicOutputWidthL may be calculated by subtracting the consistency window left offset and right offset from the picture width (e.g., see equation (7-43))

[0227] b) Let fCurWidth=PicOutputWidthL

[0228] c) Then, in step 625, fCurWidth is adjusted by subtracting the ScaledRefLayer left and right offsets (e.g., see (8-x3)); however, according to equation (7-xx), these two offsets are based on the scaled_ref_layer left and right offsets. For example, in simplified notation (i.e., by ignoring the SubWidthC scaling parameter) considering only the width of the current ROI, the width can be calculated as follows:

[0229] Image output width = image width - (consistency window left offset + consistency window right offset) (2)

[0231] ROI current width = image output width - (ROI current left offset + ROI current right offset) (3)

[0233] By combining equation (2) and equation (3),

[0234] ROI current width = image output width - ((consistency window left offset + ROI current left offset) +

[0235] (consistency window right offset + ROI current right offset))(4)

[0236] make

[0237] ROI current left total offset = consistency window left offset + ROI current left offset

[0238] ROI current right total offset = consistency window right offset + ROI current right offset

[0239] Then equation (4) can be simplified to

[0240] ROI current offset = image width - ((ROI current left total offset) + (ROI current right total offset)) (5)

[0242] The definitions of the new "sum" offsets (eg, ROI current left sum offset) correspond to those of ref_region_left_offset_sum defined earlier in equation (1).

[0243] Thus, as described above, if the scaled_ref_layer left and right offsets are redefined to include the sum of the layer's conf_win_left_offset, the steps in boxes (615) and (625) for calculating the width and height of the current ROI and the reference ROI (e.g., equations (2) and (3)) can be combined into one equation (e.g., equation (5)) (e.g., in step 630).

[0244] like Figure 6B As depicted, steps 615 and 625 can now be combined into a single step 630. Fig. 6A This approach saves some additions, but now the corrected offsets (e.g., scaled_ref_layer_left_offset_sum[]) are larger quantities, so these offsets require more bits to be encoded in the bitstream. Note that the conf_win_xxx_offset values ​​for each layer may be different, and these values ​​can be extracted from the PPS information for each layer.

[0245] In a third embodiment, the horizontal scaling factor and the vertical scaling factor (eg, hori_scale_fp and vert_scale_fp) between inter-layer pictures can be explicitly signaled. In this case, each layer needs to transmit the horizontal scaling factor and the vertical scaling factor as well as the top offset and the left offset.

[0246] Similar methods apply to embodiments of pictures containing multiple ROIs in each picture using arbitrary upsampling filters and downsampling filters.

[0247] References

[0248] Each of the references listed herein is incorporated by reference in its entirety.

[0249] [1] High efficiency video coding, H.265, H series, Coding of moving video, ITU, (02 / 2018).

[0250] [2] B. Bross, J. Chen and S. Liu, “Versatile Video Coding (Draft 6)”, JVET output document, JVET-O2001, vE, uploaded on July 31, 2019.

[0251] [3] S. Wenger et al., “AHG8: Spatial scalability using reference picture resampling,” JVET-O0045, JVET Meeting, Gothenburg, SE, July 2019.

[0252] [4] R. Skupin et al., AHG12: “On filtering of independently coded region”, JVET-O0494(v3), JVET Meeting, Gothenburg, SE, July 2019.

[0253] Example Computer System Implementation

[0254] Embodiments of the present invention may be implemented using a computer system, a system configured with electronic circuits and components, an integrated circuit (IC) device (such as a microcontroller, a field programmable gate array (FPGA) or another configurable or programmable logic device (PLD), a discrete time or digital signal processor (DSP), an application specific IC (ASIC), and / or an apparatus including one or more of such systems, devices, or components. The computer and / or IC may execute, control, or implement instructions related to canvas size scalability, such as those described herein. The computer and / or IC may calculate any of the various parameters or values ​​related to canvas size scalability described herein. The image and video embodiments may be implemented in hardware, software, firmware, and various combinations thereof.

[0255] Certain embodiments of the present invention include a computer processor that executes software instructions that cause the processor to perform the method of the present invention. For example, one or more processors in a display, an encoder, a set-top box, a transcoder, etc. can implement the method related to canvas size scalability as described above by executing software instructions in a program memory accessible to the processor. An embodiment of the present invention can also be provided in the form of a program product. The program product may include any non-transient and tangible medium that carries a set of computer-readable signals, and the set of computer-readable signals includes instructions that cause the data processor to perform the method of the present invention when executed by a data processor. The program product according to the present invention may adopt any of various non-transient and tangible forms. The program product may include, for example, a physical medium, such as a magnetic data storage medium including a floppy disk, a hard disk drive, an optical data storage medium including a CD ROM, a DVD, an electronic data storage medium including a ROM, a flash memory RAM, etc. The computer-readable signal on the program product may be optionally compressed or encrypted.

[0256] Where components (e.g., software modules, processors, components, devices, circuits, etc.) are mentioned above, unless otherwise indicated, references to the components (including references to "means") should be interpreted to include any components that perform the functions of the described components as equivalents (e.g., functionally equivalent) to the components, including components that are not structurally equivalent to the disclosed structures that perform the functions in the illustrated example embodiments of the invention.

[0257] Equivalents, extensions, alternatives and miscellaneous

[0258] Thus, example embodiments related to canvas size scalability have been described. In the foregoing specification, embodiments of the invention have been described with reference to many specific details that may vary depending on the implementation. Therefore, the sole and exclusive indicator of the invention and the applicant's inventive intent is the set of claims issuing in the specific form in which such claims are issued from this application, including any subsequent amendments. Any definitions expressly set forth herein for terms contained in such claims shall govern the meaning of such terms as used in the claims. Therefore, no limitation, element, property, feature, advantage or attribute that is not expressly recited in a claim should in any way limit the scope of such claim. Therefore, the specification and drawings should be viewed in an illustrative rather than a restrictive sense.

Claims

1. A method for decoding a bitstream encoded with a scalable canvas size, the method being executed by a processor and comprising: For the current image, Receives the current image width and the current image height as unsigned integer values; Receiving a first offset parameter for determining a rectangular area on the current picture, wherein the first offset parameter comprises a signed integer value; Calculate the current area width and the current area height of the rectangular area on the current picture based on the current picture width, the current picture height and the first offset parameter; For the reference area, access the reference area width, reference area height, reference area left offset, and reference area top offset; Calculating a horizontal scaling factor based on the current region width and the reference region width, wherein calculating the horizontal scaling factor hori_scale_fp includes calculating hori_scale_fp=((fRefWidth<<14)+(fCurWidth>>1)) / fCurWidth Wherein, fRefWidth represents the reference region width, and fCurWidth represents the current region width; The vertical scaling factor vert_scale_fp is calculated based on the current region height and the reference region height, wherein calculating the vertical scaling factor vert_scale_fp includes calculating vert_scale_fp=((fRefHeight<<14)+(fCurHeight>>1)) / fCurHeight Wherein, fRefHeight represents the reference area height, and fCurHeight represents the current area height; calculating a left offset adjustment and a top offset adjustment of the current region based on the first offset parameter; and performing motion compensation based on the horizontal scaling factor and the vertical scaling factor, the left offset adjustment, the top offset adjustment, the reference region left offset, and the reference region top offset, Among them, performing motion compensation includes calculating refxSb L =(((xSb–fCurLeftOffset)<<4)+refMvLX[0])*hori_scale_fp refx L =((Sign(refxSb L )*((Abs(refxSb L )+128)>>8)+x L *((hori_scale_fp+8) >>4))+32)>>6+(fRefLeftOffset<<4) refySb L =(((ySb–fCurTopOffset)<<4)+refMvLX[1])*vert_scale_fp dimensions L =((Sign(measureSb L )*((Abs(measureSb L )+128)>>8)+y L *((vert_scale_fp+8) >>4))+32)>>6+(fRefTopOffset<<4) Wherein, hori_scale_fp represents the horizontal scaling factor, vert_scale_fp represents the vertical scaling factor, fCurLeftOffset represents the left offset adjustment, fCurTopOffset represents the top offset adjustment, fRefLeftOffset represents the reference area left offset, fRefTopOffset represents the reference area top offset, and (refxSb L ,refySb L ) and (refx L ,refy L ) is the luminance position pointed to by the motion vector (refMvLX[0], refMvLX[1]) given in units of 1 / 16 samples.

2. The method of claim 1, wherein: The first offset parameters include a left offset, a top offset, a right offset, and a bottom offset.

3. The method of claim 2, wherein: One or more of the left offset, the top offset, the right offset, or the bottom offset includes -2 14 with 2 14 The value between .

4. The method of claim 2, wherein: Calculating the current region width includes subtracting a first sum of the left offset and the right offset from the current picture width, and calculating the current region height includes subtracting a second sum of the top offset and the bottom offset from the current picture height.

5. The method of claim 1, wherein: Accessing the reference region width and the reference region height further comprises: for a reference picture, Access reference image width and reference image height; receiving a second offset parameter for determining a rectangular area on the reference picture, wherein the second offset parameter comprises a signed integer value; and The reference region width and the reference region height of the rectangular region on the reference picture are calculated based on the reference picture width, the reference picture height, and the second offset parameter. 6 . The method of claim 5 , further comprising calculating the reference region left offset and the reference region top offset based on the second offset parameter.

7. The method of claim 1, wherein: The reference area includes a reference picture.

8. The method of claim 6, wherein: The reference region width, the reference region height, the reference region left offset and the reference region top offset are calculated based on one or more of a consistency window parameter, a reference picture width, a reference picture height or a region of interest offset parameter for the reference picture.

9. A non-transitory computer-readable storage medium having computer-executable instructions stored thereon, the computer-executable instructions being used to execute the method of any one of claims 1 to 8 using one or more processors.

10. An apparatus for decoding an encoded bitstream having a scalable canvas size, the apparatus comprising a processor and configured to perform the method of any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method, an apparatus and a computer program product for coding a 360-degree panoramic video

    US20170085917A1