Expandable drawing area size video coding

By enhancing HEVC tile concepts with cross-boundary prediction and layer-adaptive slice addressing, the solution addresses the challenge of scalable HDR video distribution, ensuring compatibility and reducing coding inefficiencies and artifacts across different resolutions.

JP2026105075APending Publication Date: 2026-06-25DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
DOLBY LABORATORIES LICENSING CORP
Filing Date
2026-04-22
Publication Date
2026-06-25

AI Technical Summary

Technical Problem

Existing video coding standards struggle to efficiently support scalable distribution of high dynamic range (HDR) video content across different resolutions, leading to limitations in playback capabilities and compatibility with existing devices.

Method used

The proposed solution involves modifying the HEVC tile concept to enable cross-boundary prediction and entropy coding, using layer-adaptive slice addressing, and incorporating SEI message communication for in-loop filtering to support scalable video coding that allows decoding of specific regions of interest, while maintaining coding efficiency and reducing boundary artifacts.

Benefits of technology

This approach enables flexible and efficient decoding of HDR video content at various resolutions, ensuring compatibility with legacy and future playback devices, and minimizing coding inefficiencies and boundary artifacts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026105075000001_ABST
    Figure 2026105075000001_ABST
Patent Text Reader

Abstract

Methods and systems for expanding the drawing surface size across the same or different bitstream layers of a video coding bitstream are described. [Solution] The offset parameter of the fitting window, the reference region of interest (ROI) of the reference layer, and the current ROI of the current layer are received. The width and height of the current ROI and reference ROI are calculated based on the offset parameter, and these are used to generate width and height scaling factors that should be used by the reference picture resampling unit to generate an output picture based on the current ROI and reference ROI.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [Related applications] This application claims priority rights to U.S. Provisional Application No. 62 / 883,195, filed 6 August 2019, No. 62 / 902,818, filed 19 September 2019, and No. 62 / 945,931, filed 10 December 2019.

[0002] [Technical field] This specification relates in general to images. More specifically, embodiments of the present invention relate to video coding with expandable drawing surface size. [Background technology]

[0003] When used in this specification, the term “dynamic range (DR)” may relate to the human visual system’s ability to perceive a range of intensity (e.g., luminance, lumens) within an image, from, for example, the darkest gray (black) to the brightest white (highlight). In this scene, DR relates to “scene-reference” intensity. DR may also relate to the ability of a display device to render a range of intensity of a particular width appropriately or approximately. In this scene, DR relates to “display-reference” intensity. In any part of the description in this specification, unless explicitly stated that a particular scene has a particular degree of importance, it should be presumed that the terms may be used synonymously in any of the scenes.

[0004] As used in this specification, the term “high dynamic range (HDR)” refers to a DR width that is 14 to 15 times or greater than the size of the human visual system (HVS). In practice, a DR in which a wide range of intensity can be perceived simultaneously by humans may be abbreviated in some way in relation to HDR.

[0005] In practice, an image contains one or more color components (e.g., lumens Y and chromins Cb and Cr), and each color component is represented with n bits per pixel (e.g., n=8). Using nonlinear luminance coding, an image where n≦8 (e.g., a 24-bit color JPEG image) can be considered a standard dynamic range (SDR) image. On the other hand, an image where n>8 can be considered an extended dynamic range image. HDR images may be stored and delivered using a high-resolution (e.g., 16-bit) floating-point format, such as the OpenEXR file format developed by Industrial Light and Magic.

[0006] Currently, the distribution of high dynamic range video content such as HDR10 in Dolby Laboratories' Dolby Vision or Blu-ray® is limited to 4K resolution (e.g., 4096×2160 or 3840×2160, etc.) and 60 frames per second (fps) due to the capabilities of many playback devices. In future versions, it is expected that content up to 8K resolution (e.g., 7680×4320) and 120fps will be available for distribution and playback. To simplify the HDR playback content ecosystem, such as Dolby Vision, it is desirable that future content types be compatible with existing playback devices. Ideally, content creators should be able to adopt and distribute future HDR technologies without having to derive and distribute special versions of content compatible with existing HDR devices (such as HDR10 or Dolby Vision). As the inventors acknowledge here, improved technologies for scalable distribution of video content, especially HDR content, are desired.

[0007] The approaches described in this chapter are pursued, but not necessarily previously conceived or pursued. Therefore, unless otherwise indicated, none of the approaches described in this chapter should be considered prior art simply by their inclusion in this chapter. Similarly, any problems identified with respect to one or more approaches should not be assumed, unless otherwise indicated, to have been recognized in any prior art based on this chapter. [Brief explanation of the drawing]

[0008] Embodiments of the present invention are described by example, not limiting them, and similar reference numerals in the accompanying figures represent similar elements.

[0009] [Figure 1] This illustrates an example of a video distribution pipeline.

[0010] [Figure 2A] An example of a picture sub-region for defining the display area of ​​input content according to the resolution of the target display is shown.

[0011] [Figure 2B] Figure 2A shows an example of a limitation that spans boundaries in tile representation according to the embodiment, specifically regarding the picture area.

[0012] [Figure 2C] An example of layer-adaptive slice addressing according to an embodiment is shown.

[0013] [Figure 3] This shows an example of the possibility of spatial expansion using conventional technology.

[0014] [Figure 4] An example of the possibility of extending the drawing surface according to the embodiment is shown.

[0015] [Figure 5] Examples of base layer and extended layer pictures, and corresponding adapted windows according to the embodiment are shown.

[0016] [Figure 6A] An exemplary processing flow supporting the scalability of the drawing screen size according to an embodiment of the present invention is shown. [Figure 6B] An exemplary processing flow supporting the scalability of the drawing screen size according to an embodiment of the present invention is shown.

Embodiments for Carrying Out the Invention

[0017] Exemplary embodiments related to the scalability of the drawing screen size for video coding are described in the specification of this application. In the following detailed description, for the purpose of explanation, numerous specific details are described to provide a complete understanding of various embodiments of the present invention. However, it is clear that various embodiments of the present invention may be implemented without some of these specific details. In other instances, well-known structures and devices are not described in exhaustive detail to avoid obscuring, obfuscating, or making unclear the embodiments of the present invention.

[0018] <Abstract> The exemplary embodiments described in the specification of this application are related to the scalability of the drawing screen size in video coding. In an embodiment, a processor receives an offset parameter of a conformity window in a first layer, accesses a reference picture width and a reference picture height of a coding region in a reference layer, receives an offset parameter for a first region of interest (ROI) in the first layer, receives an offset parameter for a second ROI in the reference layer, calculates a first picture width and a first picture height for a coding region in the first layer based on the offset parameter for the conformity window, Based on the first picture width, the first picture height, and the offset parameter for the first ROI, the second picture width and second picture height for the current ROI in the first layer are calculated. Based on the aforementioned reference picture width, the aforementioned reference picture height, and the offset parameter for the second ROI in the reference layer, the third picture width and third picture height for the reference ROI in the reference layer are calculated. Based on the second picture width and the third picture width, calculate the horizontal magnification. Based on the second picture height and the third picture height, calculate the vertical magnification. The reference ROI is scaled based on the horizontal and vertical scaling factors to generate a scaled reference ROI. Based on the current ROI and the scaled reference ROI, an output picture is generated.

[0019] In the second embodiment, the decoder is Receive the offset parameter of the fitting window in the first layer, Access the reference picture width and reference picture height of the coding area in the reference layer, The adjusted offset parameter for the first region of interest (ROI) in the first layer is received, and the adjusted offset parameter combines the offset parameter for the first ROI with the offset parameter for the fitting window in the first layer. The adjusted offset parameter for the second ROI in the reference layer is received, and the adjusted offset parameter is combined with the offset parameter for the fitting window in the reference layer. Based on the adjusted offset parameter for the first ROI in the first layer, calculate the first picture width and first picture height for the current ROI in the first layer. Based on the adjusted offset parameter for the second ROI in the reference layer, the second picture width and second picture height for the reference ROI in the reference layer are calculated. Based on the first picture width and the second picture width, calculate the horizontal magnification. Based on the first picture height and the second picture height, calculate the vertical magnification. The reference ROI is scaled based on the horizontal and vertical scaling factors to generate a scaled reference ROI. Based on the current ROI and the scaled reference ROI, an output picture is generated.

[0020] <Example video distribution processing pipeline> Figure 1 shows an exemplary process of a conventional video distribution pipeline 100, illustrating the various stages from video capture to video content display. A sequence of video frames 102 is captured or generated using an image generation block 105. The video frames 102 may be captured digitally (e.g., by a digital camera) or generated by a computer (e.g., using computer animation) to provide video data 107. Alternatively, the video frames 102 may be captured on film by a film camera. The film is converted to a digital format to provide video data 107. In the production stage 110, the video data 107 is edited to provide a video production stream 112.

[0021] The video data from production stream 112 is then provided to the processor in block 115 for post-production editing. Post-production editing in block 115 may include adjusting or modifying the color or brightness of specific areas of the image to improve image quality or achieve a specific appearance in accordance with the video producer's intentions. This is sometimes called "color timing" or "color grading." Other edits (e.g., scene selection and ordering, image cropping, addition of computer-generated visual special effects, control of severe vibration or blur, frame rate control, etc.) may be performed in block 115 to produce the final version 117 of the production for distribution. During post-production editing 115, the video image is displayed on the reference display 125. Following the production 115, the final production video data 117 may be distributed to an encoding block 120 for downstream distribution to decoding and playback devices such as television sets, set-top boxes, movie theaters, etc. In some embodiments, the coding block 120 may include audio and video encoders, such as those defined by ATSC, DVB, DVD, Blu-ray®, and other distribution formats, to generate a coded bit stream 122. At the receiver, the coded bit stream 122 is decoded by a decoding unit 130 to generate a decoded signal 132 that is identical to or a very close approximation of the signal 117. The receiver may be mounted on a target display 140, which may have entirely different characteristics from the reference display 125. In that case, a display management block 135 may be used to map the dynamic range of the decoded signal 132 to the characteristics of the target display 140 by generating a display-mapped signal 137.

[0022] <Extensible coding> Extensible coding is already part of numerous video coding standards such as MPEG-2, AVC, and HEVC. In embodiments of the present invention, extensible coding is extensible to improve performance and flexibility, particularly when related to ultra-high-resolution HDR content.

[0023] <Possibility of expanding drawing area size> As is conventionally known, spatial scalability is primarily used to enable decoders to generate content at various resolutions. In embodiments of the present invention, spatial or drawing surface scalability is designed to allow extraction of different regions of an image. For example, a content creator may choose to frame (i.e., specify a display area for) content differently for large and small displays. For example, the region framed for display may depend on the screen size or the distance from the screen to the viewer. Embodiments of the present invention divide an image into overlapping regions (typically rectangles) and encode them so that a selected number of sub-regions can be decoded for presentation independently of other sub-regions.

[0024] An example is shown in Figure 2A, where various regions encompass and / or are encompassed by other regions. For example, the smallest region 215 has a 2K resolution, and the largest region 205 has an 8K resolution. The basic layer bitstream corresponds to the smallest spatial region, while additional layers within the bitstream correspond to progressively larger image regions. Thus, a 2K display will only show content within the range of the 2K region 215. A 4K display will show content in both the 2K and 4K regions (within the range of 210), and an 8K display will show everything within the even wider 205. In another example, a 2K display may show a downsampled version of 4K content, and a 4K display may show a downsampled version of 8K content. Ideally, the basic layer region can be decoded by legacy equipment, while the other regions can be used by future equipment to expand the drawing surface size.

[0025] Existing coding standards like HEVC can enable drawing surface extensibility using tiles. In tile representation, a frame is divided into a set of non-overlapping rectangular regions. The receiver can decide to decode and display only the set of tiles necessary for display. In HEVC, coding dependencies between tiles are invalid. Specifically, entropy coding and reconstruction dependencies are not allowed to span tile boundaries. This includes motion vector prediction, intra prediction, and context selection. (In-loop filtering is the only exception, which is allowed to span boundaries but can be disabled by a flag in the bitstream.) Furthermore, encoder-side constraints on motion-constrained tiles (MCTS) are required so that the base layer can be decoded independently, and a supplemental enhancement information (SEI) message is required for the time-motion-constrained tile set. The motion-constrained tile set extraction information set SEI message is required for the purpose of bitstream extraction and adaptation. A particular flaw in the tile definition in HEVC, associated with the ability to decode independently, is the loss of coding efficiency.

[0026] In alternative implementations, HEVC enables drawing area expandability by using pan-scan rectangle SEI messages to extract regions of interest (ROIs). SEI message communication specifies a rectangular area but provides information or constraints that allow the ROI to be decoded independently of other regions. Typically, the decoder needs to decode the entire image to obtain the ROI.

[0027] In this embodiment, a novel solution is proposed by modifying the HEVC tile concept. For example, given the region shown in Figure 2A, in this embodiment, independent decoding is required only for region 2K. As shown in Figure 2B, for tiles within the 2K range, the proposed method allows cross-boundary prediction (intra / inter) and entropy coding. For 4K, it allows cross-boundary prediction (intra / inter) and entropy coding from 2k and within the 4K range. For 8K, it allows cross-boundary prediction (intra / inter) and entropy coding from 2k and 4K and within the 8K range. Here, it is proposed to assign layer_id 0 to 2K, layer_id 1 to 4K, and layer_id 2 to 8K. Given that layer_id=N is currently being decoded, cross-boundary tile prediction and entropy coding are only allowed from layer_ids less than or equal to N. In this case, the loss of coding efficiency is reduced compared to HEVC-type tiles. Exemplary syntax is shown in Tables 1 and 2 below, where new syntax elements proposed for the proposed VVC (Versatile Video Codec) draft specification in Ref[2] are highlighted in gray. Table 1: Exemplary sequence parameter set enabling drawing area resizing (RBSP syntax) [Table 1] Table 2: Exemplary picture parameters for resizing the drawing area (RBSP syntax) [Table 2]

[0028] In SPS (Table 1), the flag sps_canvas_tile_enabled_flag has been added. If sps_canvas_tile_enabled_flag is equal to 1, it indicates that the canvas tile is currently enabled in CVS. If sps_canvas_tile_enabled_flag is equal to 0, it indicates that the canvas tile is not currently enabled in CVS.

[0029] In PPS (Table 2), the new layer_id information parameter, tile_layer_id[i], specifies the layer ID of the i-th drawing plane tile. If the constraint is that the tile_layer_id values ​​start from 0 and are continuous, in the embodiment, according to the proposed VVC draft in progress (Ref.[2]), the maximum possible value of tile_layer_id may be NumTilesInPic-1.

[0030] While tiles are used for descriptive purposes, "bricks," slices, and subpictures, as defined in VVC and conventionally known, can be constructed similarly.

[0031] <Layer-adaptive slice addressing> As the inventors have recognized, certain streaming applications may have the following desirable features: (1) When using network abstraction layer (NAL) units in the video coding layer (VCL), a 2K resolution bitstream must be self-contained, and all of its NAL units must have the same nuh_layer_id value (e.g., layer0). Additional bitstreams enabling 4K resolution must also be self-contained, and their NAL units must have the same nuh_layer_id value, but different from the nuh_layer_id of the 2K layer (e.g., layer1). Finally, any additional bitstreams enabling 8K resolution must also be self-contained, and their NAL units must have the same nuh_layer_id value, but different from the nuh_layer_id values ​​of the 2K and 4K layers (e.g., layer2). Therefore, by analyzing the NAL unit header, it should be possible to extract the bitstreams having the target resolution or region of interest (e.g., 2K, 4K, or 8K) using the nuh_layer_id. (2) In non-VCL NAL units, the stream and picture parameter set headers (e.g., SPS, PPS, etc.) must be self-contained for each resolution. (3) At the target resolution, the bitstream extraction process must be able to discard NAL units that are not needed for the target resolution. After bitstream extraction for the target resolution, the bitstream follows a single-layer profile, and therefore the decoder can simply decode the single-resolution bitstream. Note that 2K, 4K, and 8K resolutions are provided only as examples and are not limiting, and the same method can be applied to any number of different spatial resolutions or regions of interest. For example, one might start with a picture at the highest possible resolution (e.g., res_layer[0]=8K) and define a sublayer or region of interest at the lowest resolution, where res_layer[i] > res_layer[i+1] for i=1, 2, ..., N-1, and N is the total number of layers. Then, one might want to decode a specific sublayer without first decoding the entire picture. This helps the decoder reduce complexity, save power, etc.

[0032] To satisfy the above requirements, the following method is proposed in the embodiment: The video parameter set (VPS) syntax can be reused to specify layer information, including the number of layers, inter-layer dependencies, layer representation formats, DPB size, and other information related to defining bitstream suitability, such as layer sets, output layer sets, profile tier levels, and timing-related parameters, using high-level syntax. Each signal parameter set (SPS) associated with a different layer, picture resolution, fitting window, sub-picture, etc., should adhere to a separate resolution (e.g., 2K, 4K, or 8K). • In picture parameter sets (PPS) associated with each different layer, tile, brick, slice, etc., the information should conform to a separate resolution (e.g., 2K, 4K, or 8K). If separate regions are set to be the same within the VCS, tile / brick / slice information may also be set in the SPS. • In the slice header, slice_address should be set to the lowest target resolution containing the slice. As mentioned above, in independent layer decoding, during prediction, a layer can only use tile / brick / slice neighbor information from lower layers and / or from the same layer.

[0033] A VVC (Ref.[2]) defines an integer number of bricks of a picture, exclusively contained within a single NAL unit, as a slice. A brick is defined as a CTU row of a rectangular region within a particular tile in the picture. A CTU (coding tree unit) is a block of samples that has luma and chroma information.

[0034] In our 2K / 4K / 8K examples, in embodiments, the slice_address value (which indicates the slice address of the slice) may need to be different for a 2K bitstream than for a 4K or 8K bitstream. Therefore, conversion of slice_address from low resolution to high resolution may be necessary. Accordingly, in embodiments, such information is provided at the VPS layer.

[0035] Figure 2C shows an example of a 4K picture having one sublayer (e.g., 2K and 4K). Consider picture 220 having 9 tiles and 3 slices. Suppose the gray tiles specify regions of 2K resolution. In a 2K bitstream, the slice_address of the gray regions must be 0, but in a 4K bitstream, the slice_address of the gray regions must be 1. The proposed new syntax allows specifying slice_address according to the resolution layer. For example, in a VPS, slice_address conversion information may be added to specify that for nul_layer_id=1, the slice_address should be changed to 1 in the case of 4K. To simplify the implementation, embodiments may wish to constrain that the slice information for each resolution must be kept the same within the coded video stream (CVS). An exemplary syntax in a VPS based on the HEVC video parameter set RBSP syntax is shown in Table 3 (Ref. [1] Section 7.3.2.1). Information can also be transmitted through other layers of high-level syntax (HLS), such as SPS, PPS, slice headers, and SEI messages. Table 3: Exemplary syntax in a VPS supporting layer-adaptive slice addressing [Table 3] A value of 1 for `vps_layer_slice_info_present_flag` indicates that slice information exists within the VPS() syntax structure. A value of 0 for `vps_layer_slice_info_present_flag` indicates that slice information does not exist within the VPS() syntax structure. Specifying +1 for num_slices_in_layer_minus1[i] specifies the number of slices in the i-th layer. The value of num_slices_in_layer_minus1[i] is equal to num_slices_in_pic_minus1 in the i-th layer. `layer_slice_address[i][j][k]` specifies the i-th layer slice address of the target for the k-th slice within the j-th layer.

[0036] For example, returning to the example in Figure 2C, picture 220 contains two layers: Layer 0 (e.g., 2K) has one slice 230 (gray) with slice address 0. Layer 1 (e.g., 4K) has three slices (225, 230, and 235) with slice addresses 0, 1, and 2. When decoding layer 1 (i=1), within layer 0 (j=0), slice 0 (k=0) 230 has slice address 1, and therefore, according to the notation in Table 3, layer_slice_address[1][0][0]=1.

[0037] <SEI message after filtering> When implementing drawing surface extensibility using bricks / tiles / slices / subpictures, a potential problem is the implementation of in-loop filtering across boundaries (e.g., deblocking, SAO, ALF). For example, Ref. [4] describes the problems when component windows are coded using independent regions (or subpictures). When encoding an entire picture using independent coding regions (which can be done, for example, by bricks / tiles / slices / subpictures, etc.), in-loop filtering across independent coding regions can result in drift and boundary artifacts. In drawing surface size applications, it is important to have good visual quality for both high-resolution and low-resolution video. In high-resolution video, boundary artifacts must be mitigated. Therefore, in-loop filtering across independent coding regions (especially deblocking filters) must be enabled. In low-resolution video, drift and boundary artifacts must also be minimized.

[0038] Ref.[4] proposes a solution for storing subpicture boundary padding for interpretation. This approach can be implemented by encoder-only constraints, such as prohibiting motion vectors that use those pixels affected by in-loop filtering. Alternatively, embodiments propose solving this problem using post-filtering communicated to the decoder via SEI message communication.

[0039] First, it is proposed that in-loop filtering spanning independent coding regions (e.g., slice boundaries within regions 225 and 230) be disabled. Filtering spanning independent coding regions for the entire picture may be performed in post-filtering processing. Post-filtering may include one or more of the following: deblocking, SAO, ALF, or other filters. Deblocking may be the most important filter for removing ROI boundary artifacts. Typically, the decoder or display / user may have their own choice of which filters should be used. Table 4 shows exemplary syntax for SEI message communication after ROI-related filtering. Table 4: Example syntax after ROI-related filtering [Table 4] For example, syntax parameters may be defined as follows:

[0040] A `deblocking_enabled_flag` equal to 1 indicates that deblocking may be applied to independent ROI boundaries of a picture reconstructed for display purposes. A `deblocking_enabled_flag` equal to 0 indicates that deblocking should not be applied to independent ROI boundaries of a picture reconstructed for display purposes.

[0041] A sao_enabled_flag equal to 1 indicates that sample adaptive offset (SAO) processing may be applied to the independent ROI boundaries of a picture reconstructed for display purposes. A sao_enabled_flag equal to 0 indicates that sample adaptive processing should not be applied to the independent ROI boundaries of a picture reconstructed for display purposes.

[0042] A value of 1 for `alf_enabled_flag` indicates that adaptive loop filtering (ALF) may be applied to independent ROI boundaries of a picture reconstructed for display purposes. A value of 0 for `alf_enabled_flag` indicates that adaptive loop filtering should not be applied to independent ROI boundaries of a picture reconstructed for display purposes.

[0043] A user_defined_filter_enabled_flag equal to 1 indicates that user-defined filtering may be applied to independent ROI boundaries of a picture reconfigured for display purposes. A user_defined_filter_enabled_flag equal to 0 indicates that user-defined filtering should not be applied to independent ROI boundaries of a picture reconfigured for display purposes.

[0044] In the embodiment, the SEI message communication in Table 4 can be simplified by removing one or more of the proposed flags. If all flags are removed, only the presence of the SEI message independent_ROI_across_boundary_filter(payloadSize){} indicates to the decoder that a post-filter should be used to mitigate ROI-related boundary artifacts.

[0045] <Expandability of Area of ​​Interest (ROI)> The latest specification for VVC (Ref.[2]), as discussed in more detail in Ref.[3], describes spatial, quality, and view extensibility using a combination of reference picture resampling (RPR) and reference picture selection (RPS). Based on single-loop decoding and block-based on-the-fly resampling, RPS is used to define predictive relationships between the base layer and one or more extension layers, or more specifically, between coding pictures assigned to either the base layer or one or more extension layers. RPR is used to code a subset of pictures, i.e., pictures in the spatial extension layers, at a higher / lower resolution than the base layer, while predicting from smaller / higher base layer pictures. Figure 3 shows an example of spatial extensibility within the RPS / RPR framework.

[0046] As shown in Figure 3, the bitstream includes two streams: a low-resolution (LR) stream 305 (e.g., standard definition (HD), 2K, etc.) and a higher-resolution (HR) stream 310 (e.g., HD, 2K, 4K, 8K, etc.). Arrows indicate possible interpredictive coding dependencies. For example, HR frame 310-P1 depends on LR frame 305-I. To predict a block in 310-P1, the decoder needs to upscale 305-I. Similarly, HR frame 310-P2 may depend on HR frame 310-P1 and LR frame 305-P1. Any prediction from LR frame 305-P1 requires spatial upscaling from LR to HR. In other embodiments, the order of LR and HR frames may be reversed, so that the base layer is the HR stream and the extended layer is the LR stream. Note that scaling of the base layer picture is not performed explicitly as in SHVC. Instead, it is incorporated into interlayer motion compensation and calculated on the fly. In Ref.[2], the scalability ratio is implicitly derived using a cropping window.

[0047] ROI extensibility is supported in HEVC (Ref.[1]) as part of Annex H, "Scalable high efficiency video coding," commonly known as SHVC. For example, Section F.7.3.2.3.4 defines the syntax elements related to scaled_ref_layer_offset_present_flag[i] and ref_region_offset_present_flag[i]. The relevant parameters are derived from equations (H-2) to (H-21) and (H-67) to (H-68). VVC does not yet support region of interest (ROI) extensibility. As the inventors have recognized, support for ROI extensibility could enable drawing plane size extensibility using the same single-loop VVC decoder without requiring an extension of extensibility as in SHVC.

[0048] For example, given three data layers (e.g., 2K, 4K, and 8K) as shown in Figure 2B, Figure 4 shows an exemplary embodiment of a bitstream that supports drawing surface size scalability using an existing RPS / RPR framework.

[0049] As shown in Figure 4, the bitstream assigns its pictures to three layers or streams: 2K stream 402, 4K stream 405, and 8K stream 410. The arrows illustrate examples of possible interpredictive coding dependencies. For example, it may depend on pixel blocks in 8K frame 410-P2, blocks in 8K frame 410-P1, 4K frame 405-P2, and 2K frame 402-P1. Compared to conventional extensibility schemes using multiple loop decoders, the proposed ROI extensibility scheme has the following advantages and disadvantages. - Advantages: Requires a single-loop decoder and no other tools. The decoder does not need to consider how to handle brick / tile / slice / subpicture boundary issues. - Disadvantages: Decoding the extension layer requires decoded pictures of both the base and extension layers in the decoded picture buffer (DPB), thus requiring a larger DPB size than non-extendable solutions. Since both the base and extension layers need to be decoded, a faster decoder speed may also be required.

[0050] The main difference in enabling ROI extensibility support between the embodiments proposed for SHVC and VVC is that in SHVC, the picture resolution must be the same for all pictures within the same layer. However, in VVC, with RPR support, pictures within the same layer may have different resolutions. For example, in Figure 3, SHVC, 305-I, 305-P1, and 305-P2 must have the same spatial resolution. However, in VVC, with RPR support, 305-I, 305-P1, and 305-P2 can have different resolutions. For example, 305-I and 305-P1 may have a first lower resolution (e.g., 720p), and 305-P2 may have a second lower resolution (e.g., 480p). Embodiments of the present invention aim to support both ROI extensibility between different layers and RPR for pictures within the same layer. Another key difference is that in SHVC, the motion vector from inter-layer predictions is constrained to be zero. However, in VVC, such constraints do not exist, and motion vectors can be zero or non-zero. This reduces the constraints for identifying interlayer correspondence.

[0051] The VVC coding tree only allows coding of entire coding units (CUs). Most standard formats code the picture area in multiples of 4 or 8 pixels, but non-standard formats may need padding in the encoder to fit the minimum CTU size. The same problem existed in HEVC. It was solved by generating a "conformance window" that specifies the picture range to be considered in order to conform to the picture output. The conformance window was also added to VVC (Ref.[2]) and is specified by four variables: conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset. For ease of reference, the following section is copied from Ref.[2].

[0052] conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset specify a sample of the picture in the CVS output from the decoding process, in terms of the rectangular area specified in the picture coordinates for the output. When conformance_window_flag is equal to 0, the values ​​of conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset are assumed to be equal to 0. The fitted cropping window constrains the luma samples by horizontal picture coordinates from SubWidthC*conf_win_left_offsettopic_width_in_luma_samples-(SubWidthC*conf_win_right_offset+1), including both ends, and vertical coordinates from SubHeightC*conf_win_top_offsettopic_height_in_luma_samples-(SubHeightC*conf_win_bottom_offset+1). The value of SubWidthC*(conf_win_left_offset+conf_win_right_offset) must be less than pic_width_in_luma_samples, and the value of SubHeightC*(conf_win_top_offset+conf_win_bottom_offset) must be less than pic_height_in_luma_samples. The variables PicOutputWidthL and PicOutputHeightL are derived as follows:

number

[0053] In the first embodiment, the newly defined ROI offset is combined with the existing offset of the fitting window to derive the magnification. An exemplary embodiment of the proposed syntax elements is shown in Figure 5, which shows the base layer picture 520 and the extended layer picture 502, along with their corresponding fitting windows. The following ROI syntax elements are defined: Base Layer (BL)

number

number

[0054] The width 512 and height 514 of EL picture 502 can be calculated using the above-described equations (7-43) and (7-44) and the adapted window parameters of the extended layer. (For example, pic_width_in_luma_samples may correspond to a width of 512, and PicOutputWidth may correspond to the width of the dotted window 518.)

[0055] For example, Table 5 shows how pic_parameter_set_rbsp(), defined in Section 7.3.2.4 of Ref.[2], is modified to support the new syntax element (edits are highlighted in gray). Table 5: Exemplary syntax for supporting ROI extensibility in VVC [Table 5] num_ref_loc_offsets specifies the number of reference layer position offsets present in the PPS. The value of num_ref_loc_offsets should be in the range of 0 to vps_max_layers_minus1, including both ends. ref_loc_offset_layer_id[i] specifies the nuh_layer_id value to which the i-th reference layer position offset parameter is specified. Note: ref_loc_offset_layer_id[i] does not need to exist between direct reference layers, for example, when a spatial correspondence between an auxiliary picture and its associated primary picture is specified. The i-th reference layer position offset parameter includes the i-th scaled reference layer offset parameter and the i-th reference region offset parameter. A value of 1 for `scaled_ref_layer_offset_present_flag[i]` indicates that the i-th scaled reference layer offset parameter exists within the PPS. A value of 0 for `scaled_ref_layer_offset_present_flag[i]` indicates that the i-th scaled reference layer offset parameter does not exist within the PPS. When it does not exist, the value of `scaled_ref_layer_offset_present_flag[i]` is assumed to be equal to 0. The i-th scaled reference layer offset parameter specifies the spatial correspondence of the picture referencing this PPS to the reference region in the composite picture that has a nuh_layer_id equal to ref_loc_offset_layer_id[i]. The sum of scaled_ref_layer_left_offset[ref_loc_offset_layer_id[i]] and conf_win_left_offset specifies the horizontal offset between the sample in the current picture that is at the same position as the top-left luma sample of the reference region in the decoded picture having a nuh_layer_id equal to ref_loc_offset_layer_id[i], and the top-left luma sample of the current picture within the unit of the subWC luma sample, where subWC is equal to the SubWidthC of the picture referencing this PPS. The value of the sum of scaled_ref_layer_left_offset[ref_loc_offset_layer_id[i]] and conf_win_left_offset is -2 including both ends. 14 ~2 14It should be in the range of -1. If it does not exist, the value of scaled_ref_layer_left_offset[ref_loc_offset_layer_id[i]] is assumed to be equal to 0. The sum of scaled_ref_layer_top_offset[ref_loc_offset_layer_id[i]] and conf_win_top_offset specifies the vertical offset between the sample in the current picture that is at the same position as the top-left luma sample in the reference region of the decoded picture having a nuh_layer_id equal to ref_loc_offset_layer_id[i], and the top-left luma sample of the current picture within the unit of the subHC luma sample, where subHC is equal to the SubHeightC of the picture referencing this PPS. The value of the sum of scaled_ref_layer_top_offset[ref_loc_offset_layer_id[i]] and conf_win_top_offset is -2 including both ends. 14 ~2 14 It should be in the range of -1. If it does not exist, the value of scaled_ref_layer_top_offset[ref_loc_offset_layer_id[i]] is assumed to be equal to 0. The sum of scaled_ref_layer_right_offset[ref_loc_offset_layer_id[i]] and conf_win_right_offset specifies the horizontal offset between the sample in the current picture that is at the same position as the lower-right luma sample in the reference region of the decoded picture having a nuh_layer_id equal to ref_loc_offset_layer_id[i], and the lower-right luma sample of the current picture within the unit of the subWC luma sample, where subWC is equal to the SubWidthC of the picture referencing this PPS. The value of the sum of scaled_ref_layer_right_offset[ref_loc_offset_layer_id[i]] and conf_win_right_offset is -2 including both ends. 14 ~2 14It should be in the range of -1. If it does not exist, the value of scaled_ref_layer_right_offset[ref_loc_offset_layer_id[i]] is assumed to be equal to 0. The sum of scaled_ref_layer_bottom_offset[ref_loc_offset_layer_id[i]] and conf_win_bottom_offset specifies the vertical offset between the sample in the current picture that is at the same position as the bottom right luma sample in the reference region of the decoded picture having a nuh_layer_id equal to ref_loc_offset_layer_id[i], and the bottom right luma sample of the current picture within the unit of the subHC luma sample, where subHC is equal to the SubHeightC of the picture referencing this PPS. The value of the sum of scaled_ref_layer_bottom_offset[ref_loc_offset_layer_id[i]] and conf_win_bottom_offset is -2 including both ends. 14 ~2 14 It should be in the range of -1. If it does not exist, the value of scaled_ref_layer_bottom_offset[ref_loc_offset_layer_id[i]] is assumed to be equal to 0. Let currTopLeftSample, currBotRightSample, colRefRegionTopLeftSample, and colRefRegionBotRightSample be the top-left luma sample of the current picture, the bottom-right luma sample of the current picture, the sample in the current picture located at the same position as the top-left luma sample of the reference region in the decoded picture that has a nuh_layer_id equal to ref_loc_offset_layer_id[i], and the sample in the current picture located at the same position as the bottom-right luma sample of the reference region in the decoded picture that has a nuh_layer_id equal to ref_loc_offset_layer_id[i], respectively. When the value of (scaled_ref_layer_left_offset[ref_loc_offset_layer_id[i]]+conf_win_left_offset) is greater than 0, colRefRegionTopLeftSample is positioned to the right of currTopLeftSample. When the value of (scaled_ref_layer_left_offset[ref_loc_offset_layer_id[i]]+conf_win_left_offset) is less than 0, colRefRegionTopLeftSample is positioned to the left of currTopLeftSample. When the value of (scaled_ref_layer_top_offset[ref_loc_offset_layer_id[i]]+conf_win_top_offset) is greater than 0, colRefRegionTopLeftSample is positioned below currTopLeftSample. When the value of (scaled_ref_layer_top_offset[ref_loc_offset_layer_id[i]]+conf_win_top_offset) is less than 0, colRefRegionTopLeftSample is positioned above currTopLeftSample. When the value of (scaled_ref_layer_right_offset[ref_loc_offset_layer_id[i]]+conf_win_right_offset) is greater than 0, colRefRegionBotRightSample is positioned to the left of currBotRightSample. When the value of (scaled_ref_layer_right_offset[ref_loc_offset_layer_id[i]]+conf_win_right_offset) is less than 0, colRefRegionTopLeftSample is positioned to the right of currBotRightSample. When the value of (scaled_ref_layer_bottom_offset[ref_loc_offset_layer_id[i]]+conf_win_bottom_offset) is greater than 0, colRefRegionBotRightSample is positioned above currBotRightSample. When the value of (scaled_ref_layer_bottom_offset[ref_loc_offset_layer_id[i]]+conf_win_bottom_offset) is less than 0, colRefRegionTopLeftSample is positioned below currBotRightSample. A value of ref_region_offset_present_flag[i] equal to 1 indicates that the offset parameter of the i-th reference region exists within the PPS. A value of ref_region_offset_present_flag[i] equal to 0 indicates that the offset parameter of the i-th reference region does not exist within the PPS. When it does not exist, the value of ref_region_offset_present_flag[i] is assumed to be equal to 0. The offset parameter of the i-th reference region specifies the spatial correspondence of reference regions in the decoded picture that have a nuh_layer_id equal to ref_loc_offset_layer_id[i] for the same decoded picture. refConfLeftOffset[ref_loc_offset_layer_id[i]], refConfTopOffset[ref_loc_offset_layer_id[i]], refConfRightOffset[ref_loc_offset_layer_id[i]], and refConfBottomOffset[ref_loc_offset_layer_id[i]] are used as the values ​​for conf_win_left_offset, conf_win_top_offset, conf_win_right_offset, and conf_win_bottom_offset of the decoded picture having a nuh_layer_id equal to ref_loc_offset_layer_id[i], respectively. The sum of ref_region_left_offset[ref_loc_offset_layer_id[i]] and refConfLeftOffset[ref_loc_offset_layer_id[i]] specifies the horizontal offset between the top-left luma sample of the reference region in the decoded picture having the nuh_layer_id equal to ref_loc_offset_layer_id[i] and the top-left luma sample of the same decoded picture within the unit of subWC luma samples. Here, subWC is equal to SubWidthC of the layer having the nuh_layer_id equal to ref_loc_offset_layer_id[i]. The value of the sum of ref_region_left_offset[ref_loc_offset_layer_id[i]] and refConfLeftOffset[ref_loc_offset_layer_id[i]] should be in the range of -2 14 ~2 14 -1, inclusive. When it does not exist, the value of ref_region_left_offset[ref_loc_offset_layer_id[i]] is presumed to be equal to 0. The sum of ref_region_top_offset[ref_loc_offset_layer_id[i]] and refConfTopOffset[ref_loc_offset_layer_id[i]] specifies the vertical offset between the top-left luma sample of the reference region in the decoded picture having the nuh_layer_id equal to ref_loc_offset_layer_id[i] and the top-left luma sample of the same decoded picture within the unit of subHC luma samples. Here, subHC is equal to SubHeightC of the layer having the nuh_layer_id equal to ref_loc_offset_layer_id[i]. The value of the sum of ref_region_top_offset[ref_loc_offset_layer_id[i]] and refConfTopOffset[ref_loc_offset_layer_id[i]] should be in the range of -2 14 ~2 14It should be in the range of -1. If it does not exist, the value of ref_region_top_offset[ref_loc_offset_layer_id[i]] is assumed to be equal to 0. The sum of ref_region_right_offset[ref_loc_offset_layer_id[i]] and refConfRightOffset[ref_loc_offset_layer_id[i]] specifies the horizontal offset between the lower-right luma sample of the reference region in the decoded picture having a nuh_layer_id equal to ref_loc_offset_layer_id[i] and the lower-right luma sample of the same decoded picture within the unit of subWC luma samples, where subWC is equal to the SubWidthC of the layer having a nuh_layer_id equal to ref_loc_offset_layer_id[i]. The value of the sum of ref_layer_right_offset[ref_loc_offset_layer_id[i]] and refConfRightOffset[ref_loc_offset_layer_id[i]] is -2 including both ends. 14 ~2 14 It should be in the range of -1. If it does not exist, the value of ref_region_right_offset[ref_loc_offset_layer_id[i]] is assumed to be equal to 0. The sum of ref_region_bottom_offset[ref_loc_offset_layer_id[i]] and refConfBottomOffset[ref_loc_offset_layer_id[i]] specifies the vertical offset between the lower right luma sample of the reference region in the decoded picture having a nuh_layer_id equal to ref_loc_offset_layer_id[i] and the lower right luma sample of the same decoded picture within the unit of subHC luma samples, where subHC is equal to the SubHeightC of the layer having a nuh_layer_id equal to ref_loc_offset_layer_id[i]. The value of the sum of ref_layer_bottom_offset[ref_loc_offset_layer_id[i]] and refConfBottomOffset[ref_loc_offset_layer_id[i]] is -2 including both ends. 14 ~2 14 It should be in the range of -1. If it does not exist, the value of ref_region_bottom_offset[ref_loc_offset_layer_id[i]] is assumed to be equal to 0. Let refPicTopLeftSample, refPicBotRightSample, refRegionTopLeftSample, and refRegionBotRightSample be the top-left luma sample of the decoded picture having a nuh_layer_id equal to ref_loc_offset_layer_id[i], the bottom-right luma sample of the decoded picture having a nuh_layer_id equal to ref_loc_offset_layer_id[i], the top-left luma sample of the reference region within the decoded picture having a nuh_layer_id equal to ref_loc_offset_layer_id[i], and the bottom-right luma sample of the reference region within the decoded picture having a nuh_layer_id equal to ref_loc_offset_layer_id[i], respectively. When the value of (ref_region_left_offset[ref_loc_offset_layer_id[i]]+refConfLeftOffset[ref_loc_offset_layer_id[i]]) is greater than 0, refRegionTopLeftSample is positioned to the right of refPicTopLeftSample. When the value of (ref_region_left_offset[ref_loc_offset_layer_id[i]]+refConfLeftOffset[ref_loc_offset_layer_id[i]]) is less than 0, refRegionTopLeftSample is positioned to the left of refPicTopLeftSample. When the value of (ref_region_top_offset[ref_loc_offset_layer_id[i]]+refConfTopOffset[ref_loc_offset_layer_id[i]]) is greater than 0, refRegionTopLeftSample is located below refPicTopLeftSample. When the value of (ref_region_top_offset[ref_loc_offset_layer_id[i]]+refConfTopOffset[ref_loc_offset_layer_id[i]]) is less than 0, refRegionTopLeftSample is positioned above refPicTopLeftSample. When the value of (ref_region_right_offset[ref_loc_offset_layer_id[i]]+refConfRightOffset[ref_loc_offset_layer_id[i]]) is greater than 0, refRegionBotRightSample is located to the left of refPicBotRightSample. When the value of (ref_region_right_offset[ref_loc_offset_layer_id[i]]+refConfRightOffset[ref_loc_offset_layer_id[i]]) is less than 0, refRegionBotRightSample is located to the right of refPicBotRightSample. When the value of (ref_region_bottom_offset[ref_loc_offset_layer_id[i]]+refConfBottomOffset[ref_loc_offset_layer_id[i]]) is greater than 0, refRegionBotRightSample is positioned above refPicBotRightSample. When the value of (ref_region_bottom_offset[ref_loc_offset_layer_id[i]]+refConfBottomOffset[ref_loc_offset_layer_id[i]]) is less than 0, refRegionBotRightSample is positioned below refPicBotRightSample.

[0056] Given the proposed syntax elements, the corresponding VVC sections may, in embodiments, be modified as follows, but not limited to: Expressions denoted as (7-xx) and (8-xx) indicate new expressions that need to be inserted into the VVC specification and are renumbered as necessary.

number

number

number

[0057] To support ROI (Risk Area Size) expandability, the specification must be modified as follows:

number

number

number

number

[0058]

number

number

[0059] In another embodiment, unlike SHVC, VVC has no constraints on the size of motion vectors during interlayer coding, so when finding pixel correspondences between ROI regions, it is not necessary to consider the upper-left position of the reference layer ROI and the scaled reference layer (current picture) ROI. Accordingly, references to fRefLeftOffset, fRefTopOffset, fCurLeftOffset and fCurTopOffset can be removed from all of the above equations.

[0060] Figure 6A provides an exemplary summary of the processing flow described above. As shown in Figure 6A, in step 605, the decoder may receive syntax parameters related to the fitting window (e.g., conf_win_xxx_offset, where xxx is left, top, right, or bottom), a scaled reference layer offset for the current picture (e.g., scaled_ref_layer_xxx_offset[]), and a reference layer region offset (e.g., ref_region_xxx_offset[]). If it is not intercoding (step 610), decoding proceeds as in single-layer decoding; otherwise, in step 615, the decoder calculates the fitting window for both the reference and the current picture (e.g., using equations (7-43) and (7-44)). If interlayer coding is not performed (step 620), it is still necessary to calculate the RPR multiplier for interpretation for pictures with different resolutions on the same layer, and then decoding proceeds as in single-layer decoding; otherwise (with interlayer coding), the decoder calculates the multiplier for the current and reference pictures based on the received offset (for example, by calculating hori_scale_fp and vert_scale_fp in equations (8-753) and (8-754)).

[0061] As mentioned above (see, for example, equations (8-x1) to (8-x2)), the decoder needs to calculate the width and height of the reference ROI (e.g., fRefWidth and fRefHeight) by subtracting the left and right offset values ​​of the reference layer (e.g., RefLayerRegionLeftOffset and RefLayerRegionRightOffset) and the top and bottom offsets (e.g., RefLayerRegionTopOffset and RefLayerRegionBottomOffset) from the PicOutputWidthL and PicOutputHeightL of the reference layer picture.

[0062] Similarly (see, for example, equations (8-x3) to (8-x4)), the decoder needs to calculate the width and height (e.g., fCurWidth and fCurHeight) of the current ROI by subtracting the left and right offset values ​​(e.g., ScaledRefLayerLeftOffset and ScaledRefLayerRightOffset) and the top and bottom offsets (e.g., ScaledRefLayerTopOffset and ScaledRefLayerBottomOffset) from the PicOutputWidthL and PicOutputHeightL of the current layer picture. Given these adjusted sizes of the current and reference ROIs, with minimal additional modifications required (as shown in the gray-highlighted section above), the decoder determines the horizontal and vertical scaling (e.g., the processing in equations (8-755) to (8-766) as in existing VVC RPR blocks (see, for example, equations (8-753) and (8-754)).

[0063] In equations (8-x5) to (8-x8), adjusted left and top offsets are also calculated to determine the correct position of the reference and current ROI relative to the top-left corner of the fitting window for proper pixel mapping.

[0064] In the second embodiment, the definitions of the ref_region_xxx_offset[] and scaled_ref_region_xxx_offset[] offsets may be redefined to combine both the fitted window offset and the ROI offset (for example, by adding them together). For example, in Table 5, scaled_ref_layer_xxx_offset may be replaced with scaled_ref_layer_xxx_offset_sum, which is defined as follows:

number

[0065] A similar definition can be generated for ref_region_xxx_offset_sum, where xxx = bottom, top, left, and right. As will be explained, the processing in step 615 may be combined with the processing in step 625, so these parameters allow the decoder to skip step 615.

[0066] For example, in Figure 6A, a) In step 615, PicOutputWidthL may be calculated by subtracting the left and right offsets of the fitting window from the picture width (see, for example, equation (7-43)). b) Let fCurWidth = PicOutputWidthL. c) Next, in step 625, we adjust fCurWidth (see, for example, (8-x3)) by subtracting the ScaledRefLayer left and right offsets, but from equation (7-xx), these are based on the scaled_ref_layer left and right offsets. For example, if we now only want to adjust the width of the ROI, in simplified notation (i.e., by ignoring the SubWidthC scaling parameter), we can calculate it as follows: Picture output width =Picture width -(Fitment window left offset + Fitment window right offset) (2) ROI current width = Picture output width - (ROI current left offset + ROI current right offset) (3) By combining equations (2) and (3), ROI current width =Picture width -((Fitment window left offset + ROI current left offset) +(Fitment window right offset + ROI current right offset)) (4)

[0067] Assuming the following, ROI current left total offset = fitting window left offset + ROI current left offset ROI current right total offset = Fitting window right offset + ROI current right offset Equation (4) can be simplified as follows: Current ROI width = Picture width - ((Current ROI left total offset) + (Current ROI right total offset)) (5) The definition of the new "sum" offset (e.g., ROI current left sum offset) corresponds to the definition of ref_region_left_offset_sum defined in equation (1).

[0068] Therefore, if we redefine the scaled_ref_layer left and right offsets to include the sum of the layer's conf_win_left_offset as described above, we proceed to blocks 615 and 625 to calculate the width and height of the current and reference ROI (for example, equations (2) and (3) can be combined into one (for example, equation (5))) (for example, step 630).

[0069] As shown in Figure 6B, steps 615 and 625 can be combined into a single step 630. Compared to Figure 6A, this approach saves some additions, but the revised offsets (e.g., scaled_ref_layer_left_offset_sum[]) are now larger and therefore require more bits to be encoded in the bitstream. Note that the conf_win_xxx_offset values ​​may be different for each layer, and these values ​​can be extracted from the PPS information in each layer.

[0070] In a third embodiment, horizontal and vertical scaling (e.g., hori_scale_fp and vert_scale_fp) may be explicitly signaled between interlayer pictures. In such a scenario, the horizontal and vertical scaling, as well as the top and left offsets, must be communicated for each layer.

[0071] A similar approach is applicable to embodiments having pictures that incorporate multiple ROIs, each using arbitrary upsampling and downsampling filters. <References> Each of the references listed here is incorporated here in its entirety by reference. [1] High efficiency video coding, H.265, Series H, Coding of moving video, ITU, (02 / 2018) [2] B. Bross, J. Chen, and S. Liu, “Versatile Video Coding (Draft 6),” JVET output document, JVET-O2001, vE, uploaded July 31, 2019 [3] S. Wenger, et al., “AHG8: Spatial scalability using reference picture resampling,” JVET-O0045, JVET Meeting, Gothenburg, SE, July 2019 [4] R. Skupin et al., AHG12: “On filtering of independently coded region,” JVET-O0494 (v3), JVET Meeting, Gothenburg, SE, July 2019 <Example of a computer system implementation> Embodiments of the present invention may be implemented by computer systems, systems configured within electronic circuits and components, integrated circuit (IC) devices such as microcontrollers, FPGAs (field programmable gate arrays), or other configurable or programmable logic devices (PLDs), individual time or digital signal processors (DSPs), application-specific ICs (ASICs), and / or devices including one or more such systems, devices, or components. Computers and / or ICs may execute, control, or perform instructions related to the expandability of the drawing surface size as described in this specification. Computers and / or ICs may calculate any of the various parameters or values ​​related to the expandability of the drawing surface size as described in this specification. Embodiments of images and videos may be implemented in hardware, software, firmware, and various combinations thereof.

[0072] A particular implementation of the present invention includes a computer processor that executes software instructions causing the processor to perform the method of the present invention. For example, one or more processors among a display, encoder, set-top box, transcoder, etc., may perform the method relating to the above-described drawing surface size expandability by executing software instructions in the processor's accessible program memory. Embodiments of the present invention may be provided in the form of a program product. The program product may include any non-temporary tangible medium that carries a set of computer-readable signals, which, when executed by a data processor, causes the data processor to perform the method of the present invention. The program product according to the present invention may be any of the various non-temporary tangible forms. The program product may include a physical medium such as, for example, a magnetic data storage medium including a floppy disk, an optical data storage medium including a hard disk drive, a CD-ROM, a DVD, an electronic data storage medium including ROM, flash RAM, etc. The computer-readable signals on the program product may be optically compressed or encrypted.

[0073] Although components (e.g., software modules, processors, parts, devices, circuits, etc.) have been mentioned above, unless otherwise specified, any mention of such components (including any mention of “means”) should be interpreted as including equivalents of such components, any components that perform the function of the described components (e.g., functionally equivalent), and components that are not structurally equivalent to the structures of the disclosures that perform the function in the illustrated exemplary embodiments of the present invention.

[0074] <Equivalents, Extensions, Alternatives, and Miscellaneous> Accordingly, exemplary embodiments relating to the expandability of the drawing surface size are described. In the above specification, embodiments of the present invention have been described with reference to a number of specific details that may vary from implementation to implementation. Accordingly, the sole and exclusive indication of what the present invention is, and what the applicant intends to be the present invention, is stated in the claims issued by this application in a specific form, including any subsequent amendments. Any definitions of terms included in such claims that are expressly stated in this specification should govern the meaning of such terms as used in the claims. Accordingly, any limitations, elements, features, advantages, or attributes not expressly stated in the claims should not limit the scope of the claims in any way. The specification and drawings should therefore be considered as explanatory, not restrictive.

Claims

1. A device for encoding a picture stream into a coding bitstream, wherein the device comprises: Receive a stream of input pictures and input, A processor that encodes the aforementioned input picture into a coding bitstream, Includes, The processor, with respect to the current picture to be coded into the coding bitstream, Determine the current picture width and current picture height, which are unsigned integer values. A first offset parameter is determined to determine the rectangular area on the current picture, and the first offset parameter has a signed integer value. Based on the current picture width, the current picture height, and the first offset parameter, the current range width and current range height of the rectangular area on the current picture are calculated. Regarding the reference range, access the reference range width, reference range height, reference range left offset, and reference range top offset. Accessing the aforementioned reference range width and reference range height further involves, for reference pictures in the coding picture stream, Access the reference picture width and reference picture height, A second offset parameter is determined for determining the rectangular area within the reference picture, and the second offset parameter has a signed integer value. Based on the aforementioned reference picture width, reference picture height, and second offset parameter, the reference range width and reference range height of the rectangular area within the reference picture are calculated. Based on the second offset parameter, the left offset and the upper offset of the reference range are calculated. This includes, The horizontal scale (hori_scale_fp) and vertical scale (vert_scale_fp) are calculated as follows: [Math 1] [Math 2] fRefWidth represents the reference range width, fCurWidth represents the current range width, fRefHeight represents the reference range height, and fCurHeight represents the current range height. Based on the first offset parameter, the left offset adjustment and the upper offset adjustment for the current range are calculated. A device that performs motion compensation based on the horizontal magnification, the vertical magnification, the left offset adjustment, the upper offset adjustment, the left offset of the reference range, and the upper offset of the reference range.

2. The apparatus according to claim 1, wherein the first offset parameter includes a left offset, an upper offset, a right offset, and a lower offset.

3. One or more of the left offset, upper offset, right offset, or lower offset are -2 14 From 2 14 The apparatus according to claim 2, having a value between the following two values.

4. A method for transmitting a coding bitstream generated by a video encoding device, the method comprising the steps of transmitting the coding bitstream, and generating a coding picture in the coding bitstream, A step of determining the current picture width and current picture height, which have unsigned integer values, A step of determining a first offset parameter that determines the rectangular area on the current picture, wherein the first offset parameter has a signed integer value, A step of calculating the current range width and current range height of the rectangular area on the current picture based on the current picture width, the current picture height, and the first offset parameter, A step of accessing the reference range width, reference range height, reference range left offset, and reference range top offset, wherein the step of accessing the reference range width and the reference range height further applies to the reference picture in the bitstream of the coding picture, Access the reference picture width and reference picture height, A second offset parameter is determined for determining the rectangular area within the reference picture, and the second offset parameter has a signed integer value. Based on the aforementioned reference picture width, reference picture height, and second offset parameter, the reference range width and reference range height of the rectangular area within the reference picture are calculated. Based on the second offset parameter, the left offset and the upper offset of the reference range are calculated. Steps that include, The steps are as follows to calculate the horizontal scale (hori_scale_fp) and the vertical scale (vert_scale_fp): [Math 3] [Math 4] fRefWidth represents the reference range width, fCurWidth represents the current range width, fRefHeight represents the reference range height, and fCurHeight represents the current range height, with a step, A step of calculating the left offset adjustment and the upper offset adjustment of the current range based on the first offset parameter, A step of performing motion compensation based on the horizontal magnification, the vertical magnification, the left offset adjustment, the upper offset adjustment, the left offset of the reference range, and the upper offset of the reference range, Methods that include...

5. The method according to claim 4, wherein the first offset parameter includes a left offset, an upper offset, a right offset, and a lower offset.

6. One or more of the left offset, upper offset, right offset, or lower offset are -2 14 From 2 14 The method according to claim 5, having a value between the following two values.