Video decoder, method for decoding, video encoder, method for encoding, and non-transitory computer readable medium

The patent addresses inefficiencies in HEVC by implementing dynamic sample aspect ratio signaling and restricted reference picture resampling, enhancing parallel processing and reducing implementation costs in video codecs.

JP2025103003AActive Publication Date: 2025-07-08FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025063599
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-09-24
Filing Date
2025-04-08
Publication Date
2025-07-08
Estimated Expiration
2040-09-24

AI Technical Summary

Technical Problem

Existing video codecs like HEVC do not efficiently support parallel processing capabilities in video encoders and decoders, particularly in handling variable aspect ratios and reference picture resampling, leading to inefficiencies and increased implementation burden.

Method used

Introduces dynamic sample aspect ratio signaling and restricted reference picture resampling methods to enhance video coding efficiency, allowing for flexible region-based reference picture resampling and reduced implementation costs.

Benefits of technology

Enhances video coding efficiency by maintaining consistent aspect ratios and optimizing resampling processes, reducing hardware requirements and improving decoding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025103003000001_ABST
    Figure 2025103003000001_ABST
Patent Text Reader

Abstract

To provide a decoder and a decoding method for decoding a plurality of pictures having different SAR values.SOLUTION: A generic video decoder 151 executes operations including the steps of: receiving an encoded video signal having a sequence parameter set (SPS) with an indication that encoded picture data and sample aspect ratio (SAR) can be modified in a video sequence; receiving a supplemental enhancement information (SEI) message within the encoded video signal having a SAR value for one or more first groups of multiple pictures of the video sequence; decoding the encoded picture data for obtaining a plurality of decoded pictures; and outputting the plurality of decoded pictures and sample aspect information for one or more of the decoded pictures.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a video decoder, a method for decoding, a video encoder, a method for encoding, and a non-transitory computer-readable medium.

[0002] H.265 / HEVC (HEVC = High Efficiency Video Coding) is a video codec that already provides tools for improving or enabling parallel processing in an encoder and / or decoder. For example, HEVC supports the sub-division of a picture into an array of tiles that are encoded independently of each other. Another concept supported by HEVC relates to WPP, according to which the CTU rows or CTU lines of a picture may be processed in parallel from left to right, for example in stripes, in accordance with some minimum CTU offset in the processing of consecutive CTU lines (CTU = coding tree unit). However, it is desirable to have a video codec nearby that more efficiently supports the parallel processing capabilities of a video encoder and / or a video decoder.

[0003] Hereinafter, VCL partitioning according to the state of the art will be described (VCL = video coding layer).

[0004] Normally, in video coding, the encoding process of picture samples requires smaller partitions, and the samples are divided into several rectangular regions for combined processing such as predictive coding or transform coding. Therefore, a picture is divided into blocks of a specific size that is constant during the encoding of a video sequence. In H.264 / AVC, standard fixed-size blocks of 16×16 samples, so-called macroblocks, are used (AVC = Advanced Video Coding).

[0005] In the latest HEVC standard (see ISO / IEC, ITU-T. High efficiency video coding. ITU-T Recommendation H.265 | ISO / IEC 23008 10 (HEVC), edition 1, 2013; edition 2, 2014), there are coding tree blocks (CTBs) or coding tree units (CTUs) with a maximum size of 64×64 samples. In further descriptions of HEVC, the more general term CTU is used for such types of blocks.

[0006] CTUs are processed in raster scan order starting from the top-left CTU and proceeding down to the bottom-right CTU in picture line units.

[0007] The encoded CTU data is organized into a kind of container called a slice. Originally, in previous video coding standards, a slice meant a section that comprised one or more consecutive CTUs of a picture. Slices are used for the segmentation of the encoded data. From another perspective, a complete picture can also be defined as one large segment, and thus, historically, the term slice still applies. In addition to the encoded picture samples, a slice also includes additional information related to the encoding process of the slice itself, which is placed in the so-called slice header.

[0008] According to the state of the art, the VCL (video coding layer) also has techniques for fragmentation and spatial division. Such divisions can be applied to video coding for various reasons, among which are processing load balancing in parallelization, CTU size matching in network transmission, error reduction, etc.

[0009] Another example relates to RoI (RoI = Region of Interest) coding, in which, for example, the central region of a picture that a viewer can select with a zoom-in operation (decoding of only the RoI), or intra data (typically put into one frame of a video sequence) is temporally distributed over several consecutive frames as a series of intra blocks that locally reset the temporal prediction chain in the same manner as intra pictures do for the entire picture plane, for example, by swiping over the picture plane. In the latter case, two regions exist in each picture, one that has been recently reset and one that is potentially affected by errors and error propagation.

[0010] Reference Picture Resampling (RPR) is a technique used in video coding to adapt the quality / rate of video by not only using a coarse quantization parameter but also potentially adapting the resolution of each transmitted picture. Thus, the reference used for inter prediction may have a different size from the picture currently being predicted for coding. Basically, RPR requires resampling processes in the prediction loop, for example, defined upsampling and downsampling filters.

[0011] Depending on the flavor, RPR can result in a change in the coded picture size in any picture, or be restricted to occur only in specific pictures, such as for adaptive HTTP streaming, only at specific positions defining the boundaries of segments. SUMMARY OF THE INVENTION

[0012] An object of the present invention is to provide an improved concept for video coding and video decoding.

[0013] The object of the present invention is solved by the subject matter of the independent claims.

[0014] Preferred embodiments are provided in the dependent claims.

Brief Description of the Drawings

[0015] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings:

[0016]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5a

Figure 5b

Figure 6a

Figure 6b

Figure 7

Figure 8

Figure 9

Modes for Carrying Out the Invention

[0017] The following description of the drawings begins with an explanation of an encoder and a decoder of a block-based prediction codec for encoding a video picture in order to form an example of an encoding framework that can incorporate embodiments of the present invention. Each encoder and decoder is described with respect to FIGS. 7 through 9. Thereafter, an explanation of embodiments of the concepts of the present invention is presented along with an explanation of how to incorporate such concepts into the encoders and decoders of FIGS. 7 and 8, respectively. However, FIGS. 1 through 3 and the embodiments described below may also be used to form encoders and decoders that do not operate according to the encoding framework that underlies the encoders and decoders of FIGS. 7 and 8.

[0018] FIG. 7 shows an apparatus, designated by reference numeral 10, for predictive encoding picture 12 into data stream 14 using a video encoder, typically transform-based residual coding. FIG. 8 shows a corresponding video decoder 20, i.e., an apparatus 20 configured to predictively decode picture 12' from data stream 14, also using, e.g., transform-based residual decoding, where the apostrophe is used to indicate that picture 12', reconstructed by decoder 20, deviates from picture 12 originally encoded by apparatus 10 with respect to the coding loss introduced by quantization of the prediction residual signal. FIGS. 7 and 8 exemplarily use transform-based prediction residual coding, but embodiments of the present application are not limited to this kind of prediction residual coding. This also applies to other details described with respect to FIGS. 7 and 8, as outlined below.

[0019] Encoder 10 is configured to subject the prediction residual signal to a spatial spectral transform and encode the thus obtained prediction residual signal into data stream 14. Similarly, decoder 20 is configured to decode the prediction residual signal from data stream 14 and subject the thus obtained prediction residual signal to a spectral spatial transform.

[0020] Internally, the encoder 10 may include a prediction residual signal former 22 that generates a prediction residual 24, thereby enabling measurement of the deviation of the original signal from the prediction signal 26, for example, from picture 12. The prediction residual signal former 22 may be a subtractor that subtracts the prediction signal from the original signal, for example, from picture 12. Then, the encoder 10 further includes a transformer 28 that applies the prediction residual signal 24 to a spatial spectrum conversion to obtain a spectral domain prediction residual signal 24'. This residual signal is quantized by a quantizer 32 also constituted by the encoder 10. Thus, the quantized prediction residual signal 24 is encoded in the bitstream 14. Optionally, the encoder 10 constitutes an entropy encoder 34 that converts and quantizes the prediction residual signal into a data stream by entropy encoding. The prediction signal 26 is encoded in the data stream 14 by the prediction stage 36 of the encoder 10 or is generated based on the prediction residual signal 24 that is decodable from the data stream 14. For this purpose, as shown in FIG. 7, the prediction stage 36 may internally include a dequantizer 38 that dequantizes the prediction residual signal 24'' to obtain a spectral domain prediction residual signal 24''' corresponding to the signal 24' excluding quantization loss. Then, an inverse transformer 40 inverse-transforms the latter prediction residual signal 24''', for example, by applying a spectral-spatial conversion, to obtain a prediction residual signal 24'''' corresponding to the original prediction residual signal 24 excluding quantization loss. Then, a combiner 42 of the prediction stage 36 can recombine the prediction signal 26 and the prediction residual signal 24'''' by, for example, addition, to obtain a reconstructed signal 46, the original signal, for example, the reconstruction of picture 12. The reconstructed signal 46 may correspond to the signal 12'. And the prediction module 44 of the prediction stage 36 generates the prediction signal 26 based on the signal 46 by using, for example, spatial prediction, i.e., intra-picture prediction, and / or temporal prediction, i.e., inter-picture prediction.

[0021] Similarly, as shown in FIG. 8, the decoder 20 may be internally configured from components corresponding to the prediction stage 36 and interconnected in a corresponding manner. In particular, the entropy decoder 50 of the decoder 20 can entropy-decode the quantized spectral region prediction residual signal 24'' from the data stream, where the inverse quantizer 52, the inverse transformator 54, the combiner 56 and the prediction module 58 are interconnected and cooperate in the manner described above with respect to the modules of the prediction stage 36 to recover the reconstructed signal based on the prediction residual signal 24'', whereby, as shown in FIG. 8, the output of the combiner 56 produces the reconstructed signal, i.e., the picture 12'.

[0022] Although not specifically described above, it is readily apparent that, for example, encoder 10 may set some encoding parameters, such as a certain rate and distortion-related criteria, e.g., a method of optimizing encoding cost, e.g., according to a certain optimization scheme, e.g., prediction mode, motion parameters, etc. For example, encoder 10 and decoder 20, as well as corresponding modules 44, 58, may each support different prediction modes, such as intra encoding mode and inter encoding mode. The granularity at which the encoder and decoder switch between these prediction mode types may correspond to the subdivision of each of pictures 12 and 12’ into encoding segments or encoding blocks. In units of these encoding segments, for example, a picture may be subdivided into blocks to be intra encoded and blocks to be inter encoded. Intra encoded blocks are predicted based on the spatial, already encoded / decoded neighborhood of each block, as will be outlined in more detail below. There may be several intra encoding modes, and each intra encoding mode may be selected for each intra encoding segment, including a directional or angular intra encoding mode in which each segment is filled by extrapolating sample values of the neighborhood along a certain direction specific to each directional intra encoding mode to each intra encoding segment. The intra encoding mode may also include, for example, a DC encoding mode in which the prediction for each intra encoded block assigns a DC value to all samples within each intra encoded segment, and / or a planar intra encoding mode in which the prediction for each block is approximated or determined to be a spatial distribution of sample values described by a two-dimensional linear function having a driving slope and an offset of a plane defined by a two-dimensional linear function based on adjacent samples over the sample positions of each intra encoded block, etc., with one or more additional modes. In contrast, inter encoded blocks may be predicted temporally, for example.In the case of an inter-coded block, the motion vector may be signaled within the data stream. The motion vector indicates the spatial displacement of a portion of a previously encoded picture of the video to which picture 12 belongs. The previously encoded / decoded pictures are sampled to obtain the prediction signal for each inter-coded block. This means that in addition to the residual signal coding provided by the data stream 14, such as the entropy-coded transform coefficient levels representing the quantized spectral region prediction residual signal 24'', the data stream 14 may also encode additional optional parameters, such as coding mode parameters for assigning coding modes to various blocks, prediction parameters for some blocks such as motion parameters for inter-coded segments, and parameters for controlling and signaling the subdivision of pictures 12 and 12' into segments. The decoder 20 uses these parameters to subdivide the picture in the same way as the encoder did, assign the same prediction mode to the segments, and perform the same prediction to result in the same prediction signal.

[0023] FIG. 9 shows the relationship between, on the one hand, the reconstructed signal, e.g., the reconstructed picture 12’, and, on the other hand, the combination of the prediction residual signal 24’’’’ signaled in the data stream 14 and the prediction signal 26. As already described above, the combination may be additional. The prediction signal 26 is shown in FIG. 9 as the subdivision of the picture area into intra-coded blocks, exemplified using hatching, and inter-coded blocks, exemplified without hatching. The subdivision may be any subdivision, such as regularly subdividing the picture area into rows and columns of square or non-square blocks, or a multi-tree subdivision that subdivides picture 12 from a tree root block into a plurality of leaf blocks of various sizes, such as a quadtree subdivision. The picture area is first subdivided into rows and columns of the tree root block and then further subdivided into one or more leaf blocks according to a recursive multi-tree subdivision, and a mixture of these is shown in FIG. 9.

[0024] Also, the data stream 14 may have an intra-coding mode coded therein for the intra-coded blocks 80, which assigns to each intra-coded block 80 one of several supported intra-coding modes. For the inter-coded blocks 82, the data stream 14 may have one or more motion parameters coded therein. Generally speaking, the inter-coded blocks 82 are not limited to being temporally coded. Alternatively, the inter-coded blocks 82 may be any block predicted from a previously coded part beyond the current picture 12 itself, such as a previously coded picture of the video to which picture 12 belongs, or, if the encoder and decoder are scalable encoders and decoders respectively, another view or a hierarchically lower picture.

[0025] Also, the prediction residual signal 24'''' in FIG. 9 is illustrated as a subdivision of the picture area into blocks 84. These blocks may sometimes be referred to as transform blocks to distinguish them from the same blocks as the encoding blocks 80 and 82. In fact, FIG. 9 shows that the encoder 10 and the decoder 20 can each use two different subdivisions of the pictures 12 and 12', one subdivision into blocks, i.e., the encoding blocks 80 and 82 respectively, and another subdivision into the transform blocks 84. The two subdivisions may be the same. For example, each of the encoding blocks 80 and 82 may simultaneously form the transform block 84. However, FIG. 9 shows a case where, for example, the subdivision into the transform block 84 is such that any boundary between the two blocks 80 and 82 overlaps the boundary between the two blocks 84, or alternatively, each of the blocks 80, 82 either coincides with one of the transform blocks 84 or coincides with a cluster of the transform blocks 84, thus forming an extension of the subdivision into the encoding blocks 80, 82. However, the subdivisions may be determined or selected independently of each other so that the transform block 84 can cross the block boundaries between the blocks 80, 82. As far as the subdivision into the transform block 84 is concerned, it is the same as that presented for the subdivision into the blocks 80, 82. For example, the block 84 may be the result of a regular subdivision of the picture area into blocks (with or without an arrangement into rows and columns), the result of a recursive multi-tree subdivision of the picture area, or a combination thereof, or the result of any other type of blocking. Incidentally, it should be noted that the blocks 80, 82, and 84 are not limited to being square, rectangular, or any other shape.

[0026] FIG. 9 further shows that the combination of the prediction signal 26 and the prediction residual signal 24'''' directly results in the reconstructed signal 12'. However, it should be noted that according to an alternative embodiment, two or more prediction signals 26 may be combined with the prediction residual signal 24'''' to obtain the picture 12' as well.

[0027] In FIG. 9, the conversion block 84 has the following importance. The converter 28 and the inverse converter 54 perform their conversions in units of these conversion blocks 84. For example, many codecs use some kind of DST or DCT for all the conversion blocks 84. Some codecs allow skipping the conversion so that for some of the conversion blocks 84, the prediction residual signal is directly encoded in the spatial domain. However, according to the embodiments described below, the encoder 10 and the decoder 20 are configured to support some conversions. For example, the conversions supported by the encoder 10 and the decoder 20 can include the following: · DCT-II (or DCT-III), where DCT represents Discrete Cosine Transform · DST-IV, where DST represents Discrete Sine Transform · DCT-IV · DST-VII · Identity Transform (IT)

[0028] Of course, the converter 28 supports all of the forward conversion versions of these conversions, while the decoder 20 or the inverse converter 54 supports the corresponding reverse or inverse versions: · Inverse DCT-II (or Inverse DCT-III) · Inverse DST-IV · Inverse DCT-IV · Inverse DST-VII · Identity Transform (IT)

[0029] The following description provides further details on which conversions can be supported by the encoder 10 and the decoder 20. In any case, note that the set of supported conversions may include only one conversion, such as one spectral-spatial conversion or spatial-spectral conversion.

[0030] As already outlined above, FIGS. 7 to 9 are presented as examples in which the concepts of the present invention, which will be further described below, can be implemented to form specific examples of the encoder and decoder according to the present application. The encoders and decoders of FIGS. 7 and 8 may each represent possible implementations of the encoders and decoders described below. However, FIGS. 7 and 8 are merely examples. However, the encoder according to an embodiment of the present application is outlined in detail below. For example, it may be a still image encoder instead of a video encoder, does not support inter prediction, or the subdivision into blocks 80 is performed in a manner different from the method illustrated in FIG. 9. It may perform block-based encoding of picture 12 different from the encoder of FIG. 7. Similarly, the decoder according to an embodiment of the present invention can perform block-based decoding of picture 12' from data stream 14 using the encoding concept further outlined below. However, for example, it may be a still picture decoder instead of a video decoder, does not support intra prediction, or it subdivides picture 12' into blocks in a manner different from that described with respect to FIG. 9, and / or it may be different from the decoder 20 of FIG. 8 in that, for example, it does not derive a prediction residual from data stream 14 in the transform domain.

[0031] Hereinafter, a general video encoder according to an embodiment is described in FIG. 1, a general video decoder according to an embodiment is described in FIG. 2, and a general system according to an embodiment is described in FIG. 3.

[0032] FIG. 1 shows a general video encoder 101 according to an embodiment.

[0033] The video encoder 101 is configured to encode a plurality of pictures of video by generating an encoded video signal, and each of the plurality of pictures comprises original picture data.

[0034] The video encoder 101 includes a data encoder 110 configured to generate an encoded video signal with encoded picture data, and the data encoder is configured to encode a plurality of pictures of the video into the encoded picture data.

[0035] Furthermore, the video encoder 101 includes an output interface 120 configured to output the encoded picture data of each of the plurality of pictures.

[0036] Figure 2 shows a general video decoder 151 according to an embodiment.

[0037] The video decoder 151 is configured to decode an encoded video signal with encoded picture data to reconstruct a plurality of pictures of the video.

[0038] The video decoder 151 includes an input interface 160 configured to receive the encoded video signal.

[0039] Furthermore, the video decoder includes a data decoder 170 configured to reconstruct a plurality of pictures of the video by decoding the encoded picture data.

[0040] Figure 3 shows a general system according to an embodiment.

[0041] The system includes the video encoder 101 of FIG. 1 and the video decoder 151 of FIG. 2.

[0042] The video encoder 101 is configured to generate an encoded video signal.

[0043] The video decoder 151 is configured to decode the encoded video signal to reconstruct the pictures of the video.

[0044] The first aspect of the present invention is described in claims 1 to 33. The first aspect provides a sample aspect ratio signal.

[0045] The second aspect of the present invention is described in claims 34 to 72. The second aspect provides Reference Picture Resampling restrictions to reduce the implementation burden.

[0046] The third aspect of the present invention is described in claims 73 to 131. The third aspect provides flexible region-based reference for zooming of reference picture resampling, and in particular provides a more efficient address zoom usage case.

[0047] Hereinafter, the first aspect of the present invention will be described in detail.

[0048] In particular, the first aspect provides sample aspect ratio signalling.

[0049] The sample aspect ratio (SAR) is related to correctly presenting the encoded video to the consumer. As a result, when the aspect ratio of the encoded sample array changes over time via RPR (e.g., by subsampling in one dimension), the aspect ratio of the presented picture can be kept constant as intended.

[0050] The latest SAR-signalled video sequence in Video Usability Information (VUI) in a sequence parameter set (SPS) such as HEVC or AVC only allows setting a constant SAR for the entire video sequence. For example, SAR changes are only permitted at the start of the encoded video sequence (e.g., the sample aspect ratio is constant for each encoded video sequence).

[0051] Accordingly, as part of the present invention, a new mode of SAR signalling is introduced into video encoding. The sequence level parameter set, i.e., the SPS, includes the following instructions: ·RPR is being used (therefore, the size of the coded picture may change). ·The actual SAR is not given in the VUI. ·Instead, the SAR of the coded video is indicated as being dynamic and may change within a CVS (coded video sequence). ·The actual SAR of the coded picture is indicated via an SEI (supplemental enhancement information) message at the resolution switching point.

Table 1

[0052] Dynamic SAR information SEI message

Table 2

[0053] Similarly, vui_aspect_ratio_constant_flag may be adopted, for example.

[0054] vui_aspect_ratio_constant_flag may be an indicator indicating, for example, whether the aspect ratio of the samples is constant for the video sequence or whether the aspect ratio of the samples can be changed within the video sequence.

[0055] For example, if vui_aspect_ratio_constant_flag can be set to 0 (or can be set to FALSE, or can be set to -1, for example), this may indicate that dynamic SAR information is present, for example, in the SEI message.

[0056] In an alternative embodiment, the SAR information within the VUI (e.g., SPS) is used as the default, which is used unless the SEI message is available. The information of the SEI message is overwritten by the information of the SPS.

Table 3

[0057] In another embodiment, the SAR information is associated with the picture resolution and signaled in the PPS (Picture Parameter Set), and the picture resolution is signaled. The default SAR is signaled in the SPS. When the SAR changes for a specific picture resolution, a different SAR is signaled, overwriting the default SAR.

[0058] SPS VUI:

Table 4

[0059] Also, in the case of the SEI, the SPS can further indicate that the SAR may be changed and that the SAR is updated to the PPS (similar to the previous aspect_ratio_dynamic_SEI_present_flag). Therefore, in some applications, it is possible to impose constraints or limitations so as not to change the SAR, which can facilitate implementation or RPR / ARC.

[0060] PPS:

Table 5

[0061] When the pps_aspect_ratio_info_present_flag is set to 0, the default SAR is obtained from the SPS; otherwise, the actual SAR is provided.

[0062] Next, the second aspect of the present invention will be described in detail.

[0063] In particular, the second aspect provides signaling regarding constraints for reference picture resampling.

[0064] By restricting the RPR method in various ways, the implementation burden can be reduced. In a general RPR scheme without additional restrictions, such as the following invention, the implementer needs to over-provision its decoder hardware to perform the following: · For any current picture, worst case: resampling in all pictures · Resampling of any picture at a defined position in the DPB (decoded picture buffer), intermediate GOP (group of pictures) pair, with fewer reference pictures · Simultaneous resampling of multiple images of various resolutions to the target resolution · Cascaded resampling chain of ref pic with loss of (reference) picture quality

[0065] The following invented restrictions make it possible to reduce the implementation cost of a codec featuring such a restricted RPR scheme compared to an unrestricted RPR codec.

[0066] In one embodiment, resolution change is only permitted at the RAP (random access point). For example, the maximum number of resampled pictures is the amount of RASL (random access decodable skipped picture) pictures at this RAP. The RAP is usually at a distance of one or more GOPs, for example, dozens of pictures apart. This reduces the worst-case rate at which such resampling operations must be supported.

[0067] In another embodiment, the resolution change occurs at a key picture within a hierarchical GOP, e.g., where the picture is: · Belonging to the bottom temporal layer, and · Occurring once in all GOPs, and · All pictures following it in the coding order have a lower POC (e.g., earlier presentation time stamp), When the reference picture is resampled, none of the pictures immediately following it in the GOP of the higher temporal layer require cascade up / down sampling.

[0068] According to another embodiment, the resolution change is permitted only at a picture that follows a key picture in the presentation order, i.e., at the first picture of the next GOP in the presentation order.

[0069] In another embodiment, the time distance between successive resolution changes is restricted by the minimum POC (picture order count) distance in the level definition.

[0070] In another embodiment, the temporal distance between successive resolution changes is restricted by the minimum number of intermediate coded pictures in the level definition.

[0071] In another embodiment, the resolution change can occur only at a picture marked as non-discardable or as a reference picture where the non_reference_picture_flag is equal to 0.

[0072] In another embodiment, the rate of resolution change is restricted by the level definition.

[0073] In another embodiment, the resampling of the reference picture for the current picture is restricted to use a single resampling ratio, e.g., all reference pictures of the current picture having a different resolution than the current picture are required to have the same resolution.

[0074] In another embodiment, when one reference picture of the current picture requires resampling, all reference pictures of the current picture need to use resampling, for example, having the same original resolution as one reference picture.

[0075] In another embodiment, only one reference picture of the current picture is permitted to require resampling.

[0076] According to another embodiment, the maximum number of pictures that require resampling at the resolution change point is optionally indicated in the coded video sequence / bitstream as a guarantee of the decoder, and if no indication exists, inference or indication is made according to the level definition.

[0077] In another embodiment, after resampling, so that only the resampled reference pictures of the original (non-resampled) reference pictures are available therefrom, they are removed from the reference picture list and / or the decoded picture buffer, for example, marked as not being used for reference.

[0078] In another embodiment, the resampling ratio used within the coded video sequence is limited to the set of resampling ratios included in a parameter set (decoding parameter set, DPS; sequence parameter set, SPS) having a sequence or bitstream scope.

[0079] Hereinafter, the third aspect of the present invention will be described in detail.

[0080] In particular, the second aspect provides flexible region-based reference for zooming for reference picture resampling.

[0081] As described above, in hierarchical coders such as SHVC and SVC, two extended scalability modes, namely, RoI scalability as shown in FIG. 4 below (the area of the lower picture is enlarged in the upper layer) and extended scalability (the lower picture is extended in the upper layer via additional content) are addressed.

[0082] Extended scalability may refer to use cases such as what is called a so-called zoom-out, for example, a use case that changes over time in the sense that a video fully covers more content, such as a larger capture angle, more parts of a scene, a larger area, etc.

[0083] FIG. 4 shows the region of interest (RoI) scalability for extended scalability.

[0084] In scenarios where zoom-in and zoom-out are allowed, the zoom and movement areas used for prediction and predicted are defined. This is known as RoI scalability (typically zoom-in) or extended scalability (typically zoom-out). In RoI scalability by scalable coding, usually, the area upscaled to the dimensions of the reference picture is defined within the reference picture. However, in scalable coding, the upper and lower pictures for which prediction is performed represent the same time.

[0085] In the case of SHVC and SVC, this is done for layered coding, and in these cases, the juxtaposed base layer does not represent any motion. For example, since the corresponding samples in the base layer are known, it was possible to fully upscale the known area in the base layer and operate based on that upscaled reference.

[0086] However, in the RPR application, the two pictures between which prediction is performed do not depict the same time instance, and thus, some content outside the defined region can move to the zoom-in / zoom-out region from time instance A (low resolution) to time instance B (high resolution). Prohibiting the reference to these regions for prediction is detrimental to coding efficiency.

[0087] However, in the case of RPR, for example, by zooming an object moving into the RoI within the area, the reference can point to an area outside the corresponding reference region. This is shown in Figure 5a without actually changing the coding resolution:

[0088] Figure 5a depicts a first example of a content fragment (gray) moving within the picture over time.

[0089] In the first embodiment, a reference region is defined that includes an area larger than the area of the RoI such that the gray box in the figure entering the RoI zoom region is within the reference:

[0090] Figure 5b depicts a second example of a content fragment (gray) moving within the picture over time.

[0091] This leads to reconstructing a region slightly larger than the RoI for the picture corresponding to the RoI, and the additional region is removed by indicating a cropping window. This problem arises from the fact that the scaling factors used to upsample the reference are calculated in VVC (Versative Video Coding) from the cropped-out picture. First, assuming there is no RoI, the horizontal scale factor HorScale and the vertical scale factor VerScale are calculated as follows: HorScale = CroppedOutWidthPic / CroppedOutWidthRefPic VerScale = CroppedOutHeightPic / CroppedOutHeightRefPic

[0092] The reason for showing the ratio based on the cropped-out picture is that since the codec needs to be a multiple of the minimum size (in the case of VVC 8 samples), depending on the picture size of interest, some additional samples need to be decoded. Therefore, if either Pic or RefPic is not a multiple of 8, some samples are added to the input picture to make them a multiple of 8, which results in different ratios and leads to incorrect scaling factors. This problem may be further exacerbated when the bitstream is desired to be encoded as "mergeable", for example, when the picture size needs to be a multiple of the CTU size up to 128 and the bitstream can be merged with other bitstreams. Therefore, the correct scaling factor needs to consider the cropping window.

[0093] In the described scenario (combining RPR with RoI) where a cropping window is used to include some additional references, the use of the cropping window is insufficient.

[0094] As described above, the RoI within the reference picture can be defined slightly larger and is used for reference but is discarded together with the cropping window within the current reconstructed picture. However, when the horizontal scaling factor HorScale and the vertical scaling factor VerScale are calculated as follows: HorScale = CroppedOutWidthPic / WidthEnlargedRefRoI VerScale = CroppedOutHeightPic / HeightEnlargeRefRoI Some of the samples within the enlarged RoI actually correspond to the samples within the cropped-out region, so the result is incorrect.

[0095] The following describes the cropping window based concept according to the first group of embodiments.

[0096] Therefore, in the embodiments of the first group described above, the calculation may be, for example, as follows: HorScale = CodedPicWidth / RefRoIWidth VerScale = CodedPicHeight / RefRoIHeight Such a calculation includes samples that are cropped for the calculation of the scale factor.

[0097] Regarding signaling, in one embodiment, the signaling of the enlarged RoI indicates that the cropping window information should be ignored in the calculation of the scaling factor.

[0098] In another embodiment, whether the cropping window needs to be considered in the calculation of the scaling factor is indicated in the bitstream (e.g., parameter set or slice header). [Table 6]

[0099] The cropping window may also be called, for example, a compliant cropping window. The offsets for the cropping window / compliant cropping window may also be called, for example, pps_conf_win_left_offset, pps_conf_win_top_offset, pps_conf_win_right_offset, and pps_conf_win_botton_offset.

[0100] Instead of using the use_cropping_for_scale_factor_derivation_flag to determine whether to ignore the information in the encoded video signal on the cropping window for upscaling the region in the reference picture (or to determine whether to use the information in the encoded video signal on the cropping window for upscaling the region in the reference picture), for example, the pps_scaling_window_explicit_signalling_flag may be used.

[0101] For example, if the pps_scaling_window_explicit_signalling_flag is set to 0 (or, for example, set to FALSE, or, for example, set to -1), the information in the encoded video signal on the cropping window may be used, for example, for upscaling the region in the reference picture. And, for example, if the pps_scaling_window_explicit_signalling_flag is set to 1 (or, for example, set to TRUE), the information in the encoded video signal on the cropping window may be ignored, for example, for upscaling the region in the reference picture.

[0102] One drawback of the above approach is that the region to be decoded for the current picture becomes larger in order to enable referencing samples outside the RoI, for example, referencing samples on the enlarged RoI. More specifically, samples are decoded in regions outside the RoI and later discarded by the cropping window. This leads to additional sample overhead and reduced coding efficiency that can potentially counteract the coding efficiency gain that enables referencing outside the corresponding RoI in the reference picture.

[0103] A more efficient approach is to decode only the RoI (omitting the additional samples required to create the 8x picture or CTU as described above), enabling reference to samples within the enlarged RoI.

[0104] The following describes the bounding box based concept according to the second group of embodiments.

[0105] In the above second group of embodiments, samples that are outside the red rectangle but inside the green box (the RoI offset with an additional RoI offset) are used to determine the resampled ref pic instead of using only the red RoI.

[0106] The size of the bounding box for the MV around the red cutout is defined / signaled with the advantage of limiting memory access / line buffer requirements and enabling implementation by a pic-wise upsampling approach.

[0107] Such signaling can be included in the PPS (additional_roi_X):

Table 7

[0108] Therefore, the derivation of the scaling factors is as follows: HorScale = CroppedOutWidthPic / RefRoIWidth VerScale = CroppedOutHeightPic / RefRoIHeight

[0109] In one embodiment, the reference sample finds the samples arranged using RoI_X_offset, and is identified by applying the MV that is clipped when the reference sample is outside the enlarged RoI indicated by additional_RoI_x. Alternatively, the samples outside this enlarged RoI are padded with the last sample within the enlarged RoI.

[0110] In another embodiment, this enlarged RoI is used only as a limitation or constraint that can be used for implementation optimization. For example, if the reference picture is not on-the-fly (block-based) but is first fully upsampled as needed, only the enlarged RoI is resampled instead of the entire picture, saving a lot of processing.

[0111] A further problem is when two or more reference pictures are used simultaneously. In this case, it is necessary to identify the picture to which the RoI region information is applied. In such a case, instead of adding information to the PPS, the slice header indicates that some of the entries in the reference list refer to a part rather than the whole picture. For example,

Table 8

[0112] In a further embodiment, additional constraints are in place: · Only the reference pictures with lower POC can have RoI information. Typically, the RoI switching is applied to the open GOP switching scenario using the features described, and thus the POCs with higher POC already represent the RoI scene. · Only one reference picture can have RoI information.

[0113] In another embodiment, RoIInfo() is carried in the picture parameter set, and the slice header only carries a flag (RoI_flag) for each reference picture, indicating whether the RoI information should be applied for resampling (derivation of the scaling factor). The following figures show the principle in four coded pictures two before and two after the switching point. At the switching point, the total resolution remains constant, but upsampling of the RoI is performed. Two PPSs are defined, and the PPS of the latter two pictures indicates the RoI within the reference picture. Further, the slice headers of the latter two pictures carry RoI_flag[i] for each of their reference pictures, and the value is shown as "RoI_flag" or "RF = x" in the figure.

[0114] Furthermore, the slice header can carry, for each reference picture, not only the RoI_flag as described above, but also, when the flag is true, an additional index to the array of RoIInfo() carried in the parameter set to identify the RoI information to be applied to a specific reference picture.

[0115] Figure 6a is a diagram showing the current picture with a hybrid reference picture.

[0116] Hereinafter, the case of zooming out according to the third group of embodiments will be described.

[0117] Instead of RoI scalability, in the third group of embodiments described above, extended scalability, for example, scalability going from the RoI picture to a larger area, can be considered. In such a case, in particular, when the area of the current decoded picture is identified as an area for scalability such as zooming in and out, the cropping window of the reference picture should also be ignored.

[0118] Figure 6b shows an example of ignoring the cropping window of the reference picture in the case of the area identified within the current picture.

[0119] HorScale = IdentifiedRegionInPicWidth / CodedRefPicWidth VerScale = IdentifiedRegionInPicHeight / CodedRefPicHeight

[0120] Although some aspects have been described in the context of an apparatus, it is clear that these aspects also represent corresponding method descriptions where a block or device corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of a method step also represent corresponding block or item or feature descriptions of a corresponding apparatus. Some or all of the method steps may be performed by (or using) hardware devices such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps can be performed by such a device.

[0121] Depending on specific implementation requirements, embodiments of the present invention can be implemented in hardware, or in software, or at least partially in hardware, or at least partially in software. This implementation can be carried out using a digital storage medium, such as a floppy (registered trademark) disk, a DVD (registered trademark), a Blu-ray (registered trademark), a CD (registered trademark), a ROM, a PROM, an EPROM, an EEPROM, or a flash memory, which stores electronically readable control signals that cooperate (or can cooperate) with a programmable computer system so that the respective methods are executed. Thus, the digital storage medium may be computer-readable.

[0122] Some embodiments according to the present invention comprise a data carrier having electronically readable control signals that can cooperate with a programmable computer system so that one of the methods described herein is executed.

[0123] Generally, embodiments of the present invention can be implemented as a computer program product having program code, the program code being operable to perform one of the methods when the computer program product is executed on a computer. The program code may be stored, for example, on a machine-readable carrier.

[0124] Other embodiments include a computer program stored on a machine-readable carrier for performing one of the methods described herein.

[0125] In other words, one embodiment of the method of the present invention is thus a computer program having program code for performing one of the methods described herein when the computer program is executed on a computer.

[0126] Accordingly, a further embodiment of the method of the present invention is a data carrier (or digital storage medium, or computer-readable medium) having recorded thereon a computer program for performing one of the methods described herein. The data carrier, digital storage medium, or recorded medium is typically tangible and / or non-transitory.

[0127] Accordingly, a further embodiment of the method of the present invention is a sequence of data streams or signals representing a computer program for performing one of the methods described herein. The data stream or signal sequence may be configured to be transferred, for example, via a data communication connection, for example, via the Internet.

[0128] A further embodiment comprises processing means, for example a computer or a programmable logic device, configured or adapted to perform one of the methods described herein.

[0129] A further embodiment includes a computer having a computer program installed thereon for performing one of the methods described herein.

[0130] A further embodiment according to the present invention comprises an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may include, for example, a file server for transferring the computer program to the receiver.

[0131] In some embodiments, a programmable logic device (e.g., a field programmable gate array) can be used to perform some or all of the functions of the methods described herein.

[0132] In some embodiments, a field programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. Generally, the method is preferably performed by any hardware device.

[0133] The apparatus described herein can be implemented using a hardware device, or using a computer, or using a combination of a hardware device and a computer.

[0134] The methods described herein can be performed using a hardware device, or using a computer, or using a combination of a hardware device and a computer.

[0135] The above embodiments are merely for explaining the principles of the present invention. It is understood that modifications and variations of the configurations and details described herein will be apparent to other persons skilled in the art. Accordingly, it is intended to be limited only by the impending claims and not by the specific details presented by the description and explanations of the embodiments herein.

Claims

1. A video decoder for decoding an encoded video signal comprising encoded picture data to decode a plurality of pictures of a video sequence of a video, the video decoder comprising: at least one memory; at least one processor communicatively connected to the at least one memory; comprising; the at least one processor reads instructions from the at least one memory, receiving the encoded video signal comprising encoded picture data and a sequence parameter set (SPS) comprising an indication that a sample aspect ratio (SAR) is changeable within an encoded video sequence (CVS); receiving a first supplemental enhancement information (SEI) message in the encoded video signal, the first SEI message comprising a first SAR value for a first group of one or more of the plurality of pictures within the CVS; receiving a second SEI message in the encoded video signal, the second SEI message comprising a second SAR value for a second group of one or more of the plurality of pictures within the CVS, the second SAR value being different from the first SAR value, and the second group being different from the first group; decoding the encoded picture data corresponding to the first group and the second group of one or more of the plurality of pictures to obtain a plurality of decoded pictures; outputting the plurality of decoded pictures, wherein one or more of the plurality of decoded pictures have a SAR value different from one or more other pictures of the plurality of decoded pictures; configured to perform operations comprising; a video decoder.

2. The SPS comprises information as to whether the SPS comprises a default SAR value, The video decoder according to claim 1.

3. The first group of the one or more pictures among the plurality of pictures in the CVS is the encoded picture data received after the first SEI message having the first SAR value is received and before the second SEI message having the second SAR value is received, and includes the one or more pictures among the plurality of pictures in the CVS encoded therein. The second group of the one or more pictures among the plurality of pictures in the CVS includes the one or more pictures among the plurality of pictures in the CVS encoded in the encoded picture data received after the second SEI message having the second SAR value is received. The video decoder according to claim 1.

4. A method of decoding an encoded video signal comprising encoded picture data for decoding a plurality of pictures of a video sequence of a video, the method comprising: Receiving the encoded video signal comprising the encoded picture data and a sequence parameter set (SPS) comprising an indication that a sample aspect ratio (SAR) is changeable within a coded video sequence (CVS); Receiving a first supplemental enhancement information (SEI) message in the encoded video signal, the first SEI message comprising a first SAR value for a first group of one or more pictures among the plurality of pictures in the CVS; Receiving a second SEI message in the encoded video signal, the second SEI message comprising a second SAR value for a second group of one or more pictures among the plurality of pictures in the CVS, the second SAR value being different from the first SAR value, and the second group being different from the first group; Decoding the encoded picture data corresponding to the first group and the second group of the one or more pictures among the plurality of pictures to obtain a plurality of decoded pictures. A step of outputting the plurality of decoded pictures, wherein one or more of the plurality of decoded pictures have a different SAR value from one or more other pictures of the plurality of decoded pictures, said step; A method comprising. **Claim 5** The SPS comprises information as to whether the SPS has a default SAR value. The method according to claim 4. **Claim 6** The first group of one or more of the plurality of pictures within the CVS comprises one or more of the plurality of pictures within the CVS that are encoded in the encoded picture data received after the first SEI message having the first SAR value is received and before a second SEI message having the second SAR value is received. The second group of one or more of the plurality of pictures within the CVS comprises one or more of the plurality of pictures within the CVS that are encoded in the encoded picture data received after the second SEI message having the second SAR value is received. The method according to claim 4. **Claim 7** At runtime, a non-transitory computer-readable medium comprising instructions for performing the method according to claim 4. Non-transitory computer-readable medium. **Claim 8** A video encoder for encoding a plurality of pictures of a video sequence of a video, the video encoder comprising: At least one memory; At least one processor communicatively coupled to the at least one memory; Comprising; The at least one processor reads instructions from the at least one memory and Encoding the plurality of pictures of the encoded video sequence (CVS) into encoded picture data; Generating a sequence parameter set (SPS) comprising an indication that a sample aspect ratio (SAR) is changeable within the CVS; Generating a first supplemental enhancement information (SEI) message having a first SAR value for a first group of one or more of the plurality of pictures within the CVS; Generating a second SEI message comprising a second SAR value for one or more pictures of a second group of the plurality of pictures within the CVS, wherein the second SAR value is different from the first SAR value and the second group is different from the first group; Generating the encoded video signal such that the encoded video signal comprises the encoded picture data, the SPS, the first SEI message, and the second SEI message; Transmitting the encoded video signal; configured to perform an operation comprising; a video encoder. **Claim 9** The operation further comprises generating the SPS such that the SPS comprises information as to whether the SPS comprises a default SAR value. The video encoder according to claim 8. **Claim 10** The encoded video signal is arranged such that: the SPS is before the first SEI message, the first SEI message is before encoding of the one or more pictures of the first group of the plurality of pictures of the CVS, encoding of the one or more pictures of the first group of the plurality of pictures of the CVS is before the second SEI message, the second SEI message is before encoding of the one or more pictures of the second group of the plurality of pictures of the CVS. The video encoder according to claim 8. **Claim 11** A method of encoding a plurality of pictures of a video sequence of a video, the method comprising: encoding the plurality of pictures of a coded video sequence (CVS) into coded picture data; generating a sequence parameter set (SPS) comprising an indication that a sample aspect ratio (SAR) is changeable within the CVS; generating a first supplemental enhancement information (SEI) message comprising a first SAR value for one or more pictures of a first group of the plurality of pictures within the CVS; ​ Generating a second SEI message comprising a second SAR value for one or more pictures of a second group of the plurality of pictures within the CVS, wherein the second SAR value is different from the first SAR value and the second group is different from the first group; Generating the encoded video signal such that the encoded video signal comprises the encoded picture data, the SPS, the first SEI message, and the second SEI message; Transmitting the encoded video signal; A method comprising the above.

12. The method further comprises generating the SPS such that the SPS comprises information as to whether the SPS comprises a default SAR value. The method according to claim 11.

13. The encoded video signal is arranged such that: The SPS is before the first SEI message; The first SEI message is before encoding of the first group of the one or more pictures of the plurality of pictures of the CVS; Encoding of the first group of the one or more pictures of the plurality of pictures of the CVS is before the second SEI message; The second SEI message is before encoding of the second group of the one or more pictures of the plurality of pictures of the CVS; The method according to claim 11.

14. A non-transitory computer-readable medium comprising instructions for performing the method according to claim 11 when executed. A non-transitory computer-readable medium. ​

Citation Information

Patent Citations

  • Low-latency video encoding system and its operation method

    JP2015526970A

  • Picture / video coding supporting varying resolutions and / or efficiently handling region-based packing - Patents.com

    JP2021514168A

  • Signalling Information for Consecutive Coded Video Sequences that Have the Same Aspect Ratio but Different Picture Resolutions

    US20140003539A1