Encoder, decoder, encoding method and decoding method with profile- and level-dependent coding options
By adapting motion vector wraparound techniques and optimizing affine prediction in video codecs, the challenges of parallel processing in omnidirectional content and affine prediction modes are addressed, resulting in improved efficiency and coding performance.
Patent Information
- Application Number
- JP2025040367
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-03-11
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-17
AI Technical Summary
Existing video codecs, such as H.265/HEVC, face challenges in efficiently supporting parallel processing in video encoding and decoding, particularly in handling omnidirectional content and affine prediction modes.
The proposed solution involves adapting motion vector wraparound techniques to handle rotations in omnidirectional content and optimizing affine prediction by restricting reference sub-block diffusion and using shorter tap interpolation filters based on content resolution.
This approach enhances the efficiency of parallel processing in video encoding and decoding, reduces memory bandwidth overhead, and improves coding efficiency by adapting to different content resolutions and formats.
Smart Images

Figure 2025090775000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to video encoding and video decoding, and more particularly, to an encoder and a decoder using profile and level dependent coding options, an encoding method, and a decoding method.
Background Art
[0002] H.265 / HEVC (HEVC = High Efficiency Video Coding) is a video codec that already provides tools for improving or even enabling parallel processing in an encoder and / or a decoder. For example, HEVC supports the sub-division of a picture into an array of tiles encoded independently of each other. Another concept supported by HEVC is related to WPP, according to which, in the processing of consecutive CTU lines (CTU = Coding Tree Unit), if some minimum CTU offsets are preserved, the CTU rows or CTU lines of a picture can be processed in parallel from left to right, for example, in stripes. However, it is preferable to have at hand a video codec that more efficiently supports the parallel processing function of a video encoder and / or a video decoder. Hereinafter, the introduction of state-of-the-art VCL partitioning will be described (VCL = Video Coding Layer).
[0003] Normally, in video coding, the coding process of picture samples requires smaller partitions, and the samples are divided into several rectangular regions for joint processing such as prediction or transform coding. Therefore, a picture is partitioned into blocks of a specific size that is constant during the encoding of a video sequence. In the H.264 / AVC standard (AVC = Advanced Video Coding), fixed-size blocks of 16×16 samples, so-called macroblocks, are used.
[0004] In the state-of-the-art HEVC standard (see [1]), there are coded tree blocks (CTBs) or coding tree units (CTUs) of a maximum size of 64×64 samples. For further description of HEVC, the more general term CTU is used for such types of blocks. CTUs start from the top-left CTU, process the CTUs in the picture in a line direction, and are processed in raster scan order up to the bottom-right CTU.
[0005] The coded CTU data is organized into a kind of container called a slice. Originally, in conventional video coding standards, a slice means a segment that contains one or more consecutive CTUs in a picture. Slices are used for the segmentation of coded data. From another perspective, a complete picture can also be defined as one large segment, and thus, historically, the term slice still applies. In addition to the coded picture samples, a slice also contains additional information related to the coding process of the slice itself, which is placed in the so-called slice header.
[0006] According to the state-of-the-art, the VCL (Video Coding Layer) also includes techniques for fragmentation and spatial partitioning. Such partitioning can be applied to video coding for various reasons, including load distribution in parallelization, CTU size matching in network transmission, error mitigation, etc. SUMMARY OF THE INVENTION PROBLEMS TO BE SOLVED BY THE INVENTION
[0007] An object of the present invention is to provide an improved concept for video encoding and video decoding. MEANS FOR SOLVING THE PROBLEMS
[0008] The object of the present invention is solved by the subject matter of the independent claims. Preferred embodiments are provided in the dependent claims. Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Modes for Carrying Out the Invention
[0010] The following description of the drawings begins with the presentation of an explanation of an encoder and a decoder of a block-based predictive codec for coding a video picture in order to form an example of a coding framework into which embodiments of the present invention can be incorporated. Each encoder and decoder will be described with respect to FIGS. 9 through 11. Below, an explanation of embodiments of the concepts of the present invention will be presented along with an explanation of how such concepts can be incorporated into the encoders and decoders of FIGS. 9 and 10, respectively, although the embodiments described in FIGS. 1 through 3 and below can also be used to form encoders and decoders that do not operate according to the coding framework underlying the encoders and decoders of FIGS. 9 and 10.
[0011] FIG. 9 shows a video encoder, which is a device for predictively coding picture 12 into data stream 14, exemplarily using transform-based residual coding. The device or encoder is indicated using reference numeral 10. FIG. 10 shows a corresponding video decoder 20, i.e., a device 20 configured to predictively decode picture 12' from data stream 14, also using transform-based residual decoding, where the apostrophe is used to indicate that picture 12' reconstructed by decoder 20 deviates from picture 12 initially encoded by device 10 with respect to coding losses introduced by quantization of the prediction residual signal. FIGS. 9 and 10 exemplarily use transform-based prediction residual coding, but embodiments of the present application are not limited to this type of prediction residual coding. This also applies to other details described with respect to FIGS. 9 and 10, as outlined below.
[0012] Encoder 10 is configured to spatially spectrally transform the prediction residual signal and encode the thus obtained prediction residual signal into data stream 14. Similarly, decoder 20 is configured to decode the prediction residual signal from data stream 14 and spectrally spatially transform the thus obtained prediction residual signal.
[0013] Internally, the encoder 10 can include a prediction residual signal former 22 that generates a prediction residual 24 to measure the deviation of the original signal, i.e., the prediction signal 26 from the picture 12. The prediction residual signal former 22 may be, for example, a subtractor that subtracts the prediction signal from the original signal, i.e., from the picture 12. Next, the encoder 10 further includes a converter 28 that spatially spectrally transforms the prediction residual signal 24 to obtain a spectral region prediction residual signal 24' that is also quantized by a quantizer 32 included in the encoder 10. The prediction residual signal 24'' thus quantized is coded into the bitstream 14. For this purpose, the encoder 10 can include an entropy coder 34 that entropy-codes the prediction residual signal that is converted and quantized into the data stream 14, if necessary. The prediction signal 26 is generated by a prediction stage 36 of the encoder 10 based on the prediction residual signal 24'' that is decodable from the data stream 14. For this purpose, as shown in FIG. 9, the prediction stage 36 includes an inverse quantizer 38 that inverse-quantizes the prediction residual signal 24'' to obtain a spectral region prediction residual signal 24''' corresponding to the signal 24' other than the quantization loss, and an inverse converter 40 that inverse-transforms, i.e., spectrally spatially transforms, the latter prediction residual signal 24''' to obtain a prediction residual signal 24'''' corresponding to the original prediction residual signal 24 other than the quantization loss. Next, a combiner 42 of the prediction stage 36 recombines the prediction signal 26 and the prediction residual signal 24'''' by addition or the like to obtain a reconstructed signal 46, i.e., a reconstruction of the original signal 12. The reconstructed signal 46 can correspond to the signal 12'. Next, a prediction module 44 of the prediction stage 36 generates the prediction signal 26 based on the signal 46, using, for example, spatial prediction, i.e., intra-picture prediction, and / or temporal prediction, i.e., inter-picture prediction.
[0014] Similarly, as shown in FIG. 10, the decoder 20 may be internally configured from components corresponding to the prediction stage 36 and interconnected in a manner corresponding to the prediction stage. In particular, the entropy decoder 50 of the decoder 20 can entropy-decode the quantized spectral region prediction residual signal 24'' from the data stream. At this time, the inverse quantizer 52, the inverse transformer 54, the combiner 56, and the prediction module 58 are interconnected and cooperate in the manner described above with respect to the modules of the prediction stage 36 to recover the signal reconstructed based on the prediction residual signal 24'', whereby, as shown in FIG. 10, the output of the combiner 56 results in the reconstructed signal, i.e., the picture 12'.
[0015] Although not specifically described above, it is readily apparent that the encoder 10 can set several coding parameters, such as prediction mode, motion parameters, etc., according to several optimization methods, such as several rate and distortion related criteria, i.e., methods for optimizing coding cost. For example, the encoder 10, the decoder 20, and the corresponding modules 44, 58 can each support different prediction modes, such as intra coding mode and inter coding mode. The granularity at which the encoder and decoder switch between these prediction mode types can respectively correspond to the subdivision into coding segments or coding blocks of pictures 12 and 12'. In units of these coding segments, for example, a picture can be subdivided into intra-coded blocks and inter-coded blocks. Intra-coded blocks are predicted based on the spatial already-coded / decoded neighborhood of each block, as outlined in more detail below. There are several intra coding modes, including a directional intra coding mode or an angular intra coding mode, and according to those modes, each intra coding segment can be selected such that each intra coding segment is filled by extrapolating sample values in the neighborhood along a specific direction unique to each directional intra coding mode to each intra coding segment. The intra coding mode can also include one or more additional modes, such as a DC coding mode in which the prediction of each intra-coded block assigns a DC value to all samples within each intra-coded segment, and / or a plane intra coding mode in which the prediction of each block is approximated or determined to be a spatial distribution of sample values described by a two-dimensional linear function having a driving slope and an offset of a plane defined by a two-dimensional linear function based on adjacent samples over the sample positions of each intra-coded block. In comparison, inter-coded blocks can be predicted, for example, temporally.In the case of an inter-coded block, motion vectors can be signaled within the data stream, and the motion vectors indicate the spatial displacement of a portion of a previously coded picture of the video to which picture 12 belongs. The previously coded / decoded pictures are sampled to obtain the prediction signal for each inter-coded block. This means that in addition to the coding of the residual signal included in the data stream 14, such as the entropy-coded transform coefficient levels representing the quantized spectral region prediction residual signal 24'', the data stream 14 can encode therein any further parameters, such as coding mode parameters for assigning coding modes to various blocks, some prediction parameters of the blocks such as motion parameters of the inter-coded segments, and parameters for controlling and signaling the further division of each segment of pictures 12 and 12' into. The decoder 20 uses these parameters to divide the picture in the same way as the encoder did, assign the same prediction mode to the segments, and perform the same prediction to yield the same prediction signal.
[0016] FIG. 11 shows the relationship between, on the one hand, the reconstructed signal, i.e., the reconstructed picture 12', and, on the other hand, the combination of the prediction residual signal 24'''' and the prediction signal 26 signaled in the data stream 14. As already described above, the combination may be an addition. In FIG. 11, the prediction signal 26 is shown as dividing the picture area into an intra-coded block exemplarily shown using hatching and an inter-coded block exemplarily shown without using hatching. The division may be any division, such as a regular division of the picture area into rows and columns of square or non-square blocks, or a multiple-tree division of picture 12 from a tree-root block into a plurality of leaf blocks of various sizes, such as a quadtree division, and a mixture thereof is shown in FIG. 11, where the picture area is first divided into rows and columns of the tree-root block and then further divided into one or more leaf blocks according to a recursive multiple-tree division.
[0017] Here too, the data stream 14 can have an intra-coding mode coded for the intra-coding block 80 that assigns one of several supported intra-coding modes to each intra-coded block 80. In the case of the inter-coding block 82, the data stream 14 can have one or more motion parameters coded therein. Generally speaking, the inter-coding block 82 is not limited to being coded temporally. Alternatively, the inter-coding block 82 can be any block predicted from a previously coded portion beyond the current picture 12 itself, such as a previously coded picture of the video to which the picture 12 belongs, or, if the encoder and decoder are respectively scalable encoders and decoders, a picture of another view or a hierarchically lower layer.
[0018] The prediction residual signal 24'''' of FIG. 11 is also shown as a subdivision of the block 84 in the picture area. These blocks may sometimes be called transform blocks to distinguish them from the coding blocks 80 and 82. In fact, FIG. 11 shows that the encoder 10 and the decoder 20 can use two different subdivisions of the blocks of the pictures 12 and 12', namely, one subdivision into the coding blocks 80 and 82, and the other subdivision into the transform blocks 84. The two subdivisions may be the same, that is, each coding block 80 and 82 may simultaneously form a transform block 84. However, FIG. 11 shows that, for example, when the subdivision into the transform blocks 84 forms an extension of the subdivision into the coding blocks 80 and 82 such that any boundary between the two blocks 80 and 82 covers the boundary between the two blocks 84, or when each block 80 and 82 coincides with one of the transform blocks 84 or coincides with a cluster of the transform blocks 84. However, the subdivisions may also be determined or selected independently of each other so that the transform block 84 can cross or alternatively cross the block boundaries between the blocks 80 and 82. Therefore, as far as the subdivision into the transform blocks 84 is concerned, the same description as presented for the subdivision into the blocks 80 and 82 applies, that is, the block 84 can be the result of a regular subdivision of the picture area into blocks (regardless of the arrangement into rows and columns), the result of a recursive multi-tree subdivision of the picture area, or a combination thereof, or any other kind of blocking. Incidentally, it should be noted that the blocks 80, 82, and 84 are not limited to quadratic, rectangular, or any other shape.
[0019] FIG. 11 further shows that the combination of the prediction signal 26 and the prediction residual signal 24'''' directly results in the reconstructed signal 12'. However, it should be noted that according to an alternative embodiment, a plurality of prediction signals 26 can be combined with the prediction residual signal 24'''' to form the picture 12'.
[0020] In FIG. 11, the conversion block 84 shall have the following importance. The converter 28 and the inverse converter 54 perform conversion in units of these conversion blocks 84. For example, many codecs use some kind of DST or DCT for all the conversion blocks 84. Some codecs allow skipping the conversion so that for some of the conversion blocks 84, the prediction residual signal is directly coded in the spatial domain. However, according to the embodiments described later, the encoder 10 and the decoder 20 are configured such that they support some conversions. For example, the conversions supported by the encoder 10 and the decoder 20 include the following: · DCT-II (or DCT-III), where DCT represents the discrete cosine transform · DST-IV, where DST represents the discrete sine transform · DCT-IV · DST-VII · Identity transform (IT)
[0021] Of course, the converter 28 supports all the forward conversion versions of these conversions, while the decoder 20 or the inverse converter 54 supports the corresponding reverse direction or reverse version: · Inverse DCT-II (or inverse DCT-III) · Inverse DST-IV · Inverse DCT-IV · Inverse DST-VII · Identity transform (IT)
[0022] The following description provides further details regarding the fact that the conversions can be supported by the encoder 10 and the decoder 20. In any case, note that the set of supported conversions can include only one conversion, such as a conversion from one spectrum to space or from space to spectrum.
[0023] As already outlined above, FIGS. 9 through 11 are presented as examples in which the inventive concepts to be further described below can be implemented to form specific examples of the encoder and decoder according to the present application. To that extent, the encoders and decoders of FIGS. 9 and 10 can each represent possible implementations of the encoders and decoders described later in this specification. However, FIGS. 9 and 10 are merely examples. However, the encoder according to the embodiment of the present application is outlined in more detail below. For example, it is a still image encoder rather than a video encoder, does not support inter prediction, or the subdivision into blocks 80 is performed in a manner different from the method illustrated in FIG. 11. Using a concept different from the encoder of FIG. 9, block-based encoding of picture 12 can be performed. Similarly, the decoder according to the embodiment of the present application can perform block-based decoding of picture 12' from data stream 14 using the encoding concept to be further outlined below. However, for example, the same is a still image decoder rather than a video decoder, the same does not support intra prediction, or the same subdivides picture 12' into blocks in a manner different from that described with respect to FIG. 11, and / or the same derives a prediction residual from data stream 14 in the transform domain but not, for example, in the spatial domain. It can be different from the decoder 20 of FIG. 10.
[0024] Hereinafter, a general-purpose video encoder according to an embodiment is described in FIG. 1, a general-purpose video decoder according to an embodiment is described in FIG. 2, and a general-purpose system according to an embodiment is described in FIG. 3.
[0025] FIG. 1 shows a general-purpose video encoder 101 according to an embodiment. The video encoder 101 is configured to encode a plurality of pictures of video by generating an encoded video signal, and each of the plurality of pictures includes original picture data.
[0026] The video encoder 101 includes a data encoder 110 configured to generate an encoded video signal including encoded picture data, and the data encoder is configured to encode a plurality of pictures of the video into the encoded picture data.
[0027] Furthermore, the video encoder 101 includes an output interface 120 configured to output the encoded picture data of each of the plurality of pictures.
[0028] Figure 2 shows a general-purpose video decoder 151 according to an embodiment. The video decoder 151 is configured to decode an encoded video signal including encoded picture data in order to reconstruct a plurality of pictures of the video. The video decoder 151 includes an input interface 160 configured to receive the encoded video signal.
[0029] Furthermore, the video decoder includes a data decoder 170 configured to reconstruct a plurality of pictures of the video by decoding the encoded picture data.
[0030] Figure 3 shows a general-purpose system according to an embodiment. The system includes the video encoder 101 of FIG. 1 and the video decoder 151 of FIG. 2. The video encoder 101 is configured to generate an encoded video signal. The video decoder 151 is configured to decode the encoded video signal to reconstruct the pictures of the video.
[0031] The first aspect of the present invention is claimed in claims 1 to 15. The second aspect of the present invention is claimed in claims 16 to 30. The third aspect of the present invention is claimed in claims 31 to 45. The fourth aspect of the present invention is claimed in claims 46 to 58. The fifth aspect of the present invention is claimed in claims 59 to 71. The sixth aspect of the present invention is claimed in claims 72 to 90.
[0032] Hereinafter, the first aspect of the present invention will be described in detail. In particular, the first aspect relates to motion vector wraparound. When omnidirectional content is coded using a specific projection format such as a widely deployed ERP format (ERP = equirectangular projection), the samples on the left and right boundaries of the projected image are actually adjacent in the 3D space after reprojection. In state-of-the-art video coding, to utilize this property and prevent visual artifacts, when motion compensation block-based inter-prediction reaches outside the picture boundary, a typical process of extrapolating the image sample values on the left and right picture boundaries is adjusted.
[0033] Instead of creating conventional boundary padding sample values by vertical extrapolation of the last pixel columns on the left and right sides of the picture, a clip operation is used in the derivation of the reference sample positions that result in wraparound of the motion vectors and reference blocks around the sample columns defined as follows. xInt i = sps_ref_wraparound_enabled_flag? ClipH(sps_ref_wraparound_offset, picW, (xInt L + i - 3)) : Clip3(0, picW - 1, xInt L + i - 3) yInt i = Clip3(0, picH - 1, yInt L + i - 3) where ClipH(o, W, x) is defined as mod(x, o); for x < 0 mod((x - W), o)+W - o; for x > W - 1, or X; otherwise Clip3 is defined as follows in [1]:
Number
[0034] Problems occur when such content is further used in omnidirectional video services, such as those based on the OMAF viewport-dependent profile that uses tiles (OMAF = omnidirectional media format). In such services, the low-resolution deformations of the ERP content version are often mixed with other high-resolution fragments of the full content to create viewport-dependent streams. The need to place fragments of the high-resolution full scene and the low-resolution variant of the content on a rectangular video frame can result in the rotation of the low-resolution ERP content, whereby the coding cannot utilize the above-described characteristics to avoid visual artifacts. For example, in FIG. 4, two arrangements are shown featuring a low-resolution ERP overview and high-resolution tiles. It can be seen that the ERP tiles are rotated or adjacent to the tiles on the right side. Therefore, the two tile arrangements shown in FIG. 4 result in requirements for MV wraparound that are not met by state-of-the-art designs. FIG. 4 shows a viewport-dependent tile arrangement having an overview according to an embodiment.
[0035] According to this aspect, the MV wraparound can be adapted in the following manner to enable, for example, the above tile arrangement. In one embodiment, the signaling sps_ref_wraparound_idc is used to distinguish MV wraparound options to handle rotations such as those on the left side of FIG. 4 as follows: If (sps_ref_wraparound_idc == 0) { xInt i=ClipH(sps_ref_wraparound_offset,picW,(xInt L +i - 3)) yInt i =Clip3(0,picH - 1,yInt L +i - 3) } else if(sps_wraparound_idc == 1) { xInt i =Clip3(0,picW - 1,xInt L +i - 3) yInt i =ClipH(sps_ref_wraparound_offset_second,picH,(yInt L +i - 3)) } else if(sps_wraparound_idc == 2) { xInt i =ClipH(sps_ref_wraparound_offset,picW,(xInt L +i - 3)) yInt i =ClipH(sps_ref_wraparound_offset_second,picH,(yInt L +i - 3)) }
[0036] In another embodiment, the wraparound region is defined to cover the scenario shown on the right side of FIG. 4.
[0037] Hereinafter, the second aspect of the present invention will be described in detail. In particular, the second aspect relates to the affine prediction runtime cost.
[0038] State-of-the-art video codecs are typically designed to cover a wide range of applications, use cases, and scenarios. For example, codecs such as HEVC include techniques targeted at streaming, broadcasting, real-time conversations, or surveillance services. Each of these services typically operates at different resolutions depending on the following service constraints: broadcasting typically uses a higher resolution than conversations, etc. The numerous individual tools that a codec has are tuned to fit this wide range of relevant requirements, but there can be a loss of efficiency or an increase in complexity at each individual operating point.
[0039] For example, video coding has conventionally used only a translational motion model, i.e., rectangular blocks are displaced according to a two-dimensional motion vector to form the motion compensation predictor of the current block. Such a model cannot represent common rotational or zoom motions in many video sequences, and thus efforts have been made to slightly extend the conventional translational motion model, called affine motion in, e.g., JEM (JEM = Joint Exploration Model). As shown in FIG. 5, in this affine motion model, two (or three) so-called control point motion vectors are used for each block (v0, v1, ···) to model rotational and zoom motions. The control point motion vectors are selected from the MV candidates of adjacent blocks. The resulting motion compensation predictor is derived by calculating a sub-block-based motion vector field and depends on the conventional translational motion compensation of rectangular sub-blocks.
[0040] FIG. 5 shows affine motion by sub-block division. In particular, FIG. 5 illustrates the sub-block reference or predictor (dashed red) resulting from the upper-right sub-block (solid red).
[0041] FIG. 5 represents a more or less simple rotational motion according to the motion vectors of the control points. However, for example, when representing a zoom motion, it may happen that the individual sub-block references spread over a wide distance within the picture. In such a case, the resulting sub-block references can be positioned as shown in FIG. 6, illustratively shown with a block size smaller than that resulting in, for example, 2×2 sub-blocks. In particular, FIG. 6 shows a zoom motion by affine prediction.
[0042] In such a case, the decoder needs to fetch individual sub-block references from a larger overall picture area and then be covered by the complete current block. In order to implement such an affine prediction mode cost-effectively, it is necessary to adjust the fetching of individual sub-block references on the currently available hardware. Such an adjustment can be, for example, a simultaneous fetch of a continuous reference picture area containing all sub-block references. Thus, extreme zoom motion can be costly to implement on the decoder side because an efficient reference fetch strategy is precluded.
[0043] Techniques of the art have been proposed to limit the spread of prediction blocks within the reference picture in order to limit worst-case reference picture access without sacrificing substantial coding gain in the affine prediction mode.
[0044] According to this aspect, such a limitation can be derived, for example, according to the maximum memory bandwidth allowed. This can be specified with respect to the overall maximum memory bandwidth (e.g., in the level limit specification) to substantially limit the memory bandwidth overhead (defining the maximum increase in reference memory access required compared to normal translational prediction), or can be explicitly signaled by the encoder in high-level parameters. These methods may be mixed.
[0045] In one embodiment, the restriction of reference sub-block diffusion is implemented by ensuring that the bounding box around the resulting sub-block reference block does not exceed a threshold. Such a threshold can be derived from the block size, picture resolution, maximum allowable bandwidth, or can be explicitly signaled in high-level parameters.
[0046] In another embodiment, the restriction of reference sub-block diffusion depends on the coded picture resolution or is defined by level restrictions.
[0047] Hereinafter, a third aspect of the present invention will be described in detail. In particular, the third aspect relates to affine prediction chroma handling.
[0048] The coding efficiency of the above-described affine prediction model is inversely proportional to the sub-block size in order to enable a better approximation of the affine motion by a translational block-based motion model with smaller sub-blocks as a basis. However, fine-grained reference picture access incurs a significant runtime cost.
[0049] For example, when the luminance and chroma components of a video are sampled non-uniformly, such as in the dominant consumer video format 4:2:0, the prediction / reference blocks are usually sampled at the same ratio. For example, when a 16×16 luminance block is predicted by a conventional method such as translational motion compensation block-based prediction, each 8×8 chroma block is predicted using the same luminance block motion vector.
[0050] In the context of the above affine prediction model, since the luminance sub-blocks are already at a very fine granularity or small size, state-of-the-art solutions for processing chroma blocks are based on chroma processing with the same luminance block size regardless of format subsampling. Thus, a 4×4 sub-block within the luminance channel in affine prediction results in a 4×4 sub-block within the chroma channel, such that, as shown in FIG. 7, there are fewer numbers within the chroma channel than within the luminance channel. FIG. 7 shows the sub-block partitioning for affine prediction in different channels.
[0051] Therefore, a method for deriving the motion information of the corresponding chroma blocks is required. There are different features for the solution. For example, the simplest state-of-the-art solution is to average the motion information mvLX[] of all four corresponding luminance sub-blocks as follows: mvAvgLX=(mvLX[(xSbIdx>>1<<1)][(ySbIdx>>1<<1)]+ mvLX[(xSbIdx>>1<<1)+1][(ySbIdx>>1<<1)]+ mvLX[(xSbIdx>>1<<1)][(ySbIdx>>1<<1)+1]+ mvLX[(xSbIdx>>1<<1)+1][(ySbIdx>>1<<1)+1]+2)>>2
[0052] Another option is to average only the corresponding upper-left and lower-right luminance motion vectors as follows. mvAvgLX=(mvLX[(xSbIdx>>1<<1)][(ySbIdx>>1<<1)]+ mvLX[(xSbIdx>>1<<1)+1][(ySbIdx>>1<<1)+1]+1)>>1
[0053] Furthermore, symmetric rounding can be introduced as follows: mvAvgLX=(mvLX[(xSbIdx>>1<<1)][(ySbIdx>>1<<1)]+ mvLX[(xSbIdx>>1<<1)+1][(ySbIdx>>1<<1)+1]) mvAvgLX[0]=(mvAvgLX[0]>=0? (mvAvgLX[0]+1)>>1: -((-mvAvgLX[0]+1)>>1)) mvAvgLX[1]=(mvAvgLX[1]>=0? (mvAvgLX[1]+1)>>1: -((-mvAvgLX[1]+1)>>1))
[0054] However, if an additional averaging method beyond the above state-of-the-art features is desired, problems occur when a chroma format different from 4:2:0 is used. For example, when the 4:4:4 chroma format is used, each 4×4 chroma sub-block exactly corresponds to one 4×4 luma sub-block, and for the 4:2:2 chroma format, since averaging is only required in one dimension, averaging should be avoided. Therefore, according to this aspect, the correct luma block included in the averaging operation is derived by introducing the chroma subsampling ratio into the derivation of mvAvgLX. In one embodiment, the derivation is realized as follows.
[0055] mvAvgLX=(mvLX[(xSbIdx>>SubWidthC<<SubWidthC)][(ySbIdx>>SubHeightC<<SubHeightC)]+ mvLX[(xSbIdx>>SubWidthC<<SubWidthC)+SubWidthC][(ySbIdx>>SubHeightC<<SubHeightC)+SubHeightC]) mvAvgLX[0]=(mvAvgLX[0]>=0? (mvAvgLX[0]+1)>>1: -((-mvAvgLX[0]+1)>>1)) mvAvgLX[1]=(mvAvgLX[1]>=0? (mvAvgLX[1]+1)>>1: -((-mvAvgLX[1]+1)>>1))
[0056] Here, SubWidthC and SubHeightC are derived from the chroma format applied as follows:
[0057] [Table 1]
[0058] Hereinafter, a fourth aspect of the present invention will be described in detail. In particular, the fourth aspect relates to interpolation filter overhead.
[0059] Motion-compensated inter prediction can use sample values at sub-sample positions (e.g., half-pixels, quarter-pixels) interpolated via a convolution filter from sample values at integer sample positions. State-of-the-art codecs generally use a fixed-tap filter kernel, such as an 8-tap filter kernel, to generate such sub-pixel sample values.
[0060] However, when sub-block-based inter prediction, such as affine prediction or DMVR (Decoder-side Motion Vector Refinement), results in relatively small luminance and chroma sub-blocks, such as size 4×4, a memory bandwidth problem occurs. As the (sub)block size decreases relative to the fixed interpolation filter size, the overhead regarding the total number of accessed samples of the generated reference luminance sub-pixel samples of the (sub)block increases. Therefore, it has been proposed to use a relatively short tap filter kernel for sub-pixel interpolation of small sub-blocks, such as 4×4, which exhibits a minimum coding penalty that decreases as the content resolution increases.
[0061] According to this aspect, the selection of a shorter tap interpolation filter for sub-pixel interpolation gains benefits in terms of reducing peak memory bandwidth for high-resolution content without sacrificing the coding efficiency associated with lower-resolution content, and is adapted depending on the resolution of the coded content.
[0062] In one embodiment, several sub-pixel interpolation filters are available (e.g., bilinear, 4-tap, 6-tap, 8-tap), and the content resolution and block size are used to determine the applicable filter.
[0063] Implementations become expensive as the amount of separate filters that need to be implemented and held locally within the encoder or decoder increases. Therefore, it is advantageous to reduce the amount of separate filters implemented in parallel in the encoder or decoder.
[0064] For example, when the chroma component of a video signal uses a chroma format such as 4:2:0, it is state-of-the-art to use spatial sub-sampling. In such cases, the chroma component also typically uses a shorter tap sub-panel interpolation filter than the luminance component. Therefore, such shorter tap sub-pixel interpolation filters need to be prepared by the decoder / encoder implementation in addition to the conventional longer tap sub-pixel interpolation filter kernels used for the luminance component. In an embodiment, sub-block processing introduces the reuse of shorter tap sub-pixel interpolation filters when reducing the peak memory bandwidth requires a reduced tap filter, for example, depending on the coded picture size or by the selection of the encoder and the indication in the bitstream.
[0065] Hereinafter, the fifth aspect of the present invention will be described in detail. In particular, the fifth aspect relates to bitstream property description and profile set operation.
[0066] Typically, video streams involve signaling that describes their characteristics and enables identification of the types of functions required by a video decoder to be able to decode the bitstream.
[0067] State-of-the-art video coding standards such as AVC (Advanced Video Coding) and HEVC include profile and level indications in parameter sets within the bitstream, enabling identification of which tools need to be supported and, for example, the resolution limitations of a video coding sequence, and as a result, easily identifying whether a decoder is capable of decoding a video sequence.
[0068] Regarding profiles, it is possible to identify profiles that support only intra coding tools, or only intra and inter tools, but support only an 8-bit configuration, or a 10-bit configuration, hierarchical or non-hierarchical coding, etc.
[0069] Regarding levels (related by bitrate, resolution, and / or frame rate), it is possible to distinguish between 1080p@30Hz, 4Kp@30Hz, 4Kp@60Hz, 8Kp@30Hz, etc.
[0070] That information is typically used in a video streaming system for function exchange, and as a result, the characteristics are made public and different receivers can match their functions to the requirements for decoding the provided bitstream.
[0071] For example, in DASH (Dynamic Adaptive Streaming over HTTP), the Media Presentation Description (MPD) can include several videos provided for the same content with different corresponding parameters. This is done by indicating the codec, profile, level, etc. in an attribute called @codec='''xxx''' that it shows. Similarly, in RTP streaming (RTP = Real-Time Transport Protocol), there is a Session Description Protocol (SDP), and this is used to negotiate the operating points of the session. Then, the parties involved in the session agree on which codec, profile, and level should be used to establish the session. Also, in a transport stream, MPEG2-TS (MPEG = Moving Picture Experts Group; TS = Transport Stream) (widely used in broadcast transmission), there is a means to expose this value to the end device. This is done by including a video descriptor corresponding to a program table that identifies the codec, profile, and level of a program (i.e., for example, a TV channel), so that an appropriate channel that can be decoded by the end device can be selected.
[0072] Currently, the properties included in the bitstream are only applicable to the coded video sequence (CVS). The CVS starts with the first IRAP (Intra Random Access Point) and ends with an IRAP having one of the following properties: · It is an IDR (Instantaneous Decoder Refresh) · It is a BLA (Broken Link Access) (only related to HEVC) · It is a sequence end (EOS) SEI message (SEI = Supplementary Enhancement Information)
[0073] In most video services, a "session" can include several temporally consecutive CVSs. Therefore, the parameters shown in the AVC and HEVC parameter sets are not sufficient because they only describe the first CVS of all the CVSs included in the "session". This problem is solved in MPEG2-TS by the following statement:
[0074] "Note - In one or more sequences within an AVC video stream, the level may be lower than the level signaled in the AVC video descriptor, and a profile that is a subset of the profile signaled in the AVC video descriptor may occur. However, for the entire AVC video stream, it is assumed that only the tools included in the profile signaled in the AVC video descriptor are used if they exist. For example, if the main profile is signaled, the baseline profile can be used in some sequences, but only the tools within the main profile are used. If the sequence parameter sets within an AVC video stream signal different profiles and no additional constraints are signaled, the stream may need to be inspected to determine which profile the entire stream conforms to if it exists. If an AVC video descriptor is to be associated with an AVC video stream that does not conform to a single profile, the AVC video stream must be split into two or more sub-streams so that the AVC video descriptor can signal a single profile for each such sub-stream."
[0075] In other words, the profile and level shown in the descriptor are profiles that are a superset of any profile used in the CVS of the AVC video stream, and the level is the highest level of any CVS.
[0076] Similarly, the profiles and levels indicated in the MPD (Media Presentation Descriptor) or SDP can be "higher" than those within the video bitstream (meaning a profile superset or more tools required than the profile of each CVS).
[0077] As is clear from the start, the bitstream itself desirably contains such signaling, which is a set of tools necessary to decode all CVSs of a video "session" (where the term video session also includes, for example, a TV channel program). The signaling consists of several CVSs, i.e., parameter sets applied to the bitstream. There are two particularly interesting cases: · A CVS having a profile where the profiles that are supersets of both do not exist simultaneously. · A CVS whose number of layers may change over time.
[0078] One option is to signal all the tools necessary to decode all CVSs by tool - specific flags, but this leads to an undesirable solution as it is usually more convenient to operate based on profiles, which are easier to compare than the multiple tools that can result in a very large number of combinations.
[0079] Multiple profiles active within the bitstream can lead to the following four cases. The tools required to decode the bitstream correspond to the following: · The intersection of two profiles where there is no profile corresponding to the intersection. · The intersection of two profiles where one profile is the intersection of the two profiles. · The union of two profiles where there is no profile corresponding to the union. · The union of two profiles where there is a profile corresponding to the union. If there is a profile corresponding to all necessary tools, a single profile display is sufficient.
[0080] However, if there is no profile for describing all necessary tools, multiple profiles should be signaled, and an intersection or combination should be signaled to indicate whether support for a single profile is sufficient to decode the bitstream (intersection), or whether both profiles need to be supported to be able to decode the bitstream. Hereinafter, a sixth aspect of the present invention will be described in detail. In particular, the sixth aspect relates to a buffering period SEI for parsing dependencies and a picture timing SEI message.
[0081] In existing designs such as HEVC, the syntax elements of the buffering period SEI message (SEI = supplementary enhancement information), the picture timing SEI message, and the decoding unit SEI message depend on the syntax elements carried in the SPS (sequence parameter set). This dependency also includes the number of bits used to encode a particular syntax element and the presence of the syntax element. However, for the following reasons, it is not desirable to put all buffering period / picture timing / decoding unit related information in the same syntax position:
[0082] 1) The current syntax elements in the SPS are useful for session / function negotiation performed by exchanging parameter sets. Therefore, they cannot be signaled alone in the buffering period SEI message.
[0083] 2) Since activating a new SPS is only possible at an IRAP picture (IRAP = Intra Random Access Point), the syntax elements within the buffering period SEI message cannot be moved to the SPS, and the buffering period can start at any picture associated with temporal sublayer 0. It is also not desirable to require activation of a new parameter set for each buffering period.
[0084] The above-mentioned syntax analysis dependency leads to startup problems: There is no active SPS before the first VCL NAL unit of the video bitstream is syntax-analyzed. Since the SEI message is sent before the first VCL NAL unit (NAL = Network Abstraction Layer), it arrives before the SPS is activated.
[0085] Also, in out-of-band signaling of parameter sets, since the relevant parameter sets are not included in the bitstream, intermediate devices such as splicers that are interested in understanding the HRD parameters of the bitstream, for example, to splice the bitstream, may not be able to syntax-analyze each SEI message. Also, in the case of simple intermediate devices such as splicers, it may be difficult to syntax-analyze the entire SPS until reaching the HRD parameters in the VUI.
[0086] Also, the order of NAL units is not explicitly restricted to require that the SPS be sent before the buffering period SEI message that references it. In such a situation, the decoder may have to store each SEI message in the bitstream format (without being able to syntax-analyze it) until the SPS arrives and is activated. Only in that case can the SEI message be syntax-analyzed. Figure 8 shows an exemplary state-of-the-art bitstream order for such a case. In an embodiment, the syntax analysis of each SEI message is performed independently of the SPS in the following manner: · Specify the lengths of other syntax elements within each SEI message, or · Not required for session / function negotiation (sub-picture clock tick divisor), All syntax elements are moved to the SEI message during the buffering period.
[0087] The lengths of all syntax elements within the buffering period SEI are completely independent of other syntax structures. The lengths of syntax elements within the timing SEI message and the decoding unit SEI message can depend on the syntax elements within the buffering period SEI message, but cannot depend on any parameter set value.
[0088] The presence flags of specific HRD parameters (NAL, VCL, DU, etc.), the counts of CPB parameters (vui_cpb_cnt_minus1[i] and bp_cpb_cnt_minus1), and other syntax sections are replicated in the parameter set and the buffering period SEI message, and this mode includes restricting the values of these replicated syntax elements to be equal at both positions.
[0089] The picture timing SEI message and the decoding unit SEI message are made independent by moving or replicating all syntax elements indicating the presence of other syntax elements to their respective SEI messages (e.g., frame_field_info_present_flag). As part of this mode, to achieve bitstream compatibility, the values of the above syntax elements are made equal to their respective values within the parameter set.
[0090] As part of this aspect, the tick_devisors (indicating sub-picture timing granularity) within the buffering period should have the same value unless the concatenation_flag is equal to 1 (i.e., splicing has occurred).
[0091] As already outlined, in the embodiments, it is proposed to add HRD parameters, buffering period, picture timing, and decoder unit information SEI messages to the VVC specification. The general syntax and semantics are based on the corresponding HEVC structures. The syntax parsing of the SEI messages is made possible independent of the presence of the SPS.
[0092] As described above, a buffering model is required to check bitstream compliance. The VVC buffering model can be based on, for example, the Hypothetical Reference Decoder (HRD) specification used in HEVC with some simplifications.
[0093] The HRD parameters are present in the SPS VUI. The buffering period SEI, picture timing SEI, and decoder unit information SEI messages are added to the access unit they reference.
[0094] In HEVC, there is a syntax parsing dependency between the SPS and the HRD-related SEI messages. The semantics only allowed referring to the active SPS within the same access unit, but the bitstream order allows for situations where the SEI can arrive before the SPS. In such cases, the coded SEI messages had to be stored until the correct SPS arrived.
[0095] Another situation where this parsing dependency becomes an issue is in network elements and splicers, especially when parameter sets are exchanged out-of-band (e.g., in session negotiation). In such cases, the parameter set may not be included in the bitstream at all, and as a result, the network element or splicer may not be able to decode any buffering / timing information. It can still be interesting for these elements to recognize values such as decoding time and frame / field information.
[0096] In an embodiment, it is proposed to decouple the parsing of buffering period, picture timing, and decoder unit information SEI messages from the SPS VUI / HRD parameters as follows: - Move all syntax elements indicating the length or presence of other syntax elements from the VUI to the buffering period SEI message - Duplicate the necessary syntax elements at both positions. The values of these syntax elements are constrained to be the same.
[0097] Design decisions: - Since parameter sets are usually exchanged for session negotiation, not all syntax elements should be moved from the VUI to the buffering period SEI. Therefore, parameters important for session negotiation are kept in the VUI.
[0098] - Since the buffering period can start from a non-IRAP picture, it is not possible to move all buffering period parameters to the SPS and activate a new SPS.
[0099] - The dependency between the buffering period SEI and (picture timing SEI or decoder unit SEI) is not considered harmful because there is also a logical dependency between these messages. Syntax examples based on the HEVC syntax are shown below.
[0100] The following syntax diagrams are based on the HEVC syntax with the changes marked in red. The semantics are generally moved to the new positions of the syntax elements (to conform to the syntax element names). Whenever a syntax element is replicated, bitstream constraints are added and these syntax elements shall have the same value. VUI / HRD parameters:
[0101]
Table 2
[0102]
Table 3-1
Table 3-2
[0103]
Table 4-1
Table 4-2
[0104]
Table 5
[0105]
Table 6
[0106] Although several aspects have been described in the context of an apparatus, it will be apparent that these aspects also represent descriptions of corresponding methods, and that a block or apparatus corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of a method step also represent descriptions of corresponding blocks or items or functions of a corresponding apparatus. Some or all of the method steps can be performed by (or using) a hardware apparatus such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps can be performed by such an apparatus.
[0107] Depending on specific implementation requirements, embodiments of the present invention can be implemented in hardware or software, or at least partially in hardware or at least partially in software. The implementation can be carried out using a digital storage medium such as a floppy disk, a DVD, a Blu-ray, a CD, a ROM, a PROM, an EPROM, an EEPROM, a flash memory, etc., which stores electronically readable control signals and cooperates (or can cooperate) with a programmable computer system so that respective methods are executed. Thus, the digital storage medium can be made computer-readable.
[0108] Some embodiments according to the present invention comprise a data carrier having electronically readable control signals that can cooperate with a programmable computer system so that one of the methods described herein is executed.
[0109] In general, embodiments of the present invention can be implemented as a computer program product comprising program code, the program code operative to perform one of the methods when the computer program product is executed on a computer. The program code can be stored, for example, on a machine-readable carrier. Other embodiments comprise a computer program for performing one of the methods described herein, stored on a machine-readable carrier.
[0110] In other words, an embodiment of the method of the present invention is thus a computer program having program code for performing one of the methods described herein when the computer program is executed on a computer.
[0111] Thus, a further embodiment of the method of the present invention is a data carrier (or digital storage medium, or computer-readable medium) having recorded thereon a computer program for performing one of the methods described herein. The data carrier, digital storage medium, or recorded medium is typically tangible and / or non-transitory.
[0112] Thus, a further embodiment of the method of the present invention is a sequence of a data stream or signal representing a computer program for performing one of the methods described herein. The sequence of the data stream or signal may be configured to be transferred via a data communication connection such as the Internet, for example.
[0113] A further embodiment includes processing means, such as a computer or a programmable logic device, configured or adapted to perform one of the methods described herein. A further embodiment comprises a computer having installed thereon a computer program for performing one of the methods described herein.
[0114] A further embodiment according to the present invention includes an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver can be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system can include, for example, a file server for transferring the computer program to the receiver.
[0115] In some embodiments, a programmable logic device (e.g., a field programmable gate array) can be used to perform some or all of the functions of the methods described herein. In some embodiments, the field programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by any hardware device.
[0116] The apparatus described herein can be implemented using a hardware device, or using a computer, or using a combination of a hardware device and a computer.
[0117] The methods described herein can be performed using a hardware device, or using a computer, or using a combination of a hardware device and a computer.
[0118] The above-described embodiments merely illustrate the principles of the present invention. It is understood that changes and modifications to the configurations and details described herein will be apparent to other those skilled in the art. Therefore, it is intended to be limited only by the immediate claims, rather than by the specific details presented as descriptions and explanations of the embodiments herein.
[0119] <Item 1> A video encoder (101) for encoding a plurality of pictures of video by generating an encoded video signal, each of the plurality of pictures including original picture data, A data encoder (110) configured to generate the encoded video signal including encoded picture data, the data encoder (110) being configured to encode the plurality of pictures of the video into the encoded picture data, And an output interface (120) configured to output the encoded picture data of each of the plurality of pictures. A video session includes a plurality of encoded video sequences, and CVS profiles of a plurality of video profiles are assigned to each of the plurality of encoded video sequences. The CVS profile of the encoded video sequence defines one or more tools among a plurality of tools required to decode the encoded video sequence, and each of the plurality of video profiles defines at least one tool required to decode at least one of the plurality of encoded video sequences. The data encoder (110) is configured to determine a session profile of the video session according to the CVS profiles of the plurality of encoded video sequences of the video session, and the session profile defines at least one tool among the plurality of tools required to decode the plurality of encoded video sequences of the video session. The data encoder (110) is configured to generate the encoded video signal such that the encoded video signal includes an encoding of an indication indicating the session profile. The data encoder (110) is configured to determine the session profile from the plurality of video profiles such that the session profile includes all tools of the plurality of tools required to decode all encoded video sequences of the video session, a video encoder (101). <Item 2> If there is no video profile among the plurality of video profiles that includes all tools of the plurality of tools required to decode all encoded video sequences of the video session, the data encoder (110) is configured to determine the session profile such that the session profile is a combination of at least two of the plurality of video profiles. The video encoder (101) according to Item 1. <Item 3> A video decoder (151) for decoding an encoded video signal including encoded picture data to reconstruct a plurality of pictures of a video, an input interface (160) configured to receive the encoded video signal, and a data decoder (170) configured to reconstruct the plurality of pictures of the video by decoding the encoded picture data, wherein the data decoder (170) is configured to decode a session profile of a video session including a plurality of coded video sequences from the encoded video signal, each of the plurality of video profiles defining at least one tool from a plurality of tools required to decode at least one of the plurality of coded video sequences, the session profile defining at least one tool from the plurality of tools required to decode the plurality of coded video sequences of the video session, and the data decoder (170) is configured to reconstruct the plurality of pictures of the video according to the session profile, the video decoder (151). <Item 4> wherein the session profile is a combination of at least two of the plurality of video profiles, The video decoder (151) according to item 3. <Item 5> A method for encoding a plurality of pictures of a video by generating an encoded video signal, each of the plurality of pictures including original picture data, the method comprising: generating the encoded video signal including encoded picture data, the generating including encoding the plurality of pictures of the video into the encoded picture data, and outputting the encoded picture data of each of the plurality of pictures. A video session includes a plurality of coded video sequences, and a CVS profile of a plurality of video profiles is assigned to each coded video sequence of the plurality of coded video sequences. The CVS profile of the coded video sequence defines one or more tools among a plurality of tools required to decode the coded video sequence, and each of the plurality of video profiles defines at least one tool required to decode at least one of the plurality of coded video sequences. The method includes determining a session profile of the video session according to the CVS profiles of the plurality of coded video sequences of the video session. The session profile defines at least one tool from the plurality of tools required to decode the plurality of coded video sequences of the video session. Generating the encoded video signal is performed such that the encoded video signal includes encoding of an indication indicating the session profile. Determining the session profile from the plurality of video profiles is performed such that the session profile includes all tools of the plurality of tools required to decode all coded video sequences of the video session. A method. <Item 6> A method for decoding an encoded video signal including encoded picture data to reconstruct a plurality of pictures of a video, Receiving the encoded video signal; Decoding the encoded picture data to reconstruct the plurality of pictures of the video. The method includes decoding a session profile of a video session including a plurality of coded video sequences from the coded video signal, each of the plurality of video profiles defining at least one tool of a plurality of tools required to decode at least one of the plurality of coded video sequences, the session profile defining at least one tool of the plurality of tools required to decode the plurality of coded video sequences of the video session. A method, wherein reconstructing the plurality of pictures of the video is performed according to the session profile. <Item 7> A computer program for implementing the method according to item 5 or 6 when executed on a computer or a signal processor. <Item 8> A system, A video encoder (101) according to any one of items 1 to 2, A video decoder (151) according to any one of items 3 to 4, comprising: The video encoder (101) is configured to generate the coded video signal. A system, wherein the video decoder (151) is configured to decode the coded video signal and reconstruct the pictures of the video. <Item 9> A video encoder (101) for encoding a plurality of pictures of a video by generating a coded video signal, each of the plurality of pictures including original picture data. A data encoder (110) configured to generate the coded video signal including coded picture data, the data encoder (110) being configured to encode the plurality of pictures of the video into the coded picture data. An output interface (120) configured to output the coded picture data of each of the plurality of pictures. The data encoder (110) is configured to generate the encoded video signal by encoding a sequence parameter set and by encoding a supplemental enhancement information message, the sequence parameter set being related to two or more of the plurality of pictures and including a first group of one or more of the plurality of syntax elements, the supplemental enhancement information message being exactly related to one of the plurality of pictures and including a second group of one or more of the plurality of syntax elements, A video encoder (101) configured such that the data encoder (110) generates the sequence parameter set and the supplemental enhancement information message such that the sequence parameter set does not include at least one syntax element of the second group of syntax elements of the supplemental enhancement information message and such that the sequence parameter set does not include length information of the at least one syntax element of the second group of syntax elements of the supplemental enhancement information message. <Item 10> The data encoder (110) is configured to generate the sequence parameter set and the supplemental enhancement information message such that the sequence parameter set does not include at least one further syntax element of the second group of syntax elements of the supplemental enhancement information message and such that the sequence parameter set does not include presence information of the at least one further syntax element of the second group of syntax elements of the supplemental enhancement information message. The video encoder (101) according to item 9. <Item 11> The data encoder (110) is configured to generate the sequence parameter set such that the sequence parameter set includes at least one first syntax element of a plurality of syntax elements including session negotiation information and / or function negotiation information. The data encoder (110) is configured to generate the supplementary enhancement information message such that the supplementary enhancement information message does not include at least one first syntax element including the session negotiation information and / or the function negotiation information. The data encoder (110) is configured to generate the supplementary enhancement information message such that the supplementary enhancement information message includes at least one second syntax element among a plurality of syntax elements that do not include the session negotiation information and / or the function negotiation information. The data encoder (110) is configured to generate the sequence parameter set such that the sequence parameter set does not include at least one second syntax element that does not include the session negotiation information and / or the function negotiation information. The video encoder (101) according to item 9 or 10. <Item 12> The data encoder (110) is configured to generate the sequence parameter set such that the sequence parameter set includes all syntax elements of a plurality of syntax elements that include the session negotiation information and / or the function negotiation information. The data encoder (110) is configured to generate the supplementary enhancement information message such that the supplementary enhancement information message does not include any of a plurality of syntax elements that include the session negotiation information and / or the function negotiation information. The data encoder (110) is configured to generate the supplementary enhancement information message such that the supplementary enhancement information message includes all syntax elements of a plurality of syntax elements that do not include the session negotiation information and / or the function negotiation information. The data encoder (110) is configured to generate the sequence parameter set such that the sequence parameter set does not include any of a plurality of syntax elements that do not include the session negotiation information and / or the function negotiation information. The video encoder (101) according to any one of items 9 to 11. <Item 13> The supplementary enhancement information message is a buffering period supplementary enhancement information message, or a timing supplementary enhancement information message, or a decoder unit supplementary enhancement information message. The video encoder (101) according to any one of items 9 to 12. <Item 14> For at least one of the plurality of syntax elements, the data encoder (110) is configured to generate the sequence parameter set and the supplementary enhancement information message such that the sequence parameter set includes the at least one syntax element and the supplementary enhancement information message also includes the syntax element. The data encoder (110) is configured to generate the sequence parameter set and the supplementary enhancement information message such that, for each of the plurality of syntax elements that exist in the sequence parameter set and also exist in the supplementary enhancement information message, the value of the syntax element is equal in the sequence parameter set and the supplementary enhancement information message. The video encoder (101) according to any one of items 9 to 13. <Item 15> A video decoder (151) for decoding an encoded video signal including encoded picture data to reconstruct a plurality of pictures of a video, An input interface (160) configured to receive the encoded video signal, A data decoder (170) configured to reconstruct the plurality of pictures of the video by decoding the encoded picture data, The data decoder (170) is configured to decode a sequence parameter set and a supplementary enhancement information message from the encoded video signal, the sequence parameter set being related to two or more of the plurality of pictures and including a first group of one or more of the plurality of syntax elements, the supplementary enhancement information message being related to exactly one of the plurality of pictures and including a second group of one or more of the plurality of syntax elements, the sequence parameter set not including at least one syntax element of the second group of syntax elements of the supplementary enhancement information message, and the sequence parameter set not including length information of the at least one syntax element of the second group of syntax elements of the supplementary enhancement information message. A video decoder (151) configured such that the data decoder (170) is configured to reconstruct the plurality of pictures of the video according to the sequence parameter set and according to the supplementary enhancement information message. <Item 16> For at least one syntax element of the plurality of syntax elements, the sequence parameter set includes the at least one syntax element, and the supplementary enhancement information message also includes the syntax element. For each of the plurality of syntax elements that exists in the sequence parameter set and also exists in the supplementary enhancement information message, the value of the syntax element is equal in the sequence parameter set and the supplementary enhancement information message. The video decoder (151) according to item 15. Reference materials [1] ISO / IEC,ITU-T.High efficiency video coding.ITU-T Recommendation H.265 | ISO / IEC 23008 10(HEVC),edition 1,2013;edition 2,2014.
Claims
1. 1. An electronic device for decoding pictures from a data stream, comprising: Decoding chroma format information from the data stream; determining motion vectors corresponding to chroma sub-blocks of the picture; determining a motion vector corresponding to the chroma sub-block based on (i) selecting an upper left luma sub-block and a lower right luma sub-block in response to chroma format information specifying a format in which the chroma sub-block is sampled, and (ii) averaging motion vector information corresponding to the selected luma sub-blocks; An electronic device comprising: at least one processor configured to reconstruct a picture based at least in part on the determined motion vectors corresponding to the chroma sub-blocks.
2. the format is a first format, In response to the chroma format information specifying a second format in which the chroma sub-block is sampled, the at least one processor is further configured to determine a motion vector corresponding to the chroma sub-block based on averaging in one dimension motion vector information corresponding to two luminance sub-blocks; In response to the chroma format information specifying a third format in which the chroma sub-blocks are sampled, the at least one processor is further configured to determine the motion vector corresponding to the chroma sub-block based on motion vector information corresponding to a luma sub-block. The electronic device of claim 1 .
3. the first format is a 4:2:0 format; the second format is a 4:2:2 format; the third format is a 4:4:4 format; The electronic device of claim 2 .
4. In response to the chroma format information indicating that the chroma sub-blocks are sampled in the second format or the third format, the at least one processor: deriving a subWidthC chroma format parameter and a subHeightC chroma format parameter based on the chroma format information; selecting two luma sub-blocks based on the subWidthC chroma format parameter and the subHeightC chroma format parameter; The electronic device of claim 2 , configured to average the motion vector information corresponding to the two selected luminance sub-blocks.
5. The at least one processor deriving a subWidthC chroma format parameter and a subHeightC chroma format parameter based on the chroma format information; and determining that the chroma sub-block is sampled in the second format if the subWidthC and subHeightC chroma format parameters have different values. The electronic device of claim 2 .
6. To average the motion vector information corresponding to the selected luminance sub-blocks, the at least one processor: configured to perform symmetric rounding of the motion vector information corresponding to the selected luminance sub-block. The electronic device of claim 1 .
7. 1. A method for decoding a picture from a data stream, comprising: decoding chroma format information from the data stream; determining motion vectors corresponding to chroma sub-blocks of the picture; determining a motion vector corresponding to the chroma sub-block based on: (i) selecting an upper left luminance sub-block and a lower right luminance sub-block in response to chroma format information specifying a format in which the chroma sub-blocks are sampled; and (ii) averaging motion vector information corresponding to the selected luminance sub-blocks; reconstructing a picture based at least in part on the determined motion vectors corresponding to the chroma sub-blocks; and A method for providing the above.
8. the format is a first format, determining a motion vector corresponding to the chroma sub-block based on averaging, in one dimension, motion vector information corresponding to two luminance sub-blocks in response to the chroma format information specifying a second format in which the chroma sub-block is sampled; determining the motion vector corresponding to one luminance sub-block based on motion vector information corresponding to one luminance sub-block in response to the chroma format information specifying a third format in which the chroma sub-block is sampled; The method of claim 7 further comprising:
9. the first format is a 4:2:0 format; the second format is a 4:2:2 format; the third format is a 4:4:4 format; The method according to claim 8.
10. in response to the chroma format information indicating that the chroma sub-blocks are sampled in the second format or the third format, deriving a subWidthC chroma format parameter and a subHeightC chroma format parameter based on the chroma format information; selecting two luma sub-blocks based on the subWidthC chroma format parameter and the subHeightC chroma format parameter; averaging the motion vector information corresponding to the two selected luminance sub-blocks; and The method of claim 8 further comprising:
11. deriving a subWidthC chroma format parameter and a subHeightC chroma format parameter based on the chroma format information; determining that the chroma sub-block is sampled in the second format if the subWidthC and subHeightC chroma format parameters have different values; The method of claim 8 further comprising:
12. Averaging the motion vector information corresponding to the selected luminance sub-blocks includes: The method of claim 7 , further comprising performing symmetric rounding of the motion vector information corresponding to the selected luma sub-block.
13. 10. A non-transitory computer-readable storage medium comprising stored instructions that, when executed by at least one processor, perform the method of claim 7.
14. 1. An electronic device for encoding pictures into a data stream, comprising: Encoding chroma format information into the data stream; determining motion vectors corresponding to chroma sub-blocks of the picture; determining the motion vector corresponding to the chroma sub-block based on: (i) selecting a top left luma sub-block and a bottom right luma sub-block in response to the chroma format information specifying a format in which the chroma sub-block is a sample; and (ii) averaging motion vector information corresponding to the selected luma sub-blocks to determine the motion vector corresponding to the chroma sub-block; 11. An electronic device comprising: at least one processor configured to reconstruct the picture based at least in part on the motion vectors corresponding to the chroma sub-blocks.
15. 1. A method for encoding a picture into a data stream, comprising: encoding chroma format information into a data stream; determining motion vectors corresponding to chroma sub-blocks of the picture; determining the motion vectors corresponding to the chroma sub-blocks based on: (i) selecting a top-left luma sub-block and a bottom-right luma sub-block in response to the chroma format information specifying a format in which the chroma sub-blocks are samples; and (ii) averaging motion vector information corresponding to the selected luma sub-blocks; reconstructing the picture based at least in part on the motion vectors corresponding to the chroma sub-blocks; A method for providing the above.
16. 20. A non-transitory computer-readable storage medium comprising stored instructions that, when executed by at least one processor, performs the method of claim 15.
Citation Information
Patent Citations
Methods and apparatus of video coding for deriving affine motion vectors for chroma components
WO2020131583A1
Method and apparatus for affine based inter prediction of chroma subblocks
WO2020169114A1