Method and apparatus for affine-based interpretation of chromatic subblocks

By deriving motion vectors for affine-based interpretation of chroma subblocks based on chroma format, the method addresses the challenge of supporting multiple chroma formats, enhancing coding performance and stability in video coding systems.

JP2026077838APending Publication Date: 2026-05-13HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2026-02-18
Publication Date
2026-05-13

AI Technical Summary

Technical Problem

Current video coding technologies face challenges in supporting multiple chroma formats, particularly 4:4:4, leading to crashes in Versatile Video Coding (VTM) coders, necessitating improved methods for deriving motion vectors for affine-based interpretation of chroma subblocks.

Method used

A method and apparatus for deriving motion vectors for affine-based interpretation of chroma subblocks based on chroma format, involving determining horizontal and vertical chroma scaling factors and using these to define sets of luma subblocks, averaging motion vectors of luma subblocks to derive accurate chroma motion vectors, and optimizing processing by unifying luma and chroma operations.

Benefits of technology

This approach enhances coding performance by reducing prediction errors and improving compression efficiency, particularly for 4:4:4 chroma formats, ensuring stable video coding operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026077838000001_ABST
    Figure 2026077838000001_ABST
Patent Text Reader

Abstract

The present invention relates to a method and apparatus for deriving motion vectors for affine-based interpretation of chromatic subblocks based on chromatoromat. [Solution] Based on chroma format information, horizontal and vertical chroma scaling factors are determined, the chroma format information indicates the chroma format of the current picture to which the current image block belongs, a set of luma subblocks (S) of the luma block is determined based on the value of the chroma scaling factor, and the motion vectors for the chroma subblocks of the chroma block are determined based on the motion vectors of one or more luma subblocks within the set of luma subblocks (S).
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [Cross-references to related applications] This patent application is a divisional application of Japanese Patent Application No. 2024-23150, filed on 27 December 2024, which is a divisional application of Japanese Patent Application No. 2023-130773, filed on 10 August 2023, which is a divisional application of Japanese Patent Application No. 2021-549409, filed on 24 February 2020, which claims priority to U.S. Provisional Patent Application No. 62 / 809,551, filed on 22 February 2019, priority to U.S. Provisional Patent Application No. 62 / 823,653, filed on 25 March 2019, and priority to U.S. Provisional Patent Application No. 62 / 824,302, filed on 26 March 2019. The disclosure of the above patent application is incorporated by reference to its entirety.

[0002] [Technical field] Embodiments of this disclosure generally relate to the field of picture processing, and more particularly to affine-based interpretation (affine motion compensation), and in particular to a method and apparatus for deriving motion vectors for affine-based interpretation of chroma subblocks based on a chroma format, and a method and apparatus for affine-based interpretation of chroma subblocks. [Background technology]

[0003] Video coding (video encoding and / or decoding) is used in a wide range of digital video applications, such as broadcast digital TV, video transmission over the internet and mobile networks, real-time conversation applications like video chat, video conferencing, DVD and Blu-ray discs, video content acquisition and editing systems, and camcorders in security applications.

[0004] Even relatively short videos can require a considerable amount of video data to depict, which can cause difficulties when data is streamed or transmitted across communication networks with limited bandwidth. Therefore, video data is generally compressed before being transmitted across modern telecommunications networks. Video size can also be a concern when video is stored on a storage device, as memory resources can be limited. Video compression devices often use software and / or hardware at the source to encode video data before transmission or storage, thereby reducing the amount of data required to represent the digital video image. The compressed data is then received at the destination by a video decompression device that decodes the video data. Due to limited network resources and the ever-increasing demand for higher video quality, improved compression and decompression techniques with higher compression ratios and little to no sacrifice of picture quality are desirable.

[0005] In particular, current Versatile Video Coding and Test Model (VTM) coders primarily support the 4:2:0 chroma format as the input picture format. Using a 4:4:4 input chroma format can cause the VTM coder to crash. To avoid such problems, coders that support other chroma formats (e.g., 4:4:4 or 4:2:2) are highly desirable, and even essential for a wide range of applications. [Overview of the project]

[0006] In consideration of the above challenges, modifications to the video coding process to support multiple chroma formats are proposed in this disclosure. In particular, embodiments of the present application aim to provide a device, encoder, decoder and corresponding method for deriving motion vectors for affine-based interpretation of chroma subblocks based on one or more supported chroma formats in order to improve coding performance.

[0007] Embodiments of the present invention are defined by the features of the independent claims, and more advantageous ways of realizing the embodiments are defined by the features of the dependent claims.

[0008] Specific embodiments are outlined in the attached independent claims, and other embodiments are outlined in the dependent claims.

[0009] The above and other objectives are achieved by the subject matter of the independent claim. Further means of implementation are evident from the dependent claims, description and drawings.

[0010] According to a first aspect of this disclosure, a method for deriving chroma motion vectors used in affine-based interpretation of current image blocks including chroma blocks and chroma blocks at the same location is provided, the method is: The steps include determining the horizontal and vertical chroma scaling factors (i.e., the values ​​of the chroma scaling factors) based on chroma format information, wherein the chroma format information indicates the chroma format of the current picture to which the current image block belongs, and A step of determining the set of luma subblocks (S) of a luma block based on the value of the chroma scaling factor, A step of determining the motion vector for a chroma block of a chroma block based on the motion vectors of one or more luma subblocks (e.g., one or two luma subblocks) in a set (S) of luma subblocks. Includes.

[0011] In this disclosure, a (luma or chroma) block or subblock can be represented by its location, position, or index; therefore, selecting / determining a block or subblock means selecting or determining the location, position, or index of the block or subblock.

[0012] It should be noted that the terms “block,” “coding block,” or “image block” as used in this disclosure may refer to a transform unit (TU), a prediction unit (PU), a coding unit (CU), etc. In Versatile Video Coding (VVC), transform units and coding units are generally aligned with each other, except when TU tiling or sub-block transform (SBT) is used. Therefore, the terms “block,” “image block,” “coding block,” and “transform block” may be used interchangeably in this disclosure, and the terms “block size” and “transform block size” may also be used interchangeably in this disclosure. The terms “sample” and “pixel” may also be used interchangeably in this disclosure.

[0013] This disclosure relates to a method for considering the chroma format of a picture when obtaining chroma motion vectors from chroma motion vectors. Linear subsampling of the chroma motion field is performed by averaging the chroma motion vectors. When the chroma color plane is at the same height as the chroma plane, motion vectors are selected from horizontally adjacent chroma blocks, thereby it is found that it is more appropriate for them to have the same vertical position. Selecting chroma motion vectors based on the picture chroma format results in a more accurate chroma motion field due to more accurate chroma motion field subsampling. This dependency on the chroma format allows for the selection of the most appropriate chroma blocks when averaging the chroma motion vectors to generate the chroma motion vectors. As a result of more accurate motion field interpolation, prediction errors are reduced, which leads to a technical consequence of improved compression performance and therefore improved coding performance.

[0014] In a possible implementation of the method according to the first embodiment, the set of luma subblocks (S) is determined based on the values ​​of the horizontal and vertical chroma scaling factors. That is, one or more luma subblocks (e.g., one or two luma subblocks) are determined based on the values ​​of the horizontal and vertical chroma scaling factors.

[0015] In possible implementations of the method according to the first embodiment, the horizontal and vertical chroma scaling factors are represented by the variables SubWidthC and SubHeightC.

[0016] In a possible implementation of the method according to the first embodiment, the position of each luma subblock is represented by a horizontal subblock index and a vertical subblock index, and the position of each chroma subblock is represented by a horizontal subblock index and a vertical subblock index.

[0017] In a possible implementation of the method according to the first embodiment, the position of each of one or more Luma subblocks (e.g., one or two Luma subblocks) within a set (S) is represented by a horizontal subblock index and a vertical subblock index.

[0018] In any of the preceding implementations of the first embodiment or possible implementations of the method according to the first embodiment itself, if both variables SubWidthC and SubHeightC are equal to 1, the set of Luma subblocks (S) is: S0=(xSbIdx,ySbIdx) includes a Luma subblock, If at least one of SubWidthC and SubHeightC is not equal to 1, then the set of Luma subblocks (S) is: The first luma subblock indexed by S0=((xSbIdx>>(SubWidthC-1)<<(SubWidthC-1)),(ySbIdx>>(SubHeightC-1)<<(SubHeightC-1))), and The second Luma subblock indexed by S1 = ((xSbIdx>>(SubWidthC-1)<<(SubWidthC-1))+(SubWidthC-1),(ySbIdx>>(SubHeightC-1)<<(SubHeightC-1))+(SubHeightC-1)) Includes, SubWidthC and SubHeightC represent the horizontal and vertical chroma scaling factors, respectively. xSbIdx and ySbIdx represent the horizontal and vertical subblock indices for the Luma subblocks in set (S), respectively, where "<<" represents a left arithmetic shift and ">>" represents a right arithmetic shift, xSbIdx = 0..numSbX-1 and ySbIdx = 0..numSbY-1, where numSbX indicates the number of Luma subblocks in the Luma block along the horizontal direction and numSbY indicates the number of Luma subblocks in the Luma block along the vertical direction.

[0019] In any of the preceding implementations of the first embodiment or in possible implementations of the method according to the first embodiment itself, the number of horizontal and vertical chroma subblocks is the same as the number of horizontal and vertical luma subblocks, respectively.

[0020] In any of the preceding implementations of the first embodiment or possible implementations of the method according to the first embodiment itself, if both SubWidthC and SubHeightC are equal to 1, the set of Luma subblocks (S) is S0=(xCSbIdx,yCSbIdx) includes a Luma subblock, If at least one of SubWidthC and SubHeightC is not equal to 1, then the set of Luma subblocks (S) is: The first luma subblock indexed by S0=((xCSbIdx>>(SubWidthC-1)<<(SubWidthC-1)),(yCSbIdx>>(SubHeightC-1)<<(SubHeightC-1))), and The second Luma subblock indexed by S1 = (xCSbIdx>>(SubWidthC-1)<<(SubWidthC-1))+(SubWidthC-1),(yCSbIdx>>(SubHeightC-1)<<(SubHeightC-1))+(SubHeightC-1)) Includes, The variables SubWidthC and SubHeightC represent the horizontal and vertical chroma scaling factors, respectively, and xCSbIdx and yCSbIdx represent the horizontal and vertical subblock indices for chroma subblocks in set (S), respectively, where xCSbIdx = 0..numCSbX-1 and yCSbIdx = 0..numCSbY-1, where numCSbX indicates the number of chroma subblocks in a chroma block along the horizontal direction, and numCSbY indicates the number of chroma subblocks in a chroma block along the vertical direction.

[0021] In any prior implementation of the first embodiment or a possible implementation of the method according to the first embodiment itself, the size of each chroma subblock is the same as the size of each luma subblock. When the number of chroma subblocks is defined as equal to the number of luma subblocks, and the chroma color plane size is equal to the luma plane size (e.g., the chroma format of the input picture is 4:4:4), it is permissible to assume that the motion vectors of adjacent chroma subblocks are the same. When implementing this processing step, optimization may be performed by skipping the iterative value calculation step.

[0022] The proposed invention discloses a method for defining the size of a chroma subblock equal to the size of a luma subblock. In this case, the implementation can be simplified by unifying the luma and chroma processing, and redundant motion vector calculations are naturally avoided.

[0023] In possible implementations of any prior implementation of the first embodiment or the method according to the first embodiment itself, if the size of each chroma subblock is the same as the size of each luma subblock, the number of horizontal chroma subblocks depends on the number of horizontal luma subblocks and the value of the horizontal chroma scaling factor, and the number of vertical chroma subblocks depends on the number of vertical luma subblocks and the value of the vertical chroma scaling factor.

[0024] In any of the preceding implementations of the first embodiment or possible implementations of the method according to the first embodiment itself, xCSbIdx is obtained based on the step values ​​of xSbIdx and SubWidth, and yCSbIdx is obtained based on the step values ​​of ySbIdx and SubHeightC.

[0025] In any of the preceding implementations of the first embodiment or the method of the first embodiment itself, numCSbX=numSbX>>(SubWidthC-1) numCSbY=numSbY>>(SubHeightC-1) And, numCSbX and numCSbY represent the number of chroma subblocks in the horizontal and vertical directions, respectively. numSbX and numSbY represent the number of luma subblocks within a luma block in the horizontal and vertical directions, respectively.

[0026] In any of the preceding implementations of the first embodiment or possible implementations of the method according to the first embodiment itself, with respect to the chroma block, the set of luma subblocks (S) is: S0=(xSbIdxL,ySbIdxL) S1=(xSbIdxL,ySbIdxL+(SubHeightC-1)) S2=(xSbIdxL+(SubWidthC-1),ySbIdxL) S3=(xSbIdxL+(SubWidthC-1),ySbIdxL+(SubHeightC-1)) Includes one or more subblocks indexed by The luma block position or index S0 is represented by the horizontal subblock index xSbIdxL and the vertical subblock index ySbIdxL. Regarding the chroma block position (e.g., [xSbIdxL][ySbIdxL] in mvCLX[xSbIdxL][ySbIdxL]), The luma block position or index S1 is represented by the horizontal subblock index xSbIdxL and the vertical subblock index ySbIdxL+(SubHeightC-1), The luma block position or index S2 is represented by the horizontal subblock index xSbIdxL+(SubWidthC-1) and the vertical subblock index ySbIdxL. The luma block position or index S3 is represented by the horizontal subblock index xSbIdxL+(SubWidthC-1) and the vertical subblock index ySbIdxL+(SubHeightC-1).

[0027] In any of the preceding implementations of the first embodiment or possible implementations of the method according to the first embodiment itself, the set of luma subblocks (S) is: S0=(xSbIdxL,ySbIdxL) S1=(xSbIdxL+(SubWidthC-1),ySbIdxL+(SubHeightC-1)) It includes two Luma subblocks indexed by, The luma block position or index S0 is represented by the horizontal subblock index xSbIdxL and the vertical subblock index ySbIdxL. The luma block position or index S1 is represented by the horizontal subblock index xSbIdxL+(SubWidthC-1) and the vertical subblock index ySbIdxL+(SubHeightC-1).

[0028] In any of the preceding implementations of the first embodiment or the method of the first embodiment itself, If the chroma format is 4:4:4, then set(S) contains (consists of) one luma subblock in the same position as the chroma subblock, If the chroma format is 4:2:2, then set(S) contains two horizontally adjacent chroma subblocks. If the chroma format is 4:2:0, then set(S) contains two chroma subblocks that are diagonally opposite each other.

[0029] In any of the preceding implementations of the first embodiment or possible implementations of the method according to the first embodiment itself, if there are more than one luma subblock in the set (S), determining the motion vector for a chroma subblock based on the motion vector of at least one luma subblock in the set (S) of luma subblocks is: The motion vectors of the Luma subblocks within set S are averaged. Deriving motion vectors for chroma subblocks based on the average chroma motion vector. Includes.

[0030] In any of the preceding implementations of the first embodiment or possible implementations of the method according to the first embodiment itself, averaging the motion vectors of the luma subblocks in set S is: Average the horizontal component of the motion vector of the Luma subblock within set S, and / or Average the vertical component of the motion vector of the Luma subblock within set S. Includes.

[0031] In any of the preceding implementations of the first embodiment or possible implementations of the method according to the first embodiment itself, averaging the motion vectors of the luma subblocks in set S includes checking whether the sum of the motion vectors of the luma subblocks in set S is 0 or greater.

[0032] In any of the preceding implementations of the first embodiment or possible implementations of the method according to the first embodiment itself, averaging the motion vectors of the luma subblocks in set S is: mvAvgLX=Σ i mvLX[S i x ][S i y ] If mvAvgLX[0] is greater than or equal to 0, mvAvgLX[0]=(mvAvgLX[0]+(N>>1)-1)>>log2(N), otherwise, mvAvgLX[0] = -((-mvAvgLX[0] + (N >> 1) - 1) >> log2(N)) When mvAvgLX[1] is 0 or greater, mvAvgLX[1] = (mvAvgLX[1] + (N >> 1) - 1) >> log2(N); otherwise, mvAvgLX[1] = -((-mvAvgLX[1] + (N >> 1) - 1) >> log2(N)) including, mvAvgLX is the result of averaging, mvAvgLX[0] is the horizontal component of the motion vector mvAvgLX, mvAvgLX[1] is the vertical component of the motion vector mvAvgLX, S i x and S i y are the horizontal and vertical indices of the sub-block Si within the set (S) of luma sub-blocks in the motion vector array, mvLX[S i x [S i y is the motion vector of the luma sub-block having the indices S i x and S i y and N is the number of elements (e.g., luma sub-blocks) within the set (S) of luma sub-blocks, log2(N) represents the logarithm of N to the base 2, the power to which the number 2 is raised to obtain the value N, and ">>" is a right arithmetic shift.

[0033] In any preceding implementation of the first aspect or a possible implementation form of the method according to the first aspect itself, N is equal to 2.

[0034] In any preceding implementation of the first aspect or a possible implementation form of the method according to the first aspect itself, averaging the motion vectors of the luma sub-blocks within the set S mvAvgLX=mvLX[(xSbIdx>>(SubWidthC-1)<<(SubWidthC-1))][(ySbIdx>>(SubHeightC-1)<<(SubHeightC-1))]+mvLX[(xSbI dx>>(SubWidthC-1)<<(SubWidthC-1))+(SubWidthC-1)][(ySbIdx>>(SubHeightC-1)<<(SubHeightC-1))+(SubHeightC-1)] If mvAvgLX[0]>=0, mvAvgLX[0]=(mvAvgLX[0]+1-(mvAvgLX[0]>=0))>>1 If mvAvgLX[1] >= 0, mvAvgLX[1]=(mvAvgLX[1]+1-(mvAvgLX[1]>=0))>>1 Includes, mvAvgLX[0] is the horizontal component of the averaged motion vector mvAvgLX, mvAvgLX[1] is the vertical component of the averaged motion vector mvAvgLX, SubWidthC and SubHeightC represent the horizontal and vertical chroma scaling factors, respectively, xSbIdx and ySbIdx represent the horizontal and vertical subblock indices for the chroma subblocks in set(S), "<<" is a left arithmetic shift and ">>" is a right arithmetic shift.

[0035] In any of the preceding implementations of the first embodiment or the method of the first embodiment itself, in case 1: when mvAvgLX[0]>=0, the value of "(mvAvgLX[0]>=0)" is 1, and in case 2: when mvAvgLX[0]<0, the value of "(mvAvgLX[0]>=0)" is 0.

[0036] In any of the preceding implementations of the first embodiment or possible implementations of the method according to the first embodiment itself, averaging the motion vectors of the luminance subblocks in set S is performed. When the sum of the motion vectors of Luma subblocks in set S is 0 or greater, the sum of the motion vectors of Luma subblocks in set S includes dividing by a right shift operation, depending on the number of elements (e.g., Luma subblocks) in the set of Luma subblocks (S).

[0037] In any of the preceding implementations of the first embodiment or possible implementations of the method according to the first embodiment itself, determining the horizontal and vertical chroma scaling factors based on chroma format information is: This includes determining the horizontal and vertical chroma scaling factors based on a mapping between chroma format information and the horizontal and vertical chroma scaling factors.

[0038] Possible implementations of any prior implementation of the first embodiment or the method according to the first embodiment itself further include generating predictions of chroma subblocks based on determined motion vectors.

[0039] In any of the preceding implementations of the first embodiment or possible implementations of the method according to the first embodiment itself, the chroma format includes one of the YUV4:2:2 format, the YUV4:2:0 format, or the YUV4:4:4 format.

[0040] A prior implementation of any of the first embodiments or a method according to the first embodiment itself is implemented by an encoding device.

[0041] A preceding implementation of any of the first embodiments or a method according to the first embodiment itself is implemented by a decoding device.

[0042] According to a second aspect of the present invention, an apparatus is provided for affine-based interpretation of a current image block including lumar and chroma blocks at the same location, the apparatus, A determination module configured to determine horizontal and vertical chroma scaling factors based on chroma format information, wherein the chroma format information indicates the chroma format of the current picture to which the current image block belongs, and is configured to determine the set of chroma subblocks (S) of the chroma block based on the value of the chroma scaling factor, A motion vector derivation module configured to determine the motion vector for a chroma block of a chroma block based on the motion vectors of one or more chroma subblocks within a set (S) of chroma subblocks, and Includes.

[0043] A method according to the first aspect of the present invention can be carried out by an apparatus according to the second aspect of the present invention. Further features and implementations of the apparatus according to the second aspect of the present invention correspond to features and implementations of the method according to the first aspect of the present invention.

[0044] According to a third aspect, the present invention relates to an apparatus for decoding a video stream, comprising a processor and memory. The memory stores instructions causing the processor to perform the method according to the first aspect.

[0045] According to a fourth aspect, the present invention relates to an apparatus for encoding a video stream, comprising a processor and memory. The memory stores instructions causing the processor to perform the method according to the first aspect.

[0046] According to a fifth aspect, a computer-readable storage medium is proposed that stores instructions causing one or more processors configured to code video data when executed. The instructions cause one or more processors to perform a method according to the first aspect or any possible embodiment of the first aspect.

[0047] According to a sixth aspect, the present invention relates to a computer program that, when executed on a computer, includes program code for performing a method according to a possible embodiment of the first or second aspect or the first aspect.

[0048] Details of one or more embodiments are described in the accompanying drawings and the following description. Other features, purposes, and advantages will become apparent from the specification, drawings, and claims. [Brief explanation of the drawing]

[0049] Embodiments of the present invention will be described in more detail below with reference to the attached figures and drawings. [Figure 1A] This block diagram shows an example of a video coding system configured to implement the embodiments presented herein. [Figure 1B] This block diagram shows another example of a video coding system configured to implement the embodiments presented herein. [Figure 2] This block diagram shows an example of a video encoder configured to implement the embodiments presented herein. [Figure 3] This block diagram shows an exemplary structure of a video decoder configured to realize the embodiments presented herein. [Figure 4] This is a block of examples of encoding or decoding devices. [Figure 5] This is a block of other examples of encoding or decoding devices. [Figure 6a] An example is shown regarding the control point motion vector position for a 4-parameter affine motion model. [Figure 6b] An example is shown regarding the control point motion vector position for a 6-parameter affine motion model. [Figure 7] An example of a subblock motion vector field for an affine motion model is shown. [Figure 8] Block diagram for motion compensation using an affine motion model. [Figure 9A] This shows an example of the nominal vertical and horizontal positions of 4:2:0 luminous and chroma samples within a picture. [Figure 9B]This shows an example of the nominal vertical and horizontal positions of 4:2:2 luma and chroma samples within a picture. [Figure 9C] This shows an example of the nominal vertical and horizontal positions of 4:4:4 luma and chroma samples within a picture. [Figure 9D] Various sampling patterns are shown. [Figure 10A] This example shows the luma and chroma blocks at the same position within the current image block of the current picture, with the current picture's chroma format being 4:2:0. [Figure 10B] This example shows the luma and chroma blocks at the same position within the current image block of the current picture, with the current picture's chroma format being 4:2:2. [Figure 10C] This example shows the luma and chroma blocks at the same position within the current image block of the current picture, with the current picture's chroma format being 4:4:4. [Figure 11A] As shown in Figure 10A, this is an example showing the positions of two luma subblocks with respect to a given position of a luma subblock during the chroma motion vector derivation from the luma motion vector, when the current picture chroma format is 4:2:0. [Figure 11B] As shown in Figure 10B, this is an example showing the positions of two luma subblocks with respect to a given position of a luma subblock during the deriving of a luma motion vector from a luma motion vector, when the current chroma format of the picture is 4:2:2. [Figure 11C] As shown in Figure 10C, this is an example showing the position of a chroma subblock relative to a given position of the chroma subblock during the chroma motion vector derivation from the chroma motion vector, when the current picture chroma format is 4:4:4. [Figure 12A] This shows various examples of subset S containing the positions of luma subblocks relative to a given position of chroma subblock when the chroma format is set to 4:4:4. [Figure 12B]Various examples of subset S are shown, and chroma subblocks located on the boundary of a chroma block have corresponding positions within the chroma block, as in the fourth case "D" shown in Figure 12A. [Figure 13A] This example demonstrates the selection of two luma subblocks to derive the motion vector for a given luma subblock during the deriving of the luma motion vector from the luma motion vector, given that the current picture chroma format is 4:2:0. [Figure 13B] This example demonstrates the selection of two luma subblocks for a given luma subblock during the derivation of a luma motion vector from a luma motion vector, given that the current picture chroma format is 4:2:2. [Figure 13C] This example demonstrates the selection of a given chroma subblock for a given chroma subblock during the chroma motion vector derivation from the chroma motion vector, given that the picture's chroma format is currently 4:4:4. [Figure 14A] This shows examples of subdividing a 16x16 luma block into subblocks, and subdividing a luma block at the same location, when the chroma format is YUV4:2:0. [Figure 14B] This shows examples of subdividing a 16x16 luma block into subblocks, and subdividing a luma block at the same location, when the chroma format is YUV4:2:2. [Figure 14C] This shows examples of subdividing a 16x16 luma block into subblocks, and subdividing a luma block at the same location, when the chroma format is YUV4:4:4. [Figure 15] A flowchart illustrating an exemplary process for deriving motion vectors for affine-based interpretation of chroma subblocks based on a chroma format according to several aspects of this disclosure is provided. [Figure 16] A flowchart illustrating other exemplary processes for deriving motion vectors for affine-based interpretation of chroma subblocks based on chroma formats according to several aspects of this disclosure is shown. [Figure 17] Schematic diagrams of devices for affine-based interpretation according to several aspects of this disclosure are shown. [Figure 18] This is a block diagram illustrating an example structure of a content supply system that enables content distribution services. [Figure 19] This is a block diagram showing the structure of an example terminal device.

[0050] In the following, unless otherwise explicitly specified, the same reference numeral indicates the same or at least functionally equivalent features. [Modes for carrying out the invention]

[0051] The following description refers to the accompanying drawings, which form part of this disclosure and illustrate by example specific aspects of embodiments of the present invention or specific ways in which embodiments of the present invention may be used. It is understood that embodiments of the present invention may be used in other ways and may include structural or logical modifications not shown in the drawings. Accordingly, the following detailed description should not be taken as limiting, and the scope of the invention is defined by the appended claims.

[0052] For example, disclosure relating to a described method may also apply to a corresponding device or system configured to perform the method, and vice versa. For example, if one or more steps of a particular method are described, the corresponding device may include one or more units, e.g., functional units, for performing the steps of the described one or more methods, even if one or more such units are not explicitly described or shown in the drawings (e.g., one unit performs one or more steps, or multiple units each perform one or more of the steps). On the other hand, for example, if a particular device is described based on one or more units, e.g., functional units, the corresponding method may include one step for performing the function of one or more units, even if one or more such steps are not explicitly described or shown in the drawings (e.g., one step performs the function of one or more units, or multiple steps each perform one or more of the functions of the multiple units). Furthermore, it is understood that the various exemplary embodiments and / or features described herein may be combined with each other unless otherwise specified.

[0053] Typically, video coding refers to the processing of a sequence of pictures that make up a video or video sequence. Instead of the term “picture,” the terms “frame” or “image” may be used synonymously in the field of video coding. Video coding (or coding in general) comprises two parts: video encoding and video decoding. Video encoding is performed on the source side and typically involves processing the original video picture (e.g., by compression) to reduce the amount of data required to represent the video picture (for more efficient storage and / or transmission). Video decoding is performed on the destination side and typically involves the reverse processing compared to the encoder in order to reconstruct the video picture. Embodiments referring to “coding” a video picture (or picture in general) are understood to relate to “encoding” or “decoding” a video picture or the respective video sequence. The combination of the encoding and decoding parts is also called a codec (Coding and Decoding).

[0054] In lossless video coding, the original video picture can be reconstructed, meaning the reconstructed video picture has the same quality as the original video picture (assuming there is no transmission loss or other data loss during storage or transmission). In lossy video coding, further compression is performed, for example, by quantization, to reduce the amount of data representing the video picture, which cannot be fully reconstructed by the decoder, meaning the quality of the reconstructed video picture is lower or worse compared to the quality of the original video picture.

[0055] Several video coding standards belong to the group of “lossy hybrid video codecs” (i.e., combining spatial and temporal prediction in the sample domain with 2D transform coding to apply quantization in the transform domain). Each picture in a video sequence is typically partitioned into a set of non-overlapping blocks, and coding is typically performed at the block level. In other words, in the encoder, for example, spatial (intra-picture) prediction and / or temporal (inter-picture) prediction are used to generate predicted blocks, the predicted blocks are subtracted from the current block (the block currently being processed / to be processed) to obtain the residual block, and the residual block is transformed to quantize the residual block in the transform domain to reduce the amount of data to be transmitted (compression), so that video is typically processed, i.e., coded, at the block (video block) level. On the other hand, in the decoder, the reverse processing compared to the encoder is applied to the coded or compressed block to reconstruct the current block for representation. Furthermore, the encoder replicates the decoder processing loop so that both generate the same predictions (e.g., intra and inter-predictions) and / or reconstructions to process, i.e., code, subsequent blocks.

[0056] This disclosure relates to improvements in the interpretation process. In particular, this disclosure relates to improvements in the derivation process for chroma motion vectors. Specifically, this disclosure relates to improvements in the derivation process for motion vectors of affine chroma blocks (chroma subblocks, etc.). More specifically, this disclosure relates to improvements in the derivation process for motion vectors for affine-based interpretation of chroma subblocks based on a chroma format.

[0057] Here, an improved mechanism is disclosed to support multiple chroma formats for the derivation process of chroma motion vectors.

[0058] In the following embodiments of the video coding system 10, the video encoder 20 and video decoder 30 will be described with reference to Figures 1 to 3.

[0059] Figure 1A is a schematic block diagram showing an exemplary coding system 10, for example, a video coding system 10 (or simply coding system 10), which may utilize the technology of the present application. The video encoder 20 (or simply encoder 20) and video decoder 30 (or simply decoder 30) of the video coding system 10 represent examples of devices that may be configured to perform the various exemplary technologies described in the present application.

[0060] As shown in Figure 1A, the coding system 10 includes a source device 12 configured to provide encoded picture data 21 to, for example, a destination device 14 for decoding the encoded picture data 21.

[0061] The source device 12 includes an encoder 20 and may also include, optionally, a picture source 16, a preprocessor (or preprocessing unit) 18, for example, a picture preprocessor 18, and a communication interface or communication unit 22.

[0062] The picture source 16 may be any type of picture capture device, e.g., a camera for capturing real-world pictures, and / or any type of picture generation device, e.g., a computer graphics processor for generating computer-animated pictures, or any other type of device for acquiring and / or providing real-world pictures, computer-generated pictures (e.g., screen content, virtual reality (VR) pictures), and / or any combination thereof (e.g., augmented reality (AR) pictures). The picture source may also be any type of memory or storage for storing any of the above pictures.

[0063] In contrast to the processing performed by the preprocessor 18 and the preprocessing unit 18, the picture or picture data 17 may also be called the raw picture or raw picture data 17.

[0064] The preprocessor 18 is configured to receive (raw) picture data 17, perform preprocessing on the picture data 17, and obtain a preprocessed picture 19 or preprocessed picture data 19. The preprocessing performed by the preprocessor 18 may include, for example, cropping, color format conversion (e.g., RGB to YCbCr), color correction, or noise reduction. It is understood that the preprocessing unit 18 may be any selected component.

[0065] The video encoder 20 is configured to receive pre-processed picture data 19 and provide encoded picture data 21 (further details are described below, for example, based on Figure 2).

[0066] The communication interface 22 of the source device 12 may be configured to receive encoded picture data 21 and transmit the encoded picture data 21 (or any further processed version thereof) over the communication channel 13 to another device, for example, the destination device 14 or any other device, for storage or direct reconstruction.

[0067] The destination device 14 includes a decoder 30 (e.g., a video decoder 30), and may also include, optionally, a communication interface or communication unit 28, a post-processor 32 (or post-processing unit 32), and a display device 34.

[0068] The communication interface 28 of the destination device 14 is configured to receive encoded picture data 21 (or any further processed version thereof) from, for example, directly from the source device 12 or from any other source, such as a storage device, such as an encoded picture data storage device, and to provide the encoded picture data 21 to the decoder 30.

[0069] Communication interfaces 22 and 28 may be configured to transmit or receive encoded picture data 21 or encoded data 21 via a direct communication link between the source device 12 and the destination device 14, for example, via a direct wired or wireless connection, or via any type of network, for example, a wired or wireless network or a combination thereof, or any type of private and public network, or a combination thereof.

[0070] The communication interface 22 may be configured, for example, to package the encoded picture data 21 into an appropriate format, such as a packet, and / or process the encoded picture data using any type of transmit encoding or processing for transmission over a communication link or communication network.

[0071] The communication interface 28 that forms the counterpart to the communication interface 22 may be configured, for example, to receive the transmitted data and process the transmitted data using any of the corresponding types of transmission decoding or processing and / or depackaging to obtain the encoded picture data 21.

[0072] Both communication interfaces 22 and 28 may be configured as unidirectional or bidirectional communication interfaces, as indicated by the arrows in Figure 1A for the communication channel 13 pointing from source device 12 to destination device 14, and may be configured, for example, to send and receive messages, and to approve and exchange any other information relating to communication links and / or data transmission, such as encoded picture data transmission, for example, to establish a connection.

[0073] The decoder 30 is configured to receive encoded picture data 21 and provide decoded picture data 31 or decoded picture 31 (further details are described below, for example, based on Figure 3 or Figure 5).

[0074] The post-processor 32 of the destination device 14 is configured to post-process the decoded picture data 31 (also called reconstructed picture data), for example, the decoded picture 31, to obtain post-processed picture data 33, for example, the post-processed picture 33. The post-processing performed by the post-processing unit 32 may include, for example, color format conversion (e.g., YCbCr to RGB), color correction, cropping or resampling, or any other processing to prepare the decoded picture data 31 for display by, for example, the display device 34.

[0075] The display device 34 of the destination device 14 is configured to receive the post-processed picture data 33 and, for example, display the picture to a user or viewer. The display device 34 may be any type of display that presents the reconstructed picture, for example, an integrated or external display or monitor, or may include one such display. The display may include, for example, a liquid crystal display (LCD), an organic light-emitting diode (OLED) display, a plasma display, a projector, a micro-LED display, a liquid crystal on silicon (LCoS), a digital light processor (DLP), or any other type of display.

[0076] Figure 1A shows the source device 12 and the destination device 14 as separate devices, but the device embodiment may also include the functionality of either or both of the source device 12 or its corresponding function and the destination device 14 or its corresponding function. In such embodiments, the source device 12 or its corresponding function and the destination device 14 or its corresponding function may be implemented using the same hardware and / or software, or by separate hardware and / or software, or any combination thereof.

[0077] As will be apparent to those skilled in the art based on the description, the presence and (exact) division of different units or functions within the source device 12 and / or destination device 14 as shown in Figure 1A may vary depending on the actual device and application.

[0078] The encoder 20 (e.g., video encoder 20) or the decoder 30 (e.g., video decoder 30), or both the encoder 20 and the decoder 30, may be implemented via processing circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, dedicated video coding, or any combination thereof, as shown in Figure 1B. The encoder 20 may be implemented via processing circuits 46 to embody various modules, as described with respect to the encoder 20 in Figure 2 and / or any other encoder system or subsystem described herein. The decoder 30 may be implemented via processing circuits 46 to embody various modules, as described with respect to the decoder 30 in Figure 3 and / or any other decoder system or subsystem described herein. The processing circuits may be configured to perform various operations, as described below. As shown in Figure 5, if the technology is partially implemented in software, the device may store instructions for the software in a suitable non-temporary computer-readable storage medium, and may execute the instructions in hardware using one or more processors to perform the technology of this disclosure. Either the video encoder 20 or the video decoder 30 may be integrated as part of a combined encoder / decoder (CODEC) within a single device, for example, as shown in Figure 1B.

[0079] The source device 12 and destination device 14 may include any of a broad range of devices, including handheld or fixed devices of any kind, such as notebook or laptop computers, mobile phones, smartphones, tablets or tablet computers, cameras, desktop computers, set-top boxes, televisions, display devices, digital media players, video game consoles, video streaming devices (such as content service servers or content distribution servers), broadcast receiver devices, broadcast transmitter devices, etc., and may or may not use any kind of operating system. In some cases, the source device 12 and destination device 14 may be equipped for wireless communication. Therefore, the source device 12 and destination device 14 may also be wireless communication devices.

[0080] In some cases, the video coding system 10 shown in Figure 1A is merely an example, and the technology of the present invention may be applied to video coding configurations (e.g., video coding or video decoding) that do not necessarily involve any data communication between the coding device and the decoding device. In other examples, the data may be retrieved from local memory, streamed over a network, etc. The video coding device may code the data and store it in memory, and / or the video decoding device may retrieve the data from memory and decode it. In some examples, coding and decoding are performed by devices that do not communicate with each other but simply code the data into memory and / or retrieve the data from memory and decode it.

[0081] For the sake of explanation, embodiments of the present invention are described herein by reference to, for example, reference software for High-Efficiency Video Coding (HEVC) or Versatile Video Coding (VVC), and next-generation video coding standards developed by the Joint Collaboration Team on Video Coding (JCT-VC) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Motion Picture Experts Group (MEPEG). Those skilled in the art will understand that embodiments of the present invention are not limited to HEVC or VVC.

[0082] Encoder and encoding method Figure 2 shows a schematic block diagram of an exemplary video encoder 20 configured to realize the technology of the present invention. In the example of Figure 2, the video encoder 20 includes an input 201 (or input interface 201), a residual calculation unit 204, a transformation processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transformation processing unit 212, a reconstruction unit 214, a loop filter unit 220, a decoded picture buffer (DPB) 230, a mode selection unit 260, an entropy coding unit 270, and an output 272 (or output interface 272). The mode selection unit 260 may include an inter-prediction unit 244, an intra-prediction processing unit 254, and a partition unit 262. The inter-prediction unit 244 may include a motion estimation unit and a motion compensation unit (not shown). The video encoder 20 as shown in Figure 2 may also be called a hybrid video encoder or a video encoder with a hybrid video codec.

[0083] The residual calculation unit 204, the conversion processing unit 206, the quantization unit 208, and the mode selection unit 260 may also be referred to as forming the forward signal path of the encoder 20. On the other hand, the inverse quantization unit 210, the inverse conversion processing unit 212, the reconstruction unit 214, the buffer 216, the loop filter 220, the decoded picture buffer (DPB) 230, the inter-prediction unit 244, and the intra-prediction unit 254 may also be referred to as forming the reverse signal path of the video encoder 20, where the reverse signal path corresponds to the signal path of the decoder (see decoder 30 in Figure 3). The inverse quantization unit 210, the inverse conversion processing unit 212, the reconstruction unit 214, the loop filter 220, the decoded picture buffer (DPB) 230, the inter-prediction unit 244, and the intra-prediction unit 254 are also referred to as forming the “built-in decoder” of the video encoder 20.

[0084] Picture and picture partition (picture and block) The encoder 20 may be configured, for example, to receive a picture 17 (or picture data 17) via input 201, for example, a picture of a sequence of pictures that form a video or video sequence. The received picture or picture data may also be a preprocessing picture 19 (preprocessing picture data 19). For brevity, the following description will refer to picture 17. Picture 17 may also be called the current picture or the picture to be coded (in particular, in video coding, to distinguish the current picture from other pictures, for example, pictures encoded and / or decoded before the same video sequence, i.e., the video sequence that also contains the current picture).

[0085] A (digital) picture is or can be thought of as a two-dimensional array or matrix of samples having intensity values. A sample in an array may also be called a pixel (short for picture element) or pel. The number of samples in the horizontal and vertical directions (or axes) of the array or picture defines the size and / or resolution of the picture. For color representation, typically three color components are used; that is, a picture may or may contain three sample arrays. In the RGB format or color space, a picture contains corresponding red, green, and blue sample arrays. However, in video coding, each pixel is typically represented by a luminance and chrominance format or color space, e.g., YCbCr, which includes a luminance component represented by Y (sometimes L is used instead) and two chrominance components represented by Cb and Cr. The luminance (or abbreviated as luma) component Y represents brightness or gray level intensity (e.g., in a grayscale picture). On the other hand, the two chrominance (or chroma for short) components Cb and Cr represent chromaticity or color information components. Therefore, a picture in YCbCr format includes a luminance sample array of luminance sample values ​​(Y) and two chrominance sample arrays of chrominance values ​​(Cb and Cr). A picture in RGB format may be converted to or from YCbCr format, and vice versa; the process is also known as color conversion or transformation. If the picture is monochrome, the picture may include only a luminance sample array. Therefore, a picture may be, for example, a luminance sample array in a monochrome format, or two corresponding arrays of luminance samples and chroma samples in 4:2:0, 4:2:2, and 4:4:4 color formats.

[0086] Embodiments of the video encoder 20 may include a picture partition unit (not shown in Figure 2) configured to partition a picture 17 into a plurality of (typically non-overlapping) picture blocks 203. These blocks may also be called root blocks, macro blocks (H.264 / AVC), coding tree blocks (CTB), or coding tree units (CTU) (H.265 / HEVC and VVC). The picture partition unit may be configured to use the same block size for all pictures in the video sequence and for the corresponding grid that defines the block size, or to change the block size between pictures or subsets or groups of pictures and partition each picture into the corresponding block.

[0087] In a further embodiment, the video encoder may be configured to directly receive blocks 203 of picture 17, for example, one, some, or all of the blocks that make up picture 17. Picture block 203 may also be called the picture block now or the picture block to be coded.

[0088] Similar to picture 17, picture block 203 is, or can be thought of as, a two-dimensional array or matrix of samples having intensity values ​​(sample values), but with fewer dimensions than picture 17. In other words, block 203 may contain, for example, one sample array (e.g., a luma array in the case of monochrome picture 17, or a luma or chroma array in the case of a color picture) or three sample arrays (e.g., a luma and two chroma arrays in the case of color picture 17), or any other number and / or type of arrays depending on the applied color format. The number of samples in the horizontal and vertical directions (or axes) of block 203 defines the size of block 203. Thus, the block may be, for example, an M×N (M columns × N rows) array of samples, or an M×N array of conversion coefficients.

[0089] An embodiment of the video encoder 20 shown in Figure 2 may be configured to encode the picture 17 block by block, for example, encoding and prediction may be performed for each block 203.

[0090] An embodiment of the video encoder 20, as shown in Figure 2, may be further configured to partition and / or encode a picture by using slices (also called video slices), the picture may be partitioned into one or more slices (typically non-overlapping) or encoded using such slices, each slice may contain one or more blocks (e.g., CTUs).

[0091] Embodiments of the video encoder 20, as shown in Figure 2, may be further configured to partition and / or encode a picture by using tile groups (also called video tile groups) and / or tiles (also called video tiles), wherein the picture may be partitioned into or encoded using one or more (typically non-overlapping) tile groups, each tile group may include, for example, one or more blocks (e.g., CTUs) or one or more tiles, each tile may be, for example, rectangular in shape and may include one or more blocks (e.g., CTUs), for example, complete or partial blocks.

[0092] Residual calculation The residual calculation unit 204 is configured to calculate the residual block 205 by, for example, subtracting the sample value of the prediction block 265 from the sample value of the picture block 203 for each sample (for each pixel) based on the picture block 203 and the prediction block 265 (further details regarding the prediction block 265 are provided below), thereby obtaining the residual block 205 in the sample domain.

[0093] conversion The transformation processing unit 206 is configured to obtain transformation coefficients 207 in the transformation domain by applying a transformation, such as a discrete cosine transform (DCT) or discrete sine transform (DST), to the sample values ​​of the residual block 205. The transformation coefficients 207 are also called transformation residual coefficients and may represent the residual block 205 in the transformation domain.

[0094] The conversion processing unit 206 may be configured to apply an integer approximation of the DCT / DST, such as the conversion specified for H.265 / HEVC. Compared to the orthogonal DCT conversion, such an integer approximation is typically scaled by a specific factor. Further scaling factors are applied as part of the conversion process to maintain the norm of the residual blocks processed by the forward and inverse conversions. The scaling factors are typically selected based on specific constraints, such as the scaling factor being a power of 2 for the shift operation, the bit depth of the conversion coefficients, and the trade-off between precision and implementation cost. A specific scaling factor may be specified, for example, for the inverse conversion by the inverse conversion processing unit 212 (and, for example, the corresponding inverse conversion by the inverse conversion processing unit 312 in the video decoder 30), and a corresponding scaling factor may be specified accordingly, for example, for the forward conversion by the conversion processing unit 206 in the encoder 20.

[0095] Embodiments of the video encoder 20 (each a conversion processing unit 206) may be configured to output conversion parameters, for example, the type of conversion or multiple conversions, which are encoded or compressed, for example, directly or via the entropy coding unit 270, so that, for example, the video decoder 30 may receive and use the conversion parameters for decoding.

[0096] quantization The quantization unit 208 may be configured to quantize the transformation coefficient 207 by, for example, applying scalar quantization or vector quantization to obtain the quantized coefficient 209. The quantized coefficient 209 may also be called the quantized transformation coefficient 209 or the quantized residual coefficient 209.

[0097] The quantization process may reduce the bit depth associated with some or all of the conversion coefficients 207. For example, n-bit conversion coefficients may be truncated to m-bit conversion coefficients during quantization, where n is greater than m. The degree of quantization may be changed by adjusting the quantization parameter (QP). For example, in scalar quantization, different scaling may be applied to achieve finer or coarser quantization. Smaller quantization step sizes correspond to finer quantization, while larger quantization step sizes correspond to coarser quantization. Applicable quantization steps may be indicated by the quantization parameter (QP). The quantization parameter may be, for example, an index to a given set of applicable quantization step sizes. For example, a small quantization parameter may correspond to finer quantization (smaller quantization step size), a large quantization parameter may correspond to coarser quantization (larger quantization step size), and vice versa. Quantization may involve division by the quantization step size, and for example, the corresponding dequantization and / or dequantization by the dequantization unit 210 may involve multiplication by the quantization step size. Some standards, e.g., embodiments by HEVC, may be configured to use quantization parameters to determine the quantization step size. Generally, the quantization step size may be calculated based on the quantization parameters using a fixed-point approximation of the equations involving division. Further scaling factors for quantization and dequantization may be introduced to restore the norm of the residual block, which may be modified due to the scaling used in the fixed-point approximation of the equations for the quantization step size and quantization parameters. In one exemplary implementation, the scaling of the inverse transform and dequantization may be combined. Alternatively, a customized quantization table may be used and signaled, for example, from encoder to decoder in a bitstream. Quantization is an irreversible operation, and losses increase with increasing quantization step size.

[0098] Embodiments of the video encoder 20 (each a quantization unit 208) may be configured to output a quantization parameter (QP) which is encoded, for example, directly or via an entropy coding unit 270, so that, for example, the video decoder 30 may receive and apply the quantization parameter for decoding.

[0099] inverse quantization The inverse quantization unit 210 is configured to obtain the de-quantized coefficients 211 by applying the inverse of the quantization scheme applied by the quantization unit 208 to the quantized coefficients, for example, by applying the inverse of the quantization scheme applied by the quantization unit 208, based on or using the same quantization step size as the quantization unit 208. The de-quantized coefficients 211 are also called de-quantized residual coefficients 211 and may correspond to the transformation coefficients 207, although they are typically not identical to the transformation coefficients due to quantization losses.

[0100] Inverse Transform The inverse transformation processing unit 212 is configured to obtain a reconstructed residual block 213 (or the corresponding antiquantized coefficient 213) in the sample domain by applying the inverse transformation of the transformation applied by the transformation processing unit 206, for example, the inverse discrete cosine transform (DCT) or the inverse discrete sine transform (DST), or other inverse transformations. The reconstructed residual block 213 may also be called the transformation block 213.

[0101] Reconstruction The reconstruction unit 214 (e.g., an adder or totaler 214) is configured to obtain the reconstructed block 215 in the sample domain by adding the transformed block 213 (i.e., the reconstructed residual block 213) to the prediction block 265, for example by adding the sample values ​​of the reconstructed residual block 213 and the prediction block 265 sample by sample.

[0102] filtering The loop filter unit 220 (or simply "loop filter" 220) is configured to filter the reconstructed block 215 to obtain a filtered block 221, or more generally, to filter the reconstructed sample to obtain a filtered sample. The loop filter unit is configured, for example, to smooth pixel transitions or to improve video quality. The loop filter unit 220 may include one or more loop filters, such as a deblocking filter, a sample-adaptive offset (SAO) filter, or one or more other filters, such as a bilateral filter, an adaptive loop filter (ALF), a sharpening, smoothing filter, or a co-filter, or any combination thereof. Although the loop filter unit 220 is shown as an in-loop filter in Figure 2, in other configurations, the loop filter unit 220 may be implemented as a post-loop filter. The filtered block 221 may also be called a filtered reconstructed block 221.

[0103] Embodiments of the video encoder 20 (each a loop filter unit 220) may be configured to output loop filter parameters (such as sample-adaptive offset information) that are encoded, for example, directly or via the entropy coding unit 220, so that, for example, the decoder 30 may receive and apply the same loop filter parameters or the respective loop filters for decoding.

[0104] Decode picture buffer The decoded picture buffer (DPB) 230 may be a memory that stores a reference picture or generally reference picture data for encoding video data by the video encoder 20. The DPB 230 may be formed from any of various memory devices, such as dynamic random access memory (DRAM) including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The decoded picture buffer (DPB) 230 may be configured to store one or more filtered blocks 221. The decoded picture buffer 230 may be further configured to store the same current picture or a different picture, for example, other previously filtered blocks of a previously reconstructed picture, for example, a previously reconstructed and filtered block 221, for example, for interpretation, it may provide a fully reconstructed, i.e., decoded picture (and corresponding reference blocks and samples) and / or a partially reconstructed current picture (and corresponding reference blocks and samples). The decoded picture buffer (DPB) 230 may also be configured to store one or more unfiltered reconstructed blocks 215, or generally, for example, if the reconstructed blocks 215 are not filtered by the loop filter unit 220, unfiltered reconstructed samples, or other further processed versions of any of the reconstructed blocks or samples.

[0105] Mode selection (partition and prediction) The mode selection unit 260 includes a partition unit 262, an inter-prediction unit 244, and an intra-prediction unit 254, and is configured to receive or acquire original picture data, e.g., original block 203 (current block 203 of current picture 17), and reconstructed picture data, e.g., filtered and / or unfiltered reconstructed samples or blocks from, e.g., a decoded picture buffer 230 or other buffers (e.g., a line buffer not shown), from the same (current) picture and / or one or more previously decoded pictures. The reconstructed picture data is used as reference picture data for predictions, e.g., inter-prediction or intra-prediction, to acquire prediction blocks 265 or predictors 265. As will be described in detail below, the embodiments presented herein provide improvements to the inter-prediction unit 244 by providing more accurate motion vector predictions, e.g., affine-based inter-prediction or subblock-based inter-prediction, used by the inter-prediction unit when performing inter-prediction.

[0106] The mode selection unit 260 may be configured to determine or select a partition (including no partition) for the current block prediction mode and a prediction mode (e.g., intra or inter-prediction mode), and to generate a corresponding prediction block 265 to be used for calculating the residual block 205 and for reconstructing the reconstructed block 215.

[0107] Embodiments of the mode selection unit 260 may be configured to select partition and prediction modes (for example, from those supported or available by the mode selection unit 260) that provide the best fit, or in other words, minimum residual (minimum residual meaning better compression for transmission or storage) or minimum signaling overhead (minimum signaling overhead meaning better compression for transmission or storage), or that consider or balance both. The mode selection unit 260 may be configured to determine the partition and prediction modes based on rate distortion optimization (RDO), i.e., to select a prediction mode that provides the minimum rate distortion. In this context, terms such as “best,” “minimum,” and “optimal” do not necessarily refer to an overall “best,” “minimum,” and “optimal,” but may refer to termination or selection criteria such as a value above or below a threshold, or the satisfaction of other constraints that result in a potentially “suboptimal selection” but reduce complexity and processing time.

[0108] In other words, the partition unit 262 may be configured to partition block 203 into smaller block partitions or subblocks (which also form blocks) by repeatedly using, for example, quad-tree partitioning (QT), binary partitioning (BT), or triple-tree partitioning (TT), or any combination thereof, and to perform predictions for each of the block partitions or subblocks, for example, mode selection includes selecting the tree structure of the partitioned block 203, and prediction modes are applied to each of the block partitions or subblocks.

[0109] The following describes in more detail the partitioning (by the partition unit 260, for example) and prediction (by the inter-prediction unit 244 and intra-prediction unit 254) processes performed by the exemplary video encoder 20.

[0110] partition The partition unit 262 may now partition (or divide) block 203 into smaller partitions, for example, smaller blocks of a square or rectangular size. These smaller blocks (also called subblocks) may be further partitioned into even smaller partitions. This is also called a tree partition or hierarchical tree partition, and for example, the root block at root tree level 0 (hierarchical level 0, depth 0) may be recursively partitioned, for example, into two or more blocks at the next lower tree level, for example, a node at tree level 1 (hierarchical level 1, depth 1), and these blocks may be further partitioned into two or more blocks at the next lower tree level, for example, tree level 2 (hierarchical level 2, depth 2), until the partitioning ends, for example, because a termination criterion is met, for example, because the maximum tree depth or minimum block size has been reached. Blocks that are not further partitioned are also called leaf blocks or leaf nodes of the tree. A tree that uses partitions into two partitions is called a binary tree (BT), a tree that uses partitions into three partitions is called a ternary tree (TT), and a tree that uses partitions into four partitions is called a quad tree (QT).

[0111] As described above, the term “block” as used herein may refer to a part of a picture, particularly a square or rectangular portion. For example, with reference to HEVC and VVC, a block may be a coding tree unit (CTU), a coding unit (CU), a prediction unit (PU), and a transform unit (TU), and / or a corresponding block, such as a coding tree block (CTB), a coding block (CB), a transform block (TB), or a prediction block (PB).

[0112] For example, a coding tree unit (CTU) may be or include a CTB of a luminous sample, two corresponding CTBs of a chroma sample of a picture having three sample sequences, or a CTB of a sample of a picture coded using a syntax structure used to code a monochrome picture or three separate color planes and samples. Correspondingly, a coding tree block (CTB) may be an N×N block of samples for some value N, thereby the division of components into a CTB is a partition. A coding unit (CU) may be or include a coding block of a luminous sample, two corresponding coding blocks of a chroma sample of a picture having three sample sequences, or a coding block of a sample of a picture coded using a syntax structure used to code a monochrome picture or three separate color planes and samples. Correspondingly, a coding block (CB) may be an M×N block of samples for some values ​​M and N, thereby the division of a CTB into a coding block is a partition.

[0113] For example, in an embodiment using HEVC, a coding tree unit (CTU) may be partitioned into CUs by using a quadtree structure, which is represented as a coding tree. The decision of whether to code a picture region using interpicture (time) prediction or intrapicture (spatial) prediction is made at the CU level. Each CU can be further partitioned into one, two, or four PUs, according to the PU partitioning type. Within a single PU, the same prediction process is applied, and relevant information is sent to the decoder for each PU. After obtaining residual blocks by applying the prediction process based on the PU partitioning type, the CU can be partitioned into transform units (TUs) according to other quadtree structures similar to the coding tree for the CU.

[0114] For example, in an embodiment of the latest video coding standard currently under development called Versatile Video Coding (VVC), combined quad-tree and binary tree (QTBT) partitions are used, for example, to partition coding blocks. In the QTBT block structure, CUs can be either square or rectangular in shape. For example, a coding tree unit (CTU) is first partitioned by a quad-tree. The quad-tree leaf nodes are further partitioned by a binary tree or a ternary tree structure. The partitioned tree leaf nodes are called coding units (CUs), and their segmentation is used for prediction and transformation processing without further partitioning. This means that CUs, PUs, and TUs have the same block size in the QTBT coding block structure. In parallel, multiple partitions, such as ternary tree partitions, may be used with the QTBT block structure.

[0115] In one example, the mode selection unit 260 of the video encoder 20 may be configured to perform any combination of the partitioning techniques described herein.

[0116] As described above, the video encoder 20 is configured to determine or select the best or most optimal prediction mode from a set of prediction modes (for example, a predetermined set). The set of prediction modes may include, for example, an intra-prediction mode and / or an inter-prediction mode.

[0117] Intra Prediction The set of intra-prediction modes may include 35 different intra-prediction modes, such as DC (or mean) mode and non-directional modes like planar mode, or directional modes like those defined for HEVC, or it may include 67 different intra-prediction modes, such as DC (or mean) mode and non-directional modes like planar mode, or directional modes like those defined for VVC.

[0118] The intra-prediction unit 254 is configured to use reconfigured samples of adjacent blocks of the same current picture to generate an intra-prediction block 265 according to a certain intra-prediction mode from a set of intra-prediction modes.

[0119] The intra-prediction unit 254 (or generally the mode selection unit 260) is further configured to output intra-prediction parameters (or generally information indicating the selected intra-prediction mode for a block) to the entropy coding unit 270 in the form of syntax elements 226 for inclusion in the coded picture data 21, so that, for example, the video decoder 30 may receive and use the prediction parameters for decoding.

[0120] Interpretation The set of interpretation modes (or possible ones) depends on the available reference picture (i.e., a previous at least partially decoded picture stored in DPB230) and other interpretation parameters, such as whether the entire reference picture is used to search for the best-fitting reference block, or only a portion of the reference picture, such as the search window area around the current block's region, and / or whether pixel interpolation, such as half / semi-per and / or quarter-per interpolation, is applied.

[0121] In addition to the prediction modes described above, skip mode and / or direct mode may also be applied.

[0122] The interpretation unit 244 may include a motion estimation (ME) unit and a motion compensation (ME) unit (neither of which are shown in Figure 2). The motion estimation unit may be configured to receive or acquire, for motion estimation, a picture block 203 (the current block 203 of the current picture 17) and a decoded picture 231, or at least one or more previously reconstructed blocks, for example, one or more other / different previously reconstructed blocks of a decoded picture 231. For example, a video sequence may include a current picture and a previous decoded picture 231, or in other words, the current picture and the previous decoded picture 231 may be part of or form part of a sequence of pictures that make up the video sequence.

[0123] The encoder 20 may be configured, for example, to select a reference block from multiple reference blocks of the same or different pictures of multiple other pictures, and to provide the motion estimation unit with an offset (spatial offset) between the reference picture (or reference picture index) and / or the position (x,y coordinates) of the reference block and the position of the current block as an interpretation parameter. This offset is also called a motion vector (MV).

[0124] The motion compensation unit is configured to acquire, for example, interpretation parameters, receive them, and perform interpretation based on or using the interpretation parameters to acquire interpretation blocks 265. Motion compensation performed by the motion compensation unit may include fetching or generating prediction blocks based on motion / block vectors determined by motion estimation, and optionally performing interpolation to sub-pixel precision. Interpolation filtering may generate further pixel samples from known pixel samples, thus potentially increasing the number of candidate prediction blocks that can be used to code picture blocks. Now, upon receiving the motion vector of a picture block, the motion compensation unit may find the prediction block pointed to by the motion vector in one of the reference picture lists. Improvements to interpretation (in particular, affine-based interpretation or sub-block-based interpretation) are made by supporting multiple chroma formats and improving the affine sub-block motion vector derivation process. In particular, improved methods and apparatus for motion vector derivation for affine-based interpretation of chroma sub-blocks based on chroma formats are introduced below.

[0125] The motion compensation unit may also generate syntax elements associated with blocks and video slices for use by the video decoder 30 when decoding picture blocks of video slices. In addition to or as an alternative to slices and their respective syntax elements, tile groups and / or tiles and their respective syntax elements may be generated or used.

[0126] Entropy coding The entropy coding unit 270 is configured to obtain encoded picture data 21 that can be output via output 272 in the form of an encoded bitstream 21, for example, by applying or bypassing (uncompressing) an entropy coding algorithm or scheme (e.g., variable length coding (VLC) scheme, context adaptive VLC (CAVLC) scheme, arithmetic coding scheme, binarization, context adaptive binary arithmetic coding (CABAC) scheme, syntax-based context-adaptive binary arithmetic coding (SBAC) scheme, probability interval partitioning entropy (PIPE) coding scheme, or other entropy coding methods or techniques) to the quantized coefficients 209, inter-prediction parameters, intra-prediction parameters, loop filter parameters and / or other syntax elements, for example, an encoded picture data 21 that can be output via output 272 in the form of an encoded bitstream 21, so that, for example, a video decoder 30 may receive and use the parameters for decoding. The encoded bitstream 21 may be transmitted to the video decoder 39, or it may be stored in memory for later transmission or retrieval by the video decoder 30.

[0127] Other structural variations of the video encoder 20 can be used to encode video streams. For example, a non-conversion-based encoder 20 can directly quantize the residual signal for a given block or frame without a conversion processing unit 206. In other implementations, the encoder 20 may have a quantization unit 208 and an inverse quantization unit 210 coupled into a single unit.

[0128] Decoder and decoding method Figure 3 shows an example of a video decoder 30 configured to realize the technology of the present invention. The video decoder 30 is configured to receive encoded picture data 21 (e.g., encoded bitstream 21) encoded by, for example, the encoder 20, in order to obtain a decoded picture 331. The decoded picture data or bitstream includes information for decoding the encoded picture data, for example, data representing picture blocks and associated syntax elements of an encoded video slice (and / or tile group or tile).

[0129] In the example shown in Figure 3, the decoder 30 includes an entropy decoding unit 304, an inverse quantization unit 310, an inverse transformation processing unit 312, a reconstruction unit 314 (e.g., a summer 314), a loop filter 320, a decoded picture buffer (DPB) 330, a mode application unit 360, an interpretation unit 344, and an intraprediction unit 354. The interpretation unit 344 may be or may include a motion compensation unit. In some examples, the video decoder 30 may perform a decoding path that is generally the reverse of the encoding path described for the video encoder 100 from Figure 2.

[0130] As described with respect to encoder 20, the inverse quantization unit 210, inverse processing unit 212, reconstruction unit 214, loop filter 220, decoded picture buffer (DPB) 230, inter-prediction unit 344, and intra-prediction unit 354 may also be called the “internal decoder” of video encoder 20. Thus, the inverse quantization unit 310 may be functionally identical to the inverse quantization unit 110, the inverse processing unit 312 may be functionally identical to the inverse processing unit 212, the reconstruction unit 314 may be functionally identical to the reconstruction unit 214, the loop filter 320 may be functionally identical to the loop filter 220, and the decoded picture buffer 330 may be functionally identical to the decoded picture buffer 230. Therefore, the descriptions provided for each unit and function of video encoder 20 apply correspondingly to each unit and function of video decoder 30.

[0131] Entropy decoding The entropy decoding unit 304 is configured to parse the bitstream 21 (or generally the encoded picture data 21) and, for example, perform entropy decoding on the encoded picture data 21 to obtain, for example, quantized coefficients 309 and / or decoded coding parameters (not shown in Figure 3), such as inter-prediction parameters (e.g., reference picture index and motion vector), intra-prediction parameters (e.g., intra-prediction mode or index), transformation parameters, quantization parameters, loop filter parameters, and / or other syntax elements. The entropy decoding unit 304 may be configured to apply a decoding algorithm or scheme corresponding to an encoding scheme such as those described with respect to the entropy coding unit 270 of the encoder 20. The entropy decoding unit 304 may be further configured to provide the inter-prediction parameters, intra-prediction parameters, and / or other syntax elements to the mode application unit 360 and to provide other parameters to other units of the decoder 30. The video decoder 30 may receive syntax elements at the video slice level and / or video block level. In addition to or as an alternative to slices and their respective syntax elements, tile groups and / or tiles and their respective syntax elements may be received and / or used.

[0132] inverse quantization The inverse quantization unit 310 may be configured to receive quantization parameters (QP, quantization parameter) (or information generally related to inverse quantization) and quantized coefficients from encoded picture data 21 (for example, by parsing and / or decoding by an entropy decoding unit 304), and to apply inverse quantization to the decoded quantized coefficients 309 based on the quantization parameters to obtain de-quantized coefficients 311, which may also be called transformed coefficients 311. The inverse quantization process may include using quantization parameters determined by the video encoder 20 for each video block in a video slice (or tile or tile group) to determine the degree of quantization and the degree of inverse quantization to be applied.

[0133] Inverse Transform The inverse transform processing unit 312 may be configured to receive the dequantized coefficients 311, also called the transform coefficients 311, and to apply a transform to the dequantized coefficients 311 in order to obtain the reconstructed residual block 213 in the sample domain. The reconstructed residual block 213 may also be called the transform block 313. The transform may be an inverse transform, such as an inverse DCT, inverse DST, inverse integer transform, or a conceptually similar inverse transform process. The inverse transform processing unit 312 may be further configured to receive transform parameters or corresponding information from the encoded picture data 21 (for example, by parsing and / or decoding by the entropy decoding unit 304) to determine the transform to be applied to the dequantized coefficients 311.

[0134] Reconstruction The reconstruction unit 314 (for example, an adder or totalizer 314) may be configured to obtain the reconstructed block 315 in the sample domain by adding the reconstructed residual block 313 to the prediction block 365, for example by adding the sample value of the reconstructed residual block 313 to the sample value of the prediction block 365.

[0135] filtering The loop filter unit 320 (either within or after the coding loop) is configured to filter the reconstructed block 315 to obtain the filtered block 321, for example, to smooth pixel transitions or to improve video quality. The loop filter unit 320 may include one or more loop filters, such as a deblocking filter, a sample-adaptive offset (SAO) filter, or one or more other filters, such as a bilateral filter, an adaptive loop filter (ALF), a sharpening, smoothing filter, or a co-filter, or any combination thereof. The loop filter unit 320 is shown in Figure 3 as an in-loop filter, but in other configurations, the loop filter unit 320 may be implemented as a post-loop filter.

[0136] Decode picture buffer The decoded video block 321 of the picture is then stored in a decoded picture buffer 330, which stores the decoded picture as a reference picture for later motion compensation for other pictures and / or for output to their respective displays.

[0137] The decoder 30 is configured to output the decoded picture 331, for example, via output 312, for presentation or viewing by the user.

[0138] prediction The inter-prediction unit 344 may be identical to the inter-prediction unit 244 (in particular, the motion compensation unit), and the intra-prediction unit 354 may be functionally identical to the inter-prediction unit 254, and they perform partition or partition decisions and predictions based on the respective information received from the partition and / or prediction parameters or encoded picture data 21 (for example, by parsing and / or decoding by the entropy decoding unit 304). The mode application unit 360 may be configured to perform predictions (intra or inter-predictions) block by block based on the reconstructed picture, block or each (filtered or unfiltered) sample to obtain a predicted block 365.

[0139] When a video slice is coded as an intra-coded (I) slice, the intra-prediction unit 354 of the mode application unit 360 is configured to generate a prediction block 365 for the picture block of the current video slice based on the signaled intra-prediction mode and data from blocks decoded before the current picture. When a video picture is coded as an inter-coded (i.e., B or P) slice, the inter-prediction unit 344 (e.g., motion compensation unit) of the mode application unit 360 is configured to generate a prediction block 365 for the video block of the current video slice based on motion vectors and other syntax elements received from the entropy decoding unit 304. In inter-prediction, the prediction block may be generated from one of the reference pictures in one of the reference picture lists. The video decoder 30 may configure the reference frame list, list 0 and list 1 using default configuration techniques based on the reference pictures stored in the DPB 330. The same or similar may apply to embodiments that use tile groups (e.g., video tile groups) and / or tiles (e.g., video tiles) in addition to or as an alternative to slices (e.g., video slices), for example, video may be coded using I, P, or B tile groups and / or tiles.

[0140] The mode application unit 360 is configured to determine prediction information for video blocks in the current video slice by parsing motion vectors or related information and other syntax elements, and uses the prediction information to generate prediction blocks for the current video block being decoded. For example, the mode application unit 360 uses some of the received syntax elements to determine the prediction mode (e.g., intra or interpredict) used to code the video blocks in the video slice, the interpredict slice type (e.g., B slice, P slice, or GPB slice), configuration information for one or more of the slice's reference picture lists, the motion vector for each intercoded video block in the slice, the interpredict state for each intercoded video block in the slice, and other information for decoding the video blocks in the current video slice. The same or similar may apply to, or thereafter, embodiments that use tile groups (e.g., video tile groups) and / or tiles (e.g., video tiles) in addition to or as an alternative to slices (e.g., video slices), for example, video may be coded using I, P, or B tile groups and / or tiles.

[0141] Embodiments of the video decoder 30, as shown in Figure 3, may be configured to partition and / or decode a picture using slices (also called video slices), the picture may be partitioned into one or more slices (typically non-overlapping) or decoded using such slices, each slice may contain one or more blocks (e.g., CTUs).

[0142] Embodiments of the video decoder 30, as shown in Figure 3, may be configured to partition and / or decode a picture using tile groups (also called video tile groups) and / or tiles (also called video tiles), wherein the picture may be partitioned into or decoded using one or more (typically non-overlapping) tile groups, each tile group may, for example, contain one or more blocks (e.g., CTUs) or one or more tiles, each tile may, for example, be rectangular in shape and contain one or more blocks (e.g., CTUs), for example, complete or partial blocks.

[0143] Other variations of the video decoder 30 can be used to decode encoded picture data 21. For example, the decoder 30 can generate an output video stream without a loop filter unit 320. For example, a non-transformation-based decoder 30 can directly dequantize the residual signal for a particular block or frame without an inverse transformation unit 312. In other implementations, the video decoder 30 may have an inverse quantization unit 310 and an inverse transformation unit 312 coupled into a single unit.

[0144] It should be understood that in encoder 20 and decoder 30, the processing result of the current step may be further processed and then output to the next step. For example, after interpolation filtering, motion vector derivation, or loop filtering, further operations such as clipping or shifting may be performed on the processing result of interpolation filtering, motion vector derivation, or loop filtering.

[0145] It should be noted that further operations may be applied to the currently derived motion vectors of the block (including, but not limited to, control point motion vectors in affine mode, subblock motion vectors in affine, planar, and ATMVP modes, time motion vectors, etc.). For example, the value of a motion vector is constrained to a predetermined range according to its representation bits. If the representation bits of the motion vector are bitDepth, the range is -2^(bitDepth-1) to 2^(bitDepth-1)-1, where "^" signifies exponential calculation. For example, if bitDepth is set to equal 16, the range is -32768 to 32767, and if bitDepth is set to equal 18, the range is -131072 to 131071. For example, the value of a derived motion vector (e.g., the MV of four 4x4 subblocks in one 8x8 block) is constrained such that the maximum difference between the integer parts of the MVs of the four 4x4 subblocks is less than or equal to N pixels, which is less than or equal to 1 pixel. Here, we provide two methods for constraining the motion vector according to bitDepth.

[0146] Method 1: Remove the most significant bit (MSB) of the overflow by performing the following operation. ux=(mvx+2 bitDepth )%2 bitDepth (1) mvx=(ux>=2 bitDepth-1 )?(ux-2 bitDepth ):ux (2) uy=(mvy+2 bitDepth )%2 bitDepth (3) mvy=(uy>=2 bitDepth-1 )?(uy-2 bitDepth ):uy (4) Here, mvx is the horizontal component of the motion vector of an image block or subblock, mvy is the vertical component of the motion vector of an image block or subblock, and ux and uy represent the intermediate values.

[0147] For example, if the value of mvx is -32769, then after applying equations (1) and (2), the resulting value is 32767. Computer systems store decimal numbers as two's complement. The two's complement of -32769 is 1,0111,1111,1111,1111 (17 bits), in which case the MSB is discarded, so the resulting two's complement is 0111,1111,1111,1111 (decimal number is 32767), which is the same as the output obtained by applying equations (1) and (2). ux=(mvpx+mvdx+2 bitDepth )%2 bitDepth (5) mvx=(ux>=2 bitDepth-1 )?(ux-2 bitDepth ):ux (6) uy=(mvpy+mvdy+2 bitDepth )%2 bitDepth (7) mvy=(uy>=2 bitDepth-1 )?(uy-2 bitDepth ):uy (8) The operation may be applied between the sum of mvp and mvd, as shown in equations (5) to (8).

[0148] Method 2: Remove the overflow MSB by clipping the value. vx=Clip3(-2 bitDepth-1 ,2 bitDepth-1 -1,vx) vy=Clip3(-2 bitDepth-1 ,2 bitDepth-1 -1, vy) Here, vx is the horizontal component of the motion vector of an image block or subblock, vy is the vertical component of the motion vector of an image block or subblock, x, y, and z correspond to the three input values ​​of the MV clipping process, and the definition of the function Clip3 is as follows.

number

[0149] Figure 4 is a schematic diagram of a video coding device 400 according to an embodiment of the present disclosure. The video coding device 400 is suitable for implementing embodiments of the disclosure described herein. In the embodiment, the video coding device 400 may be a decoder such as the video decoder 30 in Figure 1A or an encoder such as the video encoder 20 in Figure 1A.

[0150] The video coding device 400 includes an inlet port 410 (or input port 410) and a receiver unit (Rx) 420 for receiving data, a processor, logic unit, or central processing unit (CPU) 430 for processing data, a transmitter unit (Tx) 440 and an exit port 450 (or output port 450) for transmitting data, and memory 460 for storing data. The video coding device 400 may also include optical-to-electrical (OE) and electrical-to-optical (EO) components coupled to the inlet port 410, receiver unit 420, transmitter unit 440, and exit port 450 for the exit or input of optical or electrical signals.

[0151] Processor 430 is implemented by hardware and software. Processor 430 may be implemented as one or more CPU chips, cores (e.g., multi-core processors), FPGAs, ASICs, and DSPs. Processor 430 communicates with an input port 410, a receiver unit 420, a transmitter unit 440, an output port 450, and a memory 460. Processor 430 includes a coding module 470. Coding module 470 implements the embodiments of the above disclosure. For example, coding module 470 realizes, processes, prepares, or provides various coding operations. Therefore, what is included in coding module 470 provides a substantial improvement to the function of video coding device 400 and brings about the conversion of video coding device 400 to different states. Alternatively, coding module 470 is realized as instructions stored in memory 460 and executed by processor 430.

[0152] Memory 460 may include one or more disks, tape drives, and solid state drives and may be used as an overflow data storage device for storing such programs when selected for execution and for storing instructions and data read during program execution. Memory 460 may be, for example, volatile and / or non-volatile and may be read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random-access memory (SRAM).

[0153] FIG. 5 is a simplified block diagram of an apparatus 500 that may be used as one or both of the source device 12 and the destination device 14 from FIG. 1 according to an exemplary embodiment.

[0154] The processor 502 within the device 500 can be a central processing unit. Alternatively, the processor 502 can be any other type of device or devices capable of manipulating or processing information that currently exists or will be developed in the future. The disclosed implementation can be carried out with a single processor, such as processor 502 as illustrated, but the advantages in terms of speed and efficiency can be achieved using more than one processor.

[0155] The memory 504 within the device 500 can be, in an implementation, a read only memory (ROM) device or a random access memory (RAM) device. Any other suitable type of storage device can be used as the memory 504. The memory 504 can include code and data 506 that are accessed by the processor 502 using the bus 512. The memory 504 can further include an operating system 508 and an application program 510, and the application program 510 can include at least one program that enables the processor 502 to execute the methods described herein. For example, the application program 510 can include applications 1 - N that further include a video coding application that executes the methods described herein.

[0156] The device 500 can also include one or more output devices, such as a display 518. The display 518 can be, in one example, a touch - sensitive display that combines a touch - sensitive element operable to sense touch input with a display. The display 518 can be coupled to the processor 502 via the bus 512.

[0157] Although shown here as a single bus, the bus 512 of device 500 can consist of multiple buses. Furthermore, the secondary storage 514 can be directly coupled to other components of device 500 or accessed via a network, and may include a single integrated unit such as a memory card or multiple units such as multiple memory cards. Thus, device 500 can be implemented in a wide range of configurations.

[0158] The embodiments presented herein will be described in more detail below. A video source represented by a bitstream may include a sequence of pictures in decoding order. Each picture (which may be a source picture or a decoding picture) includes one or more of the following sample sequences. - Luma (Y) only (monochrome) - Luma and two chroma (YCbCr or YCgCo) - Green, blue, and red (also known as GBR or RGB) - An array representing a sample of monochrome or tristimulus color that is not otherwise specified (e.g., also known as YZX, XYZ, etc.)

[0159] For convenience of notation and terminology in this disclosure, the variables and terms related to these sequences are referred to as luma (or L or Y) and chroma, and the two chroma sequences are referred to as Cb and Cr.

[0160] Figure 9A shows the positions of the chroma components in the case of a 4:2:0 sampling scheme. Examples of other sampling schemes are shown in Figures 9B and 9C.

[0161] As shown in Figure 9A, in the 4:2:0 sampling scheme, a shift may exist between the grid of the luminous component and the grid of the chroma component. In a 2x2 pixel block, the chroma component is actually vertically shifted by half a pixel compared to the luminous component (see Figure 9A). Such a shift may affect the interpolation filter when downsampling or upsampling a picture. Figure 9D shows various sampling patterns for interlaced images. This means that parity, i.e., whether a pixel is in the top field or bottom field of the interlaced image, is also taken into consideration.

[0162] According to the Versatile Video Coding (VVC) specification draft, a special flag, "sps_cclm_colocated_chroma_flag," is signaled at the sequence parameter level. A flag equal to 1 for "sps_cclm_colocated_chroma_flag" indicates that the top-left downsampled chroma sample in the cross-component linear model intra-prediction is in the same position as the top-left chroma sample. A flag equal to 0 for "sps_cclm_colocated_chroma_flag" indicates that the top-left downsampled chroma sample in the cross-component linear model intra-prediction coexists horizontally with the top-left chroma sample, but is vertically shifted by 0.5 units relative to the top-left chroma sample.

[0163] Affine motion compensation prediction In the real world, many types of motion exist, such as zooming in / out, rotation, perspective, translation, and other irregular motions. In HEVC (ITU-T H.265), only translational motion models are used for motion compensation prediction (MCP). In VVC, affine transform motion compensation prediction is applied. The affine motion field of a block is described by two or three control point motion vectors (CPMVs) corresponding to the 4-parameter affine motion model and the 6-parameter affine motion model, respectively. The CPMV positions for the 4-parameter affine motion model are shown in Figure 6a, and the CPMV positions for the 6-parameter affine motion model are shown in Figure 6b.

[0164] In the case of a 4-parameter motion model, the motion vector field (MVF) of a block is described by the following equation.

number

[0165] The CPMV can be derived based on the motion information of adjacent blocks (for example, in a subblock merge mode process). Alternatively, or further, the CPMV can be derived by deriving a CPMV predictor (CPMVP) and obtaining the difference between the CPMV and the CPMVP from the bitstream.

[0166] To simplify motion compensation prediction, block-based affine transformation prediction is applied. For example, to derive the motion vector for each 4x4 subblock, as shown in Figure 7, the motion vector of the center sample of each subblock is calculated according to equation (1) above and rounded to a fractional precision of 1 / 16. A motion compensation interpolation filter is applied to generate predictions for each subblock using the derived motion vectors.

[0167] After motion compensation prediction (MCP), the higher-precision motion vectors of each subblock are rounded and stored with 1 / 4 the precision, the same precision as the normal motion vectors.

[0168] Figure 8 shows an example flowchart illustrating process 800 for affine-based interprediction (i.e., motion compensation using an affine motion model). Process 800 may include the following blocks:

[0169] In block 810, the control point motion vector derivation is performed in order to generate the control point motion vector cpMvLX[cpIdx].

[0170] In block 830, motion vector array derivation is performed to generate the luma subblock motion vector array mvLX[xSbIdx][ySbIdx] and the chroma subblock motion vector array mvCLX[xSbIdx][ySbIdx]. Block 830 may also include the following:

[0171] Block 831: Luma subblock. To generate the Luma motion vector array mvLX[xSbIdx][ySbIdx], the Luma motion vector array derivation is performed.

[0172] Block 833: Chroma subblock motion vector array derivation is performed to generate the chroma motion vector array mvCLX[xSbIdx][ySbIdx].

[0173] In block 850, an interpolation process is executed to generate a prediction of each sub-block with the derived motion vectors, i.e., an array of predicted samples predSamples.

[0174] The embodiments presented herein mainly focus on block 833 (this block is shown in bold in FIG. 8) for the derivation of the chroma motion vector array.

[0175] The details of the derivation process for the chroma motion vectors in previous designs (in conventional methods) are described as follows.

[0176] The inputs for this process (derivation of the chroma motion vector array) include the following: - The luma sub-block motion vector array mvLX[xSbIdx][ySbIdx], where xSbIdx = 0..numSbX - 1, ySbIdx = 0..numSbY - 1 and X is 0 or 1 - The horizontal chroma sampling ratio SubWidthC - The vertical chroma sampling ratio SubHeightC Output: - The chroma sub-block motion vector array mvCLX[xSbIdx][ySbIdx], where xSbIdx = 0..numSbX - 1, ySbIdx = 0..numSbY - 1 and X is 0 or 1

[0177] This process is realized as follows. - The average luma motion vector mvAvgLX is derived as follows. mvAvgLX = mvLX[(xSbIdx >> 1 << 1)][(ySbIdx >> 1 << 1)] + mvLX[(xSbIdx >> 1 << 1) + 1][(ySbIdx >> 1 << 1) + 1] (2) mvAvgLX[0] = (mvAvgLX[0] >= 0? (mvAvgLX[0] + 1) >> 1 : -((-mvAvgLX[0] + 1) >> 1)) (3) mvAvgLX[1]=(mvAvgLX[1]>=0?(mvAvgLX[1]+1)>>1:-((-mvAvgLX[1]+1)>>1)) (4) - Scale mvAvgLX according to the reference index value refIdxLX. Specifically, - If the reference picture corresponding to refIdxLX for the current coding unit is not the current picture, then the following applies: mvCLX[0]=mvLX[0]*2 / SubWidthC mvCLX[1]=mvLX[1]*2 / SubHeightC - Otherwise (if the reference picture corresponding to refIdxLX for the current coding unit is the current picture), then the following applies: mvCLX[0]=((mvLX[0]>>(3+SubWidthC))*32 mvCLX[1]=((mvLX[1]>>(3+SubHeightC))*32

[0178] In the above design, the calculation of mvAvgLX does not take chroma subsampling into account. This leads to an inaccurate estimate of the motion field when one of the variables SubWidthC or SubHeightC is equal to 1.

[0179] Embodiments of the present invention solve this problem by subsampling the luma motion field based on the chroma format of the picture, thereby improving the accuracy of the chroma motion field. More specifically, embodiments of the present invention disclose a method for considering the chroma format of the picture when obtaining chroma motion vectors from luma motion vectors. Linear subsampling of the luma motion field is performed by averaging the luma motion vectors. Selecting luma motion vectors based on the picture chroma format results in a more accurate chroma motion field for more accurate luma motion field subsampling. This dependence on the chroma format allows for the selection of the optimal luma block when averaging the luma motion vectors. As a result of more accurate motion field interpolation, prediction errors are reduced, which has the technical consequence of improved compression performance.

[0180] In one exemplary implementation, Table 1-1 shows the chroma formats that can be supported in this disclosure. Chroma format information such as chroma_format_idc and / or separate_colour_plane_flag may be used to determine the values ​​of the variables SubWidthC and SubHeightC. [Table 1]

[0181] `chroma_format_idc` specifies the chromasampling method for lumasampling. The value of `chroma_format_idc` must be in the range of 0 to 3.

[0182] A separate_colour_plane_flag equal to 1 specifies that the three color components of the 4:4:4 chroma format are coded separately. A separate_colour_plane_flag equal to 0 specifies that the color components are not coded separately. When separate_colour_plane_flag is not present, it is assumed to be equal to 0. When separate_colour_plane_flag is equal to 1, the coded picture consists of three separate components, each of which consists of a coded sample of one color plane (Y, Cb, or Cr) and uses the monochrome coding syntax.

[0183] The chroma format determines the priority and subsampling of the chroma sequence.

[0184] In monochromatic sampling, there is only one sample sequence that can be nominally considered to be a luma sequence.

[0185] In 4:2:0 sampling, each of the two chroma sequences has half the height and half the width of the luma sequence, as shown in Figure 9A.

[0186] In 4:2:2 sampling, each of the two chroma sequences has the same height and half the width of the luma sequence, as shown in Figure 9B.

[0187] In 4:4:4 sampling, the following applies depending on the value of separate_colour_plane_flag: - If separate_colour_plane_flag is equal to 0, each of the two chroma sequences has the same height and width as the luma sequence, as shown in Figure 9C. - Otherwise (if separate_colour_plane_flag is equal to 1), the three color planes are processed separately as monochrome sampled pictures.

[0188] In other exemplary implementations, Table 1-2 also shows chroma formats that can be supported in this disclosure. Chroma format information such as chroma_format_idc and / or separate_colour_plane_flag may be used to determine the values ​​of the variables SubWidthC and SubHeightC. [Table 2]

[0189] The number of bits required to represent each sample in the luma and chroma sequences in a video sequence is in the range of 8 to 16, and the number of bits used in the luma sequence may differ from the number of bits used in the chroma sequence.

[0190] When the value of chroma_format_idc is equal to 1, the nominal vertical and horizontal relative positions of luma and chroma samples within the picture are shown in Figure 9A. The relative positions of alternative chroma samples may be shown in the video usability information.

[0191] When the value of chroma_format_idc is equal to 2, the chroma sample coexists with the corresponding luma sample, and its nominal position within the picture is as shown in Figure 9B.

[0192] When the value of chroma_format_idc is equal to 3, all sequence samples coexist in all cases of the picture, and their nominal positions within the picture are as shown in Figure 9C.

[0193] In one exemplary implementation, the variables SubWidthC and SubHeightC are specified in Table 1-1 or Table 1-2, depending on the chroma format sampling structure specified via chroma_format_idc and separate_colour_plane_flag. It can be understood that chroma format information, such as the chroma format sampling structure, is specified via chroma_format_idc and separate_colour_plane_flag.

[0194] Unlike previous designs, in this disclosure, the derivation of the position within the chroma subblock motion vector array may be applicable to different chroma formats and depends on the values ​​of the chroma scaling factors (e.g., SubWidthC and SubHeightC). It should be understood that the “horizontal and vertical chroma scaling factors” can also be called the “horizontal and vertical chroma sampling ratios.”

[0195] Alternatively, in other exemplary implementations, SubWidthC and SubHeightC are given by SubWidthC = (1 + log2(w luma )-log2(w chroma )) and SubHeightC=(1+log2(h luma )-log2(h chroma It may also be defined as )), w luma and h luma These are the width and height of the luma array, respectively, w chroma and h chroma These represent the width and height of the chroma array, respectively.

[0196] In possible implementations of some embodiments of the present disclosure, for a given chroma format, the process of determining the position or index in the chroma motion vector array for a given index of a chroma subblock at the same position may be carried out as follows:

[0197] First, the values ​​of SubWidthC and SubHeightC are determined based on the chroma format of the picture (or frame) currently being coded or decoded.

[0198] Next, for each chroma space position specified by indices xSbIdx=0..numSbX-1, ySbIdx=0..numSbY-1, and where X is 0 or 1, the value of the chroma motion vector is stored as the corresponding mvCLX element. The steps for determining the chroma motion vector are as follows:

[0199] The first step is to perform rounding to determine the x and y indices of the same Luma subblock. xSbIdx L =(xSbIdx>>(SubWidthC-1))<<(SubWidthC-1); ySbIdx L =(ySbIdx>>(SubHeightC-1))<<(SubHeightC-1)

[0200] The second step is to determine an additional set of chroma subblock positions when determining the chroma motion vector. A possible example of defining such a set S may be written as follows: S0=(xSbIdx L ,ySbIdx L ) S1=(xSbIdx L +(SubWidthC-1),ySbIdx L +(SubHeightC-1))

[0201] The third step is to calculate the mean vector mvAvgLX.

[0202] When a set S contains N elements and N is a power of 2, in one exemplary realization, the motion vector mvAvgLX is determined as follows: - mvAvgLX=Σ i mvLX[Si x ][S i y ] - mvAvgLX[0]=(mvAvgLX[0]+N>>1)>>log2(N) - mvAvgLX[1]=(mvAvgLX[1]+N>>1)>>log2(N) Here, S i x and S i y is position S i These are the x and y coordinates.

[0203] In short, in one exemplary implementation, the determination of the mean vector mvAvgLX for averaging can be formulated as follows: mvAvgLX=mvLX[(xSbIdx>>(SubWidthC-1)<<(SubWidthC-1))] [(ySbIdx>>(SubHeightC-1)<<(SubHeightC-1))]+ mvLX[(xSbIdx>>(SubWidthC-1)<<(SubWidthC-1))+(SubWidthC-1)] [(ySbIdx>>(SubHeightC-1)<<(SubHeightC-1))+(SubHeightC-1)] (Formula 1) - mvAvgLX[0]=(mvAvgLX[0]+N>>1)>>log2(N) (Formula 2) - mvAvgLX[1]=(mvAvgLX[1]+N>>1)>>log2(N) (Equation 3)

[0204] The methods for averaging Luma motion vectors presented herein are not limited to those described above, and it should be noted that the averaging function in this disclosure can be implemented in different ways.

[0205] Although the above describes the process as a three-step process, it should be understood that the determination of the mean vector mvAvgLX formulated above in equations 1-3 can be performed in any order.

[0206] In other exemplary implementations, the third step can also be implemented as follows: - mvAvgLX=Σ i mvLX[S i x ][S i y ] - mvAvgLX[0]=(mvAvgLX[0]>=0?(mvAvgLX[0]+N>>1)>>log2(N): -((-mvAvgLX[0]+N>>1)>>log2(N)) (5) - mvAvgLX[1]=(mvAvgLX[1]>=0?(mvAvgLX[1]+N>>1)>>log2(N): -((-mvAvgLX[1]+N>>1)>>log2(N))) (6) Here, S i x and S i y is position S i These are the x and y coordinates.

[0207] The methods for averaging Luma motion vectors presented herein are not limited to those described above, and it should be noted that the averaging function in this disclosure can be implemented in different ways.

[0208] The next step is to scale mvAvgLX according to the reference index value refIdxLX. In some examples, the scaling process is carried out by replacing mvLX with mvAvgLX as follows (i.e., mvLX[0] is replaced with mvAvgLX[0], and mvLX[1] is replaced with mvAvgLX[1]): - If the reference picture corresponding to refIdxLX for the current coding unit is not the current picture, then the following applies: mvCLX[0]=mvLX[0]*2 / SubWidthC mvCLX[1]=mvLX[1]*2 / SubHeightC - Otherwise (if the reference picture corresponding to refIdxLX for the current coding unit is the current picture), then the following applies: mvCLX[0]=((mvLX[0]>>(3+SubWidthC))*32 mvCLX[1]=((mvLX[1]>>(3+SubHeightC))*32

[0209] Similarly, the derivation process for chroma motion vectors in Section 8.5.2.13, described below, can be invoked with mvAvgLX and refIdxLX as inputs and the chroma motion vector array mvCLXSub[xCSbIdx][yCSbIdx] as output. In the process described in Section 8.5.2.13, mvLX is replaced with mvAvgLX, in particular mvLX[0] is replaced with mvAvgLX[0] and mvLX[1] is replaced with mvAvgLX[1].

[0210] Details of possible implementations of the mean vector mvAvgLX calculation in the derivation process for the chroma motion vector of the proposed method are described below in the format of the revised VVC draft specification. Multiple variations of the process exist.

[0211] 1. One of the variations of the mean vector mvAvgLX calculation in the derivation process for the chroma motion vector of the proposed method may be described in the format of the modification of the VVC draft specification as follows: ... [Table 3]

[0212] Note: The above formula illustrates an example of selecting a Luma motion vector for the calculation of the average motion vector (e.g., the selection of Luma subblock positions for a given chroma subblock position). The selected Luma subblocks (and therefore each of their positions) are represented by their respective subblock indices in the horizontal and vertical directions. For example, as above, for a given chroma subblock (xSbIdx, ySbIdx), if xSbIdx and ySbIdx are the subblock indices of the horizontal and vertical chroma subblocks, respectively, then two Luma subblocks (and therefore each of their positions, e.g., their respective subblock indices) can be selected. One of the two selected Luma subblocks can be represented by its horizontal subblock index as [(xSbIdx>>(SubWidthC-1)<<(SubWidthC-1))] and its vertical subblock index as [(ySbIdx>>(SubHeightC-1)<<(SubHeightC-1))]. Other selected luma subblocks can be represented by a horizontal subblock index as [(xSbIdx>>(SubWidthC-1)<<(SubWidthC-1))+(1>>(2-SubWidthC))] and a vertical subblock index as [(ySbIdx>>(SubHeightC-1)<<(SubHeightC-1))+(1>>(2-SubHeightC))]). Therefore, the selection of luma blocks for averaging luma motion vectors depends on the picture chroma format. In particular, the selection of luma blocks for averaging luma motion vectors depends on the chroma scaling factors SubWidthC and SubHeightC, which are determined based on the picture chroma format.

[0213] The mvAvgLX obtained above can be further processed as follows. mvAvgLX[0]=(mvAvgLX[0]>=0?(mvAvgLX[0]+1)>>1:-((-mvAvgLX[0]+1)>>1)) mvAvgLX[1]=(mvAvgLX[1]>=0?(mvAvgLX[1]+1)>>1:-((-mvAvgLX[1]+1)>>1))

[0214] The methods for averaging Luma motion vectors presented herein are not limited to those described above, and it should be noted that the averaging function in this disclosure can be implemented in different ways. - The derivation process for the chroma motion vector in Section 8.5.2.13, which will be presented later in this disclosure, is invoked with mvAvgLX and refIdxLX as inputs and the chroma motion vector mvCLX[xSbIdx][ySbIdx] as output. ...

[0215] 2. Other variations of the calculation of the mean vector mvAvgLX in the derivation process for the chroma motion vector of the proposed method may be described in the format of the modifications to the VVC draft specification as follows: ... [Table 4]

[0216] / / Note: The above formula illustrates an example of selecting a Luma motion vector for the calculation of the average motion vector (e.g., the selection of Luma subblock positions for a given chroma subblock position). The selected Luma subblocks (and therefore each of their positions) are represented by their respective subblock indices in the horizontal and vertical directions. For example, as above, for a given chroma subblock (xSbIdx, ySbIdx), if xSbIdx and ySbIdx are the subblock indices of the horizontal and vertical chroma subblocks, respectively, then two Luma subblocks (and therefore each of their positions, e.g., their respective subblock indices) can be selected. One of the two selected Luma subblocks can be represented by its horizontal subblock index as [(xSbIdx>>(SubWidthC-1)<<(SubWidthC-1))] and its vertical subblock index as [(ySbIdx>>(SubHeightC-1)<<(SubHeightC-1))]. Other selected luma subblocks can be represented by a horizontal subblock index as [(xSbIdx>>(SubWidthC-1)<<(SubWidthC-1))+(1>>(4-SubWidthC-SubHeightC))] and a vertical subblock index as [(ySbIdx>>(SubHeightC-1)<<(SubHeightC-1))+(1>>(4-SubWidthC-SubHeightC))]). Therefore, the selection of luma blocks for averaging luma motion vectors depends on the picture chroma format. In particular, the selection of luma blocks for averaging luma motion vectors depends on the chroma scaling factors SubWidthC and SubHeightC, which are determined based on the picture chroma format.

[0217] Compared to the first variation, this variation uses a different method for determining the set of luma subblocks. In particular, 1>>(4-SubWidthC-SubHeightC) is used in this variation to determine the index of the second luma subblock (the luma subblock adjacent to the first luma subblock), whereas 1>>(2-SubWidthC) and 1>>(2-SubHeightC) are used in the first variation. In the first variation, the first luma subblock itself, its diagonal adjacents, or its horizontal adjacents may be used as the second subblock. In the second variation, either the first luma subblock itself or its diagonal adjacents may be used as the second luma subblock, depending on the value of the chroma scaling factor.

[0218] The mvAvgLX obtained above can be further processed as follows. mvAvgLX[0]=(mvAvgLX[0]>=0?(mvAvgLX[0]+1)>>1:-((-mvAvgLX[0]+1)>>1)) mvAvgLX[1]=(mvAvgLX[1]>=0?(mvAvgLX[1]+1)>>1:-((-mvAvgLX[1]+1)>>1)) - The derivation process for the chroma motion vector in Section 8.5.2.13 is invoked with mvAvgLX and refIdxLX as inputs and the chroma motion vector mvCLX[xSbIdx][ySbIdx] as output. ...

[0219] The methods for averaging Luma motion vectors presented herein are not limited to those described above, and it should be noted that the averaging function in this disclosure can be implemented in different ways.

[0220] 3. Other variations of the derivation process for chroma motion vectors of the proposed method may be described in the format of the VVC draft specification modification as follows: ... [Table 5]

[0221] - The derivation process for the chroma motion vector in Section 8.5.2.13 is invoked with mvAvgLX and refIdxLX as inputs and the chroma motion vector mvCLX[xSbIdx][ySbIdx] as output. ...

[0222] 4. Other variations of the derivation process for chroma motion vectors of the proposed method may be described in the format of the VVC draft specification modification as follows: ... [Table 6]

[0223] / / Note: The above formula illustrates an example of selecting a Luma motion vector for the calculation of the average motion vector (e.g., the selection of Luma subblock positions for a given chroma subblock position). The selected Luma subblocks (and therefore each of their positions) are represented by their respective subblock indices in the horizontal and vertical directions. For example, as above, for a given chroma subblock (xSbIdx, ySbIdx), if xSbIdx and ySbIdx are the subblock indices of the horizontal and vertical chroma subblocks, respectively, then two Luma subblocks (and therefore each of their positions, e.g., their respective subblock indices) can be selected. One of the two selected Luma subblocks can be represented by its horizontal subblock index as [(xSbIdx>>(SubWidthC-1)<<(SubWidthC-1))] and its vertical subblock index as [(ySbIdx>>(SubHeightC-1)<<(SubHeightC-1))]. Other selected luma subblocks can be represented by a horizontal subblock index as [(xSbIdx>>(SubWidthC-1)<<(SubWidthC-1))+(SubWidthC-1)] and a vertical subblock index as [(ySbIdx>>(SubHeightC-1)<<(SubHeightC-1))+(SubHeightC-1)]). Therefore, the selection of luma blocks for averaging luma motion vectors depends on the picture chroma format. In particular, the selection of luma blocks for averaging luma motion vectors depends on the chroma scaling factors SubWidthC and SubHeightC, which are determined based on the picture chroma format.

[0224] The mvAvgLX obtained above can be further processed as follows. mvAvgLX[0]=mvAvgLX[0]>=0?mvAvgLX[0]>>1:-((-mvAvgLX[0])>>1) mvAvgLX[1]=mvAvgLX[1]>=0?mvAvgLX[1]>>1:-((-mvAvgLX[1])>>1)

[0225] The methods for averaging Luma motion vectors presented herein are not limited to those described above, and it should be noted that the averaging function in this disclosure can be implemented in different ways.

[0226] Further details regarding the determination of chroma subblock positions when determining chroma motion vectors for different chroma formats will be explained in conjunction with Figures 10A-10C and 11A-11C below.

[0227] Figure 10A shows an example of luma and chroma blocks at the same position contained within the current image block (e.g., coding block) of the current picture, where the chroma format of the current picture is 4:2:0. As shown in Figure 10A and Table 1-1, when the chroma format of the current picture is 4:2:0, SubWidthC=2 and SubHeightC=2. If the width of a luma block is W and the height of a luma block is H, then the width of the corresponding chroma block is W / SubWidthC and the height of the corresponding chroma block is H / SubHeightC. Specifically, for a current image block containing luma and chroma blocks at the same position, the luma block generally contains four times the number of samples as the corresponding chroma block.

[0228] Figure 10B is an example showing luma and chroma blocks at the same position within the current image block of the current picture, where the chroma format of the current picture is 4:2:2. As shown in Figure 10B and Table 1-1 or 1-2, when the chroma format of the current picture is 4:2:2, SubWidthC=2 and SubHeightC=1. If the width of a luma block is W and the height of a luma block is H, then the width of the corresponding chroma block is W / SubWidthC and the height of the corresponding chroma block is H / SubHeightC. Specifically, for a current image block containing luma and chroma blocks at the same position, the luma block generally contains twice the number of samples as the corresponding chroma block.

[0229] Figure 10C is an example showing luminous and chroma blocks at the same position within the current image block of the current picture, where the chroma format of the current picture is 4:4:4. As shown in Figure 10C and Table 1-1 or 1-2, when the chroma format of the current picture is 4:4:4, SubWidthC=1 and SubHeightC=1. If the width of a luminous block is W and the height of a luminous block is H, then the width of the corresponding chroma block is W / SubWidthC and the height of the corresponding chroma block is H / SubHeightC. Specifically, for a current image block containing luminous and chroma blocks at the same position, the luminous block generally contains the same number of samples as the corresponding chroma block.

[0230] Figure 11A is an example showing the positions of two luma subblocks with respect to a given position of a luma subblock during the chroma motion vector derivation from the luma motion vector, when the current picture chroma format is 4:2:0, as shown in Figure 10A.

[0231] The x and y indices of luma subblocks at the same location are determined using the corresponding x and y indices of the chroma subblock (denoted as xSbIdx and ySbIdx). xSbIdx L=(xSbIdx>>(SubWidthC-1))<<(SubWidthC-1); ySbIdx L =(ySbIdx>>(SubHeightC-1))<<(SubHeightC-1)

[0232] Two Affin Ruma subblocks are selected for further averaging of these motion vectors. The positions of these two subblocks are defined as follows: - (xSbIdxL,ySbIdxL) and - (xSbIdxL+(SubWidthC-1),ySbIdxL+(SubHeightC-1))

[0233] As shown in Figure 11A, in the YUV4:2:0 format, the motion vectors of two luma subblocks of diagonally opposite luma blocks (1010 (8x8 luma size)) are used for averaging, and the averaged MV is used in the affine subblock motion vector derivation process for the chroma subblock. In particular, Luma subblock or Chroma subblock index 0: xSbIdx=0, ySbIdx=0 Luma subblock or Chroma subblock index 1: xSbIdx=1, ySbIdx=0 Luma subblock or Chroma subblock index 2: xSbIdx=0, ySbIdx=1 Luma subblock or Chroma subblock index 3: xSbIdx=1, ySbIdx=1

[0234] According to the design of Modification 4, the motion vector of each chroma subblock is derived based on the mean value, which is obtained based on the motion vectors of diagonal luma subblock 0 (mvLX[0][0]) and luma subblock 3 (mvLX[1][1]).

[0235] Figure 11B is an example showing the positions of two luma subblocks relative to a given position of a chroma subblock during chroma motion vector derivation from luma motion vectors, when the current picture chroma format is 4:2:2, as shown in Figure 10B. As shown in Figure 11B, in the case of the YUV4:2:2 format, the motion vectors of two horizontally adjacent luma subblocks of luma block 1010 (8x8 luma size) are used for averaging, and the averaged MV is used in the affine subblock motion vector derivation process for the chroma subblock. In particular, Luma subblock or Chroma subblock index 0: xSbIdx=0, ySbIdx=0 Luma subblock or Chroma subblock index 1: xSbIdx=1, ySbIdx=0 Luma subblock or Chroma subblock index 2: xSbIdx=0, ySbIdx=1 Luma subblock or Chroma subblock index 3: xSbIdx=1, ySbIdx=1

[0236] According to the design of Variation 4 described above, the motion vector of each chroma subblock on the first row of chroma block 920 is derived based on the mean value, which is obtained based on the motion vectors of the horizontally adjacent luma subblock 0 (mvLX[0][0]) and luma subblock 1 (mvLX[1][0]). The motion vector of each chroma subblock on the second row of chroma block 920 is derived based on the mean value, which is obtained based on the motion vectors of the horizontally adjacent luma subblock 2 (mvLX[0][1]) and luma subblock 3 (mvLX[1][1]).

[0237] Figure 11C is an example showing the position of a chroma subblock relative to a given position of the chroma subblock during the chroma motion vector derivation from the chroma motion vector, when the picture's chroma format is currently 4:4:4, as shown in Figure 10C.

[0238] As shown in Figure 11C, in the YUV4:4:4 format, the motion vectors of luma subblocks at the same position in luma block 1010 (8x8 luma size) are used for each chroma subblock to perform affine prediction; that is, the affine subblock motion vector derivation process for chroma is the same as for luma.

[0239] Averaging is not required; that is, the motion vector can be determined by using the motion vectors of luma subblocks at the same position, or it should be understood that this averaging operation takes the same MV twice as input and produces the same motion vector as output. The luma or chroma subblock size may be 4x4.

[0240] Figure 12A shows several examples of subset S containing the positions of luma subblocks relative to a given position of chroma subblock when the chroma format is set to 4:4:4. In this example, four cases of subset S derivation are considered.

[0241] In the first case, chroma position "A" (1201) has a corresponding adjacent chroma block located at position "A" (1202).

[0242] In the second case, chroma position "B" (1203) is selected on the lower boundary of the chroma block. In this case (except for the lower right position), the corresponding position 1204 of the luma subblock belonging to S is selected to be horizontally adjacent.

[0243] In the third case, chroma position "C" (1205) is selected on the right-hand boundary of the chroma block. In this case (excluding the lower-right position), the corresponding position 1206 of the luma subblock belonging to S is selected to be vertically adjacent.

[0244] In the fourth case, chroma position "D" (1207) is selected at the lower right corner of the chroma block. In this case, set S contains a single luma subblock located at the lower right corner of the luma block.

[0245] Figure 12B shows another embodiment in which subset S is obtained. In this embodiment, chroma blocks located on the boundary of a chroma block have corresponding positions within the chroma block, as in the fourth case "D" shown in Figure 12A.

[0246] In the embodiments of the present disclosure described above, when chroma subsampling is used, the number of chroma subblocks is the same as the number of luma subblocks at the same location. In particular, the number of luma subblocks in the horizontal direction, numSbX, is the same as the number of chroma subblocks in the horizontal direction, numSbX, and the number of luma subblocks in the vertical direction, numSbY, is the same as the number of chroma subblocks in the vertical direction, numSbY. Therefore, when chroma subsampling is used, the size of the chroma subblocks is different from the size of the luma subblocks at the same location.

[0247] In other scenarios, the sizes of chroma subblocks and luma subblocks are kept the same regardless of the chroma format. In these scenarios, the number of chroma subblocks in a chroma block may be different from the number of luma subblocks in a luma block at the same location. The following embodiments concern the derivation of chroma motion vectors for subblocks of the same size for chroma and luma components. That is, for a current picture containing a current image block containing luma and chroma blocks at the same location, the luma block of the current picture contains a set of luma subblocks of equal size, and the chroma block of the current picture contains a set of chroma subblocks of equal size, with the size of the chroma subblocks set to be equal to the size of the luma subblocks. When chroma subsampling is used, it can be understood that the number of chroma subblocks will be different from the number of luma subblocks at the same location. In particular, the number of luma subblocks along the horizontal direction, numSbX, differs from the number of chroma subblocks along the horizontal direction, numSbX, and the number of luma subblocks along the vertical direction, numSbY, differs from the number of chroma subblocks along the vertical direction, numSbY.

[0248] As shown in Figure 13A, in the YUV4:2:0 format, for example, the current picture's luma block has a size of 8x8, and the luma block contains four luma subblocks of equal size, and the current picture's chroma block contains a set of chroma subblocks of equal size (chroma subblock = chroma block), and the size of the chroma subblock is set to be equal to the size of the luma subblock. The motion vectors of two diagonally opposite luma subblocks are averaged, and the averaged MV is used in the affine subblock motion vector derivation process for the chroma subblock. In particular, since the number of subblocks is different, xSbIdx (subblock index of the horizontal luma subblock) changes with the step size of SubWidth, and ySbIdx (subblock index of the vertical luma subblock) changes with the step size of SubHeightC. For example, for the 4:2:0 format, xSbIdx = 0, 2, 4, 6... and ySbIdx = 0, 2, 4, 6...

[0249] As shown in Figure 13B, in the case of the YUV4:2:2 format, the motion vectors of two horizontally adjacent chroma subblocks are used to average out according to the modified equation 4 above in order to generate the motion vector for the chroma subblock. The averaged MV is used in the process of deriving the affine subblock motion vector for the chroma subblock. In particular, since the number of subblocks is different, xSbIdx changes with the step size of SubWidth and ySbIdx changes with the step size of SubHeightC. For example, for 4:2:2, xSbIdx = 0, 2, 4, 8... and ySbIdx = 0, 1, 2, 3...

[0250] As shown in Figure 13C, in the YUV4:4:4 format, the number and size of luma subblocks are equal for both chroma subblocks. In this case, for each chroma subblock, the motion vector of the luma subblock at the same position is used to perform affine prediction. In other words, the process for deriving the affine subblock motion vector for chroma blocks is the same as that for luma blocks. For example, for the 4:4:4 format, xSbIdx = 0, 1, 2, 3... and ySbIdx = 0, 1, 2, 3... In this case, it should be noted that the averaging operation specified in the above transformation 4 does not need to be performed because the two subblocks used for averaging are the same and the averaging operation outputs the same value as the input. Therefore, in this case, the motion vector of the luma subblock can be selected for the chroma subblock without going through the averaging operation formulated in any of the previous transformations.

[0251] From the above, it can be seen that when SubWidthC is greater than 1 or SubHeightC is greater than 1, the block has a different number of chroma subblocks than the number of luma subblocks. Figure 14A shows an example of subdivision of a 16x16 luma block and subdivision of chroma blocks at the same position for the YUV4:2:0 chroma format. In this example, the luma block is divided into 16 subblocks, each having a size of 4x4. The chroma block has a size of 8x8 samples and is subdivided into a total of 4 subblocks, each having a size of 4x4 samples. These 4 chroma subblocks are grouped into 2 rows, each having 2 subblocks. The letters "A", "B", "C", and "D" indicate which luma subblock is used to derive the motion vector for the corresponding chroma subblock indicated by the same letter.

[0252] Figure 14B shows a 16x16 chroma block and its subdivision into chroma blocks at the same location within a picture having the YUV4:2:2 chroma format. In this case, the chroma block has a size of 8x16 samples and is subdivided into a total of 8 subblocks, grouped into 4 rows, each row having 2 subblocks. Each chroma subblock and chroma subblock has a size of 4x4 samples. The letters "A", "B", "C", "D", "E", "F", "G", and "H" indicate which chroma subblock is used to derive the motion vector of the corresponding chroma subblock indicated by the same letter.

[0253] Assuming that a Luma block is subdivided into numSbY row subblocks, each row having numSbX subblocks, and that motion vectors are specified or obtained for each Luma subblock, the embodiments presented herein may be specified as follows:

[0254] 1. The first step is to determine the values ​​of SubWidthC and SubHeightC based on chroma format information indicating the chroma format of the picture (or frame) currently being coded or decoded. For example, the chroma format information may include the information presented in Table 1-1 or Table 1-2 above.

[0255] 2. The second step may include obtaining the number of chroma subblocks along the horizontal direction, numCSbX, and the number of chroma subblocks along the vertical direction, numCSbY, as follows: numCSbX = numSbX >> (SubWidthC-1), where numSbX is the number of luma subblocks within the luma block along the horizontal direction; numCSbY = numSbY >> (SubHeightC-1), where numSbY is the number of Luma subblocks within the Luma block along the vertical direction.

[0256] The Luma block can be the currently encoded or decoded block of the picture currently being coded or decoded.

[0257] 3. Using the spatial index (xCSbIdx, yCSbIdx), we can identify the chroma subblock located at row yCSbIdx and column xCSbIdx, where xCSbIdx = 0, ..., numCSbX-1 and yCSbIdx = 0, numCSbY-1. The value of the chroma motion vector for the chroma subblock can be determined as follows.

[0258] Spatial index of Luma subblocks at the same location (xSbIdx L, ySbIdx L ) is determined as follows: xSbIdx L =xCSbIdx<<(SubWidthC-1); ySbIdx L =yCSbIdx<<(SubHeightC-1)

[0259] The spatial position (sbX, sbY) of a chroma subblock within a chroma block may be derived using the spatial index (xCSbIdx, yCSbIdx) as follows: sbX = xCSbIdx * sbX; sbY = yCSbIdx * sbY

[0260] The same applies to determining the spatial position of a luma subblock within the luma subblock's luma block spatial index (xSbIdx, ySbIdx). sbX = xSbIdx * sbX; sbY = ySbIdx * sbY

[0261] Luma subblock (xSbIdx L ,ySbIdx L The determined spatial index of ) can be further used in determining the chroma motion vector. For example, a set of chroma subblocks can be defined as follows: S0=(xSbIdx L ,ySbIdx L ) S1=(xSbIdx L +(SubWidthC-1),ySbIdx L +(SubHeightC-1))

[0262] In this example, the set of luma sub-blocks includes two sub-blocks indexed by S0 and S1 calculated above. Each of S0 and S1 includes a pair of spatial indices that define the sub-block positions.

[0263] The set of luma sub-blocks is used to calculate the average motion vector mvAvgLX. Here, in the motion vector notation used below, X can be either 0 or 1, and correspondingly, it indicates that the reference list index for the motion vector is either L0 or L1. L0 indicates reference list 0 and L1 indicates reference list 1. It is assumed that the calculation of the motion vector is performed by independently applying the corresponding formula to the horizontal component mvAvgLX[0] and the vertical component mvAvgLX[1] of the motion vector.

[0264] If the luma motion vector of the sub-block having the spatial index (xSbIdx L ,ySbIdx L ) is shown as mvLX[xSbIdx L [ySbIdx L , the average motion vector mvAvgLX may be obtained as follows. mvAvgLX=Σ i mvLX[S i x [S i y mvAvgLX=mvAvgLX>=0?mvAvgLX>>1:-((-mvAvgLX)>>1) Here, as described above, S i x and S i y are the elements S i ​These are the horizontal and vertical space indices, where i = 0, 1...

[0265] The motion vector of the chroma subblock mvCLX, which has a spatial index (xCSbIdx, yCSbIdx), is obtained from the average motion vector mvAvgLX as follows: mvCLX[0]=mvAvgLX[0]*2 / SubWidthC mvCLX[1]=mvAvgLX[1]*2 / SubHeightC

[0266] Details of the derivation process for luma and chroma motion vectors according to the embodiments presented herein are described in part of the VVC draft specification as follows: [Table 7] JPEG2026077838000011.jpg127170

[0267] It should be noted that this embodiment differs from previous embodiments. In previous embodiments, the number of luma subblocks within a luma block is the same as the number of chroma subblocks within a chroma block at the same location. However, in this embodiment, for chroma formats 4:2:0 and 4:2:2, the number of luma subblocks within a luma block is different from the number of chroma subblocks within a chroma block at the same location. Since the number of subblocks differs within the chroma block and luma block, as can be seen below, xSbIdx changes with the step size of SubWidth, and ySbIdx changes with the step size of SubHeightC. [Table 8]

[0268] It should be noted that the averaging operation presented above is for illustrative purposes only and should not be interpreted as limiting. Various other methods can be used to perform the averaging operation in determining the chroma motion vector from the luma motion vector. [Table 9]

[0269] The behavior of this derivation process depends on how it is invoked. For example, it was previously stated that "the derivation process for chroma motion vectors in Section 8.5.2.13 is invoked with mvAvgLX and refIdxLX as inputs and the chroma motion vector array mvCLXSub[xCSbIdx][yCSbIdx] as output." In this example, the mvLX in Section 8.5.2.13 described therein is replaced with mvAvgLX to perform the operation when invoked.

[0270] There are further embodiments relating to a configuration that takes into account the offset between the subsampling positions of the chroma sample and the position of the luma sample.

[0271] An exemplary embodiment is to define or determine a set of chroma subblocks S according to the value of "sps_cclm_colocated_chroma_flag". Specifically, - When SubHeightC=1 and SubWidthC=2, and sps_cclm_colocated_chroma_flag is set to equal 1, set S is a single element S0=(xSbIdx L ,ySbIdx L It consists of ). - Otherwise, set S includes the following: S0=(xSbIdx L ,ySbIdx L ) S1=(xSbIdx L +(SubWidthC-1),ySbIdx L+(SubHeightC-1))

[0272] Other exemplary embodiments introduce a dependency between the determination of the mean motion vector and the value of "sps_cclm_colocated_chroma_flag". In particular, weights may be introduced into the averaging operation and specified differently for different chroma subblocks. An exemplary process for deriving the mean motion vector is as follows: - xSbIdx L =xCSbIdx<<(SubWidthC-1) - ySbIdx L =yCSbIdx<<(SubHeightC-1) - sps_cclm_colocated_chroma_flag is set to equal to 1, and the weight coefficients w0 and w1 are set as follows, i.e., w0=5, w1=3. mvAvgLX=w0*mvLX[xSbIdx L ][ySbIdx L ]+ +w1*mvLX[xSbIdx L +(SubWidthC-1)][ySbIdx L +(SubHeightC-1)] mvAvgLX[0]=mvAvgLX[0]>=0?(mvAvgLX[0]+3)>>3:-((-mvAvgLX[0]+3)>>3) mvAvgLX[1]=mvAvgLX[1]>=0?(mvAvgLX[1]+3)>>3:-((-mvAvgLX[1]+3)>>3)

[0273] It should be noted that the averaging operation presented in this example is for illustrative purposes only and should not be interpreted as limiting. Various other methods can be used to perform the averaging operation in determining the chroma motion vector from the luma motion vector.

[0274] Figure 15 is a flowchart of an exemplary method 1300 for affine-based interpretation of chroma subblocks, which includes the following:

[0275] In step 1501, the horizontal and vertical chroma scaling factors are determined based on the chroma format information, which indicates the chroma format of the current picture to which the current image block belongs.

[0276] In step 1503, the set of luma subblocks (S) of the luma block is determined based on the value of the chroma scaling factor.

[0277] In step 1505, the motion vector for the chroma subblock of the chroma block is determined based on the motion vectors of one or more luma subblocks in the set of luma subblocks (S).

[0278] Figure 16 is a flowchart of another exemplary method 1300 for affine-based interpretation of chroma subblocks, which includes the following:

[0279] In step 1601, the horizontal and vertical chroma scaling factors are determined based on the chroma format information, which indicates the chroma format of the current picture to which the current image block belongs.

[0280] In step 1603, the values ​​of the motion vectors for each luma subblock within the multiple luma subblocks are determined, and N luma subblocks are included in the luma block.

[0281] In step 1605, the motion vectors of the luma subblocks within the set S of luma subblocks are averaged, and the set (S) is determined based on the chroma scaling factor.

[0282] In step 1607, for chroma subblocks within multiple chroma subblocks, motion vectors for the chroma subblocks are derived based on the average chroma motion vector, and the chroma subblocks are included in the chroma block.

[0283] This invention discloses a method for considering the chroma format of a picture when obtaining chroma motion vectors from chroma motion vectors. Linear subsampling of the chroma motion field is performed by averaging the chroma motion vectors. It is found that selecting motion vectors from horizontally adjacent chroma blocks is more appropriate when the chroma color plane is at the same height as the chroma plane. The selection of chroma motion vectors that depends on the picture chroma format results in a more accurate chroma motion field due to more accurate chroma motion vector field subsampling. This dependency on the chroma format allows for the selection of the optimal chroma block position when averaging the chroma motion vectors. As a result of more accurate motion field interpolation, prediction errors are reduced, which has technical consequences in improving compression performance.

[0284] Furthermore, when the number of chroma subblocks is defined as equal to the number of luma subblocks, and the chroma color plane size is not equal to the luma plane size, the motion vectors of adjacent chroma subblocks may be assumed to be the same. When implementing this processing step, optimization may be performed by skipping the iterative value calculation step. The proposed invention discloses a method for defining the size of chroma subblocks equal to the size of luma subblocks. In this case, the implementation may be simplified by unifying the luma and chroma processing, and redundant motion vector calculations are naturally avoided.

[0285] Figure 17 shows a device for affine-based interprediction according to another aspect of the present invention. The device 1700 is A determination module 1701 is configured to determine horizontal and vertical chroma scaling factors based on chroma format information, wherein the chroma format information indicates the chroma format of the current picture to which the current image block belongs, and the set of luma subblocks (S) of the luma block is determined based on the value of the chroma scaling factor. A motion vector derivation module 1703 is configured to determine the motion vector for a chroma subblock of a chroma block based on the motion vectors of one or more luma subblocks in a set (S) of luma subblocks. Includes.

[0286] In one example, the motion vector derivation module 1703 is: A Luma motion vector derivation module 1703a configured to determine the motion vector values ​​for each Luma subblock within multiple Luma subblocks, wherein multiple Luma subblocks are included in a Luma block, A chroma motion vector derivation module 1703b is configured to determine the motion vector for a chroma subblock based on the motion vector of at least one luma subblock in a set (S) of luma subblocks, where the set (S) is determined based on a chroma scaling factor, and the chroma subblocks are contained within a chroma block. It may include. In designs of equal size, multiple chroma subblocks may contain only one chroma subblock.

[0287] Device 1700 further includes a motion compensation module 1705 configured to generate predictions of chroma subblocks based on determined motion vectors.

[0288] Correspondingly, in one example, the exemplary structure of device 1700 may correspond to the encoder 200 in Figure 2. In another example, the exemplary structure of device 1700 may correspond to the decoder 300 in Figure 3.

[0289] In other examples, the exemplary structure of device 1700 may correspond to the interpretation unit 244 in Figure 2. In other examples, the exemplary structure of device 1700 may correspond to the interpretation unit 344 in Figure 3.

[0290] This disclosure provides the following further aspects:

[0291] According to a first aspect of the present invention, a method is provided for deriving a chroma motion vector used in affine motion compensation of an interpretation unit PU, wherein the PU includes chroma and chroma blocks at the same location, and the method is as follows: The steps include determining the horizontal and vertical (SubWidthC and SubHeightC) chroma scaling factors based on the chroma format of the current picture (e.g., the picture currently being coded or decoded), Currently, the step is to divide the picture luma block into the first set of luma subblocks, The steps include obtaining the motion vector values ​​for each Luma subblock in the first set of Luma subblocks, Currently, the chroma block of the picture (in one example, the chroma block and luma block are included in the same PU) is divided into a set of chroma subblocks, This step involves determining a second set (S) of luma subblocks for each chroma subblock within a set of chroma subblocks, where the position of the luma subblocks within the second set is determined by the current picture's chroma format. The second step is to derive the motion vector for the chroma subblock based on the motion vector of the luma subblock in the second set S. Includes.

[0292] In a possible implementation of the method according to the first embodiment, the second set (S) of Luma subblocks is the following subblocks, namely, S0=(xSbIdxL,ySbIdxL) S1=(xSbIdxL,ySbIdxL+(SubHeightC-1)) S2=(xSbIdxL+(SubWidthC-1),ySbIdxL) S3=(xSbIdxL+(SubWidthC-1),ySbIdxL+(SubHeightC-1)) Includes any combination of the above.

[0293] In any of the preceding implementations of the first embodiment or possible implementations of the method according to the first embodiment itself, the second set (S) of Luma subblocks consists of two subblocks, i.e., S0=(xSbIdxL,ySbIdxL) S1=(xSbIdxL+(SubWidthC-1),ySbIdxL+(SubHeightC-1)) Includes.

[0294] In any prior implementation of the first embodiment or a possible implementation of the method according to the first embodiment itself, deriving the motion vector for the chroma subblock based on the motion vector of the luma subblock in the second set S includes averaging the motion vectors of the luma subblock in the second set S.

[0295] In any of the preceding implementations of the first embodiment or in possible implementations of the method according to the first embodiment itself, averaging the motion vectors of the luma subblocks in the second set S is done by the following steps, namely, mvAvgLX=Σ i mvLX[S i x ][S i y ] mvAvgLX[0]=(mvAvgLX[0]+N>>1)>>log2(N) mvAvgLX[1]=(mvAvgLX[1]+N>>1)>>log2(N) mvAvgLX is the result of averaging, mvAvgLX[0] is the horizontal component of the motion vector mvAvgLX, and mvAvgLX[1] is the vertical component of the motion vector mvAvgLX, S i x and S i y is the horizontal and vertical index of the subblock Si in the motion vector array, mvLX[S i x ][S iy ] is index S i x and S i y This is the motion vector of a Luma subblock, where N is the number of elements in the second set (S) of the Luma subblock, log2(N) is the exponentiation to which the number 2 must be raised to obtain the value N, and ">>" is the right arithmetic shift.

[0296] In any of the preceding implementations of the first embodiment or in possible implementations of the method according to the first embodiment itself, averaging the motion vectors of the luma subblocks in the second set S is done by the following steps, namely, mvAvgLX=Σ i mvLX[S i x ][S i y ] If mvAvgLX[0] is greater than or equal to 0, mvAvgLX[0]=(mvAvgLX[0]+N>>1)>>log2(N), otherwise, mvAvgLX[0]=-((-mvAvgLX[0]+N>>1)>>log2(N)) If mvAvgLX[1] is greater than or equal to 0, mvAvgLX[1]=(mvAvgLX[1]+N>>1)>>log2(N), otherwise, mvAvgLX[1]=-((-mvAvgLX[1]+N>>1)>>log2(N)) mvAvgLX is the result of averaging, mvAvgLX[0] is the horizontal component of the motion vector mvAvgLX, and mvAvgLX[1] is the vertical component of the motion vector mvAvgLX, S i x and S i y is the horizontal and vertical index of the subblock Si in the motion vector array, mvLX[S i x ][S i y ] is index S i x and Si y This is the motion vector of a Luma subblock, where N is the number of elements in the second set (S) of the Luma subblock, log2(N) is the exponentiation to which the number 2 must be raised to obtain the value N, and ">>" is the right arithmetic shift.

[0297] A second aspect of the present invention provides a method for deriving a chroma motion vector used in affine motion compensation of an interpretation unit PU, wherein the PU includes chroma and chroma blocks at the same location, and the method is as follows: The first step is to obtain a first set of Luma subblocks, the first set of Luma subblocks being contained in the Luma blocks of the current picture (e.g., the picture currently being coded or decoded), and the second step is to obtain a first set of Luma subblocks, the first set of Luma subblocks being contained in the Luma blocks of the current picture (e.g., the picture currently being coded or decoded), The steps include obtaining the motion vector values ​​for each Luma subblock in the first set of Luma subblocks, This is a step to obtain a set of chroma subblocks, where the set of chroma subblocks is currently contained within the picture's chroma block (in one example, the chroma block and luma block are contained within the same PU), and the next step is... The step of deriving motion vectors for chroma subblocks based on motion vectors for luma subblocks in a second set (S) of luma subblocks, wherein for chroma subblocks in a set of chroma subblocks, the second set (S) of luma subblocks (S) is determined from the first set of luma subblocks according to the current picture's chroma format, and Includes.

[0298] In a possible implementation of the method by the second embodiment itself, the second set (S) of Luma subblocks is the following subblocks, namely, S0=(xSbIdxL,ySbIdxL) S1=(xSbIdxL,ySbIdxL+(SubHeightC-1)) S2=(xSbIdxL+(SubWidthC-1),ySbIdxL) S3=(xSbIdxL+(SubWidthC-1),ySbIdxL+(SubHeightC-1)) Includes any combination of the above.

[0299] In any of the preceding implementations of the second embodiment or possible implementations of the method according to the second embodiment itself, the second set (S) of Luma subblocks consists of two subblocks, i.e., S0=(xSbIdxL,ySbIdxL) S1=(xSbIdxL+(SubWidthC-1),ySbIdxL+(SubHeightC-1)) Includes.

[0300] In any of the preceding implementations of the second embodiment or possible implementations of the method according to the second embodiment itself, deriving the motion vector for the chroma subblock based on the motion vector of the luma subblock in the second set S is: This includes averaging the motion vectors of the luma subblocks within the second set S.

[0301] In any of the preceding implementations of the second embodiment or possible implementations of the method according to the second embodiment itself, averaging the motion vectors of the luma subblocks in the second set S is: Averaging the horizontal component of the motion vector of the luma subblock in the second set S, and / or This involves averaging the vertical component of the motion vector of the luma subblock in the second set S.

[0302] In any of the preceding implementations of the second embodiment or possible implementations of the method according to the second embodiment itself, averaging the elements is performed on the element (mvLX[S i x ][S i y This includes checking whether the sum of (etc.) is 0 or greater, If the sum of the elements is 0 or greater, the sum of the elements is divided by a shift operation that depends on the number of elements. Otherwise, the absolute value of the sum of the elements is divided by a shift operation that depends on the number of elements, and a negative value is obtained from the shift result.

[0303] In any of the preceding implementations of the second embodiment or possible implementations of the method according to the second embodiment itself, the step of division by shift operation is: Rounding to 0, Rounding away from 0, Rounding away from infinity, or Rounding to infinity Includes.

[0304] In any of the preceding implementations of the second embodiment or possible implementations of the method according to the second embodiment itself, the averaging step includes averaging away from zero or rounding.

[0305] In any of the preceding implementations of the second embodiment or possible implementations of the method according to the second embodiment itself, the averaging step includes averaging toward zero or rounding.

[0306] In any of the preceding implementations of the second embodiment or possible implementations of the method according to the second embodiment itself, the averaging step includes averaging away from infinity or rounding.

[0307] In any of the preceding implementations of the second embodiment or possible implementations of the method according to the second embodiment itself, the averaging step includes averaging towards infinity or rounding.

[0308] In any of the preceding implementations of the second embodiment or the method of the second embodiment itself, averaging the motion vectors of the luma subblocks in the second set S is done by the following steps, namely, mvAvgLX=Σ i mvLX[S i x ][S i y ] mvAvgLX[0]=(mvAvgLX[0]+N>>1)>>log2(N) mvAvgLX[1]=(mvAvgLX[1]+N>>1)>>log2(N) mvAvgLX is the result of averaging, mvAvgLX[0] is the horizontal component of the motion vector mvAvgLX, and mvAvgLX[1] is the vertical component of the motion vector mvAvgLX, S i x and S i y is the horizontal and vertical index of the subblock Si in the motion vector array, mvLX[S i x ][S i y ] is index S i x and S i y This is the motion vector of a Luma subblock, where N is the number of elements in the second set (S) of the Luma subblock, log2(N) is the exponentiation to which the number 2 must be raised to obtain the value N, and ">>" is the right arithmetic shift.

[0309] In any of the preceding implementations of the second embodiment or the method of the second embodiment itself, averaging the motion vectors of the luma subblocks in the second set S is done by the following steps, namely, mvAvgLX=Σ i mvLX[S i x ][S i y ] If mvAvgLX[0] is greater than or equal to 0, mvAvgLX[0]=(mvAvgLX[0]+N>>1)>>log2(N), otherwise, mvAvgLX[0]=-((-mvAvgLX[0]+N>>1)>>log2(N)) If mvAvgLX[1] is greater than or equal to 0, mvAvgLX[1]=(mvAvgLX[1]+N>>1)>>log2(N), otherwise, mvAvgLX[1]=-((-mvAvgLX[1]+N>>1)>>log2(N)) mvAvgLX is the result of averaging, mvAvgLX[0] is the horizontal component of the motion vector mvAvgLX, and mvAvgLX[1] is the vertical component of the motion vector mvAvgLX, S i x and S i y is the horizontal and vertical index of the subblock Si in the motion vector array, mvLX[S i x ][S i y ] is index S i x and S i y This is the motion vector of a Luma subblock, where N is the number of elements in the second set (S) of the Luma subblock, log2(N) is the exponentiation to which the number 2 must be raised to obtain the value N, and ">>" is the right arithmetic shift.

[0310] In any of the preceding implementations of the second embodiment or the method of the second embodiment itself, averaging the motion vectors of the luma subblocks in the second set S is done by the following steps, namely, mvAvgLX=Σ i mvLX[S i x ][S i y ] If mvAvgLX[0] is greater than or equal to 0, mvAvgLX[0]=(mvAvgLX[0]+(N>>1)-1)>>log2(N), otherwise, mvAvgLX[0]=-((-mvAvgLX[0]+(N>>1)-1)>>log2(N)) If mvAvgLX[1] is greater than or equal to 0, mvAvgLX[1]=(mvAvgLX[1]+(N>>1)-1)>>log2(N), otherwise, mvAvgLX[1]=-((-mvAvgLX[1]+(N>>1)-1)>>log2(N)) mvAvgLX is the result of averaging, mvAvgLX[0] is the horizontal component of the motion vector mvAvgLX, and mvAvgLX[1] is the vertical component of the motion vector mvAvgLX, S i x and S i y is the horizontal and vertical index of the subblock Si in the motion vector array, mvLX[S i x ][S i y ] is index S i x and S i y This is the motion vector of a Luma subblock, where N is the number of elements in the second set (S) of the Luma subblock, log2(N) is the exponentiation of the number 2 to obtain the value N, and ">>" is the right arithmetic shift.

[0311] According to a third aspect of the present invention, a method for compensating for affine motion of a picture is provided, the method is: Currently, the steps involve dividing the picture's Luma block into sets of Luma subblocks of equal size, The current step is to divide the picture's chroma block into a set of chroma subblocks of equal size, where the size of the chroma subblock is set to be equal to the size of the chroma subblock. Includes.

[0312] In a possible implementation of the method by the third aspect itself, the method is The process further includes determining the number of horizontal and vertical chroma subblocks based on the value of the chroma scaling factor.

[0313] In any prior implementation of the third embodiment or a possible implementation of the method according to the third embodiment itself, the horizontal and vertical chroma scaling factors (SubWidthC and SubHeightC) are determined based on the chroma format of the current picture (e.g., the picture currently being coded or decoded).

[0314] In any of the preceding implementations of the third embodiment or the method of the third embodiment itself, numCSbX=numSbX>>(SubWidthC-1), and numCSbY = numSbY >> (SubHeightC-1), numCSbX and numCSbY represent the number of chroma subblocks in the horizontal and vertical directions, respectively.

[0315] In any of the preceding implementations of the third embodiment or possible implementations of the method by the third embodiment itself, currently dividing the chroma blocks of a picture is, Currently, the chroma block of the picture (in one example, the chroma block and luma block are contained in the same PU) is divided into numCSbY rows, and each row has numCSbX chroma subblocks.

[0316] In any of the preceding implementations of the third aspect or possible implementations of the method by the third aspect itself, the method is: This step determines the set (S) of luma subblocks for each chroma subblock, where the position of the luma subblock within the set is determined by the current picture's chroma format. A step of deriving the motion vector for the chroma subblock based on the motion vector of the luma subblock in set S. It also includes.

[0317] In any of the preceding implementations of the third embodiment or in possible implementations of the method according to the third embodiment itself, motion vectors are defined for each of the luma subblocks belonging to set S.

[0318] In any of the preceding implementations of the third embodiment or possible implementations of the method according to the third embodiment itself, the set of Luma subblocks (S) is the following subblocks, namely, S0=(xSbIdxL,ySbIdxL) S1=(xSbIdxL,ySbIdxL+(SubHeightC-1)) S2=(xSbIdxL+(SubWidthC-1),ySbIdxL) S3=(xSbIdxL+(SubWidthC-1),ySbIdxL+(SubHeightC-1)) Includes any combination of the above.

[0319] In any of the preceding implementations of the third embodiment or possible implementations of the method according to the third embodiment itself, the set of Luma subblocks (S) consists of two subblocks, i.e., S0=(xSbIdxL,ySbIdxL) S1=(xSbIdxL+(SubWidthC-1),ySbIdxL+(SubHeightC-1)) Includes.

[0320] In any of the preceding implementations of the third embodiment or possible implementations of the method according to the third embodiment itself, deriving the motion vector for a chroma subblock based on the motion vector of a luma subblock in set S includes averaging the motion vectors of the luma subblocks in set S.

[0321] In any of the preceding implementations of the third embodiment or possible implementations of the method according to the third embodiment itself, averaging the motion vectors of the luma subblocks in set S is: The horizontal component of the motion vector of the Luma subblock in set S is averaged. This includes averaging the vertical component of the motion vector of the Luma subblock within set S.

[0322] In any of the preceding implementations of the third embodiment or possible implementations of the method according to the third embodiment itself, averaging the elements includes checking whether the sum of the elements is 0 or greater.

[0323] In any of the preceding implementations of the third aspect or in possible implementations of the method by the third aspect itself, averaging means averaging away from zero.

[0324] In any of the preceding implementations of the third aspect or in any possible implementation of the method according to the third aspect itself, averaging means averaging toward zero.

[0325] In any of the preceding implementations of the third aspect or in possible implementations of the method by the third aspect itself, averaging means averaging away from infinity.

[0326] In any of the preceding implementations of the third aspect or in any possible implementation of the method of the third aspect itself, averaging means averaging toward infinity.

[0327] In any of the preceding implementations of the third embodiment or possible implementations of the method according to the third embodiment itself, averaging the motion vectors of the luma subblocks in set S is done by the following steps, namely: mvAvgLX=Σ i mvLX[S i x ][S i y ] mvAvgLX[0]=(mvAvgLX[0]+N>>1)>>log2(N) mvAvgLX[1]=(mvAvgLX[1]+N>>1)>>log2(N) mvAvgLX is the result of averaging, mvAvgLX[0] is the horizontal component of the motion vector mvAvgLX, and mvAvgLX[1] is the vertical component of the motion vector mvAvgLX, S i x and S i y is the horizontal and vertical index of the subblock Si in the motion vector array, mvLX[S i x ][S i y ] is index Si x and S i y This is the motion vector of a Luma subblock, where N is the number of elements in the second set (S) of the Luma subblock, log2(N) is the exponentiation to which the number 2 must be raised to obtain the value N, and ">>" is the right arithmetic shift.

[0328] In any of the preceding implementations of the third embodiment or possible implementations of the method according to the third embodiment itself, averaging the motion vectors of the luma subblocks in set S is done by the following steps, namely: mvAvgLX=Σ i mvLX[S i x ][S i y ] If mvAvgLX[0] is greater than or equal to 0, mvAvgLX[0]=(mvAvgLX[0]+N>>1)>>log2(N), otherwise, mvAvgLX[0]=-((-mvAvgLX[0]+N>>1)>>log2(N)) If mvAvgLX[1] is greater than or equal to 0, mvAvgLX[1]=(mvAvgLX[1]+N>>1)>>log2(N), otherwise, mvAvgLX[1]=-((-mvAvgLX[1]+N>>1)>>log2(N)) mvAvgLX is the result of averaging, mvAvgLX[0] is the horizontal component of the motion vector mvAvgLX, and mvAvgLX[1] is the vertical component of the motion vector mvAvgLX, S i x and S i y is the horizontal and vertical index of the subblock Si in the motion vector array, mvLX[S i x ][S i y ] is index S i x and S i yThis is the motion vector of a Luma subblock, where N is the number of elements in the second set (S) of the Luma subblock, log2(N) is the exponentiation to which the number 2 must be raised to obtain the value N, and ">>" is the right arithmetic shift.

[0329] In any of the preceding implementations of the third embodiment or possible implementations of the method according to the third embodiment itself, averaging the motion vectors of the luma subblocks in set S is done by the following steps, namely: mvAvgLX=Σ i mvLX[S i x ][S i y ] If mvAvgLX[0] is greater than or equal to 0, mvAvgLX[0]=(mvAvgLX[0]+(N>>1)-1)>>log2(N), otherwise, mvAvgLX[0]=-((-mvAvgLX[0]+(N>>1)-1)>>log2(N)) If mvAvgLX[1] is greater than or equal to 0, mvAvgLX[1]=(mvAvgLX[1]+(N>>1)-1)>>log2(N), otherwise, mvAvgLX[1]=-((-mvAvgLX[1]+(N>>1)-1)>>log2(N)) mvAvgLX is the result of averaging, mvAvgLX[0] is the horizontal component of the motion vector mvAvgLX, and mvAvgLX[1] is the vertical component of the motion vector mvAvgLX, S i x and S i y is the horizontal and vertical index of the subblock Si in the motion vector array, mvLX[S i x ][S i y ] is index S i x and S i yThis is the motion vector of a Luma subblock, where N is the number of elements in the second set (S) of the Luma subblock, log2(N) is the exponentiation to which the number 2 must be raised to obtain the value N, and ">>" is the right arithmetic shift.

[0330] In any of the preceding implementations of the third embodiment or the method of the third embodiment itself, N is equal to 1.

[0331] In any of the preceding implementations of the third embodiment or possible implementations of the method according to the third embodiment itself, set S is determined based on the value of sps_cclm_colocated_chroma_flag.

[0332] In any of the preceding implementations of the third embodiment or possible implementations of the method according to the third embodiment itself, the mean motion vector derivation is performed based on the value of sps_cclm_colocated_chroma_flag.

[0333] In any of the preceding implementations of the third embodiment or possible implementations of the method according to the third embodiment itself, the chroma format is defined as YUV4:2:2.

[0334] In any of the preceding implementations of the third embodiment or the method of the third embodiment itself, the chroma block and the luma block are included in the same PU and are located in the same position.

[0335] A fourth aspect of the encoder (20) includes a processing circuit for performing a method according to any one of the first to third aspects.

[0336] A fifth aspect of the decoder (30) includes a processing circuit for performing a method according to any one of the first to third aspects.

[0337] A sixth aspect of the computer program product includes program code for performing a method according to any one of the first to third aspects.

[0338] A seventh aspect of the decoder is: The system includes one or more processors and a non-temporary computer-readable storage medium coupled to the processors and storing a program for execution by the processors, wherein the program, when executed by the processors, configures the decoder to perform a method according to any one of the first to third embodiments.

[0339] The eighth aspect of the encoder is: The encoder includes one or more processors and a non-temporary computer-readable storage medium coupled to the processors and storing a program for execution by the processors, wherein the program, when executed by the processors, configures the encoder to perform a method according to any one of the first to third embodiments.

[0340] Based on the above, this disclosure concerns a method for considering the chroma format of a picture when obtaining chroma motion vectors from chroma motion vectors. Linear subsampling of the chroma motion field is performed by averaging the chroma motion vectors. When the chroma color plane is at the same height as the chroma plane, motion vectors are selected from horizontally adjacent chroma blocks, thereby it is found that it is more appropriate for them to have the same vertical position. Selecting chroma motion vectors depending on the picture chroma format results in a more accurate chroma motion field due to more accurate chroma motion vector field subsampling. This dependency on the chroma format allows for the selection of the most appropriate chroma blocks when averaging the chroma motion vectors to generate the chroma motion vectors. As a result of more accurate motion field interpolation, prediction errors are reduced, which leads to a technical consequence of improved compression performance.

[0341] The following describes the application of the encoding and decoding methods shown in the above embodiments, as well as the systems using them.

[0342] Figure 18 is a block diagram showing a content supply system 3100 for realizing a content distribution service. This content supply system 3100 includes a capture device 3102 and a terminal device 3106, and optionally includes a display 3126. The capture device 3102 communicates with the terminal device 3106 over a communication link 3104. The communication link may include the communication channel 13 described above. The communication link 3104 includes, but is not limited to, Wi-Fi, Ethernet, cable, wireless (3G / 4G / 5G), USB, or any combination thereof.

[0343] The capture device 3102 may generate data and encode the data using an encoding method as shown in the above embodiment. Alternatively, the capture device 3102 may distribute the data to a streaming server (not shown in the drawings), which encodes the data and transmits the encoded data to the terminal device 3106. The capture device 3102 includes, but is not limited to, a camera, a smartphone or tablet, a computer or laptop, a video conferencing system, a PDA, an in-vehicle device, or any combination thereof. For example, the capture device 3102 may include the source device 12 as described above. When the data includes video, the video encoder 20 included in the capture device 3102 may actually perform video encoding. When the data includes audio (i.e., voice), the audio encoder included in the capture device 3102 may actually perform audio encoding. In some practical scenarios, the capture device 3102 distributes the encoded video and audio data by multiplexing them together. In other practical scenarios, for example in a video conferencing system, the encoded audio data and encoded video data are not multiplexed. The capture device 3102 separately distributes encoded audio data and encoded video data to the terminal device 3106.

[0344] In the content supply system 3100, the terminal device 3106 receives and plays back encoded data. The terminal device 3106 may be any device capable of receiving and restoring data, such as a smartphone or tablet 3108, a computer or laptop 3110, a network video recorder (NVR) / digital video recorder (DVR) 3112, a TV 3114, a set-top box (STB) 3116, a video conferencing system 3118, a video surveillance system 3120, a personal digital assistant (PDA) 3122, an in-vehicle device 3124, or any combination thereof, that can decode the above encoded data. For example, the terminal device 3106 may include the destination device 14 as described above. When the encoded data includes video, the video decoder 30 included in the terminal device is prioritized to perform video decoding. When the encoded data includes audio, the audio decoder included in the terminal device is prioritized to perform audio decoding.

[0345] In terminal devices with their own displays, such as smartphones or tablets 3108, computers or laptops 3110, network video recorders (NVRs) / digital video decoders (DVRs) 3112, TVs 3114, personal digital assistants (PDAs) 3122, or in-vehicle devices 3124, the terminal device can supply decoded data to its own display. In terminal devices without a display, such as STBs 3116, video conferencing systems 3118, or video surveillance systems 3120, an external display 3126 is made contact with the terminal device to receive and display the decoded data.

[0346] When each device in this system performs encoding or decoding, a picture encoding device or a picture decoding device can be used, as shown in the embodiment described above.

[0347] Figure 19 shows the structure of an example terminal device 3106. After the terminal device 3106 receives a stream from the capture device 3102, the protocol processing unit 3202 analyzes the transmission protocol of the stream. The protocol includes, but is not limited to, Real Time Streaming Protocol (RTSP), Hyper Text Transfer Protocol (HTTP), HTTP Live Streaming Protocol (HLS), MPEG-DASH, Real-time Transport Protocol (RTP), Real Time Messaging Protocol (RTMP), or any combination thereof.

[0348] After the protocol processing unit 3202 processes the stream, a stream file is generated. The file is output to the demultiplexing unit 3204. The demultiplexing unit 3204 can separate the multiplexed data into encoded audio data and encoded video data. As described above, in some real-world scenarios, for example in a video conferencing system, the encoded audio data and encoded video data are not multiplexed. In this situation, the encoded data is sent to the video decoder 3206 and audio decoder 3208 without passing through the demultiplexing unit 3204.

[0349] Through demultiplexing, a video elementary stream (ES), an audio ES, and an optional subtitle are generated. The video decoder 3206, including the video decoder 30 as described in the above embodiment, decodes the video ES using the decoding method shown in the above embodiment to generate video frames and supplies this data to the synchronization unit 3212. The audio decoder 3208 decodes the audio ES to generate audio frames and supplies this data to the synchronization unit 3212. Alternatively, the video frames may be stored in a buffer (not shown in Figure Y) before being supplied to the synchronization unit 3212. Similarly, the audio frames may be stored in a buffer (not shown in Figure Y) before being supplied to the synchronization unit 3212.

[0350] The synchronization unit 3212 synchronizes video frames and audio frames and supplies video / audio to the video / audio display 3214. For example, the synchronization unit 3212 synchronizes the presentation of video and audio information. The information may be coded in the syntax using timestamps for the presentation of coded audio and visual data and timestamps for the delivery of the data stream itself.

[0351] If subtitles are included in the stream, the subtitle decoder 3210 decodes the subtitles, synchronizes them with the video and audio frames, and supplies the video / audio / subtitle to the video / audio / subtitle display 3216.

[0352] The present invention is not limited to the system described above, and either the picture encoding device or the picture decoding device in the above embodiment can be incorporated into other systems, such as a vehicle system.

[0353] Mathematical operators The mathematical operators used in this application are similar to those used in the C programming language. However, the results of integer division and arithmetic shift operations are more precisely defined, and further operators such as exponential calculations and real-valued division are defined. The numbering and counting rules generally start from 0, for example, "1st" is equivalent to the 0th, "2nd" is equivalent to the 1st, and so on.

[0354] Logical operators The following logical operators are defined as follows: [Table 10]

[0355] Logical operators The following logical operators are defined as follows: x&&y: Boolean "product" of x and y Boolean "union" of x||yx and y ! Boolean logic "negation" If x?y:zx is true or not equal to 0, it evaluates to the value of y; otherwise, it evaluates to the value of z.

[0356] Relational operators The following relational operators are defined as follows: > Larger than >= Above < Less than <= Below == equal != Not equal

[0357] When a relational operator is applied to a syntax element or variable assigned the value "na" (not applicable), the value "na" is treated as a separate value of the syntax element or variable. The value "na" is considered not to be equal to any other value.

[0358] Bitwise operators The following bitwise operators are defined as follows: & Bitwise "product". When operating on integer arguments, the operation is performed on the two's complement representation of the integer values. When operating on a binary argument that contains fewer bits than the other argument, the shorter argument is extended by adding higher-order bits equal to 0. | Bitwise "sum". When operating on integer arguments, the operation is performed on the two's complement representation of the integer values. When operating on a binary argument that contains fewer bits than the other argument, the shorter argument is extended by adding higher-order bits equal to 0. ^ Bitwise "exclusive sum". When operating on integer arguments, the operation is performed on the two's complement representation of the integer values. When operating on a binary argument that contains fewer bits than the other argument, the shorter argument is extended by adding higher-order bits equal to 0. x>>y Arithmetic right shift of the two's complement integer representation of x by y binary digits. This function is defined only for non-negative integer values of y. The bit shifted into the most significant bit (MSB) as a result of the right shift has a value equal to the MSB of x before the shift operation. x<<y Arithmetic left shift of the two's complement integer representation of x by y binary digits. This function is defined only for non-negative integer values of y. The bit shifted into the least significant bit (LSB) as a result of the left shift has a value equal to 0.

[0359] Assignment operators The following assignment operators are defined as follows. = Assignment operator ++ Increment. That is, x++ is equal to x=x+1. When used in an array index, it is evaluated to the value of the variable before the increment operation. -- Decrement. That is, x-- is equal to x=x-1. When used in an array index, it is evaluated to the value of the variable before the decrement operation. += increments by the specified amount. That is, x+=3 is equal to x=x+3, and x+=(-3) is equal to x=x+(-3). -= Decrement by the specified amount. That is, x-=3 is equal to x=x-3, and x-=(-3) is equal to x=x-(-3).

[0360] Range notation The following notation is used to specify a range of values. x = y . . zx takes integer values ​​between y and z (inclusive), where x, y, and z are integers, and z is greater than y.

[0361] Mathematical functions The following mathematical functions are defined.

number

number

number

number

number

number

number

[0362] Order of operations When the precedence of an expression is not explicitly indicated by the use of parentheses, the following rules apply: - Higher-priority operations are evaluated before any lower-priority operations. - Operations with the same priority are evaluated sequentially from left to right.

[0363] The table below specifies the order of operations from highest to lowest, with higher positions in the table indicating higher priority.

[0364] For operators also used in the C programming language, the precedence used herein is the same as that used in the C programming language. [Table 11]

[0365] Text description of logical operations In text, the format is as follows: if (condition 0) Statement 0 else(condition 1) Statement 1 ... else / *Reference notes regarding the remaining conditions* / statement n Statements of logical operations that are mathematically described may also be written in the following manner. ...as follows / ...the following applies: -If condition 0, statement 0 - If not, and condition 1 is true, then statement 1 -... -Otherwise (see note regarding the remaining conditions), statement n

[0366] Each "if..., otherwise..., otherwise..." statement in the text is introduced by "if..." followed immediately by "...as follows" or "...as follows." The final condition in "if..., otherwise..., otherwise..." is always "otherwise...." Alternating "if..., otherwise..., otherwise..." statements can be identified by matching them with "...as follows" or "...as follows" ending with "otherwise...."

[0367] In text, the following format: if(condition 0a &&condition 0b) Statement 0 else if(condition 1a||condition 1b) Statement 1 ... else statement n Statements of logical operations that are mathematically described may also be written in the following manner. ...as follows / ...the following applies: -Statement 0 if all of the following conditions are true: -Condition 0a -condition 0b -Otherwise, if one or more of the following conditions are true, then statement 1: -Condition 1a -Condition 1b -… -Otherwise, statement n

[0368] In text, the following format: if (condition 0) Statement 0 if (condition 1) Statement 1 Statements of logical operations that are mathematically described may also be written in the following manner. When condition 0, statement 0 When condition 1 is met, statement 1

[0369] The present invention has been described herein in relation to various embodiments. However, other variations to the embodiments disclosed can be understood and implemented by those skilled in the art in carrying out the claimed invention, from the study of the drawings, disclosure and appended claims. In the claims, the term “including” does not exclude other elements or steps, and the indefinite article “a” or “an” does not exclude plurals. A single processor or other unit may perform the functions of several items described in the claims. The mere fact that certain means are usually described in different dependent claims does not imply that combinations of these means cannot be used advantageously. Computer programs may be stored / distributed on suitable media such as optical or solid media supplied together with or as part of other hardware, or they may be distributed in other forms such as via the Internet or other wired or wireless communication systems.

[0370] Those skilled in the art will understand that the “blocks” (“units”) of various drawings (methods and apparatus) represent or describe the function of embodiments of the present invention (not necessarily individual “units” in hardware or software), and therefore, the function or features of apparatus embodiments are equally described (units = steps) along with the embodiments of methods.

[0371] The term "unit" is used solely for the purpose of describing the function of the encoder / decoder embodiment and is not intended to limit this disclosure.

[0372] In some embodiments provided herein, it should be understood that the disclosed systems, apparatuses, and methods may be implemented in other ways. For example, the embodiments of the described apparatuses are merely illustrative. For example, unit divisions are merely logical functional divisions, and other divisions may be used in actual implementations. For example, multiple units or components may be combined or integrated into other systems, or some features may be ignored or not performed. Furthermore, mutual coupling, direct coupling, or communication connections indicated or discussed may be implemented using some interfaces. Indirect coupling or communication connections between apparatuses or units may be implemented electronically, mechanically, or in other forms.

[0373] Units described as separate parts may or may not be physically separated, and parts shown as units may or may not be physical units, may be located in one location, or may be distributed across multiple network units. Some or all of the units may be selected according to the actual needs in order to achieve the objectives of the solution of the embodiment.

[0374] Furthermore, the functional units in the embodiments of the present invention may be integrated into a single processing unit, or each unit may exist physically independently, or two or more units may be integrated into a single unit.

[0375] Embodiments of the present invention may further include an apparatus comprising a processing circuit configured to perform any of the methods and / or processes described herein, such as an encoder and / or decoder.

[0376] While embodiments of the present invention have been described primarily in terms of video coding, it should be noted that embodiments of the coding system 10, encoder 20, and decoder 30 (and correspondingly system 10), as well as other embodiments described herein, may also be configured for still image picture processing or coding, i.e., processing or coding of individual pictures independent of any preceding or consecutive pictures, as in video coding. Generally, when picture processing coding is limited to a single picture 17, only the interpretation units 244 (encoder) and 344 (decoder) may be available. All other functions (also called tools or techniques) of the video encoder 20 and video decoder 30 may be used equivalently for still image picture processing, such as residual calculation 204 / 304, transformation 206, quantization 208, inverse quantization 210 / 310, (inverse) transformation 212 / 312, partitioning 262 / 362, intra prediction 254 / 354, and / or loop filtering 220, 320, as well as entropy coding 270 and entropy decoding 304.

[0377] For example, embodiments of the encoder 20 and decoder 30, and the functions described herein with respect to the encoder 20 and decoder 30, may be implemented in hardware, software, firmware, or a combination thereof. If implemented in software, the functions may be stored as one or more instructions or codes on a computer-readable medium or transmitted over a communication medium and executed by a hardware-based processing unit. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including, for example, any medium that facilitates the transfer of computer programs from one location to another according to a communication protocol. Thus, the computer-readable medium may generally correspond to (1) a non-transient tangible computer-readable storage medium, or (2) a communication medium such as a signal or carrier wave. The data storage medium may be any available medium accessible by one or more computers or one or more processors to retrieve instructions, codes and / or data structures for the implementation of the technology described herein. The computer program product may include a computer-readable medium.

[0378] Such computer-readable storage media may include, but are not limited to, computer-readable storage media, such as RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and is accessible by a computer. Any connection is also appropriately called a computer-readable medium. For example, if instructions are transmitted from a website, server or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of a medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carriers, signals, or other temporary media, but instead refer to non-temporary tangible storage media. When used herein, "disk" and "disc" include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy discs, and Blu-ray discs, where a "disk" typically reproduces data magnetically, and a "disc" reproduces data optically using a laser. Any combination of the above should also be included in the scope of computer-readable media.

[0379] Instructions may be executed by one or more processors, such as digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Therefore, when used herein, the term "processor" may refer to any of the above structures or any other structure suitable for the implementation of the technology described herein. Furthermore, in some embodiments, the functions described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into a combined codec. The technology may also be fully implemented in one or more circuits or logic elements.

[0380] The technology of this disclosure may be implemented in a wide range of devices or apparatus, including wireless handsets, integrated circuits (ICs), or sets of ICs (e.g., chipsets). Various components, modules, or units are described in this disclosure to highlight the functional aspects of devices configured to perform the technology of this disclosure, but implementation by different hardware units is not necessarily required. Rather, as described above, various units may be coupled to a codec hardware unit in combination with appropriate software and / or firmware, or they may be provided by a collection of interoperable hardware units including one or more processors as described above.

Claims

1. A method for deriving chroma motion vectors used in affine-based interpretation of current image blocks including chroma blocks and chroma blocks at the same location, The steps include determining horizontal and vertical chroma scaling factors based on chroma format information, wherein the chroma format information indicates the chroma format of the current picture to which the current image block belongs, and The steps include determining a set of luma subblocks (S) of the luma block based on the value of the chroma scaling factor, The steps of determining the motion vector for the chroma subblock of the chroma block based on the motion vectors of one or more chroma subblocks in the set (S) of chroma subblocks, and A method that includes this.

2. The method according to claim 1, wherein each of the one or more Luma subblocks in the set (S) is represented by a horizontal subblock index and a vertical subblock index.

3. When both SubWidthC and SubHeightC are equal to 1, the set (S) of Luma subblocks is S 0 Includes a Luma subblock indexed by =(xSbIdx,ySbIdx), When at least one of SubWidthC and SubHeightC is not equal to 1, the set (S) of Luma subblocks is S 0 The first luma subblock indexed by =((xSbIdx>>(SubWidthC-1)<<(SubWidthC-1)),(ySbIdx>>(SubHeightC-1)<<(SubHeightC-1))), and S 1 The second Luma subblock is indexed by =((xSbIdx>>(SubWidthC-1)<<(SubWidthC-1))+(SubWidthC-1),(ySbIdx>>(SubHeightC-1)<<(SubHeightC-1))+(SubHeightC-1)) Includes, SubWidthC and SubHeightC represent the chroma scaling factors in the horizontal and vertical directions, respectively. xSbIdx and ySbIdx represent the horizontal subblock index and the vertical subblock index, respectively, for the Luma subblock in the set (S). "<<" represents a left arithmetic shift. ">>" represents a right arithmetic shift. xSbIdx=0..numSbX-1 and ySbIdx=0..numSbY-1, numSbX indicates the number of luma subblocks within the luma block along the horizontal direction, The method according to claim 1 or 2, wherein numSbY indicates the number of luma subblocks in the luma block along the vertical direction.

4. The method according to claim 3, wherein the number of chroma subblocks in the chroma block along the horizontal and vertical directions is the same as the number of luma subblocks in the luma block along the horizontal and vertical directions, respectively.

5. When both SubWidthC and SubHeightC are equal to 1, the set (S) of Luma subblocks is S 0 Includes a Luma subblock indexed by =(xCSbIdx,yCSbIdx), When at least one of SubWidthC and SubHeightC is not equal to 1, the set (S) of Luma subblocks is S 0 The first Luma subblock indexed by =((xCSbIdx>>(SubWidthC-1)<<(SubWidthC-1)),(yCSbIdx>>(SubHeightC-1)<<(SubHeightC-1))), and S 1 The second Luma subblock is indexed by =(xCSbIdx>>(SubWidthC-1)<<(SubWidthC-1))+(SubWidthC-1),(yCSbIdx>>(SubHeightC-1)<<(SubHeightC-1))+(SubHeightC-1)) Includes, SubWidthC and SubHeightC represent the chroma scaling factors in the horizontal and vertical directions, respectively. xCSbIdx and yCSbIdx represent the horizontal subblock index and the vertical subblock index, respectively, for the Luma subblock in the set (S). The method according to claim 1 or 2, wherein xCSbIdx = 0..numCSbX-1 and yCSbIdx = 0..numCSbY-1, where numCSbX represents the number of chroma subblocks in the horizontal direction and numCSbY represents the number of chroma subblocks in the vertical direction.

6. The method according to claim 5, wherein the size of each of the chroma subblocks is the same as the size of each of the luma subblocks.

7. The number of chroma subblocks within the chroma block along the horizontal direction depends on the number of luma subblocks within the luma block along the horizontal direction and the value of the chroma scaling factor in the horizontal direction. The method according to claim 5, wherein the number of chroma subblocks in the chroma block along the vertical direction depends on the number of luma subblocks in the luma block along the vertical direction and the value of the chroma scaling factor in the vertical direction.

8. With respect to the chroma subblocks indexed by the horizontal subblock index xSbIdxL and the vertical subblock index ySbIdxL, the set (S) of luma subblocks is: S 0 =(xSbIdxL,ySbIdxL)、 S 1 =(xSbIdxL,ySbIdxL+(SubHeightC-1))、 S 2 =(xSbIdxL+(SubWidthC-1),ySbIdxL), or S 3 =(xSbIdxL+(SubWidthC-1),ySbIdxL+(SubHeightC-1)) Includes one or more subblocks indexed by The aforementioned Luma subblock index S 0 This is represented by the horizontal subblock index xSbIdxL and the vertical subblock index ySbIdxL, The aforementioned Lumablock Index S 1 This is represented by the horizontal subblock index xSbIdxL and the vertical subblock index ySbIdxL+(SubHeightC-1), The aforementioned Lumablock Index S 2 This is represented by the horizontal subblock index xSbIdxL+(SubWidthC-1) and the vertical subblock index ySbIdxL, The aforementioned Lumablock Index S 3 The method according to any one of claims 1 to 7, wherein is represented by the horizontal subblock index xSbIdxL+(SubWidthC-1) and the vertical subblock index ySbIdxL+(SubHeightC-1).

9. The set (S) of the Lumasa subblock is, S 0 The first Luma subblock is indexed by =(xSbIdxL,ySbIdxL), S 1 The second Luma subblock is indexed by =(xSbIdxL+(SubWidthC-1),ySbIdxL+(SubHeightC-1)) and Includes, The aforementioned luma block position S 0 This is represented by the horizontal subblock index xSbIdxL and the vertical subblock index ySbIdxL, The aforementioned luma block position S 1 The method according to any one of claims 1 to 8, wherein is represented by the horizontal subblock index xSbIdxL+(SubWidthC-1) and the vertical subblock index ySbIdxL+(SubHeightC-1).

10. When the chroma format is 4:4:4, the set (S) includes one luma subblock in the same position as the chroma subblock, When the chroma format is 4:2:2, the set (S) includes two chroma subblocks that are horizontally adjacent to each other. The method according to any one of claims 1 to 9, wherein, when the chroma format is 4:2:0, the set (S) includes two diagonally aligned chroma subblocks.

11. When there are more than one luma subblock in the set (S), determining the motion vector for the chroma subblock based on the motion vectors of one or more luma subblocks in the set (S) is: To generate an averaged Luma motion vector, the motion vectors of the Luma subblocks in the set S are averaged, The motion vector for the chroma subblock is derived based on the averaged luma motion vector. The method according to any one of claims 1 to 10, including

12. Averaging the motion vectors of the luma subblocks within the set S is, Average the horizontal component of the motion vector of the luma subblock within the set S, or Average the vertical component of the motion vector of the luma subblock within the set S. The method according to claim 11, comprising one or more of the above.

13. To generate an averaged Luma motion vector, the motion vectors of the Luma subblocks in the set S are averaged. mvAvgLX=Σ i mvLX[S i x ][S i y ] When mvAvgLX[0] is greater than or equal to 0, then mvAvgLX[0] = (mvAvgLX[0] + (N>>1)-1)>>log2(N), Otherwise, mvAvgLX[0]=-((-mvAvgLX[0]+(N>>1)-1)>>log2(N)) When mvAvgLX[1] is greater than or equal to 0, then mvAvgLX[1] = (mvAvgLX[1] + (N>>1)-1)>>log2(N), Otherwise, mvAvgLX[1]=-((-mvAvgLX[1]+(N>>1)-1)>>log2(N)) Includes, mvAvgLX is the motion vector resulting from the averaging, mvAvgLX[0] is the horizontal component of the resulting motion vector mvAvgLX, and mvAvgLX[1] is the vertical component of the resulting motion vector mvAvgLX. S i x and S i y These are the horizontal and vertical indices of the subblock Si in the set (S) of Luma subblocks in the motion vector array, mvLX[S i x ][S i y ] is index S i x and S i y The method according to claim 11 or 12, wherein the motion vector of a luma subblock is such that N is the number of elements in the set (S) of the luma subblock, log2(N) represents the base 2 logarithm of N, where the number 2 is raised to a power to obtain the value N, and ">>" represents a right arithmetic shift.

14. Averaging the motion vectors of the luma subblocks within the set S is, mvAvgLX=mvLX[(xSbIdx>>(SubWidthC-1)<<(SubWidthC-1))][(ySbIdx>>(SubHeightC-1)<<(SubHeightC-1))]+mvLX[(xSbI dx>>(SubWidthC-1)<<(SubWidthC-1))+(SubWidthC-1)][(ySbIdx>>(SubHeightC-1)<<(SubHeightC-1))+(SubHeightC-1)] If mvAvgLX[0]>=0, mvAvgLX[0]=(mvAvgLX[0]+1-(mvAvgLX[0]>=0))>>1 If mvAvgLX[1] >= 0, mvAvgLX[1]=(mvAvgLX[1]+1-(mvAvgLX[1]>=0))>>1 Includes, The method according to any one of claims 11 to 13, wherein mvAvgLX[0] is the horizontal component of the averaged motion vector mvAvgLX, mvAvgLX[1] is the vertical component of the averaged motion vector mvAvgLX, SubWidthC and SubHeightC represent the horizontal and vertical chroma scaling factors, respectively, and xSbIdx and ySbIdx represent the horizontal subblock index and the vertical subblock index, respectively, for the chroma subblock in the set (S), where "<<" is a left arithmetic shift and ">>" is a right arithmetic shift.

15. Determining the horizontal and vertical chroma scaling factors based on chroma format information is: The method according to any one of claims 1 to 14, comprising determining the horizontal and vertical chroma scaling factors based on a mapping between the chroma format information and the pair of horizontal and vertical chroma scaling factors.

16. The method according to any one of claims 1 to 15, further comprising the step of generating a prediction of the chroma subblock based on the determined motion vector.

17. The method according to any one of claims 1 to 16, wherein the chroma format includes one of the YUV4:2:2 format, YUV4:2:0 format, or YUV4:4:4 format.

18. The method according to any one of claims 1 to 17, as implemented by an encoding device.

19. The method according to any one of claims 1 to 17, as implemented by a decoding device.

20. An encoder (20) including a processing circuit for performing the method according to any one of claims 1 to 18.

21. A decoder (30) including a processing circuit for performing the method described in any one of claims 1 to 17 and 19.

22. A computer program product comprising program code for performing the method described in any one of claims 1 to 19.

23. One or more processors, A non-temporary computer-readable storage medium coupled to the processor and storing a program for execution by the processor. A decoder that includes, A decoder wherein, when the programming is executed by the processor, the decoder is configured to perform the method according to any one of claims 1 to 17 and 19.

24. One or more processors, A non-temporary computer-readable storage medium coupled to the processor and storing a program for execution by the processor. An encoder that includes, An encoder wherein, when the programming is executed by the processor, the encoder is configured to perform the method according to any one of claims 1 to 18.

25. A non-temporary computer-readable medium that, when executed by a computer device, carries program code that causes the computer device to execute the method according to any one of claims 1 to 19.

26. A bitstream for a video signal that includes multiple syntax elements, The plurality of syntax elements include a first flag, and the method according to any one of claims 1 to 19 is performed on a bitstream based on the value of the first flag.

27. An apparatus for affine-based interpretation of current image blocks, including luma and chroma blocks at the same location, A determination module configured to determine horizontal and vertical chroma scaling factors based on chroma format information, wherein the chroma format information indicates the chroma format of the current picture to which the current image block belongs, and is configured to determine a set of chroma subblocks (S) of the chroma block based on the value of the chroma scaling factor, A motion vector derivation module configured to determine the motion vector for the chroma subblock of the chroma block based on the motion vectors of one or more luma subblocks in the set (S) of luma subblocks. A device that includes this.

28. The apparatus according to claim 27, wherein each of the one or more Luma subblocks in the set (S) is represented by a horizontal subblock index and a vertical subblock index.

29. When both SubWidthC and SubHeightC are equal to 1, the set (S) of Luma subblocks is S 0 Includes a Luma subblock indexed by =(xSbIdx,ySbIdx), When at least one of SubWidthC and SubHeightC is not equal to 1, the set (S) of Luma subblocks is S 0 The first luma subblock indexed by =((xSbIdx>>(SubWidthC-1)<<(SubWidthC-1)),(ySbIdx>>(SubHeightC-1)<<(SubHeightC-1))), and S 1 The second Luma subblock is indexed by =((xSbIdx>>(SubWidthC-1)<<(SubWidthC-1))+(SubWidthC-1),(ySbIdx>>(SubHeightC-1)<<(SubHeightC-1))+(SubHeightC-1)) Includes, SubWidthC and SubHeightC represent the chroma scaling factors in the horizontal and vertical directions, respectively. xSbIdx and ySbIdx represent the horizontal subblock index and the vertical subblock index, respectively, for the Luma subblock in the set (S). "<<" represents a left arithmetic shift. ">>" represents a right arithmetic shift. xSbIdx=0..numSbX-1 and ySbIdx=0..numSbY-1, numSbX indicates the number of luma subblocks within the luma block along the horizontal direction, The apparatus according to claim 27 or 28, wherein numSbY indicates the number of luma subblocks in the luma block along the vertical direction.

30. The apparatus according to claim 29, wherein the number of chroma subblocks in the chroma block along the horizontal and vertical directions is the same as the number of luma subblocks in the luma block along the horizontal and vertical directions, respectively.

31. When both SubWidthC and SubHeightC are equal to 1, the set (S) of Luma subblocks is S 0 Includes a Luma subblock indexed by =(xCSbIdx,yCSbIdx), When at least one of SubWidthC and SubHeightC is not equal to 1, the set (S) of Luma subblocks is S 0 The first Luma subblock indexed by =((xCSbIdx>>(SubWidthC-1)<<(SubWidthC-1)),(yCSbIdx>>(SubHeightC-1)<<(SubHeightC-1))), and S 1 The second Luma subblock is indexed by =(xCSbIdx>>(SubWidthC-1)<<(SubWidthC-1))+(SubWidthC-1),(yCSbIdx>>(SubHeightC-1)<<(SubHeightC-1))+(SubHeightC-1)) Includes, SubWidthC and SubHeightC represent the chroma scaling factors in the horizontal and vertical directions, respectively. xCSbIdx and yCSbIdx represent the horizontal subblock index and the vertical subblock index, respectively, for the Luma subblock in the set (S). The apparatus according to claim 27 or 28, wherein xCSbIdx = 0..numCSbX-1 and yCSbIdx = 0..numCSbY-1, where numCSbX represents the number of chroma subblocks in the horizontal direction and numCSbY represents the number of chroma subblocks in the vertical direction.

32. The apparatus according to claim 31, wherein the size of each of the chroma subblocks is the same as the size of each of the luma subblocks.

33. The number of chroma subblocks within the chroma block along the horizontal direction depends on the number of luma subblocks within the luma block along the horizontal direction and the value of the chroma scaling factor in the horizontal direction. The apparatus according to claim 31, wherein the number of chroma subblocks in the chroma block along the vertical direction depends on the number of luma subblocks in the luma block along the vertical direction and the value of the chroma scaling factor in the vertical direction.

34. With respect to the chroma subblocks indexed by the horizontal subblock index xSbIdxL and the vertical subblock index ySbIdxL, the set (S) of luma subblocks is: S 0 =(xSbIdxL,ySbIdxL)、 S 1 =(xSbIdxL,ySbIdxL+(SubHeightC-1))、 S 2 =(xSbIdxL+(SubWidthC-1),ySbIdxL), or S 3 =(xSbIdxL+(SubWidthC-1),ySbIdxL+(SubHeightC-1)) Includes one or more subblocks indexed by The aforementioned Luma subblock index S 0 This is represented by the horizontal subblock index xSbIdxL and the vertical subblock index ySbIdxL, The aforementioned Lumablock Index S 1 This is represented by the horizontal subblock index xSbIdxL and the vertical subblock index ySbIdxL+(SubHeightC-1), The aforementioned Lumablock Index S 2 This is represented by the horizontal subblock index xSbIdxL+(SubWidthC-1) and the vertical subblock index ySbIdxL, The aforementioned Lumablock Index S 3 The apparatus according to any one of claims 27 to 33, wherein the subblock index is represented by the horizontal subblock index xSbIdxL+(SubWidthC-1) and the vertical subblock index ySbIdxL+(SubHeightC-1).

35. The set (S) of the Lumasa subblock is, S 0 The first Luma subblock is indexed by =(xSbIdxL,ySbIdxL), S 1 The second Luma subblock is indexed by =(xSbIdxL+(SubWidthC-1),ySbIdxL+(SubHeightC-1)) and Includes, The aforementioned luma block position S 0 This is represented by the horizontal subblock index xSbIdxL and the vertical subblock index ySbIdxL, The aforementioned luma block position S 1 The apparatus according to any one of claims 27 to 34, wherein the subblock index is represented by the horizontal subblock index xSbIdxL+(SubWidthC-1) and the vertical subblock index ySbIdxL+(SubHeightC-1).

36. When the chroma format is 4:4:4, the set (S) includes one luma subblock in the same position as the chroma subblock, When the chroma format is 4:2:2, the set (S) includes two chroma subblocks that are horizontally adjacent to each other. The apparatus according to any one of claims 27 to 35, wherein, when the chroma format is 4:2:0, the set (S) includes two diagonally aligned chroma subblocks.

37. When there are more than one Luma subblock in the set (S), the motion vector derivation module is, To generate an averaged Luma motion vector, the motion vectors of the Luma subblocks in the set S are averaged, The apparatus according to any one of claims 27 to 36, configured to derive the motion vector for the chroma subblock based on the averaged luma motion vector.

38. The motion vector derivation module described above is: The horizontal component of the motion vector of the luma subblock in the set S is averaged, or The apparatus according to claim 37, configured to average the vertical component of the motion vector of the luma subblock in the set S.

39. The motion vector derivation module is as follows, that is, mvAvgLX=Σ i mvLX[S i x ][S i y ] When mvAvgLX[0] is greater than or equal to 0, then mvAvgLX[0] = (mvAvgLX[0] + (N>>1)-1)>>log2(N), Otherwise, mvAvgLX[0]=-((-mvAvgLX[0]+(N>>1)-1)>>log2(N)) When mvAvgLX[1] is greater than or equal to 0, then mvAvgLX[1] = (mvAvgLX[1] + (N>>1)-1)>>log2(N), Otherwise, mvAvgLX[1]=-((-mvAvgLX[1]+(N>>1)-1)>>log2(N)) It is configured to average the motion vectors of the luma subblocks in the set S in order to generate an averaged luma motion vector, mvAvgLX is the motion vector resulting from the averaging, mvAvgLX[0] is the horizontal component of the resulting motion vector mvAvgLX, and mvAvgLX[1] is the vertical component of the resulting motion vector mvAvgLX. S i x and S i y These are the horizontal and vertical indices of the subblock Si in the set (S) of Luma subblocks in the motion vector array, mvLX[S i x ][S i y ] is index S i x and S i y The apparatus according to claim 37 or 38, wherein the motion vector of the luma subblock is such that N is the number of elements in the set (S) of the luma subblock, log2(N) represents the base 2 logarithm of N, where the number 2 is raised to a power to obtain the value N, and ">>" represents a right arithmetic shift.

40. The motion vector derivation module is as follows, that is, mvAvgLX=mvLX[(xSbIdx>>(SubWidthC-1)<<(SubWidthC-1))][(ySbIdx>>(SubHeightC-1)<<(SubHeightC-1))]+mvLX[(xSbI dx>>(SubWidthC-1)<<(SubWidthC-1))+(SubWidthC-1)][(ySbIdx>>(SubHeightC-1)<<(SubHeightC-1))+(SubHeightC-1)] If mvAvgLX[0]>=0, mvAvgLX[0]=(mvAvgLX[0]+1-(mvAvgLX[0]>=0))>>1 If mvAvgLX[1] >= 0, mvAvgLX[1]=(mvAvgLX[1]+1-(mvAvgLX[1]>=0))>>1 It is configured to average the motion vectors of the luma subblocks in the set S in order to generate an averaged luma motion vector, The apparatus according to any one of claims 37 to 39, wherein mvAvgLX[0] is the horizontal component of the averaged motion vector mvAvgLX, mvAvgLX[1] is the vertical component of the averaged motion vector mvAvgLX, SubWidthC and SubHeightC represent the horizontal and vertical chroma scaling factors, respectively, xSbIdx and ySbIdx represent the horizontal subblock index and the vertical subblock index, respectively, for the chroma subblock in the set (S), and "<<" is a left arithmetic shift and ">>" is a right arithmetic shift.

41. The aforementioned decision module is The apparatus according to any one of claims 27 to 40, configured to determine the horizontal and vertical chroma scaling factors based on a mapping between the chroma format information and the pairs of horizontal and vertical chroma scaling factors.

42. The apparatus according to any one of claims 27 to 41, further comprising a motion compensation module configured to generate predictions of the chroma subblock based on the determined motion vector.

43. The apparatus according to any one of claims 27 to 42, wherein the chroma format includes one of the YUV4:2:2 format, YUV4:2:0 format, or YUV4:4:4 format.