Motion Vector (MV) Constraints and Transform Constraints in Video Coding
By imposing constraints on motion vectors and transformations in video encoding, the problems of large amount of video data and low video quality in the prior art are solved, the memory bandwidth requirement for affine transformation prediction is reduced, and high-quality video transmission is achieved.
Patent Information
- Application Number
- CN202110206726.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-24
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2041-02-24
AI Technical Summary
The existing video encoding technology is difficult to improve video quality while reducing the amount of video data, and the affine transformation prediction mode increases the memory bandwidth requirement.
By imposing constraints on motion vectors (MVs) and transforms in video encoding, the memory bandwidth required to perform affine transform predictions is reduced. Specific measures include generating a list of candidate MVs, selecting the final MV, and imposing a threshold constraint on the final MV or transform to obtain the constraint MV.
It effectively reduces the memory bandwidth required for affine transformation prediction in video encoding, improves video quality, and reduces the amount of data, meeting the needs of high-quality video transmission.
Smart Images

Figure CN114979627B_ABST
Abstract
Description
Background Art
[0001] Video uses a relatively large amount of data, so the amount of bandwidth used for video transmission is relatively large. However, many networks operate at or near their bandwidth capacity. In addition, customers demand high video quality, which requires the use of more data. Therefore, there is a need to reduce the amount of data used by video and improve video quality. One solution is to compress the video during the encoding process and decompress the video during the decoding process. Improving compression and decompression technology is a focus of research and development. Summary of the invention
[0002] In one embodiment, the present invention includes a device comprising: a memory; and a processor, coupled to the memory and used to: obtain candidate MVs corresponding to neighboring blocks adjacent to a current block in a video frame, generate a candidate list of the candidate MVs, select a final MV from the candidate list, and impose constraints on the final MV or the transformation to obtain a constrained MV. In some embodiments, the constraint stipulates that: a first absolute value of a first difference between a first final MV and a second final MV is less than or equal to a first number, and the first number is determined by the threshold and the width of the current block; the constraint further stipulates that: a second absolute value of a second difference between a third final MV and a fourth final MV is less than or equal to a second number, and the second number is determined by the threshold and the height of the current block; the constraint stipulates that: a first square root of the first number is less than or equal to a second number, wherein the first number is determined by the first final MV, the second final MV, the third final MV, the fourth final MV and the width of the current block, and the second number is determined by the width and the threshold; the constraint further stipulates that: a second square root of the third number is less than or equal to a fourth number, wherein the third number is determined by the fifth final MV, the sixth final MV, the seventh final MV, the eighth final MV and the height of the current block, and the fourth number is determined by the height and the threshold; the constraint stipulates that: the first number is less than or equal to the second number, wherein the first number is determined by the The first width of the transform block, the length of the interpolation filter and the first height of the transformed block are determined, and the second number is determined by a threshold, a second width of the current block and a second height of the current block; the threshold is for single prediction or dual prediction; the constraint stipulates that: the number is less than or equal to the threshold, wherein the number is directly proportional to the first memory bandwidth and the second memory bandwidth, and the number is indirectly proportional to the width of the current block and the height of the current block; the constraint stipulates that: the first difference between the first width of the transformed block and the second width of the current block is less than or equal to the threshold, and the constraint also stipulates that: the second difference between the first height of the transformed block and the second height of the current block is less than or equal to the threshold; the processor is also used to: calculate the MVF based on the constrained MV; perform MCP based on the MVF to generate a prediction block for the current block; and encode the MV index; the processor is also used to: decode the MV index; calculate the MVF based on the constrained MV; and perform MCP based on the MVF to generate a prediction block for the current block.
[0003] In another embodiment, the present invention includes a method comprising: obtaining candidate MVs corresponding to neighboring blocks adjacent to a current block in a video frame, generating a candidate list of the candidate MVs, selecting a final MV from the candidate list, and applying constraints to the final MV or the transformation to obtain a constrained MV.
[0004] In another embodiment, the present invention includes a device comprising: a memory; and a processor, which is coupled to the memory and is used to: obtain candidate MVs corresponding to neighboring blocks adjacent to a current block in a video frame, generate a candidate list of the candidate MVs, impose constraints on the candidate MVs or transforms to obtain constrained MVs, and select a final MV from the constrained MVs. In some embodiments, the constraint stipulates that a first absolute value of a first difference between a first final MV and a second final MV is less than or equal to a first quantity, the first quantity is determined by the threshold and the width of the current block, and the constraint further stipulates that a second absolute value of a second difference between a third final MV and a fourth final MV is less than or equal to a second quantity, the second quantity is determined by the threshold and the height of the current block; the constraint stipulates that a first square root of the first quantity is less than or equal to a second quantity, wherein the first quantity is determined by the first final MV, the second final MV, the third final MV, the fourth final MV and the width of the current block, and the second quantity is determined by the width and the threshold; the constraint further stipulates that a second square root of the third quantity is less than or equal to a fourth quantity, wherein the third quantity is determined by the fifth final MV, the sixth final MV, the seventh final MV, the eighth final MV and the height of the current block, and the fourth quantity is determined by the height and the threshold; the constraint stipulates that the first quantity is less than or equal to the second quantity, wherein the The first number is determined by a first width of a transformed block, a length of an interpolation filter, and a first height of the transformed block, and the second number is determined by a threshold, a second width of the current block, and a second height of the current block; the constraint stipulates that the number is less than or equal to the threshold, wherein the number is directly proportional to a first memory bandwidth and a second memory bandwidth, and the number is indirectly proportional to a width of the current block and a height of the current block; the constraint stipulates that a first difference between the first width of the transformed block and the second width of the current block is less than or equal to the threshold, and the constraint further stipulates that a second difference between the first height of the transformed block and the second height of the current block is less than or equal to the threshold; the processor is further configured to: calculate an MVF based on the constrained MV; perform an MCP based on the MVF to generate a prediction block for the current block; and encode an MV index; the processor is further configured to: decode an MV index; calculate an MVF based on the constrained MV; and perform an MCP based on the MVF to generate a prediction block for the current block.
[0005] Any of the above embodiments can be combined with any other of the above embodiments to create a new embodiment. These and other features will be more clearly understood in the following detailed description taken in conjunction with the accompanying drawings and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] For a more thorough understanding of the present invention, reference is now made to the following brief description taken in conjunction with the accompanying drawings and detailed description, wherein like reference numerals represent like parts.
[0007] Figure 1 A schematic diagram of the coding system.
[0008] Figure 2 Schematic diagram of the MV of the current block.
[0009] Figure 3 Schematic diagram of the transformation for the current block.
[0010] Figure 4 is a flowchart of an encoding method according to an embodiment of the present invention.
[0011] Figure 5 is a flowchart of a decoding method according to an embodiment of the present invention.
[0012] Figure 6 is a flowchart of an encoding method according to another embodiment of the present invention.
[0013] Figure 7 is a flowchart of a decoding method according to another embodiment of the present invention.
[0014] Figure 8 is a schematic diagram of a device according to an embodiment of the present invention.
[0015] Fig. 9 Schematic diagram of the transformation for the current block. DETAILED DESCRIPTION
[0016] First, it should be understood that although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods may be implemented using any number of techniques, whether currently known or in existence. The present invention should in no way be limited to the illustrative implementations, drawings, and techniques described below, including the exemplary designs and implementations illustrated and described herein, but may be modified within the scope of the appended claims and their full scope of equivalents.
[0017] The following abbreviations and acronyms apply:
[0018] AF_inter: affine interpolation
[0019] AF_merge: affine merge
[0020] ASIC: application-specific integrated circuit
[0021] CPU: Central Processing Unit
[0022] CTU: coding tree unit
[0023] DSP: digital signal processor
[0024] EO: electrical-to-optical
[0025] FPGA: field-programmable gate array
[0026] ITU: International Telecommunication Union
[0027] ITU-T: ITU Telecommunication Standardization Sector
[0028] LCD: Liquid crystal display
[0029] MB: memory bandwidth
[0030] MCP: motion compensation prediction
[0031] MV: motion vector
[0032] MVF: motion vector field
[0033] NB: neighboring block
[0034] OE: optical-to-electrical
[0035] PPS: Picture Parameter Set
[0036] RAM: Random-access memory
[0037] RF: radio frequency
[0038] ROM: read-only memory
[0039] RX:receiver unit
[0040] SPS: sequence parameter set
[0041] SRAM: static RAM
[0042] TCAM: ternary content-addressable memory
[0043] TH:threshold
[0044] TX: transmitter unit.
[0045] Figure 1 1 is a schematic diagram of an encoding system 100. The encoding system 100 includes a source device 110, a medium 150, and a target device 160. The source device 110 and the target device 160 are mobile phones, tablet computers, desktop computers, laptop computers, or other suitable devices. The medium 150 is a local network, a wireless network, the Internet, or other suitable media.
[0046] Source device 110 includes video generator 120, encoder 130 and output interface 140. Video generator 120 is a camera or other device suitable for generating video. Encoder 130 may be referred to as a codec. The encoder performs encoding according to a set of rules such as "High Efficiency Video Coding" in ITU-T H.265 ("H.265") released in December 2016, which are incorporated by introduction. Output interface 140 is an antenna or other component suitable for transmitting data to target device 160. Alternatively, video generator 120, encoder 130 and output interface 140 may be in any suitable combination of devices.
[0047] The target device 160 includes an input interface 170, a decoder 180, and a display 190. The input interface 170 is an antenna or other component suitable for receiving data from the source device 110. The decoder 180 may also be referred to as a codec. The decoder 180 performs decoding according to a set of rules such as those described in H.265. The display 190 is an LCD screen or other component suitable for displaying video. Alternatively, the input interface 170, the decoder 180, and the display 190 may be in any suitable combination of devices.
[0048] In operation, the video generator 120 in the source device 110 captures the video, the encoder 130 encodes the video to create an encoded video, and the output interface 140 transmits the encoded video to the target device 160 through the medium 150. The source device 110 stores the video or the encoded video locally, or the source device 110 instructs the video or the encoded video to be stored on another device. The encoded video includes data defined at different levels such as slices and blocks. A slice is a spatially distinct area in a video frame that the encoder 130 encodes separately from any other area in the video frame. A block is a group of pixels arranged in a rectangle, such as an 8-pixel x 8-pixel square. Blocks are also called units or coding units. Other levels include regions, CTUs, and coding blocks. In the target device 160, the input interface 170 receives the encoded video from the source device 110, the decoder 180 decodes the encoded video to obtain a decoded video, and the display 190 displays the decoded video. Decoder 180 may decode the encoded video in the reverse manner in which encoder 130 encoded the video. Target device 160 stores the encoded video or the decoded video locally, or target device 160 directs that the encoded video or the decoded video be stored on another device.
[0049] Coding and decoding are collectively referred to as coding. Coding is divided into intra-frame coding and inter-frame coding, which are also called intra-frame prediction and inter-frame prediction respectively. Intra-frame prediction implements spatial prediction to reduce spatial redundancy in video frames. Inter-frame prediction implements temporal prediction to reduce temporal redundancy between consecutive video frames. MCP is one type of inter-frame prediction.
[0050] In October 2015, Huawei Technologies Co., Ltd. proposed "Affine transform prediction for next generation video coding", which was incorporated by introduction and described two MCP coding modes that simulate affine transforms: AF_inter and AF_merge. In this case, the affine transform is a way of transforming a first video block or other unit into a second video block or other unit while maintaining lines and parallelism to some extent. Therefore, AF_inter and AF_merge simulate translation, rotation, scaling, shear mapping and other features. However, both AF_inter and AF_merge increase the required memory bandwidth. Memory bandwidth is the rate at which data is read or stored from memory, usually in bytes per second. In this case, memory bandwidth refers to the rate at which samples are read from memory during the MCP encoding process. Since AF_inter and AF_merge increase the required memory bandwidth, it is necessary to reduce this memory bandwidth.
[0051] The present invention discloses embodiments of MV constraints and transform constraints in video coding. These embodiments constrain MV or transform to reduce the memory bandwidth required for executing AF_inter and AF_merge. These embodiments are applicable to transforms with two, three, four or more control points. To implement the constraints, encoders and decoders use thresholds. The encoders and decoders store the thresholds as static default values, or the encoder dynamically indicates the thresholds to the decoder in SPS, PPS, slice headers or other appropriate forms.
[0052] For a bitstream that complies with the standard, if the coding block selects the AF_inter or AF_merge mode, the motion vectors of its control points must meet the specified constraints; otherwise, the bitstream does not comply with the standard.
[0053] Figure 2 2 is a schematic diagram 200 showing the MV of the current block. Schematic diagram 200 includes a current block 210 and NB 220, NB 230, NB 240, NB 250, NB 260, NB 270, and NB 280. The current block 210 is called the current block because the encoder 130 is currently encoding it. The current block 210 includes a width w, a height h, and an upper left control point and an upper right control point represented by black dots. For example, w and h are 8 pixels. The upper left control point has an MV v 0 , expressed as (v 0x ,v 0y ), where v 0x Yes 0 The horizontal component, v 0y Yes 0 The upper right control point has MV v 1 , expressed as (v 1x ,v 1y ), where v 1x Yes 1 The horizontal component, v 1y Yes 1 The current block 210 may include other control points, such as a control point located at the center of the current block 210. NBs 220 to NB 280 are referred to as NBs because they are adjacent to the current block 210.
[0054] For an affine model with two control points and four parameters, the MVF of the current block 210 is expressed as follows:
[0055]
[0056] v x is the horizontal component of an MV of the entire current block 210; v 1x 、v0x 、w,v 1y and v 0y As described above; x is the horizontal position measured from the center of the current block 210; y is the vertical position measured from the center of the current block 210; v y is the vertical component of an MV of the entire current block 210. For an affine model with three control points and six parameters, the MVF of the current block 210 is expressed as follows:
[0057]
[0058] v x 、v 1x 、v 0x , w, x, h, y, v y 、v 1y and v 0y As mentioned above; 2x is the horizontal component of the lower left control point MV; v 2y is the vertical component of the lower left control point MV. For an affine model with four control points and eight parameters, the MVF of the current block 210 is expressed as follows:
[0059]
[0060] v x 、v 1x 、v 0x , w, x, v 2x , h, y, v x 、v 1y 、v 0y and v 2y As mentioned above; 3x is the horizontal component of the lower right control point MV; v 3y is the vertical component of the lower right control point MV.
[0061] Figure 3 Schematic diagram 300 showing the transformation of a current block. Schematic diagram 300 includes a current block 310 and a transformed block 320. The current block 310 is located in a current video frame, and the transformed block 320 is the transformed current block 310 in a video frame immediately following the current video frame. As shown in the figure, in order to transform from the current block 310 to the transformed block 320, the current block 310 is rotated and scaled at the same time.
[0062] In the current block 310, the positions of the upper left control point, the upper right control point, the lower left control point and the lower right control point in the current block 310 are respectively represented as a coordinate set (x 0 ,y 0 )、(x 1 ,y 1 )、(x 2,y 2 ) and (x 3 ,y 3 ); w is the width; h is the height. In the transformed block 320, the positions of the upper left control point, the upper right control point, the lower left control point and the lower right control point in the transformed block 320 are respectively represented by the coordinate set (x 0 ',y 0 ')、(x 1 ',y 1 ')、(x 2 ',y 2 ') and (x 3 ',y 3 '); w' is the width; h' is the height. The motion vector v is represented by the coordinate set (v 0 ,v 1 v 2 ,v 3 ), and describes the motion of the current block 310 to the transformed block 320. Therefore, the control point of the transformed block 320 can be represented by the motion vector v and the control point of the current block 310 as follows:
[0063] (x 0 ',y 0 ')=(x 0 +vx 0 ,y 0 +vy 0 )
[0064] (x 1 ',y 1 ')=(x 1 +vx 1 ,y 1 +vy 1 )
[0065] (x 2 ',y 2 ')=(x 2 +vx 2 ,y 2 +vy 2 )
[0066] (x 3 ',y 3 ')=(x 3 +vx 3 ,y 3 +vy 3 ). (4)
[0067] The size of the transformed block 320 can be expressed as:
[0068] w'=max(x 0 ',x 1 ',x2 ',x 3 ')–min(x 0 ',x 1 ',x 2 ',x 3 ')+1
[0069] h'=max(y 0 ',y 1 ',y 2 ',y 3 ')–min(y 0 ',y 1 ',y 2 ',y 3 ',)+1. (5)
[0070] The function max() selects the maximum value from its operands, and the function min() selects the minimum value from its operands. Since the positions of the control points of the current block 310 are known, v describes the positions of the control points of the transformed block 320. If v is a fraction, then v points to a subsample position that does not exist in the reference frame. In this case, the encoder 130 uses an interpolation filter with integer samples as input to generate the subsample positions, and T' is the length of the interpolation filter in affine mode, which is equal to the number of integer samples. For example, T' is 2, 4, or 8. Finally, the memory bandwidth MB required to encode the current block 310 can be expressed as:
[0071] MB=(w'+T'–1)*(h'+T'–1). (6)
[0072] w', T' and h' are as described above.
[0073] In particular, if the current block 310 is a four-parameter affine coding block, and its control points are the upper left control point and the upper right control point; or it is a six-parameter affine coding block, and its control points are the upper left control point, the upper right control point and the lower left control point, then the memory bandwidth MB required for encoding the current block 310 is calculated as follows:
[0074] a) Calculate the variables dHorX, dVerX, dHorY and dVerY:
[0075] dHorX=(vx 1 -vx 0 )<<(7-Log(w)) (7)
[0076] dHorY=(vy 1 -vy 0 )<<(7-Log(w)) (8)
[0077] If the current block 310 is a six-parameter affine coded block:
[0078] dVerX=(vx 2 -vx 0 )<<(7-Log(h)) (9)
[0079] dVerY=(vy 2 -vy 0 )<<(7-Log(h)) (10)
[0080] otherwise:
[0081] dVerX=XdHorY (11)
[0082] dVerY=dHorX (12)
[0083] b) If the current block 310 is a four-parameter affine coded block, calculate the motion vector vx of the lower left control point 2 ,vy 2 :
[0084] vx 2 =Clip3(-131072,131071,Rounding((vx 0 <<7)+dVerX*h,7)) (13)
[0085] vy 2 =Clip3(-131072,131071,Rounding((vy 0 <<7)+dVerY*h,7)) (14)
[0086] c) Calculate the motion vector vx of the lower right control point 3 ,vy 3 :
[0087] vx 3 =Clip3(-131072,131071,Rounding((vx 0 <<7)+dHorX*w+dVerX*h,7)) (15)
[0088] vx 3 =Clip3(-131072,131071,Rounding((vy 0 <<7)+dHorY*w+dVerY*h,7)) (16)
[0089] d) Calculate the following variables:
[0090] max_x=max((vx 0 >>4),max((vx 1>>4)+w,max((vx 2 >>4),(vx 3 >>4)+w))) (17)
[0091] min_x=min((vx 0 >>4),min((vx 1 >4)+w,min((vx 2 >>4),(vx 3 >>4)+w))) (18)
[0092] max_y=max((vy 0 >>4),max((vy 1 >>4),max((vy 2 >>4)+h,(vy 3 >>4)+h))) (19)
[0093] min_y=min((vy 0 >>4),min((vy 1 >>4),min((vy 2 >>4)+h,(vy 3 >>4)+h))) (20)
[0094] e) Calculate the size of the transform block:
[0095] w'=max_x–min_x (21)
[0096] h'=max_y–min_y (22)
[0097] f) Calculate MB:
[0098] MB=(w'+T'-1)*(h'+T'-1) (23)
[0099] In the above formula, Clip3 is defined as:
[0100]
[0101] It means that the input operand x is constrained to be between the interval [i,j].
[0102] Log(x) is defined as:
[0103] Log(x)=log 2 x........................................(25)
[0104] It takes the base 2 logarithm of the input operand x.
[0105] Rounding is defined as:
[0106]
[0107] Indicates that the input x is rounded based on the right shift number s.
[0108] Figure 4 4 is a flowchart showing an encoding method 400 according to an embodiment of the present invention. The encoder 130 may perform the method 400 when performing AF_inter or AF_merge. At step 410, a candidate MV is obtained. In inter-frame prediction, the encoder 130 obtains the candidate MV from the MVs of NB 220 to NB 280, and the MVs of NB 220 to NB 280 are already known when the encoder 130 encodes the current block 210. Specifically, for AF_inter, the encoder 130 regards the MVs in NB 250, NB 260, and NB 240 in this order as v 0 The encoder 130 regards the MVs in NB 270 and NB 280 as candidate MVs in this order. 1 For AF_merge, the encoder 130 considers the MVs in NB 230, NB 270, NB 280, NB 220, and NB 250 as candidate MVs in this order. 0 and v 1 Candidate MV for .
[0109] At step 420, a candidate list of candidate MVs is generated. To this end, the encoder 130 adopts known rules. For example, the encoder 130 removes duplicate candidate MVs and fills unused candidate list gaps with zero MVs.
[0110] At step 430, a final MV is selected from the candidate list. The encoder 130 is an affine model with two control points and four parameters. 0 and v 1 Select the final MV; for v in an affine model with three control points and six parameters 0 、v 1 and v 2 Select the final MV; or v in an affine model with four control points and eight parameters 0 、v 1 、v 2 and v 3 The final MV is selected. To this end, the encoder 130 uses known rules. For example, the encoder uses a rate-distortion cost check. For AF_inter, the encoder 130 may also perform motion estimation to obtain the final MV.
[0111] At step 440, it is determined whether the final MV satisfies the constraint conditions. The encoder 130 directly constrains the final MV using a set of MV constraints described below, or the encoder 130 constrains the transformation using a set of transformation constraints described below, thereby indirectly constraining the final MV. By applying the constraints, the encoder 130 calculates v according to the affine model used when constraining the final MV. 0x 、v 0y 、v 1x 、v 1y 、v 2x 、v 2y 、v 3x and v 3y or the encoder 130 calculates w', h', and T' when constraining the transform. Constraining the MV reduces memory bandwidth compared to the final MV. If the encoder 130 cannot constrain the final MV, the encoder 130 can use one of two alternatives. In the first alternative, the encoder 130 skips the remaining steps of method 400 and implements a coding mode other than AF_inter or AF_merge. In the second alternative, the encoder 130 modifies the final MV until the encoder 130 is able to obtain a final MV that satisfies the constraints. Alternatively, steps 410, 420, 430 can be replaced by a single step of determining the MV, and step 440 can be replaced by a step of applying constraints to the MV or transform to obtain a constrained MV.
[0112] At step 450, the MVF is calculated based on the MV that satisfies the constraints. The encoder 130 calculates the MVF according to equation set (1), (2) or (3). At step 460, MCP is performed based on the MVF to generate a prediction block for the current block. The prediction block includes a prediction value for each pixel in the current block 210.
[0113] Finally, in step 470, the MV index is encoded. The encoder 130 encodes the MV index as part of the encoded video. The MV index indicates the order of the candidate MV in the candidate list. For AF_inter, the encoder 130 can also encode the MV difference, which indicates the difference between the final MV and the candidate MV.
[0114] Figure 5 1 is a flowchart showing a decoding method 500 according to an embodiment of the present invention. The decoder 180 may perform the method 500 when executing AF_inter or AF_merge. At step 510, the MV index is decoded. The MV index may be Figure 4 For AF_inter, encoder 130 may also decode the MV difference. Steps 520, 530, 540, 550, 560, and 570 are respectively the same as those in step 470. Figure 4410, 420, 430, 440, 450, and 460 are similar to those in AF_inter. For AF_inter, encoder 130 may obtain the final MV and add the MV difference to the candidate MV. Alternatively, steps 520, 530, and 540 may be replaced by a single step of determining the MV, and step 550 may be replaced by a step of applying constraints to the MV or transform to obtain a constrained MV.
[0115] Figure 6 6 is a flowchart showing an encoding method 600 according to another embodiment of the present invention. The encoder 130 may perform the method 600 when performing AF_merge. Figure 4 The method 400 in is similar. Specifically, Figure 6 Step 610, step 620, step 650, step 660, and step 670 in Figure 4 410, 420, 450, 460, and 470 are similar to those in step 410, 420, 450, 460, and 470. However, compared with step 430 and step 440, step 630 and step 640 are opposite. In addition, unlike method 400 which is applicable to both AF-inter and AF-merge, method 600 is applicable to AF-merge.
[0116] Figure 7 FIG. 7 is a flowchart showing a decoding method 700 according to another embodiment of the present invention. The decoder 180 may perform the method 700 when performing AF_merge merging. At step 710, the MV index is decoded. The MV index may be Figure 6 Step 720, step 730, step 740, step 750, step 760, and step 770 are respectively the same as Figure 6 Step 610, step 620, step 630, step 640, step 650, and step 660 are similar to those in FIG. MV constraint
[0117] In a first embodiment of MV constraints, the encoder 130 uses two control points to impose the following MV constraints on the initial MV:
[0118] |v 1x –v 0x |≤w*TH
[0119] |v 1y –v 0y |≤h*TH. (27)
[0120] v 1x 、v 0x ,w,v 1y and v 0yAs described above; TH is the threshold; h is the height of the current block 210 in pixels, for example 8 pixels. TH is an arbitrary number suitable for providing sufficient MCP but also suitable for reducing memory bandwidth. For example, if only scaling transformation is used to transform the current block 310 into the transformed block 320, and if the maximum scaling factor is 2, then |v 1x –v 0x |+w, i.e., the width of the transformed block 320, should not be greater than 2*w, and TH should be set to 1. The encoder 130 and the decoder 180 store TH as a static default value. Alternatively, the encoder 130 dynamically indicates TH to the decoder 180 in the form of SPS, PPS, slice header, or other appropriate form.
[0121] In the second embodiment of the MV constraint, the encoder 130 uses three control points to impose the following MV constraints on the initial MV:
[0122] |v 1x –v 0x |≤w*TH
[0123] |v 1y –v 0y |≤h*TH
[0124] |v 2x –v 0x |≤w*TH
[0125] |v 2y –v 0y |≤h*TH. (28)
[0126] v 1x 、v 0x ,w,TH,v 1y 、v 0y and h as above; v 2x is the horizontal component of the third control point MV; v 2y is the vertical component of the third control point MV. The third control point is located at any appropriate position of the current block 210, for example, at the center of the current block 210.
[0127] In a third embodiment of MV constraints, the encoder 130 uses two control points to impose the following MV constraints on the initial MV:
[0128]
[0129] v 1x 、v 0x ,w,v 1y 、v 0y and TH as above.
[0130] In a fourth embodiment of MV constraints, the encoder 130 applies the following MV constraints to the initial MV using three control points:
[0131]
[0132] v 1x 、v 0x ,w,v 1y 、v 0y ,TH,v 2y , h and v 2x As mentioned above.
[0133] Transform constraints
[0134] In a first embodiment of transform constraints, encoder 130 applies the following transform constraints:
[0135] (w'+T'–1)*(h'+T'–1)≤TH*w*h. (31)
[0136] w', T', h', w, and h are as described above. TH is an arbitrary number suitable to provide sufficient MCP but also suitable to reduce memory bandwidth. For example, if the maximum memory bandwidth of the sample is defined as the memory access cost of a 4×4 block, then TH is as follows:
[0137]
[0138] T is the length of the interpolation filter in the translation mode. The encoder 130 and the decoder 180 store TH as a static default value, or the encoder 130 dynamically indicates TH to the decoder 180 in SPS, PPS, slice header or other appropriate forms.
[0139] In a second embodiment of transform constraints, the encoder 130 applies the following transform constraints when using uni-prediction:
[0140] (w'+T'–1)*(h'+T'–1)≤TH UNI *w*h.(33)
[0141] Single prediction means that the encoder 130 uses one reference block to determine the prediction value of the current block 210. w', T', h', w and h are as described above. TH UNI is an arbitrary number that is suitable to provide sufficient MCP but also suitable to reduce the memory bandwidth of uni-prediction. For example, if the maximum memory bandwidth of uni-predicted samples is defined as the memory access cost of a 4×4 block, then TH UNI as follows:
[0142]
[0143] Similarly, the encoder 130 applies the following transform constraints when using bi-prediction:
[0144] (w'+T'–1)*(h'+T'–1)≤TH BI *w*h. (35)
[0145] Bi-prediction means that the encoder 130 uses two reference blocks to determine the prediction value of the current block 210. w', T', h', w and h are as described above. TH BI is an arbitrary number suitable to provide sufficient MCP but also suitable to reduce the memory bandwidth of bi-prediction. TH UNI Less than or equal to TH BI For example, if the maximum memory bandwidth of bi-predicted samples is defined as the memory access cost of an 8×4 block, then TH BI as follows:
[0146]
[0147] Different from equation (11), the thresholds in equations (13) and (15) are for uni-prediction and bi-prediction, respectively.
[0148] In the third embodiment of the transform constraint, the encoder 130 uses bi-prediction, so there are two memory bandwidths, as shown below:
[0149] MB 0 =(w 0 '+T'–1)*(h 0 '+T'–1)
[0150] MB 1 =(w 1 '+T'–1)*(h 1 '+T'–1). (37)
[0151] MB 0 The first width w is used 0 ' and the first height h 0 'The memory bandwidth for the first transformation is MB 1 The second width w is used 1 ' and the second height h 1 'Memory bandwidth for the second transform. Encoder 130 uses these two memory bandwidths to apply the following transform constraints:
[0152]
[0153] MB 0 、MB 1 , w, h and TH are as described above.
[0154] In a fourth embodiment of transform constraints, encoder 130 applies the following transform constraints:
[0155] w'–w≤TH
[0156] h'–h≤TH. (39)
[0157] w', w, h', h and TH are as described above.
[0158] In a fifth embodiment of the transform constraint, the encoder 130 determines the position of the transformed block 420 as follows:
[0159] (x 0 ',y 0 ')=(x 0 +vx 0 ,y 0 +vy 0 )
[0160] (x 1 ',y 1 ')=(x 0 +w+vx 1 ,y 0 +vy 1 )
[0161] (x 2 ',y 2 ')=(x 0 +vx 2 ,y 0 +h+vy 2 )
[0162] (x 3 ',y 3 ')=(x 0 +w+vx 3 ,y 0 +h+vy 3 ). (40)
[0163] x 0 ',y 0 ', x 0 , v, y 0 、x 1 ',y 1 ', w, x 1 ,y 1 、x 2 ',y 2 ', x 2 ,h,y 2 、x 3 ',y 3 ', x 3 and 3As described above, encoder 130 uses equation set (17) and equation set (7) to determine w' and h'.
[0164] In a sixth embodiment of the transform constraints, the encoder 130 applies the following transform constraints:
[0165] (w'+T'–1)*(h'+T'–1)≤(w+T–1+w / ext)*(h+T–1+h / ext). (41)
[0166] w', T', h', w and h are as described above. T is the length of the interpolation filter in the translation mode. ext is the ratio of the number of pixels that can be diffused, which can be 1, 2, 4, 8, 16, etc.
[0167] In a seventh embodiment of the transform constraints, the encoder 130 applies the following transform constraints:
[0168] (w'+T'–1)≤(w+T–1+w / ext) (42)
[0169] (h'+T'–1)≤(h+T–1+h / ext) (43)
[0170] w', T', h', w and h are as described above. T is the length of the interpolation filter in the translation mode. ext is the ratio of the number of pixels that can be diffused, which can be 1, 2, 4, 8, 16, etc.
[0171] In the eighth embodiment of the transformation constraint, the unconstrained candidate MV is discarded. The unconstrained candidate MV refers to a candidate MV that is not subject to the above or other constraints. Therefore, the unconstrained candidate MV is a candidate MV retained in its original form.
[0172] In one embodiment, Fig. 9 As shown, according to the motion vector of the affine control point, the motion vectors of its four corner points can be calculated, thereby calculating the number of reference pixels of the entire affine coding block. If the number exceeds the set threshold, the control point motion vector is considered illegal.
[0173] The threshold is set to (width+7+width / 4)*(height+7+height / 4), which allows translational motion compensation to diffuse width / 4 and height / 4 pixels in the horizontal and vertical directions respectively.
[0174] The specific calculation method is as follows
[0175] 1) First, the motion vectors of the lower left corner (mv2) and the lower right corner (mv3) of the current CU are calculated based on the control point motion vector (mv0, mv1) or (mv0, mv1, mv2) of the original CU.
[0176] If there are 2 motion vectors in mvAffine, mv2 is calculated by the following formula:
[0177]
[0178]
[0179] mv3 is calculated by the following formula:
[0180]
[0181] 2) Then, based on the motion vectors of the four vertices, the rectangular area formed by the position coordinates of the current CU in the reference frame is calculated.
[0182]
[0183] The motion vector is shifted right by 4 because the accuracy of the motion vector is 1 / 16.
[0184] 3) Next, the number of reference pixels can be calculated, where 7 is the number of extended pixels required for pixel-by-pixel interpolation.
[0185] memory_access=(max_x–min_x+7)*(max_y–min_y+7)
[0186] 4) If memory_access is greater than (width+7+width / 4)*(height+7+height / 4), the bitstream is illegal.
[0187] In one example, in the derivation of the affine motion unit sub-block motion vector array in Section 9.20 of the text, the following code stream restriction is added:
[0188] 9.20 Derivation of Affine Motion Unit Sub-Block Motion Vector Arrays
[0189] If there are 3 motion vectors in the affine control point motion vector group, the motion vector group is represented as mvsAffine(mv0,mv1,mv2); otherwise (there are 2 motion vectors in the affine control point motion vector group), the motion vector group is represented as mvsAffine(mv0,mv1).
[0190] a) Calculate the variables dHorX, dVerX, dHorY and dVerY:
[0191]
[0192] If there are 3 motion vectors in mvsAffine, then:
[0193]
[0194] Otherwise (2 motion vectors in mvsAffine):
[0195]
[0196] (xE, yE) is the position of the upper left corner sample of the brightness prediction block of the current prediction unit in the brightness sample matrix of the current image. The width and height of the current prediction unit are width and height respectively, and the width and height of each subblock are subwidth and subheight respectively. The subblock where the upper left corner sample of the brightness prediction block of the current prediction unit is located is A, the subblock where the upper right corner sample is located is B, and the subblock where the lower left corner sample is located is C.
[0197] The code stream that complies with this standard needs to meet the following conditions:
[0198] memory_access<=(width+7+width / 4)*(height+7+height / 4)
[0199] Among them, memory_access is calculated in the following way:
[0200] If there are 2 motion vectors in mvAffine, mv2 is calculated by the following formula:
[0201]
[0202] mv3 is calculated by the following formula:
[0203]
[0204] Then the following variables are calculated:
[0205]
[0206] Next, memory_access is:
[0207] memory_access=(max_x–min_x+7)*(max_y–min_y+7)
[0208] Example 1: A device, comprising:
[0209] Memory; and
[0210] a processor, coupled to the memory and configured to:
[0211] Obtain candidate motion vectors (MVs) corresponding to adjacent blocks adjacent to the current block in the video frame, and generate a candidate list of the candidate MVs,
[0212] selecting a final MV from the candidate list, and
[0213] Constraints are imposed on the final MV or the transformation to obtain a constrained MV.
[0214] Example 2. The device according to claim 1 is characterized in that the constraint stipulates that: a first absolute value of a first difference between a first final MV and a second final MV is less than or equal to a first quantity, and the first quantity is determined by the threshold and the width of the current block; the constraint also stipulates that: a second absolute value of a second difference between a third final MV and a fourth final MV is less than or equal to a second quantity, and the second quantity is determined by the threshold and the height of the current block.
[0215] Example 3. An apparatus according to Example 1, characterized in that the constraint stipulates that: the first square root of the first number is less than or equal to the second number, wherein the first number is determined by the first final MV, the second final MV, the third final MV, the fourth final MV and the width of the current block, and the second number is determined by the width and a threshold.
[0216] Example 4. An apparatus according to Example 1, characterized in that the constraint stipulates that: the first number is less than or equal to the second number, wherein the first number is determined by the first width of the transformed block, the length of the interpolation filter and the first height of the transformed block, and the second number is determined by a threshold, the second width of the current block and the second height of the current block.
[0217] Example 5: The device according to Example 4 is characterized in that the threshold is specific to single prediction or double prediction.
[0218] Example 6. The device according to Example 1 is characterized in that the constraint stipulates that: the number is less than or equal to a threshold, wherein the number is directly proportional to the first memory bandwidth and the second memory bandwidth, and the number is indirectly proportional to the width of the current block and the height of the current block.
[0219] Example 7. An apparatus according to Example 1, characterized in that the constraint stipulates that a first difference between a first width of the transformed block and a second width of the current block is less than or equal to a threshold, and the constraint further stipulates that a second difference between a first height of the transformed block and a second height of the current block is less than or equal to the threshold.
[0220] Example 8. The apparatus according to Example 1, wherein the processor is further configured to:
[0221] Calculating a motion vector field (MVF) based on the constraint MV;
[0222] performing motion compensation prediction (MCP) based on the MVF to generate a prediction block of the current block; and
[0223] Encode the MV index.
[0224] Example 9. The device according to Example 1, wherein the processor is further configured to:
[0225] Decode the MV index;
[0226] Calculating a motion vector field (MVF) based on the constraint MV; and
[0227] Motion compensation prediction (MCP) is performed based on the MVF to generate a prediction block of the current block.
[0228] Example 10: A method, comprising:
[0229] Obtain candidate motion vectors (MVs) corresponding to neighboring blocks adjacent to the current block in the video frame;
[0230] Generating a candidate list of the candidate MVs;
[0231] Selecting a final MV from the candidate list; and
[0232] Constraints are imposed on the final MV or the transformation to obtain a constrained MV.
[0233] Example 11: A device, comprising:
[0234] Memory; and
[0235] a processor, coupled to the memory and configured to:
[0236] Obtain candidate motion vectors (MVs) corresponding to adjacent blocks adjacent to the current block in the video frame, and generate a candidate list of the candidate MVs,
[0237] applying constraints to the candidate MV or transformation to obtain a constrained MV, and
[0238] A final MV is selected from the constrained MVs.
[0239] Example 12. An apparatus according to Example 11, characterized in that the constraint stipulates that: a first absolute value of a first difference between a first final MV and a second final MV is less than or equal to a first number, and the first number is determined by the threshold and the width of the current block, and the constraint further stipulates that: based on the threshold and the height of the current block, a second absolute value of a second difference between a third final MV and a fourth final MV is less than or equal to a second number, and the second number is determined by the threshold and the height of the current block.
[0240] Example 13. An apparatus according to Example 11, characterized in that the constraint stipulates that: the first square root of the first number is less than or equal to the second number, wherein the first number is determined by the first final MV, the second final MV, the third final MV, the fourth final MV and the width of the current block, and the second number is based on the width and a threshold.
[0241] Example 14. An apparatus according to Example 11, characterized in that the constraint stipulates that: the first number is less than or equal to the second number, wherein the first number is determined by the first width of the transformed block, the length of the interpolation filter and the first height of the transformed block, and the second number is determined by a threshold, the second width of the current block and the second height of the current block.
[0242] Example 15. An apparatus according to Example 11, characterized in that the constraint stipulates that: the number is less than or equal to a threshold, wherein the number is directly proportional to the first memory bandwidth and the second memory bandwidth, and the number is indirectly proportional to the width of the current block and the height of the current block.
[0243] Example 16. An apparatus according to Example 11, characterized in that the constraint stipulates that a first difference between a first width of the transformed block and a second width of the current block is less than or equal to a threshold, and the constraint also stipulates that a second difference between a first height of the transformed block and a second height of the current block is less than or equal to the threshold.
[0244] Example 17. The device according to Example 11, wherein the processor is further configured to:
[0245] Calculating a motion vector field (MVF) based on the constraint MV;
[0246] performing motion compensation prediction (MCP) based on the MVF to generate a prediction block of the current block; and
[0247] Encode the MV index.
[0248] Example 18. The device according to Example 11, wherein the processor is further configured to:
[0249] Decode the MV index;
[0250] Calculating a motion vector field (MVF) based on the constraint MV; and
[0251] Motion compensation prediction (MCP) is performed based on the MVF to generate a prediction block of the current block.
[0252] Figure 8 8 is a schematic diagram of an apparatus 800 according to an embodiment of the present invention. The apparatus 800 may implement the disclosed embodiments. The apparatus 800 includes: an ingress port 810 and an RX 820 for receiving data; a processor, a logic unit, a baseband unit or a CPU 830 for processing data; a TX 840 and an egress port 850 for transmitting data; and a memory 860 for storing data. The apparatus 800 may also include an OE component, an EO component or an RF component, which are coupled to the ingress port 810, the RX 820, the TX 840 and the egress port 850 for input and output of optical signals, electrical signals or RF signals.
[0253] The processor 830 is any combination of hardware, middleware, firmware or software. The processor 830 includes any combination of one or more CPU chips, cores, FPGAs, ASICs or DSPs. The processor 830 communicates with the input port 810, RX 820, TX 840, the output port 850 and the memory 860. The processor 830 includes a coding component 870 that implements the disclosed embodiment. Therefore, including the coding component 870 significantly improves the function of the device 800 and allows the device 800 to change to different states. Alternatively, the memory 860 stores the coding component 870 as instructions, and the processor 830 executes these instructions.
[0254] The memory 860 includes any combination of a magnetic disk, a tape drive, or a solid state drive. The device 800 may use the memory 860 as an overflow data storage device for storing programs when the device 800 selects programs to execute, and for storing instructions and data read by the device 800 during the execution of these programs. The memory 860 may be volatile or non-volatile, and may be any combination of ROM, RAM, TCAM, or SRAM.
[0255] In an example embodiment, the apparatus includes: a storage element; and a processor coupled to the storage element and configured to: obtain a candidate MV corresponding to a neighboring block adjacent to a current block in a video frame, generate a candidate list of the candidate MVs, select a final MV from the candidate list, and impose constraints on the final MV or a transformation to obtain a constrained MV.
[0256] Although the present invention has been described in a number of specific embodiments, it should be understood that the disclosed systems and methods may also be embodied in a variety of other specific forms without departing from the spirit or scope of the present invention. The examples of the present invention should be considered illustrative rather than restrictive, and the present invention is not limited to the details given herein. For example, various elements or components may be combined or integrated in another system, or certain features may be omitted or not implemented.
[0257] In addition, the techniques, systems, subsystems, and methods described and illustrated as discrete or separate in the various embodiments may be combined or merged with other systems, components, techniques, or methods without departing from the scope of the present invention. Other items shown or discussed as coupled may be directly coupled or may be electrically, mechanically, or otherwise coupled or communicated indirectly through an interface, device, or intermediate component. Other variations, substitutions, and alternative examples will be apparent to those skilled in the art without departing from the spirit and scope disclosed herein.
Claims
1. A method for motion vector constraint and transform constraint in video coding, It is characterized in that include: Obtaining a motion vector of a control point of a current block, wherein the motion vector of the control point of the current block satisfies a constraint condition, including: a width of the current block, a height of the current block, a width of a transform block of the current block, a height of the transform block, a ratio of the number of diffused pixels, a length of an affine interpolation filter, and a length of a translational interpolation filter satisfying the constraint condition, wherein the transform block is obtained by the motion vector of the control point of the current block; and the ratio of the number of diffused pixels is 4; Calculating a motion vector of each sub-block in the current block according to a motion vector of a control point of the current block that satisfies a constraint condition; A prediction block of each sub-block is obtained according to the motion vector of each sub-block in the current block.
2. The method according to claim 1, It is characterized in that The width of the transform block and the height of the transform block are calculated based on the motion vectors of the four control points of the current block, the width of the current block or the height of the current block.
3. The method according to claim 1, It is characterized in that The constraint condition includes: the first number is less than or equal to the second number, wherein the first number is determined by the width of the transform block, the length of the affine interpolation filter and the height of the transform block, The second number is determined by a ratio of the number of diffused pixels, a length of a translational interpolation filter, a width of the current block, and a height of the current block.
4. The method according to claim 3, It is characterized in that The first quantity is calculated according to the following formula: w '+ T ' – 1) * ( h ' + T ' – 1), where w' is the width of the transform block, h' is the height of the transform block, and T' is the length of the affine interpolation filter.
5. The method according to claim 3 or 4, It is characterized in that The second quantity is calculated according to the following formula: w + T – 1 + w / ext) * (h + T – 1 + h / ext), wherein w is the width of the current block, h is the height of the current block, T is the length of the translational interpolation filter, and ext is the ratio of the number of pixels of the diffusion.
6. The method according to claim 1, It is characterized in that The constraints include: the third number is less than or equal to the fourth number, and the fifth number is less than or equal to the sixth number; wherein the third number is determined by the width of the transformation block and the length of the affine interpolation filter, wherein the fifth number is determined by the height of the transformation block and the length of the affine interpolation filter; the fourth number is determined by the proportion of the number of diffused pixels, the length of the translational interpolation filter, and the width of the current block; and the sixth number is determined by the proportion of the number of diffused pixels, the length of the translational interpolation filter, and the height of the current block.
7. An encoder, It is characterized in that Comprising processing circuitry for performing the method according to any one of claims 1 to 6.
8. A computer program product, It is characterized in that The method comprises a program code for executing the method according to any one of claims 1 to 6.
9. A decoder, It is characterized in that include: one or more processors; A non-transitory computer-readable storage medium, the non-transitory computer-readable storage medium being coupled to the processor and storing a program executed by the processor, and causing the decoder to perform the method according to any one of claims 1 to 6 when the processor executes the program.
10. A decoder, It is characterized in that Comprising processing circuitry for performing the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Motion vector (MV) constraints and transformation constraints in video coding
CN110291790A
Video image decoding method and device
CN111372086A