Motion vector processing methods, image encoding and decoding devices, and computer storage media
By refining the motion vectors during image encoding and decoding, and obtaining and updating the motion vector prediction list, the problem of low efficiency in motion vector processing is solved, achieving more efficient motion vector processing and storage.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG DAHUA TECH CO LTD
- Filing Date
- 2024-11-15
- Publication Date
- 2026-05-26
AI Technical Summary
The processing of motion vectors in existing technologies has limitations, which affects the efficiency of image encoding and decoding.
The original motion vectors of the block to be decoded are obtained, refined to obtain refined information, and the refined information is used to update the motion vector prediction list or store the refined information, including refined adjustment information, motion information before refined adjustment, and motion information after refined adjustment.
It improves the efficiency and effectiveness of motion vector processing, and enhances the accuracy of motion vector prediction and the reliability of storage.
Smart Images

Figure CN122093576A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image encoding and decoding technology, and in particular to a motion vector processing method, an image encoding and decoding device, and a computer storage medium. Background Technology
[0002] Video image data is relatively large, so it is usually necessary to compress the video pixel data (RGB, YUV, etc.). The compressed data is called the video stream. The video stream is transmitted to the user's end via wired or wireless network for decoding and viewing. The entire video encoding process includes block partitioning, prediction, transform, quantization, and encoding.
[0003] During the long-term research and development process, the inventors of this application discovered that there are still certain limitations in the processing of motion vectors, which also affects the efficiency of encoding and decoding to a certain extent. Summary of the Invention
[0004] To address the aforementioned technical problems, this application proposes a motion vector processing method, an image encoding / decoding device, and a computer storage medium.
[0005] To address the aforementioned technical problems, this application proposes a motion vector processing method, which includes:
[0006] Obtain the original motion vector of the block to be decoded;
[0007] The original motion vector is refined to obtain refinement information;
[0008] The motion vector prediction list is updated using the refined information of the block to be decoded, and / or the refined information of the block to be decoded is stored;
[0009] The refined information includes at least one of the following: refined adjustment information, motion information before refined adjustment, and motion information after refined adjustment.
[0010] To address the aforementioned technical problems, this application proposes another motion vector processing method, which includes:
[0011] Obtain the original motion vector of the block to be decoded;
[0012] The original motion vector is refined to obtain refinement information;
[0013] Obtain the predicted motion vectors from the motion vector prediction list;
[0014] Determine whether there exists any predicted motion vector whose distance to the refined information of the block to be decoded is less than a preset distance;
[0015] If not, add the refined information of the block to be decoded to the motion vector prediction list.
[0016] To address the aforementioned technical problems, this application proposes another motion vector processing method, which includes:
[0017] Obtain the original motion vector of the block to be decoded;
[0018] The original motion vector is refined to obtain refinement information;
[0019] Obtain the target motion information storage area corresponding to the level of the refined information of the block to be decoded;
[0020] The refined information of the block to be decoded is stored in the target motion information storage area;
[0021] The refinement information includes at least one of the following: coding unit-level refinement information, prediction unit sub-block-level refinement information, and prediction value refinement sub-block-level refinement information.
[0022] To address the aforementioned technical problems, this application also proposes an image encoding and decoding apparatus, which includes a memory and a processor coupled to the memory; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the motion vector processing method described above.
[0023] To address the aforementioned technical problems, this application also proposes a computer storage medium for storing program data, which, when executed by a computer, is used to implement the aforementioned motion vector processing method.
[0024] Compared with existing technologies, the beneficial effects of this application are: the image encoding / decoding device acquires the original motion vector of the block to be decoded; refines the original motion vector to obtain refinement information; updates the motion vector prediction list using the refinement information of the block to be decoded, and / or stores the refinement information of the block to be decoded; wherein the refinement information includes at least one of the following: refinement adjustment information, motion information before refinement adjustment, and motion information after refinement adjustment. Through the above motion vector processing method, the efficiency and effectiveness of motion vector processing are improved. Attached Figure Description
[0025] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0026] in:
[0027] Figure 1 This is a schematic diagram of an embodiment of the encoding / decoding process provided in this application;
[0028] Figure 2 This is a flowchart of the detailed motion information storage and updating process provided in this application;
[0029] Figure 3 This is a flowchart illustrating the first embodiment of the motion vector processing method provided in this application;
[0030] Figure 4 This is a schematic diagram of an embodiment of CU-level refined information acquisition provided in this application;
[0031] Figure 5 This is a schematic diagram of another embodiment of the CU-level refined information acquisition provided in this application;
[0032] Figure 6 This is a flowchart illustrating the second embodiment of the motion vector processing method provided in this application;
[0033] Figure 7 This is a schematic diagram illustrating the use of PU-level refined information provided in this application for HMVP updates;
[0034] Figure 8 This is a schematic diagram illustrating the use of BIO sub-block level refinement information provided in this application for HMVP updates;
[0035] Figure 9 This is a schematic diagram of an embodiment of airspace motion information storage provided in this application;
[0036] Figure 10 This is a flowchart illustrating the third embodiment of the motion vector processing method provided in this application;
[0037] Figure 11 This is a flowchart illustrating the fourth embodiment of the motion vector processing method provided in this application;
[0038] Figure 12 This is a flowchart illustrating the fifth embodiment of the motion vector processing method provided in this application;
[0039] Figure 13 This is a schematic diagram of an embodiment of the image encoding and decoding apparatus provided in this application;
[0040] Figure 14 This is a schematic diagram of the structure of an embodiment of the computer storage medium provided in this application. Detailed Implementation
[0041] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0042] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the application described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0043] Please see Figure 1 , Figure 1 This is a schematic diagram of an embodiment of the encoding / decoding process provided in this application. Figure 1 As shown, the overall encoding and decoding process is as follows: the encoding process starts from the input video frame and ends at the bitstream, while the decoding process starts from the bitstream and ends at the reconstructed frame. The dashed lines represent the common encoding and decoding processes, while the solid arrows indicating "bitstream → entropy decoding → inverse quantization & inverse transform" represent the decoding-specific processes. The remaining solid arrows represent the encoding-specific processes.
[0044] In video encoding, the most commonly used color encoding methods include YUV and RGB. This application uses the YUV color encoding method. Y represents luminance, which is the grayscale value of the image; U and V (i.e., Cb and Cr) represent chrominance, which describes the color and saturation of the image. Each Y luminance block corresponds to one Cb and one Cr chrominance block, and each chrominance block corresponds to only one luminance block. Taking a 4:2:0 sampling format as an example, an N*M block corresponds to a luminance block of size N*M, and the corresponding two chrominance blocks are both (N / 2)*(M / 2) in size, with the chrominance block being 1 / 4 the size of the luminance block. For a 4:4:4 sampling format, the luminance block and chrominance block are the same size.
[0045] Block partitioning: In video encoding, the input is a series of image frames. However, to encode a single frame, it needs to be divided into several LCUs (largest coding units). Then, each coding unit is recursively divided into CUs (coding units) of different sizes. Video encoding is performed using CUs as units. The smallest coding unit is called the SCU (smallest coding unit).
[0046] Intra-frame / Inter-frame prediction: Generally, the luminance and chrominance signal values of adjacent pixels are quite similar and have a strong correlation. If the number of samples is used directly to represent luminance and chrominance information, there is a lot of spatial redundancy in the data. If redundant data is removed before encoding, the average number of bits per pixel will decrease, which means data compression is performed to reduce spatial redundancy.
[0047] Transformation: After the prediction of the current block is completed, the true value and the predicted value of the current block are subtracted to obtain a residual block. The residual block represents the difference between the true image and the predicted image of the current block. Then, a transformation is performed on the residual block, such as using DCT (Discrete Cosine Transform) or DST (Discrete Sine Transform). Since most images have many flat regions and regions with slow content transformation, and the correlation between adjacent pixels is very strong, the transformation can reduce these correlations and transform the dispersed distribution of image energy in the spatial domain into a relatively concentrated distribution in the transform domain, thus removing spatial redundancy.
[0048] Quantization: Quantization is the process of mapping continuous signal values to multiple discrete amplitude values, achieving a many-to-one mapping of signal values. After transformation, the transform coefficients of the residual data have a large range of values. Quantization can effectively reduce the range of signal values, thus achieving better compression. Because quantization discretizes continuous values into various quantization intervals, it is the root cause of image distortion.
[0049] This application provides a refined vector storage and update encoding / decoding method, which has the same process in the encoding / decoding process.
[0050] Please refer to details. Figure 2 , Figure 2 This is a flowchart illustrating the storage and updating of detailed motion information provided in this application. For example... Figure 2As shown, firstly, the motion information of the CU or PU (Prediction Unit) is acquired. If the refinement conditions are met, refinement information is acquired, and refined prediction samples are obtained based on the refinement information; if the refinement conditions are not met, prediction samples are obtained directly based on the motion information.
[0051] Similarly, for HMVP (History-based Motion Vector Prediction) updates and motion information storage, if the refinement conditions are met, the HMVP is updated and motion information is stored based on the refined information; if the refinement conditions are not met, the HMVP is updated and motion information is stored based on the motion information.
[0052] The refinement techniques in this application include, but are not limited to, DMVR (Decoder-side motion vector refinement) and BIO (Bi-directional Optical flow) techniques. Correspondingly, Figure 6 The refinement conditions are whether to perform DMVR and BIO, and the refinement information is the motion information obtained after DMVR and BIO processing.
[0053] This application mainly involves the process of acquiring detailed information, updating HMVP, and storing motion information, specifically including the following design aspects:
[0054] (1) Acquisition of detailed information. This section mainly describes what detailed information is acquired during the refinement process.
[0055] (2) Refine the storage of motion information. This section mainly describes how to store refined motion information.
[0056] (3) HMVP update design. HMVP update includes: 1) HMVP update with refined motion information; 2) HMVP update of prediction mode containing two or more motion vectors (or two or more prediction units); 3) candidate deduplication design.
[0057] Please refer to details. Figure 3 , Figure 3 This is a flowchart illustrating the first embodiment of the motion vector processing method provided in this application.
[0058] The motion vector processing method of this application is applied to an image encoding and decoding device, which can be a server, a terminal device, or a system in which the server and the terminal device cooperate with each other. Accordingly, the various parts of the image encoding and decoding device, such as various units, sub-units, templates, and sub-templates, can all be set in the server, all in the terminal device, or separately in the server and the terminal device.
[0059] Furthermore, the aforementioned server can be either hardware or software. When the server is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When the server is software, it can be implemented as multiple software programs or software templates, such as software or software templates used to provide distributed servers, or as a single software program or software template; no specific limitations are made here.
[0060] like Figure 3 As shown, the specific steps are as follows:
[0061] Step S11: Obtain the original motion vector of the block to be decoded.
[0062] In this embodiment of the application, the image encoding and decoding device parses the transmitted bitstream information at the decoding end and obtains the original motion vector, i.e., the original MV (Motion Vector), from the bitstream information.
[0063] Step S12: Refine the original motion vector to obtain refinement information.
[0064] In the embodiments of this application, the image encoding and decoding device uses, but is not limited to, DMVR technology and BIO technology to refine the predicted value of the original MV and obtain refined information.
[0065] It should be noted that the following explanation uses DMVR and BIO technologies as possible MV refinement techniques, but is not limited to these two technologies.
[0066] The refined information in this application is defined as including, but not limited to: refined adjustment information (refined offset), motion information before refined adjustment (original MV), and motion information after refined adjustment (refined MV).
[0067] Refinement adjustment information. This includes, but is not limited to, integer-pixel refinement results, sub-pixel refinement results, or a combination of both. It also includes, but is not limited to, the BIO offset vector estimate (vx, vy) and the calculation results based on (vx, vy).
[0068] Motion information before refinement and adjustment. This includes, but is not limited to, PU motion information and calculation results based on PU motion information. Motion information before refinement and adjustment of BIO includes, but is not limited to, PU motion information, calculation results based on PU motion information, and motion information after DMVR refinement and adjustment.
[0069] Motion information after refinement and adjustment. Motion information obtained after processing the refined and adjusted information and the motion information before refinement and adjustment.
[0070] It should be noted that the motion information involved in this application includes forward motion information and / or backward motion information.
[0071] Furthermore, in addition to classifying the detailed information according to the state before and after adjustment, the detailed information in this application can also be classified according to the corresponding level, including but not limited to: detailed information at the CU level, PU level, DMVR sub-block level, BIO sub-block level, etc.
[0072] Specifically, each CU may include at least one PU, and (when performing DMVR or BIO refinement) each PU may include at least one DMVR sub-block or at least one BIO sub-block, and (when performing DMVR and BIO refinement) each DMVR sub-block may include at least one BIO sub-block.
[0073] 1) CU-level refinement information: Each CU obtains one set of refinement information, which can be obtained from the CU's PU / DMVR / BIO sub-blocks. Specific acquisition methods include, but are not limited to, obtaining the refinement information of the last PU / DMVR / BIO sub-block and obtaining the refinement information of the first PU / DMVR / BIO sub-block.
[0074] The specific process for obtaining CU-level refined information is as follows:
[0075] The image encoding / decoding device acquires several prediction unit sub-blocks (PUs) of the block to be decoded (CU); based on the original motion vector (CUMV), it acquires the prediction unit sub-block motion vector (PUMV) of one of the prediction unit sub-blocks (PU); it refines the prediction unit sub-block motion vector (PUMV) to acquire prediction unit sub-block refinement information (PU refinement information); and it uses the prediction unit sub-block refinement information (PU refinement information) as the refinement information of the block to be decoded (CU refinement information); wherein, the prediction unit sub-block refinement information includes prediction unit level refinement information (PU refinement information) and / or refinement sub-block level information (DMVR sub-block / BIO sub-block refinement information).
[0076] In this context, the PU mentioned above refers to either the last PU or the first PU of the CU; similarly, the DMVR sub-block / BIO sub-block mentioned above refers to either the last DMVR sub-block / BIO sub-block or the first DMVR sub-block / BIO sub-block of the CU.
[0077] 2) PU-level refinement information: Each PU obtains one set of refinement information, which can be obtained from the PU itself or from the PU's DMVR / BIO sub-blocks. The methods for obtaining DMVR / BIO sub-block refinement information include, but are not limited to, obtaining the refinement information of the last DMVR / BIO sub-block or obtaining the refinement information of the first DMVR / BIO sub-block.
[0078] The specific process for obtaining PU-level detailed information is as follows:
[0079] The image encoding / decoding device acquires several prediction unit sub-blocks (PUs) of the block to be decoded (CU); based on the original motion vector (CUMV), it acquires the prediction unit sub-block motion vector (PUMV) of each prediction unit sub-block (PU); it acquires several first prediction value refinement sub-blocks (DMVR sub-blocks / BIO sub-blocks) of each prediction unit sub-block (PU); based on the prediction unit sub-block motion vector (PUMV), it acquires the first prediction value refinement sub-block motion vector of one of the first prediction value refinement sub-blocks (DMVR sub-blocks / BIO sub-blocks) of the prediction unit sub-block (PU). (DMVR sub-block / BIO sub-block MV); The refined information of each prediction unit sub-block is obtained according to the following process: The motion vector of the first prediction value refined sub-block (DMVR sub-block / BIO sub-block MV) is refined to obtain the refined information of the first prediction value refined sub-block (refined information of DMVR sub-block / BIO sub-block); The refined information of the first prediction value refined sub-block (refined information of DMVR sub-block / BIO sub-block) or the refined information of the prediction unit sub-block (refined information of PU) is used as the refined information of the prediction unit sub-block (refined information of PU).
[0080] The aforementioned DMVR sub-block / BIO sub-block is either the last DMVR sub-block / BIO sub-block of the PU, or the first DMVR sub-block / BIO sub-block.
[0081] 3) DMVR sub-block level refinement information: Each DMVR sub-block obtains one set of refinement information, which can be the refinement information of the DMVR itself or the refinement information of the BIO sub-block. The methods for obtaining BIO sub-block refinement information include, but are not limited to, obtaining the refinement information of the last BIO sub-block or obtaining the refinement information of the first BIO sub-block.
[0082] The specific process for obtaining DMVR sub-block-level refined information is as follows:
[0083] The image encoding / decoding device acquires several second prediction value refinement sub-blocks (BIO sub-blocks) of the first prediction value refinement sub-block (DMVR sub-block); based on the refinement information of the first prediction value refinement sub-block (refinement information of the DMVR sub-block), it acquires the second prediction value refinement sub-block motion vector (BIO sub-block MV) of one of the second prediction value refinement sub-blocks (BIO sub-blocks); it refines the second prediction value refinement sub-block motion vector to acquire second prediction value refinement sub-block refinement information (BIO sub-block refinement information); and it uses the second prediction value refinement sub-block refinement information (BIO sub-block refinement information) or the first prediction value refinement sub-block refinement information (DMVR sub-block refinement information) as the refinement information of the first prediction value refinement sub-block (DMVR sub-block).
[0084] The aforementioned BIO sub-block is either the last BIO sub-block or the first BIO sub-block within the DMVR sub-block.
[0085] 4) BIO sub-block level refinement information: that is, each BIO sub-block obtains one set of refinement information.
[0086] The specific process for obtaining BIO sub-block level granular information is as follows:
[0087] The image encoding and decoding device uses the refinement information of the second prediction value refinement sub-block (the refinement information of the BIO sub-block) as the refinement information of the second prediction value refinement sub-block (BIO sub-block).
[0088] It is understandable that when DMVR / BIO is not performed, the relevant refined information can be obtained without acquiring it or by acquiring the motion information before refinement.
[0089] Specifically, in a specific embodiment 1 provided in this application, the detailed information scheme is as follows:
[0090] CU-level refinement information acquisition: Acquire the refinement information of the last PU in the CU (motion information before refinement adjustment), i.e., acquire motion information MV3, such as... Figure 4 The shaded area is shown. A 16*16 CU contains four 8*8 PUs, whose motion information is MV0, MV1, MV2, and MV3, respectively.
[0091] Alternatively, CU-level refinement information acquisition: acquire the refinement information (refined motion information after refinement adjustment) of the last DMVR sub-block of the CU, and... Figure 4 The shaded areas are the same (the maximum size of the DMVR sub-block is 16*16; when the width or height of the PU is less than 16, the width or height of the DMVR sub-block is the same as that of the PU). Figure 4 The DMVR sub-block is the same size as the PU.
[0092] Assume that the integer-pixel and fractional-pixel thinning results obtained from DMVR thinning are (intX, intY) and (factX, factY), respectively. Assume that the forward MV of MV3 is (mvX0, mvY0) and the backward MV is (mvX1, mvY1). Then, the forward motion information after MV3 thinning is (mvX0 + intX + factX, mvY0 + intY + factY), and the backward motion information is (mvX1 - intX - factX, mvY1 - intY - factY). It is understandable that vectors are typically represented with 1 / 4 precision; integer-pixel thinning results are represented with integer-pixel precision, while fractional-pixel thinning results are represented with 1 / 4 precision.
[0093] Therefore, the refined forward motion information of MV3 can be changed to (mvX0+(intX<<2)+factX,mvY0+(intY<<2)+factY), and the backward motion information can be changed to (mvX1-(intX<<2)-factX,mvY1-(intY<<2)-factY).
[0094] Alternatively, CU-level refinement information acquisition: acquire the refinement information (refined motion information) of the last BIO sub-block of the CU, as shown in the shaded area below. Figure 5 As shown (BIO sub-blocks are 4*4). Similarly, Figure 5 It is a 16*16 CU containing four 8*8 PUs, whose motion information is MV0, MV1, MV2, and MV3 respectively.
[0095] Assuming the CU has undergone DMVR, the motion information before BIO refinement should be the motion information after DMVR refinement. As mentioned above, the forward motion information after DMVR refinement is (mvX0+(intX<<2)+factX,mvY0+(intY<<2)+factY), and the backward motion information is (mvX1-(intX<<2)-factX,mvY1-(intY<<2)-factY).
[0096] Assuming the offset vector estimate (vx, vy) of the last BIO sub-block is (vxX, vyY), the refined forward motion information of the last BIO sub-block should be (mvX0+(intX<<2)+factX+(vxX>>scale), mvY0+(intY<<2)+factY+(vyY>>scale)), and the backward motion information should be (mvX1-(intX<<2)–factX-(vxX>>scale), mvY1-(intY<<2)–factY-(vyY>>scale)). Here, scale is the scaling factor for the BIO offset vector estimate, which can be any integer; in this embodiment, it is set to 6.
[0097] Additionally, BIO refinement information can be calculated based on the POC (Picture Order Count) distance. Assuming the POC of forward motion information is pocP, the POC of backward motion information is pocL, and the POC of the current image is pocC. Let the estimated forward offset vector of the BIO be (vxP, vyP). One way to calculate it is vxP = (pocC - pocP) / (pocL - pocP) * vxX, vyP = (pocC - pocP) / (pocL - pocP) * vxY. Therefore, the forward motion information after refinement of the last BIO sub-block should be (mvX0 + (intX << 2) + factX + (vxP >> scale), mvY0 + (intY << 2) + factY + (vyP >> scale)), and the backward motion information is (mvX1 - (intX << 2) – factX - ((vxX - vxP) >> scale), mvY1 - (intY << 2) – factY - ((vyY - vyP) >> scale)).
[0098] It is understandable that the process of obtaining detailed information at other PU level / DMVR sub-block level / BIO sub-block level is similar to the above embodiment. For example, at the PU level, detailed information of a certain DMVR sub-block or BIO sub-block in the PU can be obtained; at the DMVR sub-block level, detailed information of the DMVR or its BIO sub-blocks can be obtained; and at the BIO sub-block level, detailed information of the corresponding BIO sub-block can be obtained.
[0099] Step S13: Update the motion vector prediction list using the refinement information of the block to be decoded, and / or store the refinement information of the block to be decoded.
[0100] In this embodiment of the application, the image encoding and decoding device uses the refinement information of the block to be decoded determined in step S12, including but not limited to the refinement information of each stage and the refinement information of each level, to perform HMVP update / motion information storage.
[0101] In this application, an image encoding / decoding device acquires the original motion vector of a block to be decoded; refines the original motion vector to obtain refinement information; updates the motion vector prediction list using the refinement information of the block to be decoded, and / or stores the refinement information of the block to be decoded; wherein the refinement information includes refinement adjustment information, motion information before refinement adjustment, and / or motion information after refinement adjustment. Through the above motion vector processing method, the refined motion information is used to update the motion vector prediction list and store the motion vectors, thereby providing more accurate motion vector candidates.
[0102] Furthermore, the motion vector processing method of this application also relates to HMVP update design. The HMVP update design of this application includes: 1) HMVP update that refines motion information; 2) HMVP update of prediction patterns that includes two or more motion vectors (or two or more prediction units); and 3) candidate deduplication design.
[0103] The motion vector processing method in this application updates the HMVP using refined motion information of the current coding block. This refined information includes, but is not limited to, CU-level refined information, PU-level refined information, DMVR sub-block-level refined information, and BIO sub-block-level refined information. Notably, this application does not limit the number of PUs contained in a CU (unlike existing technologies, which only allow motion information from CUs containing one PU to be used to update the HMVP).
[0104] It is understandable that (1) CU-level and PU-level refinement information are used for HMVP updates: In the current design, the prediction mode used for updating the HMVP candidate list includes only one PU in the CU, so the obtained CU and PU-level refinement information is the same. (2) DMVR sub-block-level and BIO sub-block-level refinement information are used for HMVP updates. The refinement information of each DMVR sub-block and BIO sub-block is used for HMVP updates, so each CU will generate at least one refinement motion information for HMVP updates.
[0105] The maintenance details for Historical Motion Information Prediction (HMVP) are as follows:
[0106] HMVP candidates maintain a candidate list with a maximum length of NumOfHmvpCand (ranging from 0 to 8) using a first-in-first-out (FIFO) approach, where the value of NumOfHmvpCand is determined by the syntax in SPS.
[0107] HMVP Applications: HMVP candidates can be used for constructing candidate lists for AMVP (Adaptive Motion Vector Prediction) and direct / skip methods.
[0108] HMVP maintenance and update: If the current coding block is in inter-frame prediction mode, and is in affine, angular weighted prediction (AWP), motion vector angular prediction (MVAP), enhanced temporal motion vector prediction (ETMVP), or sub-block-based temporal motion vector prediction (SBTMVP) mode, then the HMVP list is not updated; otherwise, the HMVP list needs to be updated. The specific update process is as follows:
[0109] (1) If the MV of the current block does not exist in the HMVP list. If the HMVP list is not full, the length of the HMVP list is increased by 1, and the current MV is added to the end of the HMVP list; otherwise, the first candidate in the HMVP list is removed, all candidates in the list are moved forward by 1 position, and the current MV is added to the end of the list.
[0110] (2) If the MV of the current block already exists in the HMVP list, first remove the candidate at the position corresponding to the current MV in the HMVP list, then move the candidate at the corresponding position to the effective length of the HMVP list forward by 1 position, and finally add the current MV to the last position of the effective length.
[0111] Please continue reading for details. Figure 6 , Figure 6 This is a flowchart illustrating the second embodiment of the motion vector processing method provided in this application.
[0112] like Figure 6 As shown, the specific steps are as follows:
[0113] Step S21: Obtain the predicted motion vectors from the motion vector prediction list.
[0114] Step S22: Determine whether there exists any predicted motion vector whose distance to the refined information of the block to be decoded is less than a preset distance.
[0115] In this application embodiment, considering the increasing amount of refined motion information used for HMVP updates, the length of the HMVP candidate list can be extended or an enhanced deduplication method can be designed to reduce the number of candidates. One such enhanced deduplication method provided in this application is:
[0116] If the candidate to be added points to the same reference image as an existing candidate in the candidate list, and the distance between the reference positions is less than a preset distance, then the candidate is duplicated and will not be added to the candidate list; otherwise, the candidate needs to be added to the candidate list, i.e., proceed to step S23.
[0117] The distances mentioned above include, but are not limited to: horizontal distance, vertical distance, Euclidean distance, and the sum of horizontal and vertical distances.
[0118] It should be noted that during the plagiarism detection process, if the refined motion information is bidirectional prediction, then the plagiarism detection is performed according to the prediction direction. If both directions meet the plagiarism detection conditions, then it is considered a duplicate and the candidate is not added.
[0119] Step S23: Add the refined information of the block to be decoded to the motion vector prediction list.
[0120] In the embodiments of this application, the image encoding and decoding device can add the refined information of the block to be decoded to the motion vector prediction list according to the HMVP maintenance and update scheme of historical motion information prediction (HMVP) in the above technical description. The process will not be described in detail here.
[0121] It should be noted that the deduplication method provided in this application can be used not only for adding candidates during HMVP candidate updates, but also for the construction process of other inter-frame prediction candidate lists, which will not be listed here.
[0122] Specifically, in a specific embodiment 2 provided in this application, the HMVP update scheme is as follows:
[0123] The CU-level refinement information is used for HMVP updates. For example, the CU-level refinement information obtained in Specific Implementation 1 is used to update the HMVP candidate list. This CU-level refinement information is denoted as (mvRX0, mvRY0) for forward motion and (mvRX1, mvRY1) for backward motion.
[0124] Enhanced plagiarism detection method: Assume that there is already 1 candidate in the HMVP candidate list, where the forward motion information is (mvHX0, mvHY0) and the backward motion information is (mvHX1, mvHY1).
[0125] If (abs(mvHX0–mvRX0)>TH||abs(mvHY0–mvRY0)>TH||abs(mvHX1–mvRX1)>TH||abs(mvHY1–mvRY1)>TH), then the candidate is not repeated; otherwise, the candidate is repeated.
[0126] The abs function is a function used to calculate the absolute value of a number.
[0127] Specifically, in a specific embodiment 3 provided in this application, the HMVP update scheme is as follows:
[0128] PU-level refined information (motion information before refinement) is used for HMVP updates, such as... Figure 7As shown, there is a 16*16 CU containing four 8*8 PUs, whose motion information is MV0, MV1, MV2, and MV3 respectively. MV0, MV1, MV2, and MV3 are used to update the HMVP.
[0129] Specifically, in a specific embodiment 4 provided in this application, the HMVP update scheme is as follows:
[0130] BIO sub-block level refinement information is used for HMVP updates, assuming it is the same CU as in the previous embodiment. Figure 8 The image shows the motion information after BIO refinement, MVxy, x = {0, 1, 2, 3}, y = {0, 1, 2, 3}, representing the refined motion information of each BIO sub-block, and this refined motion information is used for HMVP updates.
[0131] Furthermore, the motion vector processing method of this application also involves refining the storage of motion information. The specific details of refining the storage of motion information are as follows:
[0132] Spatial motion vectors are stored in units of size 4*4 (scu), meaning each 4*4 motion vector stores one MV. The storage used to save the current image motion information can be called the motion field, and each CU needs to have relevant motion information at its corresponding position in the motion field.
[0133] like Figure 9 The diagram illustrates how a CU stores motion information. First, a CU can be divided into at least one prediction unit (PU). Then, each PU can be further divided into at least one 4x4 SCU block. Finally, motion prediction values (MVs) are stored in 4x4 blocks, with each PU storing the same MV. A PU is the unit that performs predictions based on MVs. For example, in sub-block-based inter-frame prediction modes such as SBTMVP, MVAP, and ETPMV, the CU can be divided into multiple PUs, and each PU can be assigned a specific MV. Figure 9 This is a schematic diagram of an embodiment of airspace motion information storage provided in this application.
[0134] The storage of temporal motion information is similar to that of the spatial domain, but the smallest unit used to store MV is 16*16. When storing, the MV corresponding to the center position of each 16*16 unit (the boundary may be less than 16*16) is used as the MV of that 16*16 storage unit.
[0135] Please refer to section 10 for details. Figure 10 This is a flowchart illustrating the third embodiment of the motion vector processing method provided in this application.
[0136] like Figure 10 As shown, the specific steps are as follows:
[0137] Step S31: Obtain the target motion information storage area corresponding to the level of the refined information of the block to be decoded.
[0138] Step S32: Store the refined information of the block to be decoded into the target motion information storage area.
[0139] In existing technology, each CU stores the motion information of the PU. Therefore, it actually stores the motion information of the PU, and then the PU is divided into 4*4 storage units, and the motion information of the PU is saved to each storage unit.
[0140] This application stores detailed information at the CU level, PU level, DMVR sub-block level, and BIO sub-block level. That is, the motion information storage area corresponding to each CU / PU / DMVR sub-block / BIO sub-block stores motion information at the CU level, PU level, DMVR sub-block level, and BIO sub-block level.
[0141] Specifically, in a specific embodiment 5 provided in this application, the HMVP update scheme is as follows:
[0142] As shown in Specific Implementation 4, detailed information of each BIO sub-block in the CU is given, and the storage unit is the same size as the BIO sub-block. Therefore, the detailed information of each BIO sub-block is stored in the corresponding storage unit (BIO sub-block motion information storage area).
[0143] Correspondingly, this application also provides corresponding solution application design and syntax design for the processes of obtaining detailed information, HMVP updating, and storing motion information.
[0144] The application design of the solution includes, but is not limited to:
[0145] (1) HMVP update scheme for refining motion information: When the CU supports refining motion information (or prediction samples), the original HMVP update method is replaced by the HMVP update method for refining motion information designed in this application.
[0146] (2) HMVP update scheme for multiple prediction units (CU): When the CU contains multiple PUs (or multiple motion vectors), the HMVP update scheme designed in this application is added.
[0147] (3) Candidate deduplication design scheme: When HMVP is updated, the candidate list of HMVP adopts the candidate deduplication design method designed in this proposal, or when CU supports the refinement of motion information (or prediction samples), its candidate list adopts the candidate deduplication design method designed in this application.
[0148] (4) Refined motion information storage scheme: When the CU supports the refinement of motion information (or prediction samples), the refined motion information storage method designed in this application is used.
[0149] Syntactic design includes, but is not limited to:
[0150] (1) Switch syntax: used to express whether the scheme proposed in this application is enabled in the codec. The switch syntax can be transmitted in syntax structures including but not limited to: Video Parameter Set (VPS), Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Picture Header (PH), coding unit, etc.
[0151] (2) Syntactic coding: Syntactic binarization methods include unary coding, truncated unary coding, truncated Rice coding, signed fixed-length coding, unsigned fixed-length coding, exponential Golomb coding, etc. Syntactic coding methods include, but are not limited to, high-entropy coding, bypass coding, etc. The coding methods are described in the standard text as descriptors. For specific meanings, refer to the corresponding standard text.
[0152] Specifically, in a specific embodiment 6 provided in this application, the application design and syntax design scheme are as follows:
[0153] This embodiment includes an HMVP update scheme for refining motion information, an HMVP update scheme for multiple prediction units (CUs), a candidate deduplication design scheme, and a storage scheme for refining motion information.
[0154] In this embodiment, the switch syntax `sps_enhanced_hmvp_enable` is used to express whether the HMVP update scheme for refining motion information and the HMVP update scheme for the multi-prediction unit (CU) prediction mode are enabled, as shown in the table below. When `sps_enhanced_hmvp_enable = 0`, it means that these schemes are not enabled; when `sps_enhanced_hmvp_enable = 1`, it means that these schemes are enabled.
[0155] The switch syntax `sps_enhanced_same_enable` is used to indicate whether the candidate plagiarism detection scheme is enabled. When `sps_enhanced_same_enable = 0`, the scheme is not enabled; when `sps_enhanced_same_enable = 1`, the scheme is enabled.
[0156] The switch syntax `sps_enhanced_store_enable` is used to indicate whether the finer-grained motion information storage scheme is enabled. When `sps_enhanced_store_enable = 0`, the scheme is disabled; when `sps_enhanced_store_enable = 1`, the scheme is enabled.
[0157]
[0158] Please continue reading. Figure 11 , Figure 11 This is a flowchart illustrating the fourth embodiment of the motion vector processing method provided in this application.
[0159] like Figure 11 As shown, the specific steps are as follows:
[0160] Step S41: Obtain the original motion vector of the block to be decoded.
[0161] Step S42: Refine the original motion vector to obtain refinement information.
[0162] Step S43: Obtain the predicted motion vectors from the motion vector prediction list.
[0163] Step S44: Determine whether there exists any predicted motion vector whose distance to the refined information of the block to be decoded is less than a preset distance.
[0164] Step S45: Add the refined information of the block to be decoded to the motion vector prediction list.
[0165] In the embodiments of this application, the image encoding and decoding device implements the HMVP update design and candidate deduplication design provided in this application based on the existing inter-frame prediction technology, which will not be described in detail here.
[0166] Please continue reading. Figure 12 , Figure 12 This is a flowchart illustrating the fifth embodiment of the motion vector processing method provided in this application.
[0167] like Figure 12 As shown, the specific steps are as follows:
[0168] Step S51: Obtain the original motion vector of the block to be decoded.
[0169] Step S52: Refine the original motion vector to obtain refinement information.
[0170] Step S53: Obtain the target motion information storage area corresponding to the level of the refined information of the block to be decoded.
[0171] Step S54: Store the refined information of the block to be decoded into the target motion information storage area.
[0172] The refinement information includes coding unit-level refinement information, prediction unit-level refinement information, and / or sub-block-level refinement information.
[0173] In the embodiments of this application, the image encoding and decoding device implements the refined motion information storage scheme provided in this application based on the existing inter-frame prediction technology, which will not be described in detail here.
[0174] This application proposes to use refined motion information for HMVP updates and motion vector storage, and can support different block levels, thereby providing more accurate motion vector candidates.
[0175] This application proposes an HMVP update method for multiple prediction units (CUs), which enriches the HMVP candidate list.
[0176] This application designs a more effective method for motion information deduplication, which can ensure that the candidates in the candidate list are more effective.
[0177] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0178] To implement the above motion vector processing method, this application also proposes an image encoding and decoding device, which can be found in the following details. Figure 13 , Figure 13 This is a schematic diagram of an embodiment of the image encoding and decoding apparatus provided in this application.
[0179] The image encoding / decoding device 700 of this embodiment includes a processor 71, a memory 72, an input / output device 73, and a bus 74.
[0180] The processor 71, memory 72, and input / output device 73 are respectively connected to the bus 74. The memory 72 stores program data, and the processor 71 is used to execute the program data to implement the motion vector processing method described in the above embodiments.
[0181] In this embodiment, processor 71 can also be referred to as a CPU (Central Processing Unit). Processor 71 may be an integrated circuit chip with signal processing capabilities. Processor 71 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor, or processor 71 can be any conventional processor.
[0182] This application also provides a computer storage medium; please refer to the following: Figure 14 , Figure 14 This is a schematic diagram of a computer storage medium according to an embodiment of the present application. The computer storage medium 600 stores a computer program 61, which, when executed by a processor, is used to implement the motion vector processing method of the above embodiment.
[0183] When the embodiments of this application are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0184] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A method of motion vector processing, the method comprising: The motion vector processing method includes: Obtain the original motion vector of the block to be decoded; The original motion vector is refined to obtain refinement information; The motion vector prediction list is updated using the refined information of the block to be decoded, and / or the refined information of the block to be decoded is stored; The refined information includes at least one of the following: refined adjustment information, motion information before refined adjustment, and motion information after refined adjustment.
2. The motion vector processing method according to claim 1, characterized in that, The refinement information includes at least one of the following: coding unit-level refinement information, prediction unit sub-block-level refinement information, and prediction value refinement sub-block-level refinement information.
3. The motion vector processing method according to claim 2, characterized in that, The process of refining the original motion vector to obtain refined information includes: Obtain at least one prediction unit sub-block of the block to be decoded; Based on the original motion vector, obtain the motion vector of one of the prediction unit sub-blocks; The motion vectors of the prediction unit sub-blocks are refined to obtain the refined information of the prediction unit sub-blocks; The refined information of the prediction unit sub-block is used as the refined information of the block to be decoded; The prediction unit sub-block refinement information includes at least one of the following: prediction unit-level refinement information and refinement sub-block-level information.
4. The motion vector processing method according to claim 3, characterized in that, The motion vector of one of the prediction unit sub-blocks is the motion vector of the first prediction unit sub-block or the motion vector of the last prediction unit sub-block.
5. The motion vector processing method according to claim 2, characterized in that, The process of refining the original motion vector to obtain refined information includes: Obtain at least one prediction unit sub-block of the block to be decoded; Based on the original motion vector, obtain the motion vector of each prediction unit sub-block; Obtain at least one first prediction value for each prediction unit sub-block to refine the sub-block; Based on the motion vector of the prediction unit sub-block, obtain one of the first prediction values of the prediction unit sub-block, refine the first prediction value of the sub-block, and refine the motion vector of the sub-block. The refinement information of each prediction unit sub-block is obtained according to the following process: the motion vector of the first prediction value refinement sub-block is refined to obtain the refinement information of the first prediction value refinement sub-block; the refinement information of the first prediction value refinement sub-block or the refinement information of the prediction unit sub-block is used as the refinement information of the prediction unit sub-block.
6. The motion vector processing method according to claim 2, characterized in that, The process of refining the original motion vector to obtain refined information includes: Obtain at least one prediction unit sub-block of the block to be decoded; Based on the original motion vector, obtain the motion vector of each prediction unit sub-block; Obtain several first predicted values for each prediction unit sub-block to refine the sub-block; Based on the motion vector of the prediction unit sub-block, obtain one of the first prediction values of the prediction unit sub-block, refine the first prediction value of the sub-block, and refine the motion vector of the sub-block. The motion vector of the first predicted value sub-block is refined to obtain the refinement information of the first predicted value sub-block; The refinement information of the first predicted value refinement sub-block is used as the refinement information of the first predicted value refinement sub-block.
7. The motion vector processing method according to claim 6, characterized in that, The motion vector processing method further includes: Obtain at least one second prediction value refinement sub-block from the first prediction value refinement sub-block; Based on the first predicted value refined sub-block information, the motion vector of the second predicted value refined sub-block of one of the second predicted value refined sub-blocks is obtained; The motion vector of the second predicted value sub-block is refined to obtain the refinement information of the second predicted value sub-block; The step of using the refined information of the first predicted value refined sub-block as the refined information of the first predicted value refined sub-block includes: The second predicted value refinement sub-block refinement information, or the first predicted value refinement sub-block refinement information, is used as the refinement information of the first predicted value refinement sub-block.
8. The motion vector processing method according to claim 7, characterized in that, The motion vector processing method further includes: The refinement information of the second predicted value refinement sub-block is used as the refinement information of the second predicted value refinement sub-block.
9. The motion vector processing method according to claim 5 or 6, characterized in that, The step of refining the motion vector of the second predicted value sub-block to obtain the refined information of the second predicted value sub-block includes: In response to the prediction unit sub-block motion vector being a bidirectional prediction motion vector, the forward and backward image sequence counting information is obtained to refine the second prediction value refined sub-block motion vector, thereby obtaining the second prediction value refined sub-block refinement information.
10. The motion vector processing method according to claim 2, characterized in that, The refinement information of the block to be decoded includes at least two prediction unit-level refinement information; The step of updating the motion vector prediction list using the refined information of the block to be decoded, and / or storing the refined information of the block to be decoded, includes: The motion vector prediction list is updated by traversing the at least two prediction unit-level refinement information and using each prediction unit-level refinement information.
11. The motion vector processing method according to claim 1 or 2, characterized in that, The step of updating the motion vector prediction list using the refined information of the block to be decoded, and / or storing the refined information of the block to be decoded, includes: Obtain the predicted motion vectors from the motion vector prediction list; Determine whether there exists any predicted motion vector whose distance to the refined information of the block to be decoded is less than a preset distance; If not, add the refined information of the block to be decoded to the motion vector prediction list.
12. The motion vector processing method according to claim 11, characterized in that, After determining whether any predicted motion vector is less than a preset distance from the refined information of the block to be decoded, the motion vector processing method further includes: If the distance between the predicted motion vector and the refined information of the block to be decoded is less than a preset distance, the refined information of the block to be decoded is not stored. Alternatively, the corresponding predicted motion vector in the motion vector prediction list can be removed, and the refined information of the block to be decoded can be stored at the end of the motion vector prediction list.
13. The motion vector processing method according to claim 11 or 12, characterized in that, The distance is one of the following: horizontal distance, vertical distance, Euclidean distance, or the sum of horizontal and vertical distances.
14. The motion vector processing method according to claim 1, characterized in that, The refinement information includes at least one of the following: coding unit level refinement information, prediction unit sub-block level refinement information, and prediction value refinement sub-block level refinement information; The step of updating the motion vector prediction list using the refined information of the block to be decoded, and / or storing the refined information of the block to be decoded, includes: Obtain the target motion information storage area corresponding to the level of the refined information of the block to be decoded; The refined information of the block to be decoded is stored in the target motion information storage area.
15. A method of motion vector processing, the method comprising: The motion vector processing method includes: Obtain the original motion vector of the block to be decoded; The original motion vector is refined to obtain refinement information; Obtain the predicted motion vectors from the motion vector prediction list; Determine whether there exists any predicted motion vector whose distance to the refined information of the block to be decoded is less than a preset distance; If not, add the refined information of the block to be decoded to the motion vector prediction list.
16. A method of motion vector processing, the method comprising: The motion vector processing method includes: Obtain the original motion vector of the block to be decoded; The original motion vector is refined to obtain refinement information; Obtain the target motion information storage area corresponding to the level of the refined information of the block to be decoded; The refined information of the block to be decoded is stored in the target motion information storage area; The refinement information includes at least one of the following: coding unit-level refinement information, prediction unit sub-block-level refinement information, and prediction value refinement sub-block-level refinement information.
17. An image coding apparatus, characterized by comprising: The image encoding / decoding device includes a memory and a processor coupled to the memory; The memory is used to store program data, and the processor is used to execute the program data to implement the motion vector processing method as described in any one of claims 1 to 16.
18. A computer storage medium, characterized in that The computer storage medium is used to store program data, which, when executed by the computer, is used to implement the motion vector processing method as described in any one of claims 1 to 16.