Encoding method, decoding method, encoder, decoder, and storage medium
Patent Information
- Application Number
- PCT/CN2025/085160
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2026-10-01
Smart Images

Figure CN2025085160_01102026_PF_FP_ABST
Abstract
Description
Encoding / decoding methods, encoders, decoders, and storage media Technical Field
[0001] This application relates to the field of video encoding and decoding technology, and in particular to an encoding and decoding method, encoder, decoder, and storage medium. Background Technology
[0002] In video coding standards, in addition to translation, video content usually includes rotation, scaling, stretching and other motions. In order to effectively represent such motions, affine motion compensation technology proposes to use the idea of affine transformation to determine the motion vectors of different pixels within a block by some known motion vectors, so as to avoid further division of the current block.
[0003] However, the affine motion compensation mode in related technologies, when constructing the control point motion vector (CPMV) using adjacent motion information, fails to consider all aspects, resulting in the inability to construct an effective CPMV in some cases and reducing the possibility of constructing an effective CPMV, which is not conducive to improving encoding and decoding performance. Summary of the Invention
[0004] This application provides an encoding / decoding method, encoder, decoder, and storage medium, which can increase the probability of constructing effective affine motion information, thereby improving encoding / decoding performance.
[0005] The technical solution of this application embodiment can be implemented as follows:
[0006] In a first aspect, embodiments of this application provide a decoding method applied to a decoder, the method comprising:
[0007] Determine the motion information of multiple adjacent blocks at the preset position of the current block;
[0008] Based on at least two motion information that meet the availability condition from multiple adjacent blocks, determine the affine motion information of the current block;
[0009] Determine the predicted value of the current block based on the affine motion information of the current block.
[0010] Secondly, embodiments of this application provide an encoding method applied to an encoder, the method comprising:
[0011] Determine the motion information of multiple adjacent blocks at the preset position of the current block;
[0012] Based on at least two motion information that meet the availability condition from multiple adjacent blocks, determine the affine motion information of the current block;
[0013] Determine the predicted value of the current block based on the affine motion information of the current block.
[0014] Thirdly, embodiments of this application provide an encoder, which includes a first determining unit and a first predicting unit, wherein:
[0015] The first determining unit is configured to determine the motion information of multiple adjacent blocks at a preset position of the current block; and to determine the affine motion information of the current block based on at least two motion information of the multiple adjacent blocks that meet the availability conditions.
[0016] The first prediction unit is configured to determine the predicted value of the current block based on the affine motion information of the current block.
[0017] Fourthly, embodiments of this application provide an encoder, which includes a first memory and a first processor, wherein:
[0018] A first memory for storing computer programs that can run on a first processor;
[0019] A first processor is configured to execute the encoding method as described in the second aspect when running a computer program.
[0020] Fifthly, embodiments of this application provide a decoder, which includes a second determining unit and a second predicting unit, wherein:
[0021] The second determining unit is configured to determine the motion information of multiple adjacent blocks at a preset position of the current block; and to determine the affine motion information of the current block based on at least two motion information of the multiple adjacent blocks that meet the availability conditions.
[0022] The second prediction unit is configured to determine the predicted value of the current block based on the affine motion information of the current block.
[0023] Sixthly, embodiments of this application provide a decoder, which includes a second memory and a second processor, wherein:
[0024] The second memory is used to store computer programs that can run on the second processor;
[0025] The second processor is used to execute the decoding method as described in the first aspect when running a computer program.
[0026] In a seventh aspect, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the decoding method as described in the first aspect or the encoding method as described in the second aspect.
[0027] Eighthly, embodiments of this application provide a computer-readable storage medium having a bitstream stored thereon, the bitstream being generated by performing the steps of the encoding method as described in the second aspect.
[0028] In a ninth aspect, embodiments of this application provide a computer program product, including a computer program or instructions, which, when executed by a processor, implement the decoding method as described in the first aspect or the encoding method as described in the second aspect.
[0029] This application provides an encoding / decoding method, encoder, decoder, and storage medium. Whether at the encoding or decoding end, the method involves determining motion information of multiple adjacent blocks at a preset position of the current block; determining affine motion information of the current block based on at least two motion information pieces that meet usability conditions from the motion information of the multiple adjacent blocks; and determining the predicted value of the current block based on the affine motion information of the current block. Thus, when constructing the affine merging list of the current block, it is no longer limited to the first motion information piece that meets the usability conditions from the motion information of multiple adjacent blocks, but can be constructed based on at least two motion information pieces that meet the usability conditions. This avoids the problem in related technologies where effective affine motion information cannot be constructed in certain cases, increases the probability of constructing effective affine motion information, makes the prediction of the current block in affine merging mode more accurate, and thus improves encoding / decoding performance. Attached Figure Description
[0030] Figure 1 is a schematic diagram of block-based motion compensation;
[0031] Figure 2 is a schematic diagram of block-based bidirectional prediction;
[0032] Figure 3 is a schematic diagram of the determination of temporal candidate MVP;
[0033] Figure 4A is a schematic diagram of the application of the 4-parameter affine model;
[0034] Figure 4B is a schematic diagram of the application of the 6-parameter affine model;
[0035] Figure 5A is a schematic diagram of determining the motion vectors of the control points corresponding to the upper left, upper right, and lower left positions of the current block;
[0036] Figure 5B is a schematic diagram of determining the motion vector of the control point corresponding to the lower right position of the current block;
[0037] Figure 6 is a schematic diagram of a video encoding and decoding network architecture provided in an embodiment of this application;
[0038] Figure 7 is a schematic block diagram of an encoder provided in an embodiment of this application;
[0039] Figure 8 is a schematic block diagram of a decoder provided in an embodiment of this application;
[0040] Figure 9 is a schematic flowchart of a decoding method provided in an embodiment of this application;
[0041] Figure 10 is a schematic flowchart of a decoding method provided in an embodiment of this application;
[0042] Figure 11 is a flowchart illustrating an encoding method provided in an embodiment of this application;
[0043] Figure 12 is a schematic diagram of the composition structure of an encoder provided in an embodiment of this application;
[0044] Figure 13 is a schematic diagram of the hardware structure of an encoder provided in an embodiment of this application;
[0045] Figure 14 is a schematic diagram of the composition structure of a decoder provided in an embodiment of this application;
[0046] Figure 15 is a schematic diagram of the hardware structure of a decoder provided in an embodiment of this application;
[0047] Figure 16 is a schematic diagram of the composition structure of an encoding / decoding system provided in an embodiment of this application. Detailed Implementation
[0048] In order to gain a more detailed understanding of the features and technical content of the embodiments of this application, the implementation of the embodiments of this application will be described in detail below with reference to the accompanying drawings. The accompanying drawings are for reference and illustration only and are not intended to limit the embodiments of this application.
[0049] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0050] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0051] It should also be noted that the terms "first, second, and third" used in the embodiments of this application are only used to distinguish similar objects and do not represent a specific order of objects. It is understood that "first, second, and third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0052] Before providing a further detailed description of the embodiments of this application, the nouns and terms used in the embodiments of this application will be explained. The nouns and terms used in the embodiments of this application shall be interpreted as follows:
[0053] Joint Video Experts Team (JVET);
[0054] Enhanced Compression Model (ECM) is a reference software test platform for enhanced compression.
[0055] Coding Unit (CU);
[0056] Coding Tree Unit (CTU);
[0057] Largest Coding Unit (LCU);
[0058] Prediction Unit (PU);
[0059] Transform Unit (TU);
[0060] Motion Vector (MV);
[0061] Sum of Absolute Difference (SAD);
[0062] Sum of Absolute Transformed Difference (SATD)
[0063] Sequence Parameter Set (SPS);
[0064] Control Point Motion Vector (CPMV);
[0065] Motion Vector Prediction (MVP);
[0066] Motion Vector Difference (MVD);
[0067] Advanced Motion Vector Prediction (AMVP);
[0068] Subblock-based Temporal Motion Vector Prediction (SbTMVP);
[0069] Picture Order Count (POC) technique.
[0070] It is understandable that video codec standards can adopt a block-based hybrid coding framework. For example, each image in a video is divided into square maximum coding units (MCUs) or coding tree units of the same size (e.g., 128×128, 64×64, etc.). Each MCU or coding tree unit can be further divided into rectangular coding units according to rules. Coding units may also be divided into prediction units, transform units, etc. The hybrid coding framework includes modules such as prediction, transform, quantization, entropy coding, and in-loop filtering. The prediction module can include intra-frame prediction and inter-frame prediction. Inter-frame prediction can include motion estimation and motion compensation. In video codec technology, intra-frame prediction can be used to eliminate spatial redundancy between adjacent samples, and inter-frame prediction can be used to eliminate temporal redundancy between adjacent images, thereby improving coding efficiency. Transformation converts the predicted image patch to the frequency domain, redistributing energy. Combined with quantization, it removes information insensitive to the human eye, thus eliminating visual redundancy. Entropy coding can eliminate character redundancy based on the current context model and the probability information of the binary bitstream.
[0071] In this embodiment, after reading a black-and-white image or a color image, the encoding end divides it into blocks. Intra-frame prediction or inter-frame prediction is applied to the current block to generate a prediction block. The prediction block is subtracted from the original block to obtain a residual block. The residual block is transformed and quantized to obtain a quantization coefficient matrix. The quantization coefficient matrix is then entropy-encoded and output to the bitstream. The decoding end applies intra-frame prediction or inter-frame prediction to the current block to generate a prediction block. The bitstream is decoded to obtain a quantization coefficient matrix. The quantization coefficient matrix is inverse-quantized and inverse-transformed to obtain a residual block. The prediction block and the residual block are added together to obtain a reconstructed block. The reconstructed block can be used to compose a reconstructed image. The decoding end performs loop filtering on the reconstructed image based on the image or based on the blocks to obtain a decoded image.
[0072] The block partitioning information, prediction, transform, quantization, entropy coding, and loop filtering mode information or parameter information determined at the encoding end need to be written into the bitstream if necessary. The decoding end determines the same block partitioning information, prediction, transform, quantization, entropy coding, and loop filtering mode information or parameter information as the encoding end by parsing and analyzing existing information, thus ensuring that the decoded image obtained by the encoding end is identical to that obtained by the decoding end. The encoding end also needs similar operations to the decoding end to obtain the decoded image. The decoded image can serve as a reference image for subsequent inter-frame prediction.
[0073] The above describes the basic flow of a video codec under a block-based hybrid coding framework. With technological advancements, some modules or steps of the framework may be optimized. The embodiments in this application apply to the basic flow of this block-based hybrid coding framework for video codecs, but are not limited to this framework and flow.
[0074] In this embodiment, the current block (CB) can be the current coding unit, the current prediction unit, or the current transform unit, etc. Due to the need for parallel processing, an image can be divided into slices, etc. Slices within the same image can be processed in parallel, meaning they have no data dependency. A "frame" is a common term, generally understood as an image. In this embodiment, the term "frame" can also be replaced with an image or a slice, etc. Furthermore, samples in this document can also be called pixels, and samples include both location information and values.
[0075] The following section introduces several prediction techniques within the relevant technologies.
[0076] (1) Inter-frame prediction.
[0077] Inter-frame prediction leverages the temporal correlation of video, using sample values from the encoded image to predict the sample values of the current image, effectively removing temporal redundancy. In video codec standards, inter-frame prediction employs block-based motion compensation techniques, as shown in Figure 1. Its basic principle is to find an optimal matching block (or reference block) in the encoded image for the current block; this process is called motion estimation (ME). The displacement from the current block to the reference block is called MV, and the process of using the reference block obtained from MV as the predicted value for the current block is called motion compensation (MC). The prediction residual is calculated from the sample values and predicted values of the current block, and MV and prediction residual are written into the bitstream. The decoder can obtain the predicted value from MV, add the transmitted prediction residual, and then obtain the reconstructed value of the current block.
[0078] When using a reference image from earlier times to predict the current image, this prediction method is called "forward prediction." Conversely, if a reference image from later times is used to predict the current image, it is called "backward prediction." If the current image is a B-frame, three inter-frame prediction methods are used: forward prediction, backward prediction, and bi-directional prediction. Thus, a coded block may have two prediction frames (MVs): one for forward prediction and one for backward prediction. In bi-directional prediction, two reference images are used simultaneously, and the weighted sum of the predicted values of the two reference blocks is used as the predicted value for the current block, as shown in Figure 2.
[0079] It should be noted that MV information accounts for a considerable proportion of the video compression bitstream. In order to improve encoding efficiency, there are many corresponding methods in modules such as MV prediction and acquisition, and motion compensation.
[0080] (a) Merge pattern.
[0081] The Merge mode directly utilizes one or a group of MVPs (each MVP includes unidirectional or bidirectional motion information) as the MV information of the current block to perform motion compensation and obtain the prediction block. The Merge mode does not require the transmission of motion vector difference (MVD), reference image index, and the current prediction direction (forward, backward, or bidirectional), making it an efficient MV encoding method. In its implementation, the Merge mode needs to establish an MVP candidate list (also called the Merge list or MergeMVP list), where each MVP contains the MV, reference image index, and prediction direction. By encoding the index value of the selected MVP in the Merge list, the motion information of the current block can be determined.
[0082] The merge list contains a maximum of MaxNumMergeCand candidates, determined by SPS-level syntax elements. The merge list can be constructed using spatial MVP candidates, temporal MVP candidates, history-based MVP candidates, average MVP candidates, zero-value MVP candidates, etc. During list construction, candidates of various types are checked sequentially and attempted to be added to the list until the list reaches its maximum value.
[0083] (b) AMVP mode.
[0084] In many cases, directly using the motion information of an already encoded block cannot effectively represent the motion information of the current block. AMVP encodes motion information by efficiently representing the motion vector difference (MVD). Similarly, the AMVP mode builds an MVP candidate list for the current block. The difference between AMVP and Merge mode is that the current block needs to encode the prediction direction and the reference image index; correspondingly, the MVP list constructed may be different when using different reference images. The MVP list of AMVP can also contain spatial domain candidate MVPs, temporal domain candidate MVPs, history-based candidate MVPs, and zero-value MVPs. When determining spatial domain candidate MVPs, an adjacent MV can only be used as an MVP candidate if the reference image corresponding to the adjacent candidate MV is the same as that of the current block. When determining temporal domain candidate MVPs, if there is motion information at the position corresponding to the current block in the co-located image, the MV of the temporal motion information is scaled according to the current image POC, the current reference image POC, the co-located image POC, and the reference image POC of the temporal motion information to obtain the temporal domain candidate MVP, as shown in Figure 3. Wherein, the POC distance between the current image and the reference image of the current image is t_b, the POC distance between the co-position image and the reference image of the co-position image is t_d, the motion information of the current block (current CU) is MVP_cur, and the motion information of the co-position block (co-position CU) is MV_col, then MVP_cur=(t_b / t_d)MV_col.
[0085] (2) Affine motion compensation.
[0086] Besides translation, video content typically includes rotation, scaling, and stretching motions. To effectively represent these motions, the image usually needs to be divided into smaller CUs (Compartment Units), and different motion information is used to perform motion compensation on different CUs. However, the motion vectors between these CUs may exhibit certain regularities. Affine motion compensation technology proposes that the motion vectors of different samples within a block can be determined by some known motion vectors through affine transformation. This avoids the need for further subdivision of the CUs.
[0087] For example, the motion vector of the top-left sample is known to be (MV 0,x ,MV 0,y The motion vector of the sample in the upper right corner is (MV). 1,x ,MV 1,y The motion vector of the sample in the lower left corner is (MV). 2,x ,MV 2,y The width and height of the current block are W and H, respectively. Using the top-left and top-right motion vectors as control points, the motion vector at position (i,j) within the block can be represented as follows:
[0088] This motion model is called a 4-parameter affine model, which consists of the control point motion vector CPMV1 in the upper left corner and CPMV2 in the upper right corner, as shown in Figure 4A.
[0089] Furthermore, the motion vector at position (i,j) within the block can also be determined by the three motion vectors mentioned above, as follows:
[0090] This motion model is called a 6-parameter affine model, which consists of the control point motion vector CPMV1 in the upper left corner, CPMV2 in the upper right corner, and CPMV3 in the lower left corner, as shown in Figure 4B.
[0091] (a) Affine merge mode.
[0092] Similar to the Merge pattern, the Affine Merge pattern obtains each CPMV by constructing a Merge List (Affine Merge List, SubBlkMergeMVP List). Each candidate in the Affine merge list can contain its affine type (e.g., 4-parameter or 6-parameter) and multiple motion vectors (4-parameter vectors contain 2 CPMVs, and 6-parameter vectors contain 3 CPMVs). Each CPMV can be a bidirectional motion prediction, but each CPMV should have the same reference image in one prediction direction. Specifically, if each CPMV is a bidirectional prediction, then these CPMVs should have the same reference image in both prediction directions.
[0093] The Affine merge list can contain two types of spatial candidates: spatially adjacent affine pattern CU inheritance candidates; and translational MV constructions of spatially and temporally adjacent CUs. Spatially adjacent affine pattern CU inheritance candidates refer to CPMVP (CPMV Prediction) candidates derived from the CPMV of adjacent blocks encoded by Affine patterns. The derived CPMVP has the same model type (4 or 6 parameters) as the adjacent blocks. For example, the current block has top-left corner coordinates (x0, y0), top-right corner coordinates (x1, y1), and bottom-left corner coordinates (x5, y5). The adjacent block has top-left corner coordinates (x2, y2), top-right corner coordinates (x3, y3), and bottom-left corner coordinates (x4, y4). The control point vectors corresponding to the top-left, top-right, and bottom-left positions of the adjacent blocks are MV2, MV3, and MV4, respectively. If the adjacent blocks are 4-parameter affine models, then based on MV2, MV3 and the relative positions of samples (x0,y0) and (x2,y2), and (x1,y1) and (x3,y3), MV0 (i.e., the MV at (x0,y0)) and MV1 can be obtained as candidate CPMVPs according to the MV calculation formula of the 4-parameter affine model described above. If the adjacent blocks are 6-parameter affine models, then based on MV2, MV3, MV4 and the relative positions of samples (x0,y0) and (x2,y2), (x1,y1) and (x3,y3), and (x4,y4) and (x5,y5), MV0, MV1, and MV5 can be obtained as candidate CPMVPs according to the MV calculation formula of the 6-parameter affine model described above.
[0094] The construction of translational MVs of spatially and temporally adjacent CUs refers to using the translational MVs of multiple spatially and temporally adjacent CUs as MVs at different positions to derive the control point vectors corresponding to the upper left, upper right, and lower left positions of the current block, thus obtaining candidate CPMVPs. First, four candidate control point MVs are set for the current block: CPMV1, CPMV2, CPMV3, and CPMV4. These candidate MVs are obtained from the MVs of spatially and temporally adjacent coding blocks, and are as follows:
[0095] As shown in Figure 5A, following the order of block B2->B3->A2, the MV of the first valid block (inter-frame prediction mode) is used as CPMV1;
[0096] As shown in Figure 5A, following the block B1->B0 order, the MV of the first valid block (inter-frame prediction mode) is used as CPMV2;
[0097] As shown in Figure 5A, following the block A1->A0 order, the MV of the first valid block (inter-frame prediction mode) is used as CPMV3;
[0098] As shown in Figure 5B, the MV of the corresponding coding block in the co-location image at the lower right corner position of the current block is adjusted according to the temporal MVP acquisition method and used as CPMV4. For example, it can also be determined based on the MV of the corresponding coding block in the co-location image at a certain offset distance from the lower right corner position of the current block. For instance, assuming the lower right corner position of the current block is (x, y), the corresponding position in the co-location image could also be (x+1, y+1).
[0099] Thus, based on the obtained CPMV1 to CPMV4, multiple candidate CPMVPs are obtained according to the combination of their subscripts, as shown in Table 1.
[0100] Table 1
[0101] Where f(CPMV1,CPMV3) represents the derivation of CPMV2 (i.e., the MV corresponding to the upper right control point) from CPMV1 and CPMV3, and may include:
[0102] CPMV2_x=(CPMV1_x<<7)+((CPMV3_y-CPMV1_y)<<(7+log2(Width)-log2(Height));
[0103] CPMV2_y=(CPMV1_y<<7)+((CPMV3_x-CPMV1_x)<<(7+log2(Width)-log2(Height)).
[0104] Here, Width and Height represent the width and height of the current block. The precision of CPMV2 can be adjusted later to maintain consistency with other CPMV values.
[0105] (3) Adaptive Reordering of Merge Candidates (ARMC).
[0106] ARMC technology proposes that in Merge mode, the predicted value of the corresponding reference block template can be obtained based on the motion information of each candidate, and compared with the template of the current block to determine the template error value of the candidate. This allows for sorting of candidates based on the template error value, placing candidates with smaller template error values at the top of the list and encoding their index values with fewer codewords. It also allows for candidate filtering, encoding only the top few candidates in the sorted list. Affine merge mode can use ARMC technology to determine the motion vectors of each sub-block in the template region based on the CPMVs of the current candidate, thereby obtaining the predicted value of the corresponding reference block template and achieving reordering based on template error value. The use of ARMC technology can be controlled by high-level syntax elements; for example, the sequence-level syntax element `sps_aml_enabled_flag` indicates whether the list is reordered based on template error value after construction of the merge list. Furthermore, ARMC technology can also be controlled by other high-level syntax elements; for example, the sequence-level syntax element `sps_tm_tools_enabled_flag` indicates whether to use the template error value-based technique.
[0107] In simple terms, the affine motion compensation mode in related technologies, when constructing the CPMV using neighboring motion information, only uses the first motion information found at each position (top left, top right, bottom left, bottom right) to predict the CPMV for the current block. In reality, more than one motion information can be obtained by checking different positions when determining the motion information at each position. Furthermore, constructing the CPMV requires that the reference images pointed to by the motion information at different positions be the same, which may lead to the inability to construct an effective CPMV. Not utilizing all available neighboring motion information reduces the likelihood of constructing an effective CPMV.
[0108] Based on this, embodiments of this application provide an encoding / decoding method that determines the motion information of multiple adjacent blocks at a preset position of the current block; determines the affine motion information of the current block based on at least two motion information that satisfy the availability condition among the motion information of the multiple adjacent blocks; and determines the predicted value of the current block based on the affine motion information of the current block.
[0109] In this way, when constructing the affine merging list for the current block, both the encoding and decoding ends are no longer limited to the first motion information that meets the usable condition among multiple adjacent blocks. Instead, they can be constructed based on at least two motion information that meet the usable condition. This avoids the problem in some related technologies that it is impossible to construct effective affine motion information in certain cases, increases the probability of constructing effective affine motion information, makes the prediction of the current block in the affine merging mode more accurate, and thus improves the encoding and decoding performance.
[0110] The embodiments of this application will now be described in detail with reference to the accompanying drawings.
[0111] Figure 6 is a schematic diagram of a video encoding and decoding network architecture provided in an embodiment of this application. As shown in Figure 6, the network architecture includes one or more electronic devices 13 to 1N and a communication network 01, wherein the electronic devices 13 to 1N can perform video interaction through the communication network 01. During implementation, the electronic devices can be various types of devices with video encoding and decoding capabilities. For example, the electronic devices may include mobile phones, tablet computers, personal computers, personal digital assistants, navigators, digital phones, video phones, televisions, sensing devices, servers, etc., and this embodiment of the application does not impose any limitations.
[0112] This application provides a network architecture for a video encoding / decoding system that includes decoding and encoding methods. The decoder or encoder in this application can be the aforementioned electronic device. That is, the electronic device in this application has video encoding / decoding capabilities and generally includes a video / image encoder (referred to as an encoder) and a video / image decoder (referred to as a decoder).
[0113] Figure 7 is a schematic block diagram of an encoder provided in an embodiment of this application. As shown in Figure 7, the encoder 100 may include a transform and quantization unit 101, an intra-frame estimation unit 102, an intra-frame prediction unit 103, a motion compensation unit 104, a motion estimation unit 105, an inverse transform and inverse quantization unit 106, a filter control and analysis unit 107, a filtering unit 108, an encoding unit 109, and a decoded image buffer unit 110, etc. Among them, the filtering unit 108 can implement deblocking filtering and sample adaptive offset (SAO) filtering, and the encoding unit 109 can implement header information encoding and context-based adaptive binary arithmetic coding (CABAC).For the input raw video signal, a video coding block can be obtained by partitioning it through a Coding Tree Unit (CTU). Then, the residual sample information obtained after intra-frame or inter-frame prediction is transformed by the transform and quantization unit 101, including transforming the residual information from the sample domain to the transform domain and quantizing the resulting transform coefficients to further reduce the bit rate. The intra-frame estimation unit 102 and the intra-frame prediction unit 103 are used to perform intra-frame prediction on the video coding block. Specifically, the intra-frame estimation unit 102 and the intra-frame prediction unit 103 are used to determine the intra-frame prediction mode to be used to encode the video coding block. The motion compensation unit 104 and the motion estimation unit 105 are used to perform inter-frame prediction coding of the received video coding block relative to one or more blocks in one or more reference frames to provide time prediction information. The motion estimation performed by the motion estimation unit 105 is a process of generating motion vectors, which can estimate the motion of the video coding block. Then, the motion compensation unit 104 is used to perform the motion estimation based on the motion vectors determined by the motion estimation unit 105. The motion compensation is performed. After determining the intra-prediction mode, the intra-prediction unit 103 is also used to provide the selected intra-prediction data to the coding unit 109, and the motion estimation unit 105 also sends the calculated motion vector data to the coding unit 109. In addition, the inverse transform and inverse quantization unit 106 is used to reconstruct the video coding block, reconstruct the residual block in the sample domain, and remove the block artifacts by the filter control analysis unit 107 and the filtering unit 108. Then, the reconstructed residual block is added to a predictive block in the frame of the decoding image buffer unit 110 to generate the reconstructed video coding block. The coding unit 109 is used to encode various coding parameters and quantized transform coefficients. In the CABAC-based coding algorithm, the context content can be based on adjacent coding blocks and can be used to encode information indicating the determined intra-prediction mode and output the bitstream of the video signal. The decoding image buffer unit 110 is used to store the reconstructed video coding block for prediction reference. As video image encoding proceeds, new reconstructed video encoding blocks are continuously generated, and these reconstructed video encoding blocks are stored in the decoding image buffer unit 110.
[0114] Figure 8 is a schematic block diagram of a decoder provided in an embodiment of this application. As shown in Figure 8, the decoder 200 includes a decoding unit 201, an inverse transform and inverse quantization unit 202, an intra-frame prediction unit 203, a motion compensation unit 204, a filtering unit 205, and a decoding image buffer unit 206, etc. The decoding unit 201 can perform header information decoding and CABAC decoding, and the filtering unit 205 can perform deblocking filtering and SAO filtering. After the input video signal undergoes the encoding processing shown in Figure 7, the bitstream of the video signal is output. This bitstream is input into the decoder 200, first passing through the decoding unit 201 to obtain the decoded transform coefficients. The transform coefficients are then processed by the inverse transform and inverse quantization unit 202 to generate residual blocks in the sample domain. The intra-frame prediction unit 203 can be used to generate prediction data for the current video decoding block based on the determined intra-frame prediction mode and data from previously decoded blocks in the current frame or image. The motion compensation unit 204 determines the prediction information for the video decoding block by analyzing motion vectors and other associated syntax elements, and uses... The prediction information is used to generate a predictive block of the video block being decoded; the decoded video block is formed by summing the residual block from the inverse transform and inverse quantization unit 202 with the corresponding predictive block generated by the intra-prediction unit 203 or the motion compensation unit 204; the decoded video signal is passed through the filtering unit 205 to remove block artifacts, which can improve video quality; then the decoded video block is stored in the decoding image buffer unit 206, which stores reference images for subsequent intra-prediction or motion compensation, and is also used for the output of the video signal, thus obtaining the recovered original video signal.
[0115] It should be noted that the method in this application embodiment affects the inter-frame prediction part in the video coding hybrid framework, for example, it mainly acts on the motion compensation unit 104 and motion estimation unit 105 shown in FIG7 and the motion compensation unit 204 shown in FIG8. That is to say, the method in this application embodiment can be applied to the encoder, the decoder, or even to both the encoder and the decoder at the same time, without any limitation.
[0116] It should also be noted that when the embodiments of this application are applied to the encoder 100, the "current block" can refer to the block to be encoded in the video image (also known as the "encoded block"); when the embodiments of this application are applied to the decoder 200, the "current block" can refer to the block to be decoded in the video image (also known as the "decoded block").
[0117] In one embodiment of this application, Figure 9 is a schematic flowchart of a decoding method provided in this application. As shown in Figure 9, the method may include:
[0118] S901, determine the motion information of multiple adjacent blocks at the preset position of the current block.
[0119] It should be noted that, in the embodiments of this application, the method is applied to a decoder. Here, the decoding method can be an inter-frame prediction method, and more specifically, it can be an extension method for constructing an affine merge list, utilizing the motion information of multiple adjacent blocks as much as possible to construct candidate affine motion information in the affine merge list.
[0120] In some embodiments, the preset position may include one of the following: the top left position of the current block, the top right position of the current block, the bottom left position of the current block, and the bottom right position of the current block.
[0121] In other words, in this embodiment of the application, the motion information of multiple adjacent blocks at a preset position may refer to the motion information of multiple adjacent blocks at the upper left position of the current block, or the motion information of multiple adjacent blocks at the upper right position of the current block, or the motion information of multiple adjacent blocks at the lower left position of the current block, or the motion information of multiple adjacent blocks at the lower right position of the current block, without any limitation.
[0122] It should also be noted that, in the embodiments of this application, the motion information of multiple adjacent blocks may include multiple motion information, and these multiple motion information are determined based on multiple adjacent blocks at a preset position, or they may be determined based on multiple adjacent positions (or including sub-blocks of the adjacent positions) within the decoding block adjacent to the preset position.
[0123] In some embodiments, multiple adjacent blocks at a preset position may include: blocks spatially adjacent to the preset position, or blocks adjacent to the preset position in a reference image. The latter is considered from a temporal perspective; therefore, "blocks adjacent to the preset position in the reference image" can also be called blocks temporally adjacent to the preset position. Here, the reference image can refer to the co-located image corresponding to the current block.
[0124] It should also be noted that, in the embodiments of this application, when the preset position is the upper left, upper right, or lower left position of the current block, the multiple adjacent blocks of the preset position may include blocks adjacent to the spatial domain of the preset position; when the preset position is the lower right position of the current block, the multiple adjacent blocks of the preset position may include blocks adjacent to the preset position in the reference image.
[0125] For example, taking Figure 5A as an example, when the preset position is the upper left position of the current block, the multiple adjacent blocks here can include blocks that are adjacent to the upper left position space of the current block, such as the blocks corresponding to B2, B3 and A2 in Figure 5A.
[0126] For example, taking Figure 5A as an example, when the preset position is the upper right position of the current block, the multiple adjacent blocks here can include blocks that are adjacent to the upper right position space of the current block, such as the blocks corresponding to B0 and B1 in Figure 5A.
[0127] For example, taking Figure 5A as an example, when the preset position is the lower left position of the current block, the multiple adjacent blocks here can include blocks that are adjacent to the lower left position space of the current block, such as the blocks corresponding to A0 and A1 in Figure 5A.
[0128] For example, taking Figure 5B as an example, when the preset position is the lower right position of the current block, the multiple adjacent blocks here can include the blocks in the reference image that are adjacent to the lower right position of the corresponding co-position block, such as the block corresponding to T in Figure 5B.
[0129] S902, determine the affine motion information of the current block based on at least two motion information that meet the availability conditions from multiple adjacent blocks.
[0130] In this embodiment of the application, taking a certain preset position as an example, after obtaining the motion information of multiple adjacent blocks, the affine motion information of the current block can be determined based on at least two motion information of the multiple adjacent blocks that meet the usability conditions.
[0131] In some embodiments, determining the affine motion information of the current block based on at least two motion information pieces that meet the availability condition from the motion information of multiple adjacent blocks may include: determining one or more candidate motion information pieces and availability identification information of one or more candidate motion information pieces from the at least two motion information pieces that meet the availability condition; and determining the affine motion information of the current block based on the one or more candidate motion information pieces that are indicated as available by the availability identification information.
[0132] In this embodiment of the application, determining one or more candidate motion information and one or more availability identification information of candidate motion information from at least two motion information that meet the availability conditions may include: determining a first candidate motion information from at least two motion information that meet the availability conditions; and marking the availability identification information of the first candidate motion information as available when the first candidate motion information does not overlap with existing candidate motion information that is already marked as available.
[0133] It should be noted that the first candidate motion information is not the same as the existing candidate motion information that is already marked as available. This can mean that the prediction direction is different, or the reference image is different in any direction, or the motion vector is different in any direction, etc. There are no specific limitations here.
[0134] It should also be noted that whether motion information meets the availability criteria can refer to whether the motion information is in inter-frame prediction mode. For example, if the motion information of a neighboring block is in inter-frame prediction mode, then it can be determined that the motion information meets the availability criteria, or in other words, the motion information is usable motion information.
[0135] In some embodiments, the method may further include: sequentially checking motion information of multiple adjacent blocks at a preset position in a first order; when the motion information of a first adjacent block at the preset position meets the availability condition, determining a list of reference images and a reference image index corresponding to the motion information of the first adjacent block; when determining that the motion information of the first adjacent block refers to a reference image corresponding to the first reference image index in the first reference image list, determining first available identifier information of candidate motion information at the preset position of the current block according to the first reference image list and the first reference image index; when the first available identifier information indicates unavailability, determining the motion information of the first adjacent block as a candidate motion information corresponding to the preset position, and marking the first available identifier information as available.
[0136] It should be noted that when the motion information of the first adjacent block at the preset position meets the availability condition, the reference image list and reference image index corresponding to the motion information of the first adjacent block are determined (which may be unidirectional or bidirectional prediction). Thus, the first reference image list and the first reference image index can be any number of reference image lists (list) and reference image indices (refIdx) from unidirectional or bidirectional prediction.
[0137] It should also be noted that when the motion information of the first adjacent block at the preset position meets the availability condition, it can also be that at least one reference image list corresponding to the current block is traversed; when it is determined that the motion information of the first adjacent block refers to the reference image corresponding to the first reference image index in the first reference image list, and the first availability identifier information of the preset position of the reference image corresponding to the first reference image index in the first reference image list indicates that it is unavailable, the motion information of the first adjacent block can be determined as a candidate motion information corresponding to the preset position, and the first availability identifier information is marked as available.
[0138] In the embodiments of this application, the first order may be a predefined inspection order, or it may be the default inspection order of related technologies, or it may be to sequentially check the motion information of multiple adjacent blocks. No limitation is made here.
[0139] It should also be noted that the first neighboring block can be one of the currently detected neighboring blocks among these multiple neighboring blocks. In some embodiments, the motion information of the first neighboring block satisfies the availability condition, which may include: the motion information of the first neighboring block is motion information in the inter-frame prediction mode. In this case, the first neighboring block is characterized as a valid block, that is, the motion information of the first neighboring block is one of the available motion information among these motion information.
[0140] It should also be noted that, assuming the number of reference images corresponding to the current block is P, for example, P equals 5; and the number of reference image lists is Q, for example, Q equals 2 (forward list L0 and backward list L1). An array isAvailable is initialized to determine whether motion information is available. This array should be equal to Q × P × 4, and can be denoted as isAvailable[2][5][4], where 4 represents the need to determine candidate motion information for 4 preset positions. Here, this array can also be called "available identifier information," used to determine whether the required prediction direction (i.e., the reference image list) and the candidate motion information at the corresponding preset positions on the reference images are available. For example, if the candidate motion information at the lower left corner position of the reference image with index 1 in the forward list L0 is unavailable, it is indicated that isAvailable[0][1][2] is false. It should be noted that, in this embodiment, the values of isAvailable[2][5][4] are all initialized to false, i.e., isAvailable[2][5][4] are all marked as unavailable.
[0141] In one possible implementation, assuming the preset position is the top-left position of the current block, the motion information of multiple adjacent blocks at the top-left position can be checked sequentially in a first order. When the motion information of the first adjacent block meets the availability condition, at least one reference image list corresponding to the current block is traversed. When it is determined that the motion information of the first adjacent block refers to the reference image corresponding to the first reference image index refIdx in the first reference image list, and the availability identifier information isAvailable[list][refIdx][0] of the top-left position of the reference image corresponding to the first reference image index refIdx in the first reference image list indicates that it is unavailable, the motion information of the first adjacent block is determined as a candidate motion information corresponding to the top-left position, and the first availability identifier information isAvailable[list][refIdx][0] is marked as available.
[0142] In this implementation, the multiple adjacent blocks in the upper left position can include the blocks corresponding to A2, B2, and B3 as shown in Figure 5A. The first order can be B2->B3->A2, or it can be a sequence check of A2->B2->B3; this is not a limitation.
[0143] In this implementation, the motion information of multiple adjacent blocks at the top-left position is traversed. If the motion information (for example, the motion information at position B2, which can be represented as Mi B2 ) is available, traverse the two reference picture lists of the current picture list=0..1; wherein, the current picture is the picture where the current block is located, that is, the current picture includes the current block. Then determine Mi B2 whether it refers to the first reference picture list (if list is 0, then Mi B2 's prediction mode shall be forward or bidirectional; otherwise, Mi B2 's prediction mode shall be backward or bidirectional, which can be expressed as Mi B2 .interDir&(1<<list)). The first reference picture index of Mi in the list is expressed as refIdx=Mi B2 refIdx[list]. If Mi B2 satisfies: Mi B2 .interDir&(1<<list)&&!isAvailable[list][refIdx][0], that is, Mi B2 refers to list, and the motion vector (MV) at the top-left position of the reference picture with the first reference picture index refIdx in list is unavailable. In this case, Mi can be B2 saved as candidate motion information (represented as Mi[list][refIdx][0]) for the top-left position of the reference picture with the first reference picture index refIdx in list, and the first available identification information isAvailable[list][refIdx][0] is marked as available. B2
[0144] In some embodiments, in the process of sequentially checking the motion information of multiple adjacent blocks, for a next adjacent block (for example, a second adjacent block), when the motion information of the second adjacent block satisfies an availability condition (for example, the motion information of the second adjacent block is motion information in inter prediction mode), the reference picture list and reference picture index corresponding to the motion information of the second adjacent block are determined; this includes the following three possible cases:
[0145] First possible case: when it is determined that the motion information of the second adjacent block refers to the reference picture corresponding to the first reference picture index in the first reference picture list, and the first available identification information indicates available, the motion information of the second adjacent block is not used as a candidate motion information corresponding to the preset position.
[0146] For example, still taking the top-left position of the current block as an example, if the motion information of the second adjacent block (for example, the motion information at position B3, which can be represented as MiB3 It is available, and Mi B3 It also references the reference image corresponding to the first reference image index refIdx in the first reference image list, but the first available identifier isAvailable[list][refIdx][0] has already been marked as available, so Mi is not included. B3 As a candidate motion information corresponding to the top left position, Mi is no longer considered. B3 The candidate motion information is saved as the top-left position of the reference image with the first reference image index refIdx in the list.
[0147] The second possible scenario is as follows: When determining the motion information of the second adjacent block by referring to the reference image corresponding to the second reference image index in the first reference image list, the second available identifier information of the candidate motion information of the current block at the preset position is determined according to the first reference image list and the second reference image index; and when the second available identifier information indicates that it is unavailable, the motion information of the second adjacent block is determined as another candidate motion information corresponding to the preset position, and the second available identifier information is marked as available.
[0148] For example, still taking the top-left position of the current block as an example, if the motion information of the second adjacent block (e.g., the motion information of position B3, can be represented as Mi) B3 It is available, and Mi B3 The reference image corresponding to the second reference image index refIdx1 in the first reference image list is used, but the second availability identifier isAvailable[list][refIdx1][0] is unavailable. Therefore, Mi can be used... B3 As another candidate motion information corresponding to the top left position, it is saved as the candidate motion information of the top left position of the reference image with the second reference image index refIdx1 in the list (represented as Mi[list][refIdx1][0]), and the second available identifier information isAvailable[list][refIdx1][0] is marked as available.
[0149] The third possible scenario: When determining the motion information of the second adjacent block by referring to the reference image corresponding to the third reference image index in the second reference image list, the third available identifier information of the candidate motion information of the current block at the preset position is determined according to the second reference image list and the third reference image index; and when the third available identifier information indicates that it is unavailable, the motion information of the second adjacent block is determined as another candidate motion information corresponding to the preset position, and the third available identifier information is marked as available.
[0150] For example, still taking the top-left position of the current block as an example, if the motion information of the second adjacent block (e.g., the motion information of position B3, can be represented as Mi) B3 It is available, and Mi B3 The reference image corresponding to the third reference image index refIdx2 in the second reference image list list1 was used, but the third available identifier isAvailable[list1][refIdx2][0] is unavailable. Therefore, Mi can be used... B3 As another candidate motion information corresponding to the top left position, it is saved as the candidate motion information of the top left position of the reference image with the index refIdx2 of the third reference image in list1 (represented as Mi[list1][refIdx2][0]), and the third available identifier information isAvailable[list1][refIdx2][0] is marked as available.
[0151] It should also be noted that, in one possible implementation of the stopping condition for the inspection, the method may include: sequentially inspecting the motion information of multiple adjacent blocks at a preset position in a first order to obtain at least one candidate motion information corresponding to the preset position and available identification information of at least one candidate motion information.
[0152] In other words, under this implementation, the motion information of all these adjacent blocks can be checked, and then at least one candidate motion information corresponding to the preset position and at least one available identifier information of the candidate motion information can be obtained.
[0153] In another possible implementation, the method may further include: when the number of available candidate motion information obtained at the preset position reaches a preset number, stopping the inspection of motion information of multiple adjacent blocks at the preset position, and determining at least one candidate motion information corresponding to the preset position and available identification information of at least one candidate motion information based on the obtained candidate motion information.
[0154] In other words, this implementation method can limit the number of available candidate motion information stored at a preset position. For example, assuming the preset number (or "maximum number") is 2, if the preset position has already obtained 2 candidate motion information, then the subsequent adjacent blocks to be checked are skipped, that is, the traversal of the motion information of these multiple adjacent blocks is stopped, and the 2 candidate motion information already obtained is directly determined as at least one candidate motion information corresponding to the preset position, along with the available identification information for determining this at least one candidate motion information.
[0155] In the embodiments of this application, for a preset position, the reference images for the at least one candidate motion information stored therein are different, that is, the reference images for the at least one candidate motion information come from different reference image lists and / or different reference image indices.
[0156] In some embodiments, the method may further include: determining temporal motion information corresponding to a block in a reference image adjacent to a preset position; scaling the temporal motion information to determine a candidate motion information corresponding to the preset position; and determining available identification information of the candidate motion information based on the reference image.
[0157] It should be noted that, in the embodiments of this application, if multiple adjacent blocks at the preset position include blocks adjacent to the preset position in the reference image, then after determining the corresponding temporal motion information, the temporal motion information can be scaled and adjusted to obtain a candidate motion information corresponding to the preset position and the available identification information of the candidate motion information.
[0158] It should also be noted that, in this embodiment, based on the above-described checking method, when the preset position is the upper right position of the current block, the motion information of multiple adjacent blocks at the upper right position can be traversed in the order of B1->B0 to obtain at least one candidate motion information and the available identifier information isAvailable corresponding to the upper right position. Similarly, when the preset position is the lower left position of the current block, the motion information of multiple adjacent blocks at the lower left position can be traversed in the order of A1->A0 to obtain at least one candidate motion information and the available identifier information isAvailable corresponding to the lower left position. Furthermore, when the preset position is the lower right position of the current block, the motion information of the corresponding block in the reference image (e.g., a co-located image) at the lower right position (or offset by a certain distance) of the current block can be checked and adjusted according to the temporal MVP acquisition method to obtain at least one candidate motion information and the available identifier information isAvailable corresponding to the lower right position.
[0159] Understandably, after obtaining at least one candidate motion information corresponding to each of the multiple preset positions such as the upper left, upper right, lower left, and lower right positions of the current block, as well as the available identifier information isAvailable, it can be used to determine the affine motion information of the current block. In some embodiments, referring to Figure 10, the method may include:
[0160] S1001, determine the candidate motion information corresponding to at least two preset positions and the available identification information of the candidate motion information.
[0161] S1002, determine at least one candidate affine motion information for the current block based on the candidate motion information corresponding to at least two preset positions and the available identifier information of the candidate motion information.
[0162] S1003, determine the affine merging list of the current block based on at least one candidate affine motion information.
[0163] S1004, Determine the affine motion information of the current block based on the affine merging list of the current block.
[0164] In this embodiment, each candidate affine motion information can be composed of candidate motion information corresponding to at least two preset positions. Here, the combination of at least two preset positions can be as shown in Table 2. For example, the at least two candidate motion information used to construct the candidate affine motion information can be determined according to the order in Table 2.
[0165] Table 2
[0166] Wherein, CPMV1 represents at least one candidate motion information corresponding to the upper left position, CPMV2 represents at least one candidate motion information corresponding to the upper right position, CPMV3 represents at least one candidate motion information corresponding to the lower left position, and CPMV4 represents at least one candidate motion information corresponding to the lower right position. Furthermore, f(CPMV1, CPMV3) represents the derivation of CPMV2 (i.e., the MV corresponding to the upper right control point) from CPMV1 and CPMV3, including:
[0167] CPMV2_x=(CPMV1_x<<7)+((CPMV3_y-CPMV1_y)<<(7+log2(Width)-log2(Height));
[0168] CPMV2_y=(CPMV1_y<<7)+((CPMV3_x-CPMV1_x)<<(7+log2(Width)-log2(Height)).
[0169] In this embodiment, when the model index is 2, it can be determined by CPMV1, CPMV2, and CPMV4. "CPMV4+CPMV1-CPMV2" indicates that the lower left control point MV can be derived from candidate motion information at preset positions such as CPMV1, CPMV2, and CPMV4. Similarly, when the model index is 3, it can be determined by CPMV1, CPMV3, and CPMV4. "CPMV4+CPMV1-CPMV3" indicates that the upper right control point MV can be derived from candidate motion information at preset positions such as CPMV1, CPMV3, and CPMV4. When the model index is 4, it can be determined by CPMV2, CPMV3, and CPMV4. "CPMV2+CPMV3-CPMV4" indicates that the upper left control point MV can be derived from candidate motion information at preset positions such as CPMV2, CPMV3, and CPMV4.
[0170] In some embodiments, determining candidate motion information corresponding to at least two preset positions may include: determining a first reference image list corresponding to the current block and a first reference image index in the first reference image list; determining at least two preset positions required for candidate affine motion information; and determining candidate motion information corresponding to at least two preset positions based on the first reference image list, the first reference image index, and the at least two preset positions.
[0171] It should be noted that, in the embodiments of this application, determining the candidate motion information corresponding to at least two preset positions can be achieved by traversing the model index, traversing the reference image list, and traversing the reference image index in any order.
[0172] For example, the construction order can be as follows: first, traverse the model index in Table 2 to determine at least two preset positions required for candidate affine motion information; then, traverse the reference image list list = 0..1 of the current image to determine the current list; then, traverse the reference image index refIdx = 0..4 to determine the current refIdx, so as to obtain the candidate motion information corresponding to each of the at least two preset positions used in the current construction.
[0173] For example, the construction order can also be adjusted as follows: first, traverse the reference image list list = 0..1 of the current image to determine the current list; then traverse the reference image index refIdx = 0..4 to determine the current refIdx; then traverse the model index in Table 2 to determine at least two preset positions required for the candidate affine motion information, so as to obtain the candidate motion information corresponding to each of the at least two preset positions used in the current construction.
[0174] For example, the construction order can also be adjusted as follows: first, traverse the reference image index refIdx = 0..4 to determine the current refIdx; then traverse the reference image list list = 0..1 of the current image to determine the current list; then traverse the model index in Table 2 to determine at least two preset positions required for the candidate affine motion information, so as to obtain the candidate motion information corresponding to each of the at least two preset positions used in the current construction.
[0175] In some embodiments, the method may further include: if the candidate motion information corresponding to at least two preset positions refers to the reference image corresponding to the first reference image index in the first reference image list and all available identification information is available, then the candidate affine motion information is determined based on the candidate motion information corresponding to at least two preset positions.
[0176] For example, taking model index 1 in Table 2 as an example, first iterate through the reference image list list = 0..1, and then iterate through the reference image indices refIdx = 0..4. Check whether the corresponding control points MV are all available. For example, when model index is 1, candidate motion information corresponding to the upper left, upper right, and lower left positions is needed. Then we can determine whether the following conditions are met: isAvailable[list][refIdx][0]&&isAvailable[list][refIdx][1]&&isAvailable[list][refIdx][2]. If they are met, then candidate affine motion information is constructed based on Mi[list][refIdx]. That is, CPMV1 = Mi[list][refIdx][0], CPMV2 = Mi[list][refIdx][1], CPMV3 = MI[list][refIdx][2], thereby determining the candidate affine motion information, that is, constructing CPMVP candidates.
[0177] It should be noted that after obtaining the candidate affine motion information, it can be compared with other candidates in the affine merging list of the current block to determine whether the obtained candidate affine motion information is redundant, i.e., a deduplication operation is performed. If the obtained candidate affine motion information is not redundant, it can be added to the affine merging list of the current block.
[0178] In some embodiments, for each candidate affine motion information, the candidate affine motion information may include an affine type and at least two candidate motion information. Specifically, when the affine type is a 4-parameter affine type, the candidate affine motion information is determined to include candidate motion information corresponding to two preset positions; when the affine type is a 6-parameter affine type, the candidate affine motion information is determined to include candidate motion information corresponding to three preset positions.
[0179] It should also be noted that, in this embodiment of the application, Table 2 is still used as an example. The model indices 1, 2, 3 and 4 are all 6-parameter affine types. The motion information of the current block (i,j) position can be determined by three candidate motion information, as shown in the aforementioned formula (2); the model indices 5 and 6 are both 4-parameter affine types. The motion information of the current block (i,j) position can be determined by two candidate motion information, as shown in the aforementioned formula (1).
[0180] In this embodiment of the application, after obtaining at least one candidate affine motion information, when constructing the affine merging list of the current block, the at least one candidate affine motion information can be reordered according to the template error value, or the reordered part of the candidate affine motion information can be added to the affine merging list of the current block.
[0181] In some embodiments, the method may further include: determining the current template of the current block; calculating the template error between the current template of the current block and the reference template corresponding to at least one candidate affine motion information, and determining the template error value corresponding to each of the at least one candidate affine motion information; sorting the at least one candidate affine motion information in ascending order of template error value, and forming an affine merging list of the current block from the top M candidate affine motion information; wherein M is a positive integer.
[0182] It should be noted that, in the embodiments of this application, the template error calculation can be SAD error calculation, SATD error calculation, or other error calculation methods, such as mean-square error (MSE), mean absolute error (MAE), etc., without any limitation.
[0183] It should also be noted that, in this embodiment, template error can be calculated based on the current template of the current block and the reference template of the reference block corresponding to each candidate affine motion information to obtain the template error value corresponding to each candidate affine motion information; then, at least one candidate affine motion information is reordered in ascending order of template error values, and the reordered at least one candidate affine motion information is added to the affine merging list. Considering the length of the affine merging list, only a portion of the candidate affine motion information with relatively small index numbers can be added to the affine merging list, for example, only the top M candidate affine motion information can be added to the affine merging list. For example, the value of M can be 6, but it is not specifically limited.
[0184] In this embodiment of the application, after obtaining the affine merging list of the current block, determining the affine motion information of the current block based on the affine merging list of the current block may include: parsing the merging candidate index in the bitstream; and determining the affine motion information of the current block based on the merging candidate index and the affine merging list.
[0185] In this embodiment, the merge candidate index can be represented by pu_merge_idx, which is a block-level syntax element used to indicate the index number (or "position") of the affine motion information of the current block in the affine merge list. If the affine merge list only includes the top M candidate affine motion information, then the number of bits corresponding to the merge candidate index in the bitstream can be saved.
[0186] In this way, based on at least two motion information that meet the availability conditions from multiple adjacent blocks, an affine merging list for the current block can be constructed. Then, by combining the merging candidate index in the bitstream, the selected affine motion information for the current block can be determined.
[0187] S903, determine the predicted value of the current block based on the affine motion information of the current block.
[0188] In this embodiment, after obtaining the affine motion information of the current block, the current block can be predicted based on the affine motion information to obtain the predicted value of the current block. For example, if the affine motion information indicates that the affine type is a 4-parameter affine type, then the motion information within the current block can be determined using the aforementioned formula (1) and affine motion compensation can be performed to obtain the predicted value of the current block; if the affine motion information indicates that the affine type is a 6-parameter affine type, then the motion information within the current block can be determined using the aforementioned formula (2) and affine motion compensation can be performed to obtain the predicted value of the current block.
[0189] It should also be noted that, in this embodiment of the application, some indication information in the form of syntax elements, or flags, can be written into the bitstream. In this way, by parsing the relevant syntax elements in the bitstream, it is possible to determine whether the current block uses affine merge mode.
[0190] In some embodiments, the method may further include: parsing a first syntax element and a second syntax element in the bitstream; and determining that the current block uses an affine merge mode if the first syntax element has a first value and the second syntax element has a first value.
[0191] In this embodiment, the first syntax element can be used to indicate whether the current block uses merge mode. The first syntax element can be represented by `pu_merge_flag`, which represents the enable flag for the block-level merge mode. For example, if the first syntax element takes a first value, it can be determined that the current block uses merge mode; if the first syntax element takes a second value, it can be determined that the current block does not use merge mode.
[0192] In this embodiment, the second syntax element can be used to indicate whether the current block uses a sub-block-based motion compensation mode. The second syntax element can be represented by `cu_affine_flag`, which represents the enable flag for the block-level sub-block-based motion compensation mode. For example, if the value of the second syntax element is a first value, it can be determined that the current block uses a sub-block-based motion compensation mode; if the value of the second syntax element is a second value, it can be determined that the current block does not use a sub-block-based motion compensation mode. It should be noted that "sub-block-based motion compensation mode" is a general term here, encompassing both the sbTMVP mode and the Affine mode.
[0193] It should be noted that the first value is different from the second value, and both the first and second values can be in parameter form or in numeric form. For example, the first value can be set to 1 and the second value can be set to 0; or, the first value can be set to true and the second value can be set to false; no restrictions are imposed here.
[0194] It should also be noted that if both pu_merge_flag and cu_affine_flag are true, it indicates that the current block uses affine merging mode, and the decoding method of this application embodiment can be executed, i.e., the decoding process shown in Figure 9.
[0195] In one possible implementation, a schematic description of the relevant syntactic elements is shown in Table 3.
[0196] Table 3
[0197] It is also understood that, in the embodiments of this application, the use of the technical solution of the embodiments of this application to determine CPMVP candidates can be controlled by high-level syntax elements such as sequence level and image level. In some embodiments, the steps include parsing a third syntax element in the bitstream; if the value of the third syntax element is a first value, performing the steps of determining the motion information of multiple adjacent blocks at a preset position of the current block; and determining the affine motion information of the current block based on at least two motion information that meet the availability conditions among the motion information of multiple adjacent blocks.
[0198] In this embodiment, if the value of the third syntax element is the first value, then it can be determined that the current block uses the technical solution of this embodiment to determine the CPMVP candidate, that is, executes the decoding process shown in FIG9; if the value of the third syntax element is the second value, then it can be determined that the current block does not use the technical solution of this embodiment to determine the CPMVP candidate, that is, does not execute the decoding process shown in FIG9, but uses existing methods in related technologies to determine the CPMVP candidate. For example, the first value can be set to 1 and the second value can be set to 0; or, the first value can be set to true and the second value can be set to false; no limitation is made here.
[0199] In one possible implementation, the third syntax element can be a sequence-level syntax element, such as represented by `sps_affine_cst_opt_enabled_flag`. If `sps_affine_cst_opt_enabled_flag` is true, it indicates that the current block uses the technical solution of this application embodiment to determine the CPMVP candidate; if `sps_affine_cst_opt_enabled_flag` is false, it indicates that the current block uses existing methods in related technologies to determine the CPMVP candidate. If this syntax element does not exist in the bitstream, it is assumed to be false by default.
[0200] In another possible implementation, the third syntax element can be other high-level syntax elements, such as the ARMC enable flag `sps_aml_enabled_flag`, or the template matching tool enable flag `sps_tm_tools_enabled_flag`. In this implementation, if the third syntax element is true, it indicates that the current block uses the adaptive merge candidate reordering method (i.e., the ARMC method), and it can be determined that the current block uses the technical solution of this application embodiment to determine CPMVP candidates. If the third syntax element is false, it indicates that the current block does not use the adaptive merge candidate reordering method (i.e., the ARMC method), and it can be determined that the current block uses existing methods in related technologies to determine CPMVP candidates.
[0201] In another possible implementation, the method may further include: determining block information of the current block; determining motion information of multiple adjacent blocks at a preset position of the current block based on the block information; and determining affine motion information of the current block based on at least two motion information of the multiple adjacent blocks that satisfy the availability condition.
[0202] In this embodiment, the block information of the current block may include information such as the block size and position of the current block. That is, the block information of the current block can also be used to determine whether the current block uses the technical solution of this embodiment to determine a CPMVP candidate. When the block information of the current block meets preset conditions, it can be determined that the current block uses the technical solution of this embodiment to determine a CPMVP candidate, i.e., the decoding process shown in Figure 9 is executed. Here, the preset conditions are the criteria used to determine whether the current block uses the technical solution of this embodiment to determine a CPMVP candidate, such as the block size of the current block being greater than a certain threshold, etc., and are not limited here.
[0203] In some embodiments, the method may further include: setting the maximum number of candidates in the affine merge list to L+N when the value of the third syntax element is the first value; where L and N are both positive integers.
[0204] In this embodiment, L represents the initial maximum number of candidates in the affine merge list. When using the technical solution of this embodiment to determine CPMVP candidates in the current block, the maximum number of candidates in the affine merge list can be increased. For example, if the value of sps_affine_cst_opt_enabled_flag is true, then the maximum number of candidates in the affine merge list increases by N, and the maximum number of candidates can be L+N. N can be set to 2, but this is not limited in any way.
[0205] It should also be noted that, in this embodiment, after obtaining the predicted value of the current block, the residual coefficients of the current block in the bitstream can be parsed, and then inverse quantization and inverse transform processing can be performed on the residual coefficients of the current block to obtain the residual value of the current block, thereby determining the reconstructed value of the current block. For example, the residual value and the predicted value of the current block can be added together to complete the reconstruction of the current block.
[0206] This application provides a decoding method that, when constructing the affine merging list of the current block, is no longer limited to the first motion information that meets the usability condition among multiple adjacent blocks, but can be constructed based on at least two motion information that meet the usability condition. This avoids the problem in related technologies where effective affine motion information cannot be constructed in some cases, increases the probability of constructing effective affine motion information, makes the prediction of the current block in the affine merging mode more accurate, and thus improves decoding performance.
[0207] In another embodiment of this application, FIG11 is a flowchart illustrating an encoding method provided in an embodiment of this application. As shown in FIG11, the method may include:
[0208] S1101, determine the motion information of multiple adjacent blocks at the preset position of the current block.
[0209] It should be noted that, in the embodiments of this application, the method is applied to an encoder. Here, the encoding method can be an inter-frame prediction method, and more specifically, it can be an extension method for constructing an affine merge list, utilizing the motion information of multiple adjacent blocks as much as possible to construct candidate affine motion information in the affine merge list.
[0210] In some embodiments, the preset position may include one of the following: the top left position of the current block, the top right position of the current block, the bottom left position of the current block, and the bottom right position of the current block.
[0211] In other words, in this embodiment of the application, the motion information of multiple adjacent blocks at a preset position may refer to the motion information of multiple adjacent blocks at the upper left position of the current block, or the motion information of multiple adjacent blocks at the upper right position of the current block, or the motion information of multiple adjacent blocks at the lower left position of the current block, or the motion information of multiple adjacent blocks at the lower right position of the current block, without any limitation.
[0212] It should also be noted that, in the embodiments of this application, the motion information of multiple adjacent blocks may include multiple motion information, and these multiple motion information are determined based on multiple adjacent blocks at a preset position, or they may be determined based on multiple adjacent positions (or including sub-blocks of the adjacent positions) within the coding block adjacent to the preset position.
[0213] In some embodiments, multiple adjacent blocks at a preset position may include: blocks spatially adjacent to the preset position, or blocks adjacent to the preset position in a reference image. The latter is considered from a temporal perspective; therefore, "blocks adjacent to the preset position in the reference image" can also be called blocks temporally adjacent to the preset position. Here, the reference image can refer to the co-located image corresponding to the current block.
[0214] It should also be noted that, in the embodiments of this application, when the preset position is the upper left, upper right, or lower left position of the current block, the multiple adjacent blocks of the preset position may include blocks adjacent to the spatial domain of the preset position; when the preset position is the lower right position of the current block, the multiple adjacent blocks of the preset position may include blocks adjacent to the preset position in the reference image.
[0215] For example, taking Figure 5A as an example, when the preset position is the upper left position of the current block, the multiple adjacent blocks here can include blocks that are adjacent to the upper left position space of the current block, such as the blocks corresponding to B2, B3 and A2 in Figure 5A.
[0216] For example, taking Figure 5A as an example, when the preset position is the upper right position of the current block, the multiple adjacent blocks here can include blocks that are adjacent to the upper right position space of the current block, such as the blocks corresponding to B0 and B1 in Figure 5A.
[0217] For example, taking Figure 5A as an example, when the preset position is the lower left position of the current block, the multiple adjacent blocks here can include blocks that are adjacent to the lower left position space of the current block, such as the blocks corresponding to A0 and A1 in Figure 5A.
[0218] For example, taking Figure 5B as an example, when the preset position is the lower right position of the current block, the multiple adjacent blocks here can include the blocks in the reference image that are adjacent to the lower right position of the corresponding co-position block, such as the block corresponding to T in Figure 5B.
[0219] S1102, determine the affine motion information of the current block based on at least two motion information that meet the availability conditions from multiple adjacent blocks.
[0220] In this embodiment of the application, taking a certain preset position as an example, after obtaining the motion information of multiple adjacent blocks, the affine motion information of the current block can be determined based on at least two motion information of the multiple adjacent blocks that meet the usability conditions.
[0221] In some embodiments, determining the affine motion information of the current block based on at least two motion information pieces that meet the availability condition from the motion information of multiple adjacent blocks may include: determining one or more candidate motion information pieces and availability identification information of one or more candidate motion information pieces from the at least two motion information pieces that meet the availability condition; and determining the affine motion information of the current block based on the one or more candidate motion information pieces that are indicated as available by the availability identification information.
[0222] In this embodiment of the application, determining one or more candidate motion information and one or more availability identification information of candidate motion information from at least two motion information that meet the availability conditions may include: determining a first candidate motion information from at least two motion information that meet the availability conditions; and marking the availability identification information of the first candidate motion information as available when the first candidate motion information does not overlap with existing candidate motion information that is already marked as available.
[0223] It should be noted that the first candidate motion information is not the same as the existing candidate motion information that is already marked as available. This can mean that the prediction direction is different, or the reference image is different in any direction, or the motion vector is different in any direction, etc. There are no specific limitations here.
[0224] It should also be noted that whether motion information meets the availability criteria can refer to whether the motion information is in inter-frame prediction mode. For example, if the motion information of a neighboring block is in inter-frame prediction mode, then it can be determined that the motion information meets the availability criteria, or in other words, the motion information is usable motion information.
[0225] In some embodiments, the method may further include: sequentially checking motion information of multiple adjacent blocks at a preset position in a first order; when the motion information of a first adjacent block at the preset position meets the availability condition, determining a list of reference images and a reference image index corresponding to the motion information of the first adjacent block; when determining that the motion information of the first adjacent block refers to a reference image corresponding to the first reference image index in the first reference image list, determining first available identifier information of candidate motion information at the preset position of the current block according to the first reference image list and the first reference image index; when the first available identifier information indicates unavailability, determining the motion information of the first adjacent block as a candidate motion information corresponding to the preset position, and marking the first available identifier information as available.
[0226] It should be noted that when the motion information of the first adjacent block at the preset position meets the availability condition, the reference image list and reference image index corresponding to the motion information of the first adjacent block are determined (which may be unidirectional or bidirectional prediction). Thus, the first reference image list and the first reference image index can be any number of reference image lists (list) and reference image indices (refIdx) from unidirectional or bidirectional prediction.
[0227] It should also be noted that when the motion information of the first adjacent block at the preset position meets the availability condition, it can also be that at least one reference image list corresponding to the current block is traversed; when it is determined that the motion information of the first adjacent block refers to the reference image corresponding to the first reference image index in the first reference image list, and the first availability identifier information of the preset position of the reference image corresponding to the first reference image index in the first reference image list indicates that it is unavailable, the motion information of the first adjacent block can be determined as a candidate motion information corresponding to the preset position, and the first availability identifier information is marked as available.
[0228] In the embodiments of this application, the first order may be a predefined inspection order, or it may be the default inspection order of related technologies, or it may be to sequentially check the motion information of multiple adjacent blocks. No limitation is made here.
[0229] It should also be noted that the first neighboring block can be one of the currently detected neighboring blocks among these multiple neighboring blocks. In some embodiments, the motion information of the first neighboring block satisfies the availability condition, which may include: the motion information of the first neighboring block is motion information in the inter-frame prediction mode. In this case, the first neighboring block is characterized as a valid block, that is, the motion information of the first neighboring block is one of the available motion information among these motion information.
[0230] It should also be noted that, assuming the number of reference images corresponding to the current block is P, for example, P equals 5; and the number of reference image lists is Q, for example, Q equals 2 (forward list L0 and backward list L1). An array isAvailable is initialized to determine whether motion information is available. This array should be equal to Q × P × 4, and can be denoted as isAvailable[2][5][4], where 4 represents the need to determine candidate motion information for 4 preset positions. Here, this array can also be called "available identifier information," used to determine whether the required prediction direction (i.e., the reference image list) and the candidate motion information at the corresponding preset positions on the reference images are available. For example, if the candidate motion information at the lower left corner position of the reference image with index 1 in the forward list L0 is unavailable, it is indicated that isAvailable[0][1][2] is false. It should be noted that, in this embodiment, the values of isAvailable[2][5][4] are all initialized to false, i.e., isAvailable[2][5][4] are all marked as unavailable.
[0231] In one possible implementation, assuming the preset position is the top-left position of the current block, the motion information of multiple adjacent blocks at the top-left position can be checked sequentially in a first order. When the motion information of the first adjacent block meets the availability condition, at least one reference image list corresponding to the current block is traversed. When it is determined that the motion information of the first adjacent block refers to the reference image corresponding to the first reference image index refIdx in the first reference image list, and the availability identifier information isAvailable[list][refIdx][0] of the top-left position of the reference image corresponding to the first reference image index refIdx in the first reference image list indicates that it is unavailable, the motion information of the first adjacent block is determined as a candidate motion information corresponding to the top-left position, and the first availability identifier information isAvailable[list][refIdx][0] is marked as available.
[0232] In this implementation, the multiple adjacent blocks in the upper left position can include the blocks corresponding to A2, B2, and B3 as shown in Figure 5A. The first order can be B2->B3->A2, or it can be a sequence check of A2->B2->B3; this is not a limitation.
[0233] In this implementation, the motion information of a plurality of adjacent blocks at the top-left position is traversed. If the motion information (e.g., the motion information at position B2, which can be denoted as Mi B2 ) is available, then two reference picture lists of the current picture where list=0..1 are traversed; wherein the current picture is the picture where the current block is located, that is, the current picture includes the current block. Then it is determined whether Mi B2 refers to the first reference picture list list (if list is 0, then Mi B2 's prediction mode shall be forward or bi-directional; otherwise, Mi B2 's prediction mode shall be backward or bi-directional, which can be expressed as Mi B2 .interDir&(1<<list)). The first reference picture index of Mi in list is denoted as refIdx=Mi B2 .refIdx[list]. If Mi B2 satisfies: Mi B2 .interDir&(1<<list)&&!isAvailable[list][refIdx][0], that is, Mi B2 refers to list, and the MV at the top-left position of the reference picture with the first reference picture index refIdx in list is unavailable. In this case, Mi B2 can be saved as candidate motion information at the top-left position of the reference picture with the first reference picture index refIdx in list (denoted as Mi[list][refIdx][0]), and the first available identification information isAvailable[list][refIdx][0] is marked as available. B2
[0234] In some embodiments, in the process of sequentially checking motion information of a plurality of adjacent blocks, for a next adjacent block (e.g., a second adjacent block), when the motion information of the second adjacent block satisfies an availability condition (e.g., the motion information of the second adjacent block is motion information in inter prediction mode), the reference picture list and reference picture index corresponding to the motion information of the second adjacent block are determined; this includes the following three possible cases:
[0235] First possible case: when it is determined that the motion information of the second adjacent block refers to the reference picture corresponding to the first reference picture index in the first reference picture list, and the first available identification information indicates available, the motion information of the second adjacent block is not taken as a candidate motion information corresponding to the preset position.
[0236] For example, still taking the top-left position of the current block as an example, if the motion information of the second adjacent block (e.g., the motion information at position B3, which can be denoted as Mi B3 It is available, and Mi B3 It also references the reference image corresponding to the first reference image index refIdx in the first reference image list, but the first available identifier isAvailable[list][refIdx][0] has already been marked as available, so Mi is not included. B3 As a candidate motion information corresponding to the top left position, Mi is no longer considered. B3 The candidate motion information is saved as the top-left position of the reference image with the first reference image index refIdx in the list.
[0237] The second possible scenario is as follows: When determining the motion information of the second adjacent block by referring to the reference image corresponding to the second reference image index in the first reference image list, the second available identifier information of the candidate motion information of the current block at the preset position is determined according to the first reference image list and the second reference image index; and when the second available identifier information indicates that it is unavailable, the motion information of the second adjacent block is determined as another candidate motion information corresponding to the preset position, and the second available identifier information is marked as available.
[0238] For example, still taking the top-left position of the current block as an example, if the motion information of the second adjacent block (e.g., the motion information of position B3, can be represented as Mi) B3 It is available, and Mi B3 The reference image corresponding to the second reference image index refIdx1 in the first reference image list is used, but the second availability identifier isAvailable[list][refIdx1][0] is unavailable. Therefore, Mi can be used... B3 As another candidate motion information corresponding to the top left position, it is saved as the candidate motion information of the top left position of the reference image with the second reference image index refIdx1 in the list (represented as Mi[list][refIdx1][0]), and the second available identifier information isAvailable[list][refIdx1][0] is marked as available.
[0239] The third possible scenario: When determining the motion information of the second adjacent block by referring to the reference image corresponding to the third reference image index in the second reference image list, the third available identifier information of the candidate motion information of the current block at the preset position is determined according to the second reference image list and the third reference image index; and when the third available identifier information indicates that it is unavailable, the motion information of the second adjacent block is determined as another candidate motion information corresponding to the preset position, and the third available identifier information is marked as available.
[0240] For example, still taking the top-left position of the current block as an example, if the motion information of the second adjacent block (e.g., the motion information of position B3, can be represented as Mi) B3 It is available, and Mi B3 The reference image corresponding to the third reference image index refIdx2 in the second reference image list list1 was used, but the third available identifier isAvailable[list1][refIdx2][0] is unavailable. Therefore, Mi can be used... B3 As another candidate motion information corresponding to the top left position, it is saved as the candidate motion information of the top left position of the reference image with the index refIdx2 of the third reference image in list1 (represented as Mi[list1][refIdx2][0]), and the third available identifier information isAvailable[list1][refIdx2][0] is marked as available.
[0241] It should also be noted that, in one possible implementation of the stopping condition for the inspection, the method may include: sequentially inspecting the motion information of multiple adjacent blocks at a preset position in a first order to obtain at least one candidate motion information corresponding to the preset position and available identification information of at least one candidate motion information.
[0242] In other words, under this implementation, the motion information of all these adjacent blocks can be checked, and then at least one candidate motion information corresponding to the preset position and at least one available identifier information of the candidate motion information can be obtained.
[0243] In another possible implementation, the method may further include: when the number of available candidate motion information obtained at the preset position reaches a preset number, stopping the inspection of motion information of multiple adjacent blocks at the preset position, and determining at least one candidate motion information corresponding to the preset position and available identification information of at least one candidate motion information based on the obtained candidate motion information.
[0244] In other words, this implementation method can limit the number of available candidate motion information stored at a preset position. For example, assuming the preset number (or "maximum number") is 2, if the preset position has already obtained 2 candidate motion information, then the subsequent adjacent blocks to be checked are skipped, that is, the traversal of the motion information of these multiple adjacent blocks is stopped, and the 2 candidate motion information already obtained is directly determined as at least one candidate motion information corresponding to the preset position, along with the available identification information for determining this at least one candidate motion information.
[0245] In the embodiments of this application, for a preset position, the reference images for the at least one candidate motion information stored therein are different, that is, the reference images for the at least one candidate motion information come from different reference image lists and / or different reference image indices.
[0246] In some embodiments, the method may further include: determining temporal motion information corresponding to a block in a reference image adjacent to a preset position; scaling the temporal motion information to determine a candidate motion information corresponding to the preset position; and determining available identification information of the candidate motion information based on the reference image.
[0247] It should be noted that, in the embodiments of this application, if multiple adjacent blocks at the preset position include blocks adjacent to the preset position in the reference image, then after determining the corresponding temporal motion information, the temporal motion information can be scaled and adjusted to obtain a candidate motion information corresponding to the preset position and the available identification information of the candidate motion information.
[0248] It should also be noted that, in this embodiment, based on the above-described checking method, when the preset position is the upper right position of the current block, the motion information of multiple adjacent blocks at the upper right position can be traversed in the order of B1->B0 to obtain at least one candidate motion information and the available identifier information isAvailable corresponding to the upper right position. Similarly, when the preset position is the lower left position of the current block, the motion information of multiple adjacent blocks at the lower left position can be traversed in the order of A1->A0 to obtain at least one candidate motion information and the available identifier information isAvailable corresponding to the lower left position. Furthermore, when the preset position is the lower right position of the current block, the motion information of the corresponding block in the reference image (e.g., a co-located image) at the lower right position (or offset by a certain distance) of the current block can be checked and adjusted according to the temporal MVP acquisition method to obtain at least one candidate motion information and the available identifier information isAvailable corresponding to the lower right position.
[0249] Understandably, after obtaining at least one candidate motion information corresponding to each of the multiple preset positions such as the top-left, top-right, bottom-left, and bottom-right positions of the current block, as well as the available identifier information isAvailable, it can be used to determine the affine motion information of the current block. In some embodiments, the method may include: determining candidate motion information corresponding to at least two preset positions and the available identifier information of the candidate motion information; determining at least one candidate affine motion information of the current block based on the candidate motion information corresponding to at least two preset positions and the available identifier information of the candidate motion information; determining an affine merging list of the current block based on the at least one candidate affine motion information; and determining the affine motion information of the current block based on the affine merging list of the current block.
[0250] In this embodiment, each candidate affine motion information can be composed of candidate motion information corresponding to at least two preset positions. Here, the combination of at least two preset positions can be as shown in Table 2 above. For example, the at least two candidate motion information used to construct the candidate affine motion information can be determined according to the order in Table 2.
[0251] In Table 2, CPMV1 represents at least one candidate motion information corresponding to the top left position, CPMV2 represents at least one candidate motion information corresponding to the top right position, CPMV3 represents at least one candidate motion information corresponding to the bottom left position, and CPMV4 represents at least one candidate motion information corresponding to the bottom right position. Furthermore, f(CPMV1, CPMV3) represents the derivation of CPMV2 (i.e., the MV corresponding to the top right control point) from CPMV1 and CPMV3, including:
[0252] CPMV2_x=(CPMV1_x<<7)+((CPMV3_y-CPMV1_y)<<(7+log2(Width)-log2(Height));
[0253] CPMV2_y=(CPMV1_y<<7)+((CPMV3_x-CPMV1_x)<<(7+log2(Width)-log2(Height)).
[0254] In this embodiment, when the model index is 2, it can be determined by CPMV1, CPMV2, and CPMV4. "CPMV4+CPMV1-CPMV2" indicates that the lower left control point MV can be derived from candidate motion information at preset positions such as CPMV1, CPMV2, and CPMV4. Similarly, when the model index is 3, it can be determined by CPMV1, CPMV3, and CPMV4. "CPMV4+CPMV1-CPMV3" indicates that the upper right control point MV can be derived from candidate motion information at preset positions such as CPMV1, CPMV3, and CPMV4. When the model index is 4, it can be determined by CPMV2, CPMV3, and CPMV4. "CPMV2+CPMV3-CPMV4" indicates that the upper left control point MV can be derived from candidate motion information at preset positions such as CPMV2, CPMV3, and CPMV4.
[0255] In some embodiments, determining candidate motion information corresponding to at least two preset positions may include: determining a first reference image list corresponding to the current block and a first reference image index in the first reference image list; determining at least two preset positions required for candidate affine motion information; and determining candidate motion information corresponding to at least two preset positions based on the first reference image list, the first reference image index, and the at least two preset positions.
[0256] It should be noted that, in the embodiments of this application, determining the candidate motion information corresponding to at least two preset positions can be achieved by traversing the model index, traversing the reference image list, and traversing the reference image index in any order.
[0257] For example, the construction order can be as follows: first, traverse the model index in Table 2 to determine at least two preset positions required for candidate affine motion information; then, traverse the reference image list list = 0..1 of the current image to determine the current list; then, traverse the reference image index refIdx = 0..4 to determine the current refIdx, so as to obtain the candidate motion information corresponding to each of the at least two preset positions used in the current construction.
[0258] For example, the construction order can also be adjusted as follows: first, traverse the reference image list list = 0..1 of the current image to determine the current list; then traverse the reference image index refIdx = 0..4 to determine the current refIdx; then traverse the model index in Table 2 to determine at least two preset positions required for the candidate affine motion information, so as to obtain the candidate motion information corresponding to each of the at least two preset positions used in the current construction.
[0259] For example, the construction order can also be adjusted as follows: first, traverse the reference image index refIdx = 0..4 to determine the current refIdx; then traverse the reference image list list = 0..1 of the current image to determine the current list; then traverse the model index in Table 2 to determine at least two preset positions required for the candidate affine motion information, so as to obtain the candidate motion information corresponding to each of the at least two preset positions used in the current construction.
[0260] In some embodiments, the method may further include: if the candidate motion information corresponding to at least two preset positions refers to the reference image corresponding to the first reference image index in the first reference image list and all available identification information is available, then the candidate affine motion information is determined based on the candidate motion information corresponding to at least two preset positions.
[0261] For example, taking model index 1 in Table 2 as an example, first iterate through the reference image list list = 0..1, and then iterate through the reference image indices refIdx = 0..4. Check whether the corresponding control points MV are all available. For example, when model index is 1, candidate motion information corresponding to the upper left, upper right, and lower left positions is needed. Then we can determine whether the following conditions are met: isAvailable[list][refIdx][0]&&isAvailable[list][refIdx][1]&&isAvailable[list][refIdx][2]. If they are met, then candidate affine motion information is constructed based on Mi[list][refIdx]. That is, CPMV1 = Mi[list][refIdx][0], CPMV2 = Mi[list][refIdx][1], CPMV3 = MI[list][refIdx][2], thereby determining the candidate affine motion information, that is, constructing CPMVP candidates.
[0262] It should be noted that after obtaining the candidate affine motion information, it can be compared with other candidates in the affine merging list of the current block to determine whether the obtained candidate affine motion information is redundant, i.e., a deduplication operation is performed. If the obtained candidate affine motion information is not redundant, it can be added to the affine merging list of the current block.
[0263] In some embodiments, for each candidate affine motion information, the candidate affine motion information may include an affine type and at least two candidate motion information. Specifically, when the affine type is a 4-parameter affine type, the candidate affine motion information is determined to include candidate motion information corresponding to two preset positions; when the affine type is a 6-parameter affine type, the candidate affine motion information is determined to include candidate motion information corresponding to three preset positions.
[0264] It should also be noted that, in this embodiment of the application, Table 2 is still used as an example. The model indices 1, 2, 3 and 4 are all 6-parameter affine types. The motion information of the current block (i,j) position can be determined by three candidate motion information, as shown in the aforementioned formula (2); the model indices 5 and 6 are both 4-parameter affine types. The motion information of the current block (i,j) position can be determined by two candidate motion information, as shown in the aforementioned formula (1).
[0265] In this embodiment of the application, after obtaining at least one candidate affine motion information, when constructing the affine merging list of the current block, the at least one candidate affine motion information can be reordered according to the template error value, or the reordered part of the candidate affine motion information can be added to the affine merging list of the current block.
[0266] In some embodiments, the method may further include: determining the current template of the current block; calculating the template error between the current template of the current block and the reference template corresponding to at least one candidate affine motion information, and determining the template error value corresponding to each of the at least one candidate affine motion information; sorting the at least one candidate affine motion information in ascending order of template error value, and forming an affine merging list of the current block from the top M candidate affine motion information; wherein M is a positive integer.
[0267] It should be noted that, in the embodiments of this application, the template error calculation can be SAD error calculation, SATD error calculation, or other error calculation methods, such as MSE, MAE, etc., without any limitation.
[0268] It should also be noted that, in this embodiment, template error can be calculated based on the current template of the current block and the reference template of the reference block corresponding to each candidate affine motion information to obtain the template error value corresponding to each candidate affine motion information; then, at least one candidate affine motion information is reordered in ascending order of template error values, and the reordered at least one candidate affine motion information is added to the affine merging list. Considering the length of the affine merging list, only a portion of the candidate affine motion information with relatively small index numbers can be added to the affine merging list, for example, only the top M candidate affine motion information can be added to the affine merging list. For example, the value of M can be 6, but it is not specifically limited.
[0269] In this embodiment of the application, after obtaining the affine merging list of the current block, determining the affine motion information of the current block based on the affine merging list can include: determining M candidate affine motion information included in the affine merging list, where M is a positive integer; calculating the encoding cost of the current block based on the M candidate affine motion information to determine the cost result corresponding to each of the M candidate affine motion information; determining the minimum cost result among the cost results corresponding to each of the M candidate affine motion information, and determining the candidate affine motion information corresponding to the minimum cost result as the affine motion information of the current block.
[0270] In this embodiment, the affine merging list includes M candidate affine motion information. The optimal candidate affine motion information, i.e., the candidate affine motion information corresponding to the minimum cost result, can be determined from these M candidate affine motion information. Furthermore, the encoding cost calculation here can be rate-distortion cost calculation, or other distortion cost calculations, such as MSE, MAE, etc. That is, the optimal candidate affine motion information can be selected based on distortion cost, bit overhead, etc., and used as the affine motion information for the current block.
[0271] In some embodiments, the method may further include: determining a candidate index for merging the current block; encoding the candidate index for merging the current block; and writing the obtained encoded bits into the bitstream.
[0272] In this embodiment, the merge candidate index can be represented by `pu_merge_idx`, which is a block-level syntax element used to indicate the index number (or "position") of the affine motion information of the current block in the affine merge list. For example, the value of `pu_merge_idx` can be an integer such as 0, 1, or 2. Since the affine merge list only includes the top M candidate affine motion information items, this also saves the number of bits corresponding to the merge candidate index in the bitstream.
[0273] In this way, based on at least two motion information that meet the availability conditions from multiple adjacent blocks, an affine merging list for the current block can be constructed. Then, by combining the merging candidate index in the bitstream, the selected affine motion information for the current block can be determined.
[0274] S1103, Determine the predicted value of the current block based on the affine motion information of the current block.
[0275] In this embodiment, after obtaining the affine motion information of the current block, the current block can be predicted based on the affine motion information to obtain the predicted value of the current block. For example, if the affine motion information indicates that the affine type is a 4-parameter affine type, then the motion information within the current block can be determined using the aforementioned formula (1) and affine motion compensation can be performed to obtain the predicted value of the current block; if the affine motion information indicates that the affine type is a 6-parameter affine type, then the motion information within the current block can be determined using the aforementioned formula (2) and affine motion compensation can be performed to obtain the predicted value of the current block.
[0276] It should also be noted that, in this embodiment, some indication information in the form of syntax elements, or flags, can be written into the bitstream. Thus, based on the relevant syntax elements written into the bitstream, the decoding end can quickly determine whether the current block uses affine merging mode by parsing the bitstream. For example, the first syntax element can be used to indicate whether the current block uses merging mode, and the second syntax element can be used to indicate whether the current block uses sub-block-based motion compensation mode.
[0277] In some embodiments, the method may further include: determining the value of a first syntax element and the value of a second syntax element; encoding the value of the first syntax element and the value of the second syntax element respectively, and writing the obtained encoded bits into a bitstream.
[0278] In this embodiment, the first syntax element can be represented by pu_merge_flag, which represents the enable flag for the block-level merge mode. For example, if the current block uses the merge mode, then the value of the first syntax element can be determined to be a first value; if the current block does not use the merge mode, then the value of the first syntax element can be determined to be a second value.
[0279] In this embodiment, the second syntax element can be represented by `cu_affine_flag`, which represents the enable flag for the block-level sub-block-based motion compensation mode. For example, if the current block uses the sub-block-based motion compensation mode, the value of the second syntax element can be determined to be a first value; if the current block does not use the sub-block-based motion compensation mode, the value of the second syntax element can be determined to be a second value. It should be noted that "sub-block-based motion compensation mode" here is a general term, encompassing both the sbTMVP mode and the Affine mode.
[0280] It should also be noted that the first value is different from the second value, and both the first and second values can be in parameter form or in numeric form. For example, the first value can be set to 1 and the second value can be set to 0; or, the first value can be set to true and the second value can be set to false; no restrictions are imposed here.
[0281] Thus, based on the high-level mode enable flag, the size and position of the current block, it can be determined whether the current block uses affine merging mode for encoding. If the current block uses affine merging mode for encoding, the values of `pu_merge_flag` and `cu_affine_flag` can be written into the bitstream as true. At this point, the encoding method of this embodiment can be executed, i.e., the encoding flow shown in Figure 11. For example, illustrative descriptions of the relevant syntax elements are shown in Table 3 above.
[0282] It is also understood that, in the embodiments of this application, the use of the technical solution of the embodiments of this application to determine CPMVP candidates can be controlled by high-level syntax elements such as sequence level and image level. In some embodiments, the value of a third syntax element is determined; the value of the third syntax element is encoded, and the obtained encoded bits are written into the bitstream. Wherein, when the value of the third syntax element is a first value, the steps of determining the motion information of multiple adjacent blocks at a preset position of the current block and determining the affine motion information of the current block based on at least two motion information that meet the availability conditions among the motion information of multiple adjacent blocks are performed.
[0283] In this embodiment, if the current block uses the technical solution of this embodiment to determine the CPMVP candidate, i.e., executes the encoding process shown in FIG11, then the value of the third syntax element can be determined to be a first value; if the current block does not use the technical solution of this embodiment to determine the CPMVP candidate, i.e., does not execute the encoding process shown in FIG11, but uses existing methods in related technologies to determine the CPMVP candidate, then the value of the third syntax element can be determined to be a second value. For example, the first value can be set to 1, and the second value can be set to 0; or, the first value can be set to true, and the second value can be set to false; no limitations are made here.
[0284] In one possible implementation, the third syntax element can be a sequence-level syntax element, such as represented by `sps_affine_cst_opt_enabled_flag`. Specifically, if the current block uses the technical solution of this application embodiment to determine the CPMVP candidate, then `sps_affine_cst_opt_enabled_flag` is set to true; if the current block uses existing methods in related technologies to determine the CPMVP candidate, then `sps_affine_cst_opt_enabled_flag` is set to false. If this syntax element does not exist in the bitstream, it is assumed to be false by default.
[0285] In another possible implementation, the third syntax element can be other higher-level syntax elements, such as the ARMC enable flag sps_aml_enabled_flag, the template matching tool enable flag sps_tm_tools_enabled_flag, etc.
[0286] In this implementation, if the current block uses the adaptive merge candidate reordering method (i.e., the ARMC method), then the value of the third syntax element is determined to be true. In this case, it can also be determined that the current block uses the technical solution of this application embodiment to determine CPMVP candidates. If the current block does not use the adaptive merge candidate reordering method (i.e., the ARMC method), then the value of the third syntax element is determined to be false. In this case, it can be determined that the current block uses existing methods in related technologies to determine CPMVP candidates. That is, when the current block uses the adaptive merge candidate reordering method, the steps of determining the motion information of multiple adjacent blocks at a preset position of the current block and determining the affine motion information of the current block based on at least two motion information from the multiple adjacent blocks that satisfy the availability condition can be executed.
[0287] In another possible implementation, the method may further include: determining block information of the current block; determining motion information of multiple adjacent blocks at a preset position of the current block based on the block information; and determining affine motion information of the current block based on at least two motion information of the multiple adjacent blocks that satisfy the availability condition.
[0288] In this embodiment, the block information of the current block may include information such as the block size and position of the current block. That is, the block information can also be used to determine whether the current block uses the technical solution of this embodiment to determine a CPMVP candidate. When the block information of the current block meets preset conditions, it can be determined that the current block uses the technical solution of this embodiment to determine a CPMVP candidate, i.e., the encoding flow shown in Figure 11 is executed. Here, the preset conditions are the criteria used to determine whether the current block uses the technical solution of this embodiment to determine a CPMVP candidate, such as the block size being greater than a certain threshold, and are not limited here.
[0289] In some embodiments, the method may further include: setting the maximum number of candidates in the affine merge list to L+N when the value of the third syntax element is the first value; where L and N are both positive integers.
[0290] In this embodiment, L represents the initial maximum number of candidates in the affine merge list. When using the technical solution of this embodiment to determine CPMVP candidates in the current block, the maximum number of candidates in the affine merge list can be increased. For example, if the value of sps_affine_cst_opt_enabled_flag is true, then the maximum number of candidates in the affine merge list increases by N, and the maximum number of candidates can be L+N. N can be set to 2, but this is not limited in any way.
[0291] It should also be noted that, in the embodiments of this application, after obtaining the predicted value of the current block, the residual value of the current block can also be determined based on the original value and the predicted value of the current block. For example, by performing a subtraction operation on the original value and the predicted value of the current block, the residual value of the current block can be obtained. Then, the residual value of the current block is transformed, quantized, and coefficient encoded to complete the encoding of the current block.
[0292] This application provides an encoding method that, when constructing the affine merging list of the current block, is no longer limited to the first motion information that meets the usability condition among multiple adjacent blocks' motion information. Instead, it can be constructed based on at least two motion information that meet the usability condition. This avoids the problem in related technologies where effective affine motion information cannot be constructed in some cases, increases the probability of constructing effective affine motion information, makes the prediction of the current block in the affine merging mode more accurate, and thus improves encoding performance.
[0293] In another embodiment of this application, based on the encoding and decoding method of the foregoing embodiments, a method is proposed to construct CPMVP candidates by utilizing motion information of spatially and temporally adjacent blocks as much as possible when constructing the merge list in Affine mode. This is because valid CPMVP candidates require that the reference images corresponding to the MVs of each control point are the same. Therefore, when traversing the motion information of spatially and temporally adjacent blocks, if the reference images corresponding to these motion information are different, these motion information should be saved to increase the probability of constructing a valid CPMVP based on this motion information.
[0294] In one possible implementation, the number of available reference images for the current image can be determined as P, for example, P equals 5; the number of reference image lists is Q, for example, Q equals 2 (forward list L0 and backward list L1). An array isAvailable is initialized to determine whether motion information is available. This array should have a size equal to Q × P × 4, and can be denoted as isAvailable[2][5][4], where 4 represents the need to determine 4 control points MV at preset positions. This array is used to determine whether the control point MV at the corresponding position in the required prediction direction and reference image is available. For example, if the control point MV at the lower left position with reference image index 1 in the forward list L0 is unavailable, it can be represented as isAvailable[0][1][2] being false.
[0295] In this embodiment of the application, the MV of adjacent blocks in the spatial and temporal domains can be checked sequentially, the available motion information can be recorded, and the corresponding judgment can be marked as true. This process may include:
[0296] For example, the motion information of the adjacent blocks at the top left position can be traversed in the order of block B2->B3->A2. If this motion information (e.g., the motion information at position B2) can be represented as Mi... B2 If available, then iterate through the list of reference images for the current image, list = 0..1;
[0297] Determine Mi B2 Did it reference a list (if list is 0, then Mi)? B2 The prediction method should be forward or bidirectional; otherwise, MiB2 the prediction mode shall be backward or bidirectional, which can be expressed as Mi B2 .interDir & (1 << list)). Mi B2 the reference picture index in list is expressed as refIdx = Mi B2 .refIdx[list]. If Mi B2 satisfies: Mi B2 .interDir & (1 << list) && !isAvailable[list][refIdx][0], that is, Mi B2 refers to list, and the control point MV at the top-left position of the reference picture with reference picture index refIdx in list is unavailable. Then Mi B2 is stored as the control point MV at the top-left position of the reference picture with reference picture index refIdx in list (expressed as Mi[list][refIdx][0]), and the corresponding isAvailable[list][refIdx][0] is marked as available;
[0298] It should be noted that the maximum number of motion information that can be stored for each control point, such as the top-left position, top-right position, bottom-left position, and bottom-right position, can also be limited herein, for example, the maximum number is 2. If two available motion information items have been obtained for the current control point, subsequent to-be-checked blocks of the current control point can be skipped, and adjacent blocks of the next control point are directly checked.
[0299] Similarly, the motion information of adjacent blocks at the top-right position can also be traversed in the order of block B1->B0; the motion information of adjacent blocks at the bottom-left position can also be traversed in the order of block A1->A0; in addition, the MV of the corresponding block in the collocated picture at the bottom-right position (or offset by a certain distance) of the current block is checked, and adjusted according to the acquisition method of temporal MVP, so as to obtain available motion information at the bottom-right position.
[0300] According to the above steps, the spatial and temporal adjacent motion information Mi of the current block and the availability identifier isAvailable can be obtained. Further, CPMVP candidates can be constructed in a certain order.
[0301] For example, the CPMV used by the current CPMVP candidate to be constructed can be determined according to Table 2 above. Assuming the model index is 1, first traverse the reference image list list = 0..1 of the current image, and then traverse the reference image indices refIdx = 0..4. Check whether the corresponding control point MVs are all available. For example, if the model index is 1 and the CPMVs of the upper left, upper right, and lower left positions are needed, then determine whether the following conditions are met: isAvailable[list][refIdx][0]&&isAvailable[list][refIdx][1]&&isAvailable[list][refIdx][2]. If the conditions are met, then the CPMVP candidate is constructed based on Mi[list][refIdx]. That is, the CPMVP candidate is determined according to CPMV1 = Mi[list][refIdx][0], CPMV2 = Mi[list][refIdx][1], and CPMV3 = Mi[list][refIdx][2].
[0302] In this way, after obtaining a CPMVP candidate, it can be compared with other candidates in the current merge list to determine whether the current candidate is redundant. If the current candidate is not redundant, it is added to the current merge list.
[0303] In some embodiments, the method proposed in this application improves the affine merge list construction process. If the decoded syntax element determines that the current block uses the Affine merge mode, an affine merge list is constructed for the current block. Specifically, when constructing CPMVP candidates using motion information from adjacent blocks in the spatial and temporal domains, the method in this application is used to record the motion information of adjacent blocks in both the spatial and temporal domains, sequentially traversing the candidate CPMVP model indexes and the motion information corresponding to different reference images, thereby constructing at least one CPMVP candidate to complete the merge list construction. Then, the affine motion information selected for the current block is determined through the decoded merge candidate index, completing subsequent motion compensation, current block reconstruction, and other processes. This completes the decoding of the current block.
[0304] For example, as shown in Table 3 above, for the syntax elements therein, `pu_merge_flag` is the enable flag for the merge mode at the coding block level. If it is true, it means that the current block uses the merge mode; if it is false, the merge mode is not used. `cu_affine_flag` is the enable flag for the sub-block-based motion compensation mode at the coding block level. If it is true, it means that the current block uses the sub-block-based motion compensation mode; if it is false, the sub-block-based motion compensation mode is not used. `pu_merge_idx` is the merge candidate index at the coding block level, used to indicate the position of the selected motion information in the merge list for the current block.
[0305] In one possible implementation, the decoding process at the decoding end is as follows:
[0306] Step 1: Decode the syntax elements. If the value of pu_merge_flag in the current block is true and the value of cu_affine_flag is true, then the current block uses the Affine merge mode.
[0307] Step 2: Construct an Affine merge list for the current block, which may include sbTMVP candidates predicted using the sbTMVP pattern, and CPMVP candidates predicted using affine motion compensation.
[0308] Step 3: When constructing CPMVP candidates using the motion information of adjacent blocks in the spatial and temporal domains (and the value of the flag sps_aml_enabled_flag of the ARMC method is true), the method described in the embodiments of this application is used to record the motion information of adjacent blocks in both the spatial and temporal domains, and sequentially traverse the CPMVP model index of the candidates and the motion information corresponding to different reference images to construct CPMVP candidates.
[0309] Step 4: Complete the construction of the Affine merge list. The candidates in the Affine merge list can be reordered according to the template error value between the current template of the current block and the reference template of the corresponding reference block of the merge candidate.
[0310] Step 5: Determine the motion information of the selected block based on the decoded pu_merge_idx. If the motion information is affine motion compensation, then according to its affine model type, perform CPMV-based affine motion compensation according to the 4-parameter affine type or 6-parameter affine type described in the previous embodiment to obtain the predicted value of the current block. Further, based on the decoding residual coefficients, inverse quantization, inverse transform, and other processes, complete the reconstruction of the current block;
[0311] Step 6: Complete the decoding of the current block according to the steps above.
[0312] In one possible implementation, the encoding process at the encoding end is as follows:
[0313] Step 7: Based on the high-level mode enable flags, the size and position of the current block, determine whether the current block can be encoded using Affine merge mode. If it is determined that the current block uses Affine merge mode, both pu_merge_flag and cu_affine_flag are set to true.
[0314] Step 8, similar to steps 2-4 above, constructs an Affine merge list for the current block, which may include sbTMVP candidates predicted using the sbTMVP model and CPMVP candidates using affine motion compensation. Specifically, when constructing CPMVP candidates using motion information from spatial and temporal adjacent blocks, the method described in this embodiment is used to record the motion information of both spatial and temporal adjacent blocks, sequentially traversing the CPMVP model indexes of the candidates and the motion information corresponding to different reference images, thereby constructing CPMVP candidates.
[0315] Step 9: Select the best motion information for the current block (i.e., the selected motion information for the current block) based on distortion, bit overhead, etc., and encode its corresponding merge candidate index pu_merge_idx in the Affine merge list. The encoding and reconstruction of the current block are completed through processes such as quantization, transformation, and coefficient encoding of the residuals.
[0316] Step 10: After traversing all coding units, the bitstream is output after passing through techniques such as loop filtering and entropy coding.
[0317] Understandably, in the embodiments of this application, for the aforementioned technical solution, "firstly, determine the CPMV used by the current CPMVP candidate to be constructed according to the aforementioned Table 2. Assuming the model index is 1, first traverse the reference image list of the current image list = 0..1, and then traverse the reference image index refIdx = 0..4. Check whether the corresponding control point MVs are all available. For example, if the model index is 1 and the CPMVs of the upper left, upper right, and lower left positions are required, then determine whether the following condition is met: isAvailable[list][refIdx][0]&&isAvailable [list][refIdx][1]&&isAvailable[list][refIdx][2]. If satisfied, then construct CPMVP candidates based on Mi[list][refIdx]. That is, determine CPMVP candidates based on CPMV1=Mi[list][refIdx][0], CPMV2=Mi[list][refIdx][1], CPMV3=MI[list][refIdx][2]. The process of “” can be implemented in any order of traversing the model index, traversing the reference list, and traversing the reference image.
[0318] For example, the construction order can be adjusted as follows:
[0319] ① Iterate through the list of reference images for the current image, list = 0..1, and determine the current list;
[0320] ② Traverse the reference image indices refIdx = 0..4 to determine the current refIdx;
[0321] ③ Determine the CPMV used by the current CPMVP candidate to be constructed according to Table 2 above.
[0322] For example, the construction order can also be adjusted as follows:
[0323] ① Traverse the reference image indices refIdx = 0..4 to determine the current refIdx;
[0324] ② Iterate through the list of reference images for the current image, list = 0..1, and determine the current list;
[0325] ③ Determine the CPMV used by the current CPMVP candidate to be constructed according to Table 2 above.
[0326] It is also understood that, in the embodiments of this application, whether to use this technical solution to construct CPMVP candidates can be controlled by high-level syntax elements such as sequence-level and image-level elements. For example, the sequence-level `sps_affine_cst_opt_enabled_flag` flag indicates that if the value is true, it means that this technical solution is used to construct CPMVP candidates; if it is false, existing methods in related technologies are used to construct CPMVP candidates. If this syntax element does not exist in the bitstream, it defaults to false. In addition, the block size of the current block, the position of the current block, and other information can also be used to determine whether to use this technical solution to construct CPMVP candidates.
[0327] It is also understood that, in the embodiments of this application, this technical solution can be used only when the ARMC method is used in the current block (i.e., after constructing the Affine merge list, the list is reordered according to the template error value) to construct CPMVP candidates; otherwise, existing methods in related technologies are used to construct CPMVP candidates. In this case, whether the current block uses the ARMC method can be determined through high-level syntax elements, such as checking the ARMC method enable flag sps_aml_enabled_flag and the template matching tool enable flag sps_tm_tools_enabled_flag.
[0328] It is also understood that, in the embodiments of this application, the maximum number of candidates in the Affine merge list can be increased when using this technical solution. For example, if the value of sps_affine_cst_opt_enabled_flag in this technical solution is true, the maximum number of candidates in the Affine merge list is increased by N, for example, N can be set to 2.
[0329] In summary, this technical solution proposes an extension method for the CPMVP construction process under the Affine merge mode. The performance improvement of applying this technical solution and expanding the maximum length L of the Affine merge list to (L+2) on the reference software ECM compared to anchor is shown in Tables 4 and 5. Specifically, the performance test results obtained under random access (RA) conditions are shown in Table 4, and the performance test results obtained under low latency (LD) conditions are shown in Table 5.
[0330] Table 4
[0331] Table 5
[0332] It should be noted that some missing test points in Tables 4 and 5 are copies of the ECM encoding results. Tables 4 and 5 show that the encoding and decoding performance of the current block has been improved.
[0333] In this application embodiment, the specific implementation of the aforementioned embodiments is described in detail through the above embodiments. It can be seen that, according to the technical solution of the aforementioned embodiments, in the construction of the merge list in Affine mode, when constructing CPMVP candidates using the motion information of adjacent blocks in the spatial and temporal domains, all available motion information is recorded and used to construct CPMVP candidates, thereby increasing the probability of constructing effective affine motion information, making the prediction of the current block in the affine merging mode more accurate, and thus improving the encoding and decoding performance.
[0334] In another embodiment of this application, based on the same inventive concept as the foregoing embodiments, FIG12 is a schematic diagram of the composition structure of an encoder provided in an embodiment of this application. As shown in FIG12, the encoder 100 may include a first determining unit 1201, a first predicting unit 1202, and an encoding unit 1203, wherein:
[0335] The first determining unit 1201 is configured to determine the motion information of multiple adjacent blocks at a preset position of the current block; and to determine the affine motion information of the current block based on at least two motion information of the multiple adjacent blocks that meet the availability conditions.
[0336] The first prediction unit 1202 is configured to determine the prediction value of the current block based on the affine motion information of the current block.
[0337] In some embodiments, the first determining unit 1201 is further configured to determine the merging candidate index of the current block; wherein the merging candidate index is used to indicate the index number of the affine motion information of the current block in the affine merging list; the encoding unit 1203 is configured to encode the merging candidate index of the current block and write the obtained encoded bits into the bit stream.
[0338] In some embodiments, the first determining unit 1201 is further configured to determine the value of the first syntax element and the value of the second syntax element; wherein the first syntax element is used to indicate whether the current block uses a merge mode, and the second syntax element is used to indicate whether the current block uses a sub-block-based motion compensation mode; the encoding unit 1203 is further configured to encode the value of the first syntax element and the value of the second syntax element respectively, and write the obtained encoded bits into the bitstream.
[0339] In some embodiments, the first determining unit 1201 is further configured to determine the value of the third syntax element; the encoding unit 1203 is further configured to encode the value of the third syntax element and write the obtained encoded bits into the code stream.
[0340] In some embodiments, the first determining unit 1201 is further configured to determine block information of the current block; and, based on the block information, perform the steps of determining motion information of a plurality of adjacent blocks at a preset position of the current block; and, based on the motion information of at least two of the motion information of the plurality of adjacent blocks that satisfy the availability condition, determine the affine motion information of the current block.
[0341] It should be noted that those skilled in the art should understand that the device description of the encoder in the embodiments of this application can be understood with reference to the relevant description of the encoding method described in the foregoing embodiments.
[0342] It should also be noted that, in the embodiments of this application, a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and can also be a module or a non-modular one. Moreover, the components in this embodiment can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit described above can be implemented in hardware or in the form of a software functional module.
[0343] In another embodiment of this application, FIG13 is a schematic diagram of the hardware structure of an encoder provided in an embodiment of this application. As shown in FIG13, the encoder 100 may include: a first communication interface 1301, a first memory 1302, and a first processor 1303; the various components are coupled together through a first bus system 1304. It is understood that the first bus system 1304 is used to realize the connection and communication between these components. In addition to a data bus, the first bus system 1304 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, all buses are labeled as the first bus system 1304 in FIG13. Wherein: the first communication interface 1301 is used for receiving and sending signals during the process of sending and receiving information with other external network elements; the first memory 1302 is used for storing computer programs that can run on the first processor 1303; the first processor 1303 is used to execute the following when running the computer program: determining the motion information of multiple adjacent blocks at a preset position of the current block; determining the affine motion information of the current block based on at least two motion information of the multiple adjacent blocks that meet the availability conditions; and determining the predicted value of the current block based on the affine motion information of the current block.
[0344] It is understood that the first memory 1302 in this embodiment can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDRSDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DRRAM). The first memory 1302 of the system and method described in this application is intended to include, but is not limited to, these and any other suitable types of memory.
[0345] The first processor 1303 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of the first processor 1303 or by instructions in software form. The first processor 1303 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules may reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in the first memory 1302. The first processor 1303 reads the information in the first memory 1302 and completes the steps of the above method in conjunction with its hardware.
[0346] It is understood that the embodiments described in this application can be implemented using hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described in this application, or combinations thereof. For software implementation, the technology described in this application can be implemented through modules (e.g., procedures, functions, etc.) that perform the functions described in this application. Software code can be stored in memory and executed by a processor. The memory can be implemented in the processor or external to the processor.
[0347] Alternatively, as another embodiment, the first processor 1303 is also configured to perform the method described in any of the foregoing embodiments when running the computer program.
[0348] This embodiment provides an encoder that, when constructing the affine merging list of the current block, is no longer limited to the first motion information that meets the usability condition among the motion information of multiple adjacent blocks, but can be constructed based on at least two motion information that meet the usability condition. This avoids the problem in related technologies that effective affine motion information cannot be constructed in some cases, increases the probability of constructing effective affine motion information, makes the prediction of the current block in the affine merging mode more accurate, and thus improves the coding performance.
[0349] In another embodiment of this application, based on the same inventive concept as the foregoing embodiments, FIG14 is a schematic diagram of the composition structure of a decoder provided in an embodiment of this application. As shown in FIG14, the decoder 200 may include a second determining unit 1401, a second predicting unit 1402, and a decoding unit 1403, wherein:
[0350] The second determining unit 1401 is configured to determine the motion information of a plurality of adjacent blocks at a preset position of the current block; and to determine the affine motion information of the current block based on at least two motion information of the plurality of adjacent blocks that meet the availability conditions.
[0351] The second prediction unit 1402 is configured to determine the prediction value of the current block based on the affine motion information of the current block.
[0352] In some embodiments, the decoding unit 1403 is configured to parse the merge candidate index in the bitstream; the second determining unit 1401 is further configured to determine the affine motion information of the current block based on the merge candidate index and the affine merge list.
[0353] In some embodiments, the decoding unit 1403 is further configured to parse a first syntax element and a second syntax element in the bitstream; wherein the first syntax element is used to indicate whether the current block uses a merging mode, and the second syntax element is used to indicate whether the current block uses a sub-block-based motion compensation mode; the second determining unit 1401 is further configured to determine that the current block uses an affine merging mode when the value of the first syntax element is a first value and the value of the second syntax element is a first value.
[0354] In some embodiments, the decoding unit 1403 is further configured to parse a third syntax element in the bitstream; the second determining unit 1401 is further configured to, when the value of the third syntax element is a first value, perform the step of determining the motion information of a plurality of adjacent blocks at a preset position of the current block; and determine the affine motion information of the current block based on at least two motion information that satisfy the availability condition among the motion information of the plurality of adjacent blocks.
[0355] It should be noted that those skilled in the art should understand that the device description of the decoder in the embodiments of this application can be understood with reference to the relevant description of the decoding method described in the foregoing embodiments.
[0356] It should also be noted that, in the embodiments of this application, a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and can also be a module or a non-modular one. Moreover, the components in this embodiment can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit described above can be implemented in hardware or in the form of a software functional module.
[0357] In another embodiment of this application, FIG15 is a schematic diagram of the hardware structure of a decoder provided in an embodiment of this application. As shown in FIG15, the decoder 200 may include: a second communication interface 1501, a second memory 1502, and a second processor 1503; the various components are coupled together through a second bus system 1504. It is understood that the second bus system 1504 is used to realize the connection and communication between these components. In addition to a data bus, the second bus system 1504 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, all buses are labeled as the second bus system 1504 in FIG15. Wherein: the second communication interface 1501 is used for receiving and sending signals during the process of sending and receiving information with other external network elements; the second memory 1502 is used for storing computer programs that can run on the second processor 1503; the second processor 1503 is used to execute the following when running the computer program: determining the motion information of multiple adjacent blocks at a preset position of the current block; determining the affine motion information of the current block based on at least two motion information of the multiple adjacent blocks that meet the availability conditions; and determining the predicted value of the current block based on the affine motion information of the current block.
[0358] Alternatively, as another embodiment, the second processor 1503 is also configured to perform the method described in any of the foregoing embodiments when running the computer program.
[0359] It is understood that the second memory 1502 has similar hardware functions to the first memory 1302, and the second processor 1503 has similar hardware functions to the first processor 1303; these will not be described in detail here.
[0360] This embodiment provides a decoder that, when constructing the affine merging list of the current block, is no longer limited to the first motion information that meets the usability condition among multiple adjacent blocks. Instead, it can be constructed based on at least two motion information that meet the usability condition. This avoids the problem in related technologies where effective affine motion information cannot be constructed in some cases, increases the probability of constructing effective affine motion information, makes the prediction of the current block in the affine merging mode more accurate, and thus improves decoding performance.
[0361] In another embodiment of this application, FIG16 is a schematic diagram of the composition structure of an encoding and decoding system provided in an embodiment of this application. As shown in FIG16, the encoding and decoding system 160 may include an encoder 1601 and a decoder 1602.
[0362] In this embodiment, encoder 1601 can be any of the encoders described in the foregoing embodiments, and decoder 1602 can be any of the decoders described in the foregoing embodiments.
[0363] In some embodiments, this application also provides a computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the method as described in any of the foregoing embodiments. Specifically, when executed by a first processor, the computer program implements the encoding method as described in any of the foregoing embodiments, or when executed by a second processor, it implements the decoding method as described in any of the foregoing embodiments.
[0364] In some embodiments, this application also provides a computer program product, including a computer program or instructions. When executed by a processor, the computer program or instructions implement the method as described in any of the foregoing embodiments. Specifically, when executed by a first processor, the computer program or instructions implement the encoding method as described in any of the foregoing embodiments, or when executed by a second processor, they implement the decoding method as described in any of the foregoing embodiments.
[0365] In some embodiments, this application also provides a computer program that, when executed by a processor, implements the method as described in any of the foregoing embodiments. Specifically, when executed by a first processor, the computer program or instructions implement the encoding method as described in any of the foregoing embodiments, or when executed by a second processor, implement the decoding method as described in any of the foregoing embodiments.
[0366] In some embodiments, this application also provides a computer-readable storage medium storing a bitstream thereon. The bitstream is generated by performing the steps of the encoding method as described in any of the foregoing embodiments.
[0367] In this embodiment of the application, the information to be encoded in the encoding method includes at least one of the following: the candidate index for merging the current block, the value of the first syntax element, the value of the second syntax element, and the value of the third syntax element. Here, the bitstream is generated based on encoding processing of this information to be encoded.
[0368] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0369] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described apparatus and unit can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0370] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0371] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0372] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes: USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, optical disks, and other media capable of storing program code.
[0373] It should be noted that, in this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0374] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0375] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0376] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0377] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method or device embodiments.
[0378] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims. Industrial applicability
[0379] In this embodiment, both the encoding and decoding ends determine the motion information of multiple adjacent blocks at a preset position of the current block; based on at least two motion information pieces that meet the usability condition from the motion information of the multiple adjacent blocks, the affine motion information of the current block is determined; and based on the affine motion information of the current block, the prediction value of the current block is determined. Thus, when constructing the affine merging list of the current block, it is no longer limited to the first motion information piece that meets the usability condition from the motion information of the multiple adjacent blocks, but can be constructed based on at least two motion information pieces that meet the usability condition. This avoids the problem in related technologies where effective affine motion information cannot be constructed in some cases, increases the probability of constructing effective affine motion information, makes the prediction of the current block in affine merging mode more accurate, and thus improves encoding and decoding performance.
Claims
1. A decoding method applied to a decoder, the method comprising: Determine the motion information of multiple adjacent blocks at the preset position of the current block; Based on at least two motion information pieces that satisfy the availability condition from the motion information of the plurality of adjacent blocks, the affine motion information of the current block is determined; The predicted value of the current block is determined based on the affine motion information of the current block.
2. The method according to claim 1, wherein, The preset position includes one of the following: the upper left position of the current block, the upper right position of the current block, the lower left position of the current block, and the lower right position of the current block; The plurality of adjacent blocks at the preset position include: blocks adjacent to the spatial domain of the preset position, or blocks adjacent to the preset position in the reference image.
3. The method according to claim 1, wherein, The step of determining the affine motion information of the current block based on at least two motion information pieces that satisfy the availability condition from the motion information of the plurality of adjacent blocks includes: From the at least two motion information that meet the availability conditions, determine one or more candidate motion information and availability identification information of the one or more candidate motion information; The affine motion information of the current block is determined based on one or more candidate motion information indicated by the available identification information.
4. The method according to claim 1, wherein determining one or more candidate motion information and availability identifier information of the one or more candidate motion information from the at least two motion information satisfying the availability condition comprises: From the at least two motion information that meet the availability conditions, determine the first candidate motion information; When the first candidate motion information does not overlap with existing candidate motion information that is already marked as available, the available identification information of the first candidate motion information is marked as available.
5. The method according to claim 1, wherein, The method further includes: When the motion information of the first adjacent block at the preset position meets the availability condition, a list of reference images and a reference image index corresponding to the motion information of the first adjacent block are determined. When determining the reference image corresponding to the first reference image index in the first reference image list when the motion information of the first adjacent block is referenced, the first available identifier information of the candidate motion information of the current block at the preset position is determined according to the first reference image list and the first reference image index. When the first available identifier indicates that the information is unavailable, the motion information of the first adjacent block is determined as a candidate motion information corresponding to the preset position, and the first available identifier is marked as available.
6. The method according to claim 5, wherein, The motion information of the first adjacent block meets the usability conditions, including: the motion information of the first adjacent block is motion information in the inter-frame prediction mode.
7. The method according to claim 5, wherein, The method further includes: When the motion information of the second adjacent block at the preset position meets the availability condition, a list of reference images and a reference image index corresponding to the motion information of the second adjacent block are determined. When the motion information of the second adjacent block is determined to be referenced by the reference image corresponding to the first reference image index in the first reference image list, and the first available identifier information indicates that it is available, the motion information of the second adjacent block is not used as a candidate motion information corresponding to the preset position.
8. The method according to claim 5, wherein, The method further includes: When the motion information of the second adjacent block at the preset position meets the availability condition, a list of reference images and a reference image index corresponding to the motion information of the second adjacent block are determined. When determining the motion information of the second adjacent block by referring to the reference image corresponding to the second reference image index in the first reference image list, the second available identifier information of the candidate motion information of the current block at the preset position is determined according to the first reference image list and the second reference image index; and when the second available identifier information indicates that it is unavailable, the motion information of the second adjacent block is determined as another candidate motion information corresponding to the preset position, and the second available identifier information is marked as available. or, When determining the motion information of the second adjacent block by referring to the reference image corresponding to the third reference image index in the second reference image list, the third available identifier information of the candidate motion information of the current block at the preset position is determined according to the second reference image list and the third reference image index; and when the third available identifier information indicates that it is unavailable, the motion information of the second adjacent block is determined as another candidate motion information corresponding to the preset position, and the third available identifier information is marked as available.
9. The method according to claim 5, wherein, The method further includes: The motion information of multiple adjacent blocks at the preset position is checked in sequence according to the first order to obtain at least one candidate motion information corresponding to the preset position and the available identification information of the at least one candidate motion information.
10. The method according to claim 9, wherein, The method further includes: When the number of available candidate motion information obtained at the preset position reaches a preset number, the inspection of motion information of multiple adjacent blocks at the preset position is stopped, and at least one candidate motion information corresponding to the preset position and the available identification information of the at least one candidate motion information are determined based on the obtained candidate motion information.
11. The method according to claim 2, wherein, The method further includes: Determine the temporal motion information of the block in the reference image that is adjacent to the preset position; The temporal motion information is scaled to determine a candidate motion information corresponding to the preset position; and the available identification information of the candidate motion information is determined based on the reference image.
12. The method according to claim 1, wherein, Determining the affine motion information of the current block includes: Determine candidate motion information corresponding to at least two preset positions and available identification information for the candidate motion information; Based on the candidate motion information corresponding to each of the at least two preset positions and the available identifier information of the candidate motion information, at least one candidate affine motion information of the current block is determined; Based on the at least one candidate affine motion information, determine the affine merging list of the current block; Based on the affine merging list of the current block, determine the affine motion information of the current block.
13. The method according to claim 12, wherein, The determination of candidate motion information corresponding to at least two preset positions includes: Determine the first reference image list corresponding to the current block and the first reference image index in the first reference image list; At least two preset positions are required to determine the candidate affine motion information; Based on the first reference image list, the first reference image index, and the at least two preset positions, determine the candidate motion information corresponding to each of the at least two preset positions.
14. The method according to claim 13, wherein, The method further includes: If the candidate motion information corresponding to each of the at least two preset positions refers to the reference image corresponding to the first reference image index in the first reference image list and all available identification information is available, then the candidate affine motion information is determined based on the candidate motion information corresponding to each of the at least two preset positions.
15. The method according to claim 14, wherein, The candidate affine motion information includes the affine type and at least two candidate motion information; the method further includes: When the affine type is a 4-parameter affine type, the candidate affine motion information is determined to include candidate motion information corresponding to two preset positions. When the affine type is a 6-parameter affine type, the candidate affine motion information is determined to include candidate motion information corresponding to three preset positions.
16. The method according to claim 12, wherein, The step of determining the affine merging list of the current block based on the at least one candidate affine motion information includes: Determine the current template of the current block; The template error is calculated between the current template of the current block and the reference template corresponding to the at least one candidate affine motion information, and the template error value corresponding to each of the at least one candidate affine motion information is determined. The at least one candidate affine motion information is sorted in ascending order of the template error value, and the top M candidate affine motion information are combined to form the affine merging list of the current block; where M is a positive integer.
17. The method according to claim 12, wherein, Determining the affine motion information of the current block based on the affine merging list of the current block includes: Parse the merge candidate index in the bitstream; Based on the merge candidate index and the affine merge list, the affine motion information of the current block is determined.
18. The method according to any one of claims 1 to 17, wherein, The method further includes: The first syntax element and the second syntax element in the parsed bitstream are used to indicate whether the current block uses a merge mode, and the second syntax element is used to indicate whether the current block uses a sub-block-based motion compensation mode. If the first syntax element has a first value and the second syntax element has a first value, then the current block is determined to use the affine merge mode.
19. The method according to any one of claims 1 to 18, wherein, The method further includes: Parse the third syntax element in the bitstream; When the value of the third syntax element is the first value, the steps of determining the motion information of multiple adjacent blocks at the preset position of the current block and determining the affine motion information of the current block based on at least two motion information of the multiple adjacent blocks that meet the availability condition are executed.
20. The method according to any one of claims 1 to 18, wherein, The method further includes: Determine the block information of the current block; Based on the block information, the steps of determining the motion information of multiple adjacent blocks at a preset position of the current block, and determining the affine motion information of the current block based on at least two motion information of the multiple adjacent blocks that meet the availability conditions.
21. The method according to claim 19, wherein, The method further includes: When the value of the third syntax element is the first value, the maximum number of candidates in the affine merge list is set to L+N; where L and N are both positive integers.
22. An encoding method applied to an encoder, the method comprising: Determine the motion information of multiple adjacent blocks at the preset position of the current block; Based on at least two motion information pieces that satisfy the availability condition from the motion information of the plurality of adjacent blocks, the affine motion information of the current block is determined; The predicted value of the current block is determined based on the affine motion information of the current block.
23. The method according to claim 22, wherein, The preset position includes one of the following: the upper left position of the current block, the upper right position of the current block, the lower left position of the current block, and the lower right position of the current block; The plurality of adjacent blocks at the preset position include: blocks adjacent to the spatial domain of the preset position, or blocks adjacent to the preset position in the reference image.
24. The method according to claim 22, wherein, The step of determining the affine motion information of the current block based on at least two motion information pieces that satisfy the availability condition from the motion information of the plurality of adjacent blocks includes: From the at least two motion information that meet the availability conditions, determine one or more candidate motion information and availability identification information of the one or more candidate motion information; The affine motion information of the current block is determined based on one or more candidate motion information indicated by the available identification information.
25. The method of claim 22, wherein determining one or more candidate motion information and availability identifier information of the one or more candidate motion information from the at least two motion information satisfying the availability condition comprises: From the at least two motion information that meet the availability conditions, determine the first candidate motion information; When the first candidate motion information does not overlap with existing candidate motion information that is already marked as available, the available identification information of the first candidate motion information is marked as available.
26. The method according to claim 22, wherein, The method further includes: When the motion information of the first adjacent block at the preset position meets the availability condition, a list of reference images and a reference image index corresponding to the motion information of the first adjacent block are determined. When determining the reference image corresponding to the first reference image index in the first reference image list when the motion information of the first adjacent block is referenced, the first available identifier information of the candidate motion information of the current block at the preset position is determined according to the first reference image list and the first reference image index. When the first available identifier indicates that the information is unavailable, the motion information of the first adjacent block is determined as a candidate motion information corresponding to the preset position, and the first available identifier is marked as available.
27. The method according to claim 26, wherein, The motion information of the first adjacent block meets the usability conditions, including: the motion information of the first adjacent block is motion information in the inter-frame prediction mode.
28. The method according to claim 26, wherein, The method further includes: When the motion information of the second adjacent block at the preset position meets the availability condition, a list of reference images and a reference image index corresponding to the motion information of the second adjacent block are determined. When the motion information of the second adjacent block is determined to be referenced by the reference image corresponding to the first reference image index in the first reference image list, and the first available identifier information indicates that it is available, the motion information of the second adjacent block is not used as a candidate motion information corresponding to the preset position.
29. The method according to claim 26, wherein, The method further includes: When the motion information of the second adjacent block at the preset position meets the availability condition, a list of reference images and a reference image index corresponding to the motion information of the second adjacent block are determined. When determining the motion information of the second adjacent block by referring to the reference image corresponding to the second reference image index in the first reference image list, the second available identifier information of the candidate motion information of the current block at the preset position is determined according to the first reference image list and the second reference image index; and when the second available identifier information indicates that it is unavailable, the motion information of the second adjacent block is determined as another candidate motion information corresponding to the preset position, and the second available identifier information is marked as available. or, When determining the motion information of the second adjacent block by referring to the reference image corresponding to the third reference image index in the second reference image list, the third available identifier information of the candidate motion information of the current block at the preset position is determined according to the second reference image list and the third reference image index; and when the third available identifier information indicates that it is unavailable, the motion information of the second adjacent block is determined as another candidate motion information corresponding to the preset position, and the third available identifier information is marked as available.
30. The method according to claim 26, wherein, The method further includes: The motion information of multiple adjacent blocks at the preset position is checked in sequence according to the first order to obtain at least one candidate motion information corresponding to the preset position and the available identification information of the at least one candidate motion information.
31. The method according to claim 30, wherein, The method further includes: When the number of available candidate motion information obtained at the preset position reaches a preset number, the inspection of motion information of multiple adjacent blocks at the preset position is stopped, and at least one candidate motion information corresponding to the preset position and the available identification information of the at least one candidate motion information are determined based on the obtained candidate motion information.
32. The method according to claim 23, wherein, The method further includes: Determine the temporal motion information of the block in the reference image that is adjacent to the preset position; The temporal motion information is scaled to determine a candidate motion information corresponding to the preset position; and the available identification information of the candidate motion information is determined based on the reference image.
33. The method according to claim 22, wherein, Determining the affine motion information of the current block includes: Determine candidate motion information corresponding to at least two preset positions and available identification information for the candidate motion information; Based on the candidate motion information corresponding to each of the at least two preset positions and the available identifier information of the candidate motion information, at least one candidate affine motion information of the current block is determined; Based on the at least one candidate affine motion information, determine the affine merging list of the current block; Based on the affine merging list of the current block, determine the affine motion information of the current block.
34. The method according to claim 33, wherein, The determination of candidate motion information corresponding to at least two preset positions includes: Determine the first reference image list corresponding to the current block and the first reference image index in the first reference image list; At least two preset positions are required to determine the candidate affine motion information; Based on the first reference image list, the first reference image index, and the at least two preset positions, determine the candidate motion information corresponding to each of the at least two preset positions.
35. The method according to claim 34, wherein, The method further includes: If the candidate motion information corresponding to each of the at least two preset positions refers to the reference image corresponding to the first reference image index in the first reference image list and all available identification information is available, then the candidate affine motion information is determined based on the candidate motion information corresponding to each of the at least two preset positions.
36. The method according to claim 35, wherein, The candidate affine motion information includes the affine type and at least two candidate motion information; the method further includes: When the affine type is a 4-parameter affine type, the candidate affine motion information is determined to include candidate motion information corresponding to two preset positions. When the affine type is a 6-parameter affine type, the candidate affine motion information is determined to include candidate motion information corresponding to three preset positions.
37. The method according to claim 33, wherein, The step of determining the affine merging list of the current block based on the at least one candidate affine motion information includes: Determine the current template of the current block; The template error is calculated between the current template of the current block and the reference template corresponding to the at least one candidate affine motion information, and the template error value corresponding to each of the at least one candidate affine motion information is determined. The at least one candidate affine motion information is sorted in ascending order of the template error value, and the top M candidate affine motion information are combined to form the affine merging list of the current block; where M is a positive integer.
38. The method according to claim 33, wherein, Determining the affine motion information of the current block based on the affine merging list of the current block includes: Determine the M candidate affine motion information included in the affine merging list, where M is a positive integer; The encoding cost of the current block is calculated based on the M candidate affine motion information, and the cost result corresponding to each of the M candidate affine motion information is determined. The minimum cost result is determined from the cost results corresponding to each of the M candidate affine motion information, and the candidate affine motion information corresponding to the minimum cost result is determined as the affine motion information of the current block.
39. The method according to claim 38, wherein, The method further includes: Determine the merge candidate index of the current block; wherein the merge candidate index is used to indicate the index number of the affine motion information of the current block in the affine merge list; The candidate index for merging the current block is encoded, and the resulting encoded bits are written into the bitstream.
40. The method according to any one of claims 22 to 39, wherein, The method further includes: Determine the values of the first syntax element and the second syntax element; wherein the first syntax element is used to indicate whether the current block uses the merge mode, and the second syntax element is used to indicate whether the current block uses the sub-block-based motion compensation mode. The values of the first syntax element and the second syntax element are encoded respectively, and the resulting encoded bits are written into the bitstream.
41. The method according to any one of claims 22 to 40, wherein, The method further includes: Determine the value of the third syntax element; The value of the third syntax element is encoded, and the resulting encoded bits are written into the bitstream.
42. The method according to any one of claims 22 to 40, wherein, The method further includes: Determine the block information of the current block; Based on the block information, the steps of determining the motion information of multiple adjacent blocks at a preset position of the current block, and determining the affine motion information of the current block based on at least two motion information of the multiple adjacent blocks that meet the availability conditions.
43. The method according to any one of claims 22 to 40, wherein, The method further includes: When the current block uses an adaptive merging candidate reordering method, the steps of determining the motion information of multiple adjacent blocks at a preset position of the current block and determining the affine motion information of the current block based on at least two motion information of the multiple adjacent blocks that meet the availability condition are executed.
44. The method according to claim 41, wherein, The method further includes: When the value of the third syntax element is the first value, the maximum number of candidates in the affine merge list is set to L+N; where L and N are both positive integers.
45. An encoder, the encoder comprising a first determining unit and a first predicting unit, wherein: The first determining unit is configured to determine the motion information of a plurality of adjacent blocks at a preset position of the current block; and to determine the affine motion information of the current block based on at least two motion information of the plurality of adjacent blocks that meet the availability condition. The first prediction unit is configured to determine the predicted value of the current block based on the affine motion information of the current block.
46. An encoder, the encoder comprising a first memory and a first processor, wherein: The first memory is used to store computer programs that can run on the first processor; The first processor is configured to perform the method as described in any one of claims 22 to 44 when running the computer program.
47. A decoder, the decoder comprising a second determining unit and a second predicting unit, wherein: The second determining unit is configured to determine the motion information of a plurality of adjacent blocks at a preset position of the current block; and to determine the affine motion information of the current block based on at least two motion information of the plurality of adjacent blocks that satisfy the availability condition. The second prediction unit is configured to determine the predicted value of the current block based on the affine motion information of the current block.
48. A decoder, the decoder comprising a second memory and a second processor, wherein: The second memory is used to store computer programs that can run on the second processor; The second processor is configured to perform the method as described in any one of claims 1 to 21 when running the computer program.
49. A computer-readable storage medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 21, or the method as described in any one of claims 22 to 44.
50. A computer-readable storage medium having a bitstream stored thereon, wherein, The bitstream is generated by performing the steps of the encoding method as described in any one of claims 22 to 44.