Inter prediction method, encoder, decoder and computer storage medium
By increasing the diversity of the motion information candidate list in GPM or AWP inter-frame prediction, and using the lower right position inside and outside the current block to determine temporal motion information, the problem of insufficient partition correlation is solved, and the encoding and decoding performance is improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-31
- Publication Date
- 2026-04-10
AI Technical Summary
During GPM or AWP inter-frame prediction, the partition of the current block is not clearly adjacent to the relevant position when constructing the motion information candidate list, which leads to a decrease in encoding and decoding performance.
By increasing the diversity of motion information in the motion information candidate list, especially by utilizing the lower right position inside and outside the current block to determine temporal motion information, a new motion information candidate list is constructed to improve relevance.
It improves the encoding and decoding performance of inter-frame prediction, especially the correlation in the lower right corner in GPM or AWP modes, thus improving encoding and decoding efficiency.
Smart Images

Figure CN115052161B_ABST
Abstract
Description
[0001] This application is a divisional application of PCT / CN2021 / 084278, which entered the Chinese national phase as Chinese Patent Application No. 202180005705.0, with the title of "Inter-frame prediction method, encoder, decoder and computer storage medium", and the filing date of March 31, 2021.
[0002] Cross-reference to Related Applications
[0003] This application claims priority to Chinese Patent Application No. 202010479444.3, filed May 29, 2020, entitled "Inter-frame prediction method, encoder, decoder and computer storage medium", the entire contents of which are incorporated herein by reference. TECHNICAL FIELD
[0004] The present application relates to the technical field of video coding, and in particular to an inter-frame prediction method, an encoder, a decoder and a computer storage medium. BACKGROUND
[0005] In the field of video coding, in addition to using an intra-frame prediction method, a current block can also be coded using an inter-frame prediction method. The inter-frame prediction can include a geometric partitioning mode (GPM) and an angular weighted prediction mode (AWP), etc. By dividing the current block into two non-rectangular partitions (or two blocks) for prediction and then fusing them by weighting, a prediction value of the current block can be obtained.
[0006] Currently, in the prediction process of GPM or AWP, although spatial motion information and temporal motion information are used to construct a motion information candidate list, the temporal motion information used is searched for a corresponding position in a specified reference frame according to the top-left corner position in the current block. In this way, for the two partitions of GPM or AWP, in a partial partition mode, some partitions are not explicitly adjacent to the positions used to construct the motion information candidate list, resulting in a weak correlation between the partitions and the positions when predicting, thereby affecting the coding performance. SUMMARY
[0007] The present application proposes an inter-frame prediction method, an encoder, a decoder and a computer storage medium, which can increase the diversity of motion information in the motion information candidate list, thereby improving the coding performance.
[0008] The technical solution of the present application is implemented as follows:
[0009] In a first aspect, an inter prediction method is provided, which is applied to a decoder and includes the following steps.
[0010] parsing a bitstream to obtain a prediction mode parameter of a current block;
[0011] when the prediction mode parameter indicates that a preset inter prediction mode is used to determine an inter prediction value of the current block, determining at least one candidate position of the current block; wherein the candidate position at least includes a lower right position inside the current block and a lower right position outside the current block;
[0012] determining at least one temporal motion information of the current block based on the at least one candidate position;
[0013] constructing a new motion information candidate list based on the at least one temporal motion information;
[0014] determining the inter prediction value of the current block according to the new motion information candidate list.
[0015] In a second aspect, an inter prediction method is provided, which is applied to an encoder and includes the following steps.
[0016] determining a prediction mode parameter of a current block;
[0017] when the prediction mode parameter indicates that a preset inter prediction mode is used to determine an inter prediction value of the current block, determining at least one candidate position of the current block; wherein the candidate position at least includes a lower right position inside the current block and a lower right position outside the current block;
[0018] determining at least one temporal motion information of the current block based on the at least one candidate position;
[0019] constructing a new motion information candidate list based on the at least one temporal motion information;
[0020] determining the inter prediction value of the current block according to the new motion information candidate list.
[0021] In a third aspect, a decoder is provided, which includes a parsing unit, a first determining unit, a first constructing unit and a first predicting unit; wherein
[0022] the parsing unit is configured to parse a bitstream to obtain a prediction mode parameter of a current block;
[0023] The first determining unit is configured to determine at least one candidate position of the current block when the prediction mode parameter indicates that a preset inter-frame prediction mode is used to determine the inter-frame prediction value of the current block, wherein the candidate position at least includes a lower right position inside the current block and a lower right position outside the current block.
[0024] The first determining unit is further configured to determine at least one time domain motion information of the current block based on the at least one candidate position.
[0025] The first constructing unit is configured to construct a new motion information candidate list based on the at least one time domain motion information.
[0026] The first predicting unit is configured to determine the inter-frame prediction value of the current block according to the new motion information candidate list.
[0027] In a fourth aspect, an embodiment of the present application provides a decoder, which comprises a first memory and a first processor; wherein,
[0028] The first memory is configured to store a computer program capable of running on the first processor.
[0029] The first processor is configured to execute the method in the first aspect when the computer program is running.
[0030] In a fifth aspect, an embodiment of the present application provides an encoder, which comprises a second determining unit, a second constructing unit and a second predicting unit; wherein,
[0031] The second determining unit is configured to determine a prediction mode parameter of a current block, and determine at least one candidate position of the current block when the prediction mode parameter indicates that a preset inter-frame prediction mode is used to determine the inter-frame prediction value of the current block, wherein the candidate position at least includes a lower right position inside the current block and a lower right position outside the current block.
[0032] The second determining unit is further configured to determine at least one time domain motion information of the current block based on the at least one candidate position.
[0033] The second constructing unit is configured to construct a new motion information candidate list based on the at least one time domain motion information.
[0034] The second predicting unit is configured to determine the inter-frame prediction value of the current block according to the new motion information candidate list.
[0035] In a sixth aspect, an embodiment of the present application provides an encoder, which comprises a second memory and a second processor; wherein,
[0036] the second memory, configured to store a computer program capable of running on the second processor;
[0037] the second processor, configured to execute the method in the second aspect when running the computer program.
[0038] In a seventh aspect, an embodiment of the present application provides a computer storage medium, which stores a computer program. The computer program is executed by a first processor to implement the method in the first aspect, or executed by a second processor to implement the method in the second aspect.
[0039] The inter prediction method, the encoder, the decoder and the computer storage medium provided by the embodiments of the present application analyze a code stream, obtain a prediction mode parameter of a current block, determine at least one candidate position of the current block when the prediction mode parameter indicates that a preset inter prediction mode is used to determine an inter prediction value of the current block, wherein the candidate position at least includes a right bottom position inside the current block and a right bottom position outside the current block, determine at least one time domain motion information of the current block based on the at least one candidate position, construct a new motion information candidate list based on the at least one time domain motion information, and determine the inter prediction value of the current block according to the new motion information candidate list. In this way, since the time domain motion information of the current block is determined based on the right bottom position inside the current block or the right bottom position outside the current block, motion information with higher correlation with the right bottom position can be supplemented in the motion information candidate list, thereby increasing the diversity of the motion information in the motion information candidate list. Especially for the GPM or AWP inter prediction mode, the correlation with the right bottom position can be improved by increasing the right bottom time domain motion information candidate position, thereby improving the coding performance. BRIEF DESCRIPTION OF DRAWINGS
[0040] Figure 1 A constituent block diagram schematic view of a video coding system provided by an embodiment of the present application is provided.
[0041] Figure 2 A constituent block diagram schematic view of a video decoding system provided by an embodiment of the present application is provided.
[0042] Figure 3 A flowchart schematic view of an inter prediction method provided by an embodiment of the present application is provided.
[0043] Figure 4 A structure schematic view of a typical group of pictures provided by an embodiment of the present application is provided.
[0044] Figure 5 A spatial position relationship schematic view of a current block and a neighboring block provided by an embodiment of the present application is provided.
[0045] Figure 6 Another schematic diagram of spatial position relationship between a current block and neighboring blocks provided by an embodiment of the present application;
[0046] Figure 7 Another schematic diagram of spatial position relationship between a current block and neighboring blocks provided by an embodiment of the present application;
[0047] Figure 8 A flowchart of another inter-prediction method provided by an embodiment of the present application;
[0048] Figure 9A A schematic diagram of weight distribution of a GPM in multiple partition modes of a 64x64 current block provided by an embodiment of the present application;
[0049] Figure 9B A schematic diagram of weight distribution of an AWP in multiple partition modes of a 64x64 current block provided by an embodiment of the present application;
[0050] Figure 10 A flowchart of another inter-prediction method provided by an embodiment of the present application;
[0051] Figure 11 A schematic diagram of component structure of a decoder provided by an embodiment of the present application;
[0052] Figure 12 A schematic diagram of hardware structure of a decoder provided by an embodiment of the present application;
[0053] Figure 13 A schematic diagram of component structure of an encoder provided by an embodiment of the present application;
[0054] Figure 14 A schematic diagram of hardware structure of an encoder provided by an embodiment of the present application. DETAILED DESCRIPTION
[0055] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. It can be understood that the specific embodiments described herein are only used to explain the related application, but not to limit the application. In addition, it should be noted that, for the convenience of description, only the parts related to the application are shown in the drawings.
[0056] In a video image, a current block (Coding Block, CB) is generally represented by a first image component, a second image component and a third image component; wherein the three image components are a luminance component, a blue chroma component and a red chroma component respectively, specifically, the luminance component is usually represented by a symbol Y, the blue chroma component is usually represented by a symbol Cb or U, and the red chroma component is usually represented by a symbol Cr or V; thus, the video image can be represented by YCbCr format or YUV format.
[0057] At present, the general video coding standards are based on a block-based hybrid coding framework. Each frame in a video image is divided into square-shaped Largest Coding Units (LCUs) of the same size (such as 128x128, 64x64, etc.), and each LCU can be further divided into rectangular Coding Units (CUs) according to a rule; and the Coding Units can be further divided into smaller Prediction Units (PUs). Specifically, the hybrid coding framework can include modules of prediction, transform, quantization, entropy coding, in loop filter, etc.; wherein the prediction module can include intra prediction and inter prediction, and the inter prediction can include motion estimation and motion compensation. Since there is a strong correlation between adjacent pixels in a frame of a video image, using the intra prediction mode in the video coding technology can eliminate the spatial redundancy between adjacent pixels; but since there is also a strong similarity between adjacent frames in a video image, using the inter prediction mode in the video coding technology can eliminate the temporal redundancy between adjacent frames, thereby improving the coding efficiency. The following detailed description of the present application will be made in detail with respect to the inter prediction.
[0058] It should be understood that the embodiments of the present application provide a video coding system, such as Figure 1As shown, the video encoding system 11 can include a transform unit 111, a quantization unit 112, a mode selection and encoding control logic unit 113, an intra prediction unit 114, an inter prediction unit 115 (including motion compensation and motion estimation), a dequantization unit 116, an inverse transform unit 117, a loop filter unit 118, an encoding unit 119, and a decoded picture buffer unit 120. For an input raw video signal, a video reconstruction block can be obtained by partitioning the raw video signal into coding tree units (CTUs). The mode selection and encoding control logic unit 113 determines an encoding mode. Then, the video reconstruction block after intra or inter prediction is transformed by the transform unit 111 and the quantization unit 112, including transforming residual information from a pixel domain to a transform domain, and quantizing the obtained transform coefficients to further reduce the bit rate. The intra prediction unit 114 is configured to perform intra prediction on the video reconstruction block. The intra prediction unit 114 is configured to determine an optimal intra prediction mode (i.e., a target prediction mode) for the video reconstruction block. The inter prediction unit 115 is configured to perform inter prediction encoding of the received video reconstruction block with respect to one or more blocks in one or more reference frames to provide temporal prediction information. The motion estimation is a process of generating a motion vector that can estimate the motion of the video reconstruction block. The motion compensation is performed based on the motion vector determined by the motion estimation. After determining the inter prediction mode, the inter prediction unit 115 is further configured to provide the selected inter prediction data to the encoding unit 119, and also send the calculated determined motion vector data to the encoding unit 119. In addition, the dequantization unit 116 and the inverse transform unit 117 are configured to reconstruct the video reconstruction block. The reconstructed residual block is removed of blocking artifacts by the loop filter unit 118. Then, the reconstructed residual block is added to a predictive block in one of the frames of the decoded picture buffer unit 120 to generate a reconstructed video reconstruction block. The encoding unit 119 is configured to encode various encoding parameters and quantized transform coefficients. The decoded picture buffer unit 120 is configured to store the reconstructed video reconstruction block for prediction reference. As the video image encoding proceeds, new reconstructed video reconstruction blocks are continuously generated and stored in the decoded picture buffer unit 120.
[0059] The embodiments of the present application also provide a video decoding system, which can include a decoding unit 121, a dequantization unit 122, an inverse transform unit 123, a loop filter unit 124, an intra prediction unit 125, an inter prediction unit 126 (including motion compensation and motion estimation), a mode selection and decoding control logic unit 127, and a decoded picture buffer unit 128. For an input encoded video signal, a video reconstruction block can be obtained by partitioning the encoded video signal into coding tree units (CTUs). The mode selection and decoding control logic unit 127 determines a decoding mode. Then, the video reconstruction block after decoding is transformed by the decoding unit 121, the dequantization unit 122, and the inverse transform unit 123, including transforming residual information from a pixel domain to a transform domain, and dequantizing the obtained transform coefficients to further reduce the bit rate. The intra prediction unit 125 is configured to perform intra prediction on the video reconstruction block. The intra prediction unit 125 is configured to determine an optimal intra prediction mode (i.e., a target prediction mode) for the video reconstruction block. The inter prediction unit 126 is configured to perform inter prediction decoding of the received video reconstruction block with respect to one or more blocks in one or more reference frames to provide temporal prediction information. The motion estimation is a process of generating a motion vector that can estimate the motion of the video reconstruction block. The motion compensation is performed based on the motion vector determined by the motion estimation. After determining the inter prediction mode, the inter prediction unit 126 is further configured to provide the selected inter prediction data to the decoding unit 121, and also send the calculated determined motion vector data to the decoding unit 121. In addition, the dequantization unit 122 and the inverse transform unit 123 are configured to reconstruct the video reconstruction block. The reconstructed residual block is removed of blocking artifacts by the loop filter unit 124. Then, the reconstructed residual block is added to a predictive block in one of the frames of the decoded picture buffer unit 128 to generate a reconstructed video reconstruction block. The decoding unit 121 is configured to decode various encoding parameters and quantized transform coefficients. The decoded picture buffer unit 128 is configured to store the reconstructed video reconstruction block for prediction reference. As the video image decoding proceeds, new reconstructed video reconstruction blocks are continuously generated and stored in the decoded picture buffer unit 128. Figure 2As shown, the video decoding system 12 can include a decoding unit 121, an inverse transformation unit 127, an inverse quantization unit 122, an intra prediction unit 123, a motion compensation unit 124, a loop filter unit 125 and a decoded picture buffer unit 126. After the input video signal is processed by the video encoding system 11, a bitstream of the video signal is output. The bitstream is input into the video decoding system 12, and first passes through the decoding unit 121 to obtain decoded transform coefficients. The inverse transformation unit 127 and the inverse quantization unit 122 process the transform coefficients to generate a residual block in the pixel domain. The intra prediction unit 123 can be used to generate prediction data of a current video decoding block based on a determined intra prediction direction and data from previously decoded blocks of the current frame or picture. The motion compensation unit 124 determines prediction information for the video decoding block by parsing motion vectors and other associated syntax elements, and uses the prediction information to generate a predictive block of the video decoding block being decoded. The decoded video block is formed by summing the residual block from the inverse transformation unit 127 and the inverse quantization unit 122 and the corresponding predictive block generated by the intra prediction unit 123 or the motion compensation unit 124. The decoded video signal passes through the loop filter unit 125 to remove blocking artifacts and improve video quality. The decoded video block is then stored in the decoded picture buffer unit 126, which stores reference pictures for subsequent intra prediction or motion compensation, and also for output of the video signal to obtain the recovered original video signal.
[0060] The inter prediction method provided by the embodiments of the present application mainly acts on the inter prediction unit 215 of the video encoding system 11 and the inter prediction unit, i.e. the motion compensation unit 124, of the video decoding system 12. That is, if a better prediction effect is obtained by the inter prediction method provided by the embodiments of the present application in the video encoding system 11, the encoding performance is improved. Correspondingly, in the video decoding system 12, the video decoding recovery quality is improved, and thus the decoding performance is improved.
[0061] Based on this, the technical solutions of the present application are further described in detail below in combination with the drawings and embodiments. Before the detailed description, it should be noted that "first", "second", "third" and the like mentioned throughout the description are only used to distinguish different features, and do not have the functions of limiting priority, sequence, size relationship and the like.
[0062] The embodiments of the present application provide an inter prediction method, which is applied to a video decoding device, i.e. a decoder. The function realized by the method can be realized by a first processor in the decoder calling a computer program, and of course the computer program can be saved in a first memory. Therefore, the decoder at least includes the first processor and the first memory.
[0063] Referring to Figure 3 , a flowchart of an inter prediction method is shown. As Figure 3 indicated, the method can include:
[0064] S301: parsing a code stream to obtain a prediction mode parameter of a current block.
[0065] It should be noted that a to-be-decoded image can be divided into a plurality of image blocks, and a current to-be-decoded image block can be referred to as a current block (which can be represented by a CU), and an image block adjacent to the current block can be referred to as a neighboring block; that is, in the to-be-decoded image, the current block and the neighboring block have an adjacent relationship. Here, each current block can include a first image component, a second image component, and a third image component, that is, the current block represents an image block in the to-be-decoded image that is currently to be predicted for the first image component, the second image component, or the third image component.
[0066] Among them, it is assumed that the current block is to be predicted for the first image component, and the first image component is a luminance component, that is, the to-be-predicted image component is a luminance component, so the current block can also be referred to as a luminance block; or, it is assumed that the current block is to be predicted for the second image component, and the second image component is a chroma component, that is, the to-be-predicted image component is a chroma component, so the current block can also be referred to as a chroma block.
[0067] It should also be noted that the prediction mode parameter indicates the prediction mode adopted by the current block and the parameters related to the prediction mode. Among them, the prediction mode usually includes an inter prediction mode, a traditional intra prediction mode, and a non-traditional intra prediction mode, and the inter prediction mode includes a normal inter prediction mode, a GPM prediction mode, and an AWP prediction mode. That is, the encoder will select the optimal prediction mode to pre-encode the current block, and in this process, the prediction mode of the current block can be determined, so that the corresponding prediction mode parameter is written into the code stream and transmitted to the decoder by the encoder.
[0068] In this way, on the decoder side, the prediction mode parameter of the current block can be directly obtained by parsing the code stream, and the prediction mode parameter obtained is used to determine whether the current block uses a preset inter prediction mode, such as a GPM prediction mode or an AWP prediction mode.
[0069] S302: When the prediction mode parameter indicates that a preset inter prediction mode is used to determine an inter prediction value of the current block, at least one candidate position of the current block is determined; wherein the candidate position at least includes a lower right position inside the current block and a lower right position outside the current block.
[0070] It should be noted that in the case that the decoder parses the code stream to obtain the prediction mode parameter indicating that the preset inter prediction mode is used to determine the inter prediction value of the current block, the inter prediction method provided in the embodiments of the present application can be used.
[0071] It should be further noted that the motion information can include motion vector (MV) information and reference frame information. Specifically, for the current block using inter prediction, the current frame in which the current block is located has one or more reference frames, and the current block can be a coding unit or a prediction unit. One motion information containing a set of motion vectors and reference frame information can be used to indicate a pixel region in a certain reference frame, which is referred to as a reference block, and the size of the reference block is the same as that of the current block. Alternatively, one motion information containing two sets of motion vectors and reference frame information can be used to indicate two reference blocks in two reference frames which can be the same or different. Then, the motion compensation can obtain the inter prediction value of the current block according to the reference blocks indicated by the motion information.
[0072] It should be understood that a P frame (Predictive Frame) is a frame that can only be predicted using reference frames before the current frame in terms of Picture Order Count (POC). In this case, the current reference frame has only one reference frame list, which is denoted as RefPicListO, and all the reference frames in RefPicListO are reference frames before the current frame in terms of POC. A B frame (Bi-directional Interpolated Prediction Frame) is a frame that can be predicted using reference frames before the current frame in terms of POC and reference frames after the current frame in terms of POC. A B frame has two reference frame lists, which are denoted as RefPicListO and RefPicListl, respectively. In RefPicListO, all the reference frames are reference frames before the current frame in terms of POC, and in RefPicListl, all the reference frames are reference frames after the current frame in terms of POC. For a current block, it can only refer to a reference block in a certain frame in RefPicListO, which is called forward prediction; or it can only refer to a reference block in a certain frame in RefPicListl, which is called backward prediction; or it can refer to a reference block in a certain frame in RefPicListO and a reference block in a certain frame in RefPicListl, which is called bi-prediction. A simple way to simultaneously refer to two reference blocks is to average the pixels in the corresponding positions in the two reference blocks to obtain the inter-prediction value (or can be called a prediction block) of each pixel in the current block. In later B frames, the restriction that all the reference frames in RefPicListO are reference frames before the current frame in terms of POC and all the reference frames in RefPicListl are reference frames after the current frame in terms of POC is removed. In other words, there can be reference frames after the current frame in terms of POC in RefPicListO, and there can be reference frames before the current frame in terms of POC in RefPicListl, i.e., the current block can simultaneously refer to reference frames before the current frame in terms of POC or simultaneously refer to reference frames after the current frame in terms of POC. However, when the current block is bi-predicted, the reference frames used must come from RefPicListO and RefPicListl, respectively. Such a B frame is also called a generalized B frame.
[0073] Because the coding order is different from the POC order in the case of random access (RA) configuration, a B frame can simultaneously refer to information before the current frame and information after the current frame, which can significantly improve the coding performance. An exemplary classic Group Of Pictures (GOP) structure of RA is shown in FIG. 1. In the exemplary GOP structure, the coding order is different from the POC order. Figure 4 Figure 4 In the middle, the arrow indicates the reference relationship. Since the I frame does not need a reference frame, after the I frame with POC 0 is decoded, the P frame with POC 4 is decoded, and the I frame with POC 0 can be referenced when the P frame with POC 4 is decoded. After the P frame with POC 4 is decoded, the B frame with POC 2 is decoded, and the I frame with POC 0 and the P frame with POC 4 can be referenced when the B frame with POC 2 is decoded, and so on. In this way, according to the POC order {0 1 2 3 4 5 6 7 8}, the decoding order is {0 3 2 4 1 7 6 8 5}. Figure 4 It can be obtained that, in the case of the POC order {0 1 2 3 4 5 6 7 8}, the corresponding decoding order is {0 3 2 4 1 7 6 8 5}.
[0074] In addition, the coding order of the Low Delay (LD) configuration is the same as the POC order, and at this time, the current frame can only refer to the information before the current frame. The Low Delay configuration is divided into Low Delay P and Low Delay B. The Low Delay P is a traditional Low Delay configuration. The typical structure is IPPP...., that is, an I frame is coded first, and then all the frames decoded are P frames. The typical structure of the Low Delay B is IBBB.... The difference between the Low Delay P and the Low Delay B is that each frame is a B frame, that is, two reference frame lists are used, and the current block can simultaneously refer to the reference block of a frame in RefPicList0 and the reference block of a frame in RefPicList1. Here, a reference frame list of the current frame can have several reference frames at most, such as 2, 3, or 4. When a current frame is coded or decoded, how many reference frames are in RefPicList0 and RefPicList1 is determined by a preset configuration or algorithm, but the same reference frame can be simultaneously present in RefPicList0 and RefPicList1, that is, the encoder or decoder allows the current block to simultaneously refer to two reference blocks in the same reference frame.
[0075] In the embodiments of the present application, the encoder or decoder can generally use an index value (denoted by index) in the reference frame list to correspond to the reference frame. If the length of a reference frame list is 4, the index has four values 0, 1, 2, and 3. For example, if the RefPicList0 of the current frame has four reference frames with POC 5, 4, 3, and 0, the index 0 of RefPicList0 is the reference frame with POC 5, the index 1 of RefPicList0 is the reference frame with POC 4, the index 2 of RefPicList0 is the reference frame with POC 3, and the index 3 of RefPicList0 is the reference frame with POC 0.
[0076] In the current Versatile Video Coding (VVC) standard, the preset inter prediction mode can be a GPM prediction mode. In the current Audio Video coding Standard (AVS) standard, the preset inter prediction mode can be an AWP prediction mode. Although the two prediction modes have different names, different specific implementation forms, but the principle is common, that is, the two prediction modes can be applied to the inter prediction method of the embodiments of the application.
[0077] Specifically, for the GPM prediction mode, if GPM is used, the prediction mode parameters under GPM, such as the specific partition mode of GPM, will be transmitted in the code stream. Generally, GPM includes 64 partition modes. For the AWP prediction mode, if AWP is used, the prediction mode parameters under AWP, such as the specific partition mode of AWP, will be transmitted in the code stream. Generally, AWP includes 56 partition modes.
[0078] In the preset prediction mode, such as GPM and AWP, two single-direction motion information needs to be used to search two reference blocks. The current implementation is to use the related information of the previously encoded / decoded part of the current block at the encoder side to construct a motion information candidate list (also called a single-direction motion information candidate list), and select single-direction motion information from the motion information candidate list. The index value (index) of the two single-direction motion information in the motion information candidate list is written into the code stream. At the decoder side, the same method is used, that is, the related information of the previously decoded part of the current block is used to construct a motion information candidate list, and this motion information candidate list must be the same as the candidate list constructed at the encoder side. In this way, the index values of the two motion information are parsed from the code stream, and then the two single-direction motion information are searched from the motion information candidate list, which are the two single-direction motion information needed by the current block.
[0079] It should be further explained that the unidirectional motion information described in the embodiments of the present application can include motion vector information, i.e. the value of (x, y), and corresponding reference frame information, i.e. the reference frame list and the reference index value in the reference frame list. One way of representation is to record the reference index values of two reference frame lists, wherein the reference index values corresponding to one reference frame list are valid, such as 0, 1, 2, etc., and the reference index values corresponding to the other reference frame list are invalid, i.e. -1. The reference frame list with valid reference index values is the reference frame list used by the motion information of the current block, and the corresponding reference frame can be found from the reference frame list according to the reference index value. Each reference frame list has a corresponding motion vector, and the motion vector corresponding to the valid reference frame list is valid, and the motion vector corresponding to the invalid reference frame list is invalid. The decoder can find the required reference frame through the reference frame information in the unidirectional motion information, find the reference block in the reference frame according to the position of the current block and the value of the motion vector (x, y), and then determine the inter prediction value of the current block.
[0080] In practical applications, the construction of the motion information candidate list not only uses spatial motion information, but also uses temporal motion information. In VVC, temporal motion information and spatial motion information are also used when constructing a merge candidate list (merge list). As shown in Figure 5 , which shows the motion information of the related positions used when constructing the merge list, the candidate positions filled with elements 1, 2, 3, 4, 5 represent spatial related positions, i.e. the motion information used by the position blocks adjacent to the current block in the current frame; the candidate positions filled with elements 6 and 7 represent temporal related positions, i.e. the motion information used by the corresponding positions in a certain reference frame, which can also be scaled. Here, for the temporal motion information, if the candidate position 6 is available, the motion information corresponding to the position 6 can be used; otherwise, the motion information corresponding to the position 7 can be used. It should be noted that the construction of the motion information candidate list in the Triangle Partition Mode (TPM) and the GPM prediction mode also involves the use of these positions; and here the size of the block is not the actual size, but is only used as an example for illustration.
[0081] For the AWP prediction mode, as shown in Figure 6 , block E is the current block, and blocks A, B, C, D, F, G are adjacent blocks of block E, and block H is the top-left position inside block E. Among them, blocks A, B, C, D, F, G can be used to determine spatial motion information, and temporal motion information is to find the corresponding position in the specified reference frame according to the top-left position inside the current block, i.e. Figure 6The gray filled positions (i.e., block H) are shown, where block H is only used to represent the position and does not represent the block size. Based on Figure 6 It can be seen that all the relevant positions of AWP are from the left, top-left, top, and top-right of the current block, but not from the bottom, bottom-right, and right of the current block.
[0082] However, GPM and AWP are different from other prediction modes in that GPM and AWP essentially divide the current block into two partitions. For these other prediction modes, Figure 6 The positions shown are all closely related to the current block, such as the positions being adjacent to at least one boundary of the current block, or the positions overlapping the current block. However, this is not the case for the two partitions of GPM and AWP, and it is particularly obvious that the partition in the lower right corner under certain division modes is not clearly adjacent to the relevant positions used to build the motion information candidate list. That is, if the current block is small, the corresponding position in the time domain can overlap the partition in the lower right corner of the current block; but if the current block is large, the partition is not adjacent to the relevant positions used to build the motion information candidate list, so the relevance of the partition to the selected positions is weak.
[0083] In addition, each partition in GPM and AWP must use different motion information, otherwise GPM and AWP are meaningless. Taking AWP as an example, in the extreme case, if the selected relevant motion information of the current block is all the same as the top-left corner of the current block, the lower right corner cannot find available motion information, resulting in that AWP cannot be used. That is, for AWP, the relevance of the lower right corner is obviously missing. For GPM, the time domain motion information is from different frames, and the spatial domain motion information is from the same frame, so the relevance of the spatial domain motion information is stronger, and the relevance of the lower right corner is weak. This deficiency in the relevance of the lower right corner has a greater impact on GPM and AWP than other prediction modes, thereby affecting the coding performance.
[0084] At this time, in order to improve the coding performance, embodiments of the present application need to enhance the relevance of the lower right corner of AWP and GPM. That is, in the AWP prediction mode, when the time domain motion information is used to build the motion information candidate list, the time domain motion information corresponding to the lower right corner of the current block can be added, or the originally used time domain motion information corresponding to the top-left corner of the current block is replaced by the time domain motion information corresponding to the lower right corner of the current block. It should be noted that if the relevance of the lower right corner is added, the order of the time domain motion information corresponding to the lower right corner of the current block can be before the time domain motion information corresponding to the top-left corner of the current block. The lower right corner position can include the lower right position inside the current block and the lower right position outside the current block.
[0085] In some embodiments, for S302, the determining the at least one candidate position of the current block can include:
[0086] obtaining a first bottom-right candidate position, a second bottom-right candidate position, a third bottom-right candidate position and a fourth bottom-right candidate position to form a candidate position set;
[0087] determining the at least one candidate position of the current block from the candidate position set;
[0088] wherein the first bottom-right candidate position represents a bottom-right position inside the current block, and the second bottom-right candidate position, the third bottom-right candidate position and the fourth bottom-right candidate position represent bottom-right positions outside the current block.
[0089] It should be noted that, for example, Figure 7 all the gray filled positions represent the bottom-right corner of the current block, which can be divided into four bottom-right candidate positions. Specifically, the bottom-right position inside the current block can be the first bottom-right candidate position, as shown in the position filled with 1 in Figure 7 ; the bottom-right position outside the current block can be the second bottom-right candidate position, as shown in the position filled with 2 in Figure 7 ; the bottom-right position outside the current block can also be the third bottom-right candidate position, as shown in the position filled with 3 in Figure 7 ; and the bottom-right position outside the current block can also be the fourth bottom-right candidate position, as shown in the position filled with 4 in Figure 7 . In other words, the bottom-right corner in the embodiments of the present application can refer to the bottom-right corner inside the current block, as shown in the position filled with 1 in Figure 7 ; or refer to the bottom-right corner outside the current block, as shown in the position filled with 2 in Figure 7 ; or refer to the bottom-right corner outside the current block, as shown in the position filled with 3 in Figure 7 ; or refer to the bottom-right corner outside the current block, as shown in the position filled with 4 in Figure 7 . Here, if consistent with the text of the AVS standard, the filled 1, 2, 3 and 4 can be replaced by H, I, J and K.
[0090] In this way, after obtaining the first bottom-right candidate position, the second bottom-right candidate position, the third bottom-right candidate position and the fourth bottom-right candidate position, a candidate position set can be formed; and then the at least one candidate position of the current block is determined from the candidate position set.
[0091] Further, in some embodiments, the obtaining the first bottom-right candidate position, the second bottom-right candidate position, the third bottom-right candidate position and the fourth bottom-right candidate position can include:
[0092] obtaining coordinate information corresponding to a top-left pixel position of the current block, width information of the current block, and height information of the current block;
[0093] performing coordinate calculation by using the coordinate information corresponding to the top-left pixel position, the width information, and the height information respectively, to obtain first coordinate information, second coordinate information, third coordinate information, and fourth coordinate information;
[0094] determining a position corresponding to the first coordinate information as the first right-down candidate position, determining a position corresponding to the second coordinate information as the second right-down candidate position, determining a position corresponding to the third coordinate information as the third right-down candidate position, and determining a position corresponding to the fourth coordinate information as the fourth right-down candidate position.
[0095] That is, the right-down corner inside the current block, i.e., the position of 1, is determined according to the pixel position of the right-down corner of the current block. Here, assuming that the top-left pixel position (i.e., the pixel position of the top-left corner) of the current block is denoted by (x, y), the width of the current block is denoted by width, and the height of the current block is denoted by height, the first coordinate information can be (x+width-1, y+height-1), and the first right-down candidate position (i.e., the position of 1) can be the position determined according to the first coordinate information; the second coordinate information can be (x+width, y+height), and the second right-down candidate position (i.e., the position of 2) can be the position determined according to the second coordinate information; the third coordinate information can be (x+width-1, y+height), and the third right-down candidate position (i.e., the position of 3) can be the position determined according to the third coordinate information; and the fourth coordinate information can be (x+width, y+height-1), and the fourth right-down candidate position (i.e., the position of 4) can be the position determined according to the fourth coordinate information.
[0096] Further, if it is considered that the four positions determined by the above-described manner are too close, it is more likely that the corresponding positions in the corresponding reference frame belong to the same block, and at this time, an offset, i.e., a first preset offset, denoted by offset, can be set. Therefore, in some embodiments, after the first coordinate information, the second coordinate information, the third coordinate information, and the fourth coordinate information are obtained, the method can further include:
[0097] modifying the second coordinate information, the third coordinate information, and the fourth coordinate information by using the first preset offset, to obtain first modified second coordinate information, first modified third coordinate information, and first modified fourth coordinate information;
[0098] The position corresponding to the first coordinate information is determined as the first bottom-right candidate position, the position corresponding to the first corrected second coordinate information is determined as the second bottom-right candidate position, the position corresponding to the first corrected third coordinate information is determined as the third bottom-right candidate position, and the position corresponding to the first corrected fourth coordinate information is determined as the fourth bottom-right candidate position.
[0099] That is, among the above four positions, the position of 1 can still be the position determined according to the first coordinate information, i.e., (x+width-1, y+height-1); the position of 2 can be the position determined according to the first corrected second coordinate information, i.e., (x+width+offset, y+height+offset); the position of 3 can be the position determined according to the first corrected third coordinate information, i.e., (x+width-1, y+height+offset); and the position of 4 can be the position determined according to the first corrected fourth coordinate information, i.e., (x+width+offset, y+height-1).
[0100] It should be further noted that the offset in the embodiments of the present application can be a preset fixed value, such as 0, 1, 2, 4, 8, 16, etc.; or, it can also be a value determined according to the size (i.e., width, height) of the current block; or, it can even be a value determined in other manners. Here, the value of offset at each position can be the same or different, which is not limited in the embodiments of the present application.
[0101] Further, since there can be a time deviation between the current frame and the reference frame of the reference motion information used to derive the temporal motion information, i.e., there can also be a motion between the block in the current frame and the reference frame used by the derived temporal motion information, a second offset can also be set, denoted as (x', y'). Therefore, in some embodiments, after the first coordinate information, the second coordinate information, the third coordinate information and the fourth coordinate information are obtained, the method can further include:
[0102] When there is a motion vector between the image block in the current frame and the candidate reference frame, a second preset offset is determined; wherein the candidate reference frame is the reference frame of the reference motion information used to determine the temporal motion information, and the image block at least includes a neighboring block, and the neighboring block is spatially adjacent to the current block in the current frame;
[0103] The first coordinate information, the second coordinate information, the third coordinate information and the fourth coordinate information are corrected by using a second preset offset, to obtain second corrected first coordinate information, second corrected second coordinate information, second corrected third coordinate information and second corrected fourth coordinate information.
[0104] The position corresponding to the second corrected first coordinate information is determined as the first right lower candidate position, the position corresponding to the second corrected second coordinate information is determined as the second right lower candidate position, the position corresponding to the second corrected third coordinate information is determined as the third right lower candidate position, and the position corresponding to the second corrected fourth coordinate information is determined as the fourth right lower candidate position.
[0105] That is, one possible implementation is that when finding the corresponding position in the reference frame of the reference motion information of the derived temporal motion information in the above manner, the motion vector between the block on the current frame and the reference frame used by the derived temporal motion information needs to be considered. Exemplarily, if the 2 position is determined according to the second coordinate information (x+width, y+height), at this time, the motion between the block (x+width, y+height) on the current frame and the reference frame used by the derived temporal motion information is considered, assuming (x', y') is the motion vector between the block (x+width, y+height) on the current frame and the reference frame used by the derived temporal motion information, then the second corrected second coordinate information is (x+width+x', y+height+y'), that is, the 2 position can be determined at (x+width+x', y+height+y') of the reference frame used by the derived temporal motion information. Wherein, the position can be (x+width+x', y+height+y'), or it can be obtained by other calculations according to (x+width+x', y+height+y'). Or, if the first preset offset is considered, the 2 position is determined according to (x+width+offset, y+height+offset), at this time, the motion between the block (x+width+offset, y+height+offset) on the current frame and the reference frame used by the derived temporal motion information is considered, assuming (x', y') is still the motion vector between the block (x+width+offset, y+height+offset) on the current frame and the reference frame used by the derived temporal motion information, then the second corrected second coordinate information is (x+width+offset+x', y+height+offset+y'), that is, the 2 position can be determined at (x+width+offset+x', y+height+offset+y') of the reference frame used by the derived temporal motion information. Wherein, the position can be (x+width+offset+x', y+height+offset+y'), or it can be obtained by other calculations according to (x+width+offset+x', y+height+offset+y'). Similarly, the 1 position, the 3 position and the 4 position can also be determined, which will not be described here.
[0106] Further, in some embodiments, the determining the second preset offset can include:
[0107] obtaining the motion vector of the preset neighboring block on the current frame to the candidate reference frame, and determining the obtained motion vector as the second preset offset; or,
[0108] scaling motion information of the preset neighboring block on the current frame to the candidate reference frame to obtain a scaled motion vector, and determining the scaled motion vector as the second preset offset.
[0109] That is, in the embodiments of the present application, for the determination of (x', y'), the motion vector of a certain neighboring block to the reference frame used for deriving the temporal motion information (or the reference frame of the motion information referred to by the temporal motion information) can be found as (x', y'), or the motion vector of the motion information of a certain neighboring block scaled to the reference frame used for deriving the temporal motion information (or the reference frame of the motion information referred to by the temporal motion information or the reference frame of the motion information referred to by the determination of the temporal motion information) can be found as (x', y'), which is not specifically limited here.
[0110] It should be understood that another possible implementation is that when finding the corresponding position in the reference frame information used for deriving the temporal motion information in the above manner, the motion vector from the block on the current frame to the reference frame used for deriving the temporal motion information is not considered. Exemplarily, if the 2 position is determined according to (x+width, y+height), the motion from the block (x+width, y+height) on the current frame to the reference frame used for deriving the temporal motion information is not considered, and the second coordinate information at this time is (x+width, y+height), that is, the 2 position can be determined according to (x+width, y+height) on the reference frame used for deriving the temporal motion information. The position can be (x+width, y+height), or can be obtained through other calculations according to (x+width, y+height). Alternatively, if the 2 position is determined according to (x+width+offset, y+height+offset) considering the first preset offset, the motion from the block (x+width+offset, y+height+offset) on the current frame to the reference frame used for deriving the temporal motion information is not considered, and the second coordinate information at this time is (x+width+offset, y+height+offset), that is, the 2 position can be determined according to (x+width+offset, y+height+offset) on the reference frame used for deriving the temporal motion information. The position can be (x+width+offset, y+height+offset), or can be obtained through other calculations according to (x+width+offset, y+height+offset). Similarly, the 1 position, the 3 position and the 4 position can also be determined, which will not be described here.
[0111] In addition, in the embodiments of the present application, one possible way is to use only one of the above four positions, and one possible way is to use a combination of some of the above four positions. Therefore, optionally, in some embodiments, the determining, from the candidate position set, at least one candidate position of the current block can include:
[0112] selecting a candidate position from the candidate position set according to a preset manner, and determining the selected candidate position as a candidate position of the current block; or
[0113] selecting a candidate position corresponding to a high priority according to a preset priority order from the candidate position set and the selected candidate position is available, and determining the selected candidate position as a candidate position of the current block.
[0114] Optionally, in some embodiments, the determining the at least one candidate position of the current block from the candidate position set can comprise:
[0115] selecting a plurality of candidate positions from the candidate position set according to a preset combination manner, and determining the selected plurality of candidate positions as the plurality of candidate positions of the current block; or
[0116] selecting a plurality of candidate positions from the candidate position set according to a preset priority order and the selected candidate positions are available, and determining the selected plurality of candidate positions as the plurality of candidate positions of the current block.
[0117] It should be noted that, for example, Figure 7 when only one temporal motion information is needed in the motion information candidate list, one of the above-mentioned positions 1, 2, 3, 4 can be selected as a candidate position according to a preset manner, or the positions can be selected according to a preset priority order, such as 2, 1, 3, 4. If the priority of 2 is the highest and 2 is available, the position 2 is selected as the candidate position. If the priority of 2 is the highest but 2 is not available, the position 1 corresponding to the highest priority is selected according to the preset priority order, and if 1 is available, the position 1 is selected as the candidate position. Or, when multiple temporal motion information is needed in the motion information candidate list, multiple positions can be selected as candidate positions from the above-mentioned positions 1, 2, 3, 4 according to a preset combination manner, or the positions can be selected according to a preset priority order, such as 2, 1, 3, 4. If the priority of 2 is the highest, the position 2 is arranged at the front, and then multiple positions are selected as candidate positions according to the preset priority order under the condition that the selected candidate positions are available.
[0118] That is, one possible way is to use only one of the above-mentioned four positions, and another possible way is to use a combination of some of the above-mentioned four positions, such as a combination of positions 2, 3, 4, a combination of positions 3, 4, etc. Here, the use of all the positions 1, 2, 3, 4 is also a combination. In other words, any permutation and combination of the above-mentioned four positions (1, 2, 3, 4) can be used as a selection method for determining candidate positions. It should be noted that when multiple positions are needed, one possible way is to arrange 2 at the first position, i.e., the position 2 has the highest priority.
[0119] Further, in some embodiments, the method can further comprise:
[0120] When selecting a candidate position from the candidate position set, if the candidate position to be selected belongs to the lower right position outside the current block and all the lower right positions outside the current block are not available, the first lower right candidate position is determined as the candidate position to be selected.
[0121] That is, still taking Figure 7 for example, the 2, 3, 4 positions are all outside the current block, and the 2, 3, 4 positions are sometimes unavailable. If the image boundary is encountered, the boundary cannot be crossed by certain inter-frame reference, etc., the 2, 3, 4 positions can all be unavailable. In the case where the 2, 3, 4 positions are all unavailable, one possible way is to use the 1 position instead of the unavailable positions. If the 1 position has already been used in the construction of the current motion information candidate list, the unavailable positions can also be skipped.
[0122] In this way, after the at least one candidate position of the current block is determined, the candidate position here is the position of the lower right corner of the current block, which can be the lower right position inside the current block or the lower right position outside the current block, and the at least one temporal motion information of the current block is determined according to the candidate position, so that when the motion information candidate list is constructed using the temporal motion information, the temporal motion information corresponding to the position of the lower right corner of the current block can be increased, thereby improving the correlation of the lower right corner.
[0123] S303: determining at least one temporal motion information of the current block based on the at least one candidate position.
[0124] It should be noted that after the at least one candidate position is obtained, the temporal motion information can be determined according to the obtained candidate position, that is, the motion information used by the temporal position in the corresponding reference frame is used as the temporal motion information of the candidate position. Here, the frame to which the current block belongs can be referred to as the current frame, and the candidate position in the current frame and the temporal position in the reference frame belong to different frames, but the positions are the same.
[0125] In some embodiments, for S303, the at least one temporal motion information of the current block is determined based on the at least one candidate position, which can include:
[0126] determining the reference frame information corresponding to each candidate position in the at least one candidate position;
[0127] for each candidate position, determining the temporal position associated with the candidate position in the corresponding reference frame information, and determining the motion information used by the temporal position as the temporal motion information corresponding to the candidate position;
[0128] corresponding to the at least one candidate position, at least one temporal motion information is obtained.
[0129] That is, the temporal motion information is determined according to the motion information used by the corresponding position in a certain reference frame information. Moreover, different temporal motion information can be obtained for different candidate positions.
[0130] Exemplarily, method one is as follows: Figure 7 Exemplarily, the steps of deriving the temporal motion information are as follows, taking the position 2 in the above table as an example:
[0131] Suppose the top-left luma sample position of the current PU is (x, y), the width of the luma prediction block is l_width, and the height of the luma prediction block is l_height; and the bottom-right luma sample position of the selected current PU is (x', y'), then x' = x + l_width, y' = y + l_height.
[0132] If the (x', y') derived above is unavailable, such as exceeding the image boundary, patch boundary, etc., then x' = x + l_width - 1, y' = y + l_height - 1.
[0133] The first step is,
[0134] If the reference frame index stored in the temporal motion information storage unit corresponding to the luma sample in the reference image in the reference image queue 1 with the reference index value of 0 and the bottom-right luma sample position of the selected current PU is -1, then the L0 reference index and the L1 reference index of the current PU are both equal to 0. Take the size and position of the coding unit in which the current PU is located as the size and position of the current PU, then take the obtained L0 motion vector predictor and L1 motion vector predictor as the L0 motion vector MvE0 and the L1 motion vector MvE1 of the current PU respectively, and let the L0 reference index RefIdxL0 and the L1 reference index RefIdxL1 of the current PU both equal to 0, and end the motion information derivation process.
[0135] Otherwise,
[0136] The L0 reference index and the L1 reference index of the current PU are both equal to 0. The distance indexes of the images corresponding to the L0 reference index and the L1 reference index of the current PU are respectively denoted as DistanceIndexL0 and DistanceIndexL1; and the BlockDistances of the images corresponding to the L0 reference index and the L1 reference index of the current PU are respectively denoted as BlockDistanceL0 and BlockDistanceL1.
[0137] The L0 motion vector of the time-domain motion information storage unit in which the luma sample corresponding to the lower right luma sample position of the selected current prediction unit in the reference image in the reference index 0 in the reference image queue 1 is recorded, is denoted as mvRef (mvRef_x, mvRef_y), the distance index of the image in which the motion information storage unit is located is denoted as DistanceIndexCol, and the distance index of the image in which the reference unit pointed to by the motion vector is located is denoted as DistanceIndexRef.
[0138] The second step is,
[0139]
[0140] The third step is,
[0141] Let the L0 reference index RefIdxL0 of the current prediction unit be equal to 0, and calculate the L0 motion vector mvE0 (mvE0_x, mvE0_y) of the current prediction unit:
[0142]
[0143] Here, mvX is mvRef, and MVX is mvE0.
[0144] Let the L1 reference index RefIdxL1 of the current prediction unit be equal to 0, and calculate the L1 motion vector mvE1 (mvE1_x, mvE1_y) of the current prediction unit:
[0145]
[0146] Here, mvX is mvRef, and MVX is mvE1.
[0147] The fourth step is that the value of interPredRefMode is equal to 'PRED_List01'.
[0148] Method two, taking the 4th position in Figure 7 as an example, the steps of deriving the time-domain motion information are as follows:
[0149] It is assumed that the upper left corner luma sample position of the current prediction unit is (x, y), the width of the luma prediction block is l_width, and the height of the luma prediction block is l_height; and the right luma sample position of the selected current prediction unit is (x', y'), x' = x + l_width, and y' = y + l_height - 1.
[0150] If the derived (x', y') is not available, such as out of image boundary, patch boundary, etc., then x' = x + l_width - 1, y' = y + l_height - 1.
[0151] The first step is to determine the reference index value of the current prediction unit.
[0152] If the reference frame index stored in the temporal motion information storage unit where the luma sample corresponding to the right luma sample position of the selected current prediction unit in the reference picture with reference index value 0 in the reference picture list 1 is -1, then the L0 reference index and the L1 reference index of the current prediction unit are both equal to 0. The size and position of the coding unit where the current prediction unit is located are taken as the size and position of the current prediction unit, and then the obtained L0 motion vector predictor and L1 motion vector predictor are taken as the L0 motion vector MvE0 and the L1 motion vector MvE1 of the current prediction unit, respectively, and the L0 reference index RefIdxL0 and the L1 reference index RefIdxL1 of the current prediction unit are both equal to 0, and the motion information derivation process is ended.
[0153] Otherwise,
[0154] The L0 reference index and the L1 reference index of the current prediction unit are both equal to 0. The distance index of the picture corresponding to the L0 reference index and the L1 reference index of the current prediction unit are denoted as DistanceIndexL0 and DistanceIndexL1, respectively; and the BlockDistance of the picture corresponding to the L0 reference index and the L1 reference index of the current prediction unit are denoted as BlockDistanceL0 and BlockDistanceL1, respectively.
[0155] The L0 motion vector of the temporal motion information storage unit where the luma sample corresponding to the right luma sample position of the selected current prediction unit in the reference picture with reference index value 0 in the reference picture list 1 is denoted as mvRef (mvRef_x, mvRef_y), the distance index of the picture where the motion information storage unit is located is denoted as DistanceIndexCol, and the distance index of the picture where the reference unit pointed to by the motion vector is located is denoted as DistanceIndexRef.
[0156] The second step is to determine the reference index value of the current prediction unit.
[0157]
[0158] The third step is to determine the reference index value of the current prediction unit.
[0159] The L0 reference index RefIdxL0 of the current prediction unit is equal to 0, and the L0 motion vector mvE0 (mvE0_x, mvE0_y) of the current prediction unit is calculated as follows:
[0160]
[0161] Here, mvX is mvRef, and MVX is mvE0.
[0162] Let the L1 reference index RefIdxL1 of the current prediction unit be equal to 0, and calculate the L1 motion vector mvE1 (mvE1_x, mvE1_y) of the current prediction unit:
[0163]
[0164] Here, mvX is mvRef, and MVX is mvE1.
[0165] In the fourth step, the value of interPredRefMode is equal to 'PRED_List01'.
[0166] Method three, taking the 3 position in Figure 7 as an example, the steps of deriving the temporal motion information are as follows:
[0167] Suppose the top-left corner luminance sample position of the current prediction unit is (x, y), the width of the luminance prediction block is l_width, and the height of the luminance prediction block is l_height; and the luminance sample position below the selected current prediction unit is (x', y'), then x' = x + l_width - 1, y' = y + l_height.
[0168] If the (x', y') derived above is not available, such as exceeding the image boundary, patch boundary, etc., then x' = x + l_width - 1, y' = y + l_height - 1.
[0169] In the first step,
[0170] If the reference frame index stored in the temporal motion information storage unit corresponding to the luminance sample in the reference image queue 1 with the reference index value of 0 and the luminance sample position below the selected current prediction unit is -1, then the L0 reference index and the L1 reference index of the current prediction unit are both equal to 0. Take the size and position of the coding unit in which the current prediction unit is located as the size and position of the current prediction unit, and then take the obtained L0 motion vector prediction value and L1 motion vector prediction value as the L0 motion vector MvE0 and L1 motion vector MvE1 of the current prediction unit, respectively, and let the L0 reference index RefIdxL0 and the L1 reference index RefIdxL1 of the current prediction unit be both equal to 0, and end the motion information derivation process.
[0171] Otherwise,
[0172] The L0 reference index and the L1 reference index of the current prediction unit are both equal to 0. The distance index of the image corresponding to the L0 reference index and the L1 reference index of the current prediction unit are denoted as DistanceIndexL0 and DistanceIndexL1 respectively; the BlockDistance of the image corresponding to the L0 reference index and the L1 reference index of the current prediction unit are denoted as BlockDistanceL0 and BlockDistanceL1 respectively.
[0173] The L0 motion vector of the time-domain motion information storage unit in which the luma sample corresponding to the lower luma sample position of the selected current prediction unit in the image with reference index 0 in the reference image queue 1 is denoted as mvRef(mvRef_x, mvRef_y), the distance index of the image in which the time-domain motion information storage unit is located is denoted as DistanceIndexCol, and the distance index of the image in which the reference unit pointed to by the motion vector is located is denoted as DistanceIndexRef.
[0174] The second step,
[0175]
[0176] The third step,
[0177] Let the L0 reference index RefIdxL0 of the current prediction unit be equal to 0, and calculate the L0 motion vector mvE0(mvE0_x, mvE0_y) of the current prediction unit:
[0178]
[0179] Here, mvX is mvRef, and MVX is mvE0.
[0180] Let the L1 reference index RefIdxL1 of the current prediction unit be equal to 0, and calculate the L1 motion vector mvE1(mvE1_x, mvE1_y) of the current prediction unit:
[0181]
[0182] Here, mvX is mvRef, and MVX is mvE1.
[0183] The fourth step, the value of interPredRefMode is equal to 'PRED_List01'.
[0184] In this way, after the time-domain motion information is derived, the obtained time-domain motion information can be filled into the motion information candidate list to obtain a new motion information candidate list.
[0185] S304: constructing a new motion information candidate list based on the at least one temporal motion information.
[0186] It should be noted that after obtaining the at least one temporal motion information, it can be filled into the motion information candidate list to obtain the new motion information candidate list. Specifically, for S304, the step can include: filling the at least one temporal motion information into the motion information candidate list to obtain the new motion information candidate list.
[0187] It should also be noted that the existing motion information candidate list only reserves one filling position of temporal motion information, in order to improve the correlation of the lower right corner, the filling position of temporal motion information in the motion information candidate list can also be increased. Specifically, in some embodiments, the method can also include:
[0188] adjusting the proportion value of the temporal motion information in the new motion information candidate list;
[0189] controlling the reserved filling positions of at least two temporal motion information in the new motion information candidate list according to the adjusted proportion value.
[0190] That is, the proportion value of temporal motion information in the motion information candidate list can be increased. If the candidate list in the AWP prediction mode reserves at least 1 position for temporal motion information, it can be adjusted to reserve at least 2 (or 3) positions for temporal motion information in the candidate list in the AWP prediction mode, so that at least two filling positions of temporal motion information are reserved in the new motion information candidate list.
[0191] It can be understood that when the prediction mode parameter indicates that the preset inter prediction mode (such as GPM or AWP) is used to determine the inter prediction value of the current block, two partitions of the current block can be determined. That is, the method can also include: when the prediction mode parameter indicates that GPM or AWP is used to determine the inter prediction value of the current block, determining two partitions of the current block; wherein the two partitions include a first partition and a second partition.
[0192] Here, in the GPM or AWP prediction mode, the candidate position used to derive the temporal motion information can also be selected according to the division mode of GPM or AWP, or the arrangement combination of the temporal motion information derived according to different positions can also be selected. In some embodiments, the method can also include:
[0193] grouping the various division modes under GPM or AWP to obtain at least two division mode sets;
[0194] determining at least one candidate position corresponding to each of the at least two sets of partition modes respectively; wherein different sets of partition modes correspond to different at least one candidate position;
[0195] For each partition mode in each set of partition modes, performing the step of determining the at least one temporal motion information of the current block based on the corresponding determined at least one candidate position.
[0196] Further, the at least two sets of partition modes include a first set of partition modes and a second set of partition modes, and the method can further include:
[0197] If the current partition mode belongs to the first set of partition modes, determining a top-left pixel position inside the current block as the at least one candidate position of the current block, and performing the step of determining the at least one temporal motion information of the current block based on the at least one candidate position.
[0198] If the current partition mode belongs to the second set of partition modes, determining a bottom-right pixel position outside the current block as the at least one candidate position of the current block, and performing the step of determining the at least one temporal motion information of the current block based on the at least one candidate position.
[0199] Further, the at least two sets of partition modes include a first set of partition modes, a second set of partition modes, a third set of partition modes and a fourth set of partition modes, and the method can further include:
[0200] If the current partition mode belongs to the first set of partition modes, determining a first bottom-right candidate position of the current block as the at least one candidate position of the current block, and performing the step of determining the at least one temporal motion information of the current block based on the at least one candidate position.
[0201] If the current partition mode belongs to the second set of partition modes, determining a second bottom-right candidate position of the current block as the at least one candidate position of the current block, and performing the step of determining the at least one temporal motion information of the current block based on the at least one candidate position.
[0202] If the current partition mode belongs to the third set of partition modes, determining a third bottom-right candidate position of the current block as the at least one candidate position of the current block, and performing the step of determining the at least one temporal motion information of the current block based on the at least one candidate position.
[0203] If the current partition mode belongs to the fourth set of partition modes, a fourth bottom-right candidate position of the current block is determined as at least one candidate position of the current block, and the step of determining at least one temporal motion information of the current block based on the at least one candidate position is performed.
[0204] That is, the partition modes of GPM or AWP can be grouped into at least two sets of partition modes. For example, in one possible implementation, if the partition modes of AWP are divided into a first set of partition modes 0-31 and a second set of partition modes 32-55, if the current partition mode of AWP is one of the first set of partition modes 0-31, the temporal motion information can be derived according to the existing method (i.e., the top-left corner position of the current block); otherwise, if the current partition mode of AWP is one of the second set of partition modes 32-55, the temporal motion information can be derived according to the method of the present application (i.e., the bottom-right corner position outside the current block). Alternatively, the partition modes of GPM or AWP can be divided into four sets of partition modes, and the four sets of partition modes correspond to four bottom-right candidate positions of the current block, for example, a first set of partition modes corresponds to a first bottom-right candidate position (e.g., position 1 in Figure 7 , a second set of partition modes corresponds to a second bottom-right candidate position (e.g., position 2 in Figure 7 , a third set of partition modes corresponds to a third bottom-right candidate position (e.g., position 3 in Figure 7 , and a fourth set of partition modes corresponds to a fourth bottom-right candidate position (e.g., position 4 in Figure 7 .
[0205] In another possible implementation, the position used to derive the temporal motion information or the arrangement combination of deriving the temporal motion information according to different positions can be selected according to the partition mode of GPM or AWP. For example, the two ways of deriving the temporal motion information using the top-left corner position of the current block and the bottom-right corner position (e.g., position 2 in Figure 8 ) of the current block can be taken as an example, that is, some partition modes of GPM or AWP can use the top-left corner position of the current block to derive the temporal motion information, and other partition modes of GPM or AWP can use the bottom-right corner position (e.g., position 2 in Figure 9A ) of the current block to derive the temporal motion information. Here, the present application is not limited to the top-left corner and the bottom-right corner of the current block, and can also be the bottom-left corner and the top-right corner, which are not limited by the present application.
[0206] It is also noted that, since the number of closely related spatial positions that can be found by GPM or AWP prediction mode for some partition mode is different for two partitions, the number can also be used to determine which part of the position derivation of the temporal motion information is used for the current block. Therefore, in some embodiments, the method can further comprise:
[0207] determining a first number corresponding to a first spatial pixel position, wherein the first spatial pixel position is adjacent to at least one boundary space of the first partition;
[0208] determining a second number corresponding to a second spatial pixel position, wherein the second spatial pixel position is adjacent to at least one boundary space of the second partition;
[0209] if the first number or the second number is less than a preset value, determining a lower right pixel position outside the current block as at least one candidate position of the current block, and performing the step of determining at least one temporal motion information of the current block based on the at least one candidate position;
[0210] if the first number and the second number are both greater than the preset value, determining an upper left pixel position of the current block as at least one candidate position of the current block, and performing the step of determining at least one temporal motion information of the current block based on the at least one candidate position.
[0211] Here, one possible implementation principle is that if the number of closely related spatial positions (i.e., the spatial positions F, G, C, A, B, D in Figure 9A adjacent to the nearest pixel position of the partition) that can be found by GPM or AWP under partitioning of a certain partition is less than N, where N is 1 or 2 or 3 or 4, etc.; at this time, the temporal motion information can use the position of the lower right corner of the current block to derive the temporal motion information, otherwise the temporal motion information uses the position of the upper left corner of the current block to derive the temporal motion information.
[0212] Another possible implementation principle is that if the number of closely related spatial positions (i.e., the spatial positions F, G, C, A, B, D in Figure 9B adjacent to the nearest pixel position of the partition) that can be found by GPM or AWP under partitioning of two partitions is greater than N, where N is 1 or 2 or 3 or 4, etc.; at this time, the temporal motion information can use the position of the upper left corner of the current block to derive the temporal motion information, otherwise the temporal motion information uses the position of the lower right corner of the current block to derive the temporal motion information.
[0213] It is important to note that for different coding block or prediction block shapes, such as square, 2N×N, N×2N, 4N×N, N×4N, etc., the same partitioning pattern of GPM or AWP can find closely related spatial locations (i.e., Figure 9B The number of spatial positions (F, G, C, A, B, D, etc.) adjacent to the nearest pixel position of the partition is not necessarily the same. One possible approach is to use the same temporal motion information derivation method for the same partitioning pattern under any shape, and which one or more positions are used to derive the temporal motion information may be determined by the case corresponding to a square. Another possible approach is to use different temporal motion information derivation methods for the same partitioning pattern under different shapes. Yet another possible approach is to specify certain partitioning patterns to use one spatial motion information derivation method, and specify other partitioning patterns to use another spatial motion information derivation method, and so on. Here, specifying which patterns use which spatial motion information derivation method is determined based on the above-mentioned correlations, and this application embodiment does not make specific limitations.
[0214] In this way, after deriving the temporal motion information, a new candidate list of motion information can be constructed. The inter-frame prediction value for the current block is then determined based on this new candidate list.
[0215] S305: Determine the inter-frame prediction value of the current block based on the new motion information candidate list.
[0216] It should be noted that when the prediction mode parameter indicates that GPM or AWP is used to determine the inter-frame prediction value of the current block, the two partitions of the current block can be determined; the two partitions can include the first partition and the second partition.
[0217] In this way, after obtaining a new list of motion information candidates, the motion information corresponding to the first partition and the motion information of the second partition of the current block can be determined; then, based on the motion information corresponding to the first partition and the motion information of the second partition, the inter-frame prediction value of the current block can be determined.
[0218] Specifically, such as Figure 9A The diagram illustrates a flowchart of another inter-frame prediction method provided in an embodiment of this application. This method may include:
[0219] S801: Parse the bitstream and determine the first motion information index value corresponding to the first partition and the second motion information index value corresponding to the second partition;
[0220] S802: Based on the new motion information candidate list, determine the motion information in the new motion information candidate list indicated by the first motion information index value as the motion information of the first partition, and determine the motion information in the new motion information candidate list indicated by the second motion information index value as the motion information of the second partition;
[0221] S803: Calculate the first predicted value of the first partition using the motion information of the first partition, and calculate the second predicted value of the second partition using the motion information of the second partition;
[0222] S804: The first predicted value and the second predicted value are weighted and fused to obtain the inter-frame predicted value of the current block.
[0223] It's important to note that traditional unidirectional prediction simply finds a reference block of the same size as the current block, while traditional bidirectional prediction uses two reference blocks of the same size. The pixel value of each point within the predicted block is the average of the corresponding positions in the two reference blocks; that is, all points in each reference block occupy a 50% proportion. Bidirectional weighted prediction allows the proportions of the two reference blocks to differ, such as all points in the first reference block occupying a 75% proportion and all points in the second reference block occupying a 25% proportion. However, the proportions of all points within the same reference block are the same. Other optimization methods, such as using Decoder-Side Motion Vector Refinement (DMVR) and Bidirectional Optical Flow (BIO), can cause some changes in the reference or predicted pixels. Furthermore, GPM or AWP also use two reference blocks of the same size as the current block, but some pixel positions use 100% of the pixel values from the corresponding positions in the first reference block, and some pixel positions use 100% of the pixel values from the corresponding positions in the second reference block. In the boundary regions, the pixel values from the corresponding positions in the two reference blocks are used in a certain proportion. How these weights are specifically allocated is determined by the prediction mode of GPM or AWP, or it can be considered that GPM or AWP uses two reference blocks that are different from the current block size, that is, taking a portion of each block as a reference block.
[0224] For example, such as Figure 9B As shown, this illustration illustrates a weight allocation diagram of multiple partitioning modes of GPM on a 64×64 current block, provided by an embodiment of this application. Figure 7 In China, GPM has 64 possible partitioning methods. For example... Figure 7 As shown, this illustration illustrates the weight allocation diagram for various partitioning modes of an AWP on a 64×64 current block, provided by an embodiment of this application. Figure 7In the AWP, there are 56 partition modes. Figure 7 Figure 7 In each partition mode, the black region represents that the weight value of the first reference block is 0%, the white region represents that the weight value of the first reference block is 100%, and the gray region represents that the weight value of the first reference block is greater than 0% and less than 100% according to the color depth, and the weight value of the second reference block is 100% minus the weight value of the first reference block.
[0225] It should be understood that in the early coding technology, only rectangular partition mode exists, whether it is CU, PU or transform unit (TU) partition. The GPM or AWP realizes non-rectangular partition, that is, a straight line can divide a rectangular block into two partitions. According to the position and angle of the straight line, the two partitions may be triangular or trapezoidal or rectangular, so that the partition is closer to the edge of the object or the edge of two different motion regions. It should be noted that the partition here is not a real partition, but more like a prediction effect partition. Because this partition only divides the weight of the two reference blocks when generating the prediction block, or it can be simply understood that part of the position of the prediction block comes from the first reference block, and the other part of the position comes from the second reference block, and the current block is not really divided into two CUs or PUs or TUs according to the partition line. In this way, the residual transformation, quantization, inverse transformation, and inverse quantization after prediction are all processed as a whole.
[0226] It should also be noted that the GPM or AWP belongs to an inter-frame prediction technology. The GPM or AWP needs to transmit a flag in the code stream to indicate whether the GPM or AWP is used. The flag can indicate whether the current block uses the GPM or AWP. If the GPM or AWP is used, the encoder needs to transmit the specific mode used in the code stream, that is, one of the 64 GPM partition modes or one of the 56 AWP partition modes; and the index values of the two single-direction motion information. That is, for the current block, the decoder can obtain the information whether the GPM or AWP is used by analyzing the code stream. If it is determined that the GPM or AWP is used, the decoder can analyze the prediction mode parameters of the GPM or AWP and the two motion information index values. For example, if the current block can be divided into two partitions, the first motion information index value corresponding to the first partition and the second motion information index value corresponding to the second partition can be analyzed.
[0227] Before calculating the inter-frame prediction value of the current block, a new motion information candidate list needs to be constructed. The following will take the AWP in AVS as an example to introduce the construction method of the motion information candidate list.
[0228] like Figure 10 As shown, block E is the current block, while blocks A, B, C, D, F, and G are all neighboring blocks of block E. Among them, block A, a neighboring block of block E, is a sample (…). , The block containing block E, and the adjacent block B of block E are samples ( , The block containing block E, and the adjacent block C of block E are samples ( , The block containing block E, and the adjacent block D of block E are samples ( , The block containing block E, and the adjacent block F of block E are samples ( , The block containing block E, and the adjacent block G of block E are samples ( , The block containing (). , ) is the coordinate of the top-left corner sample of block E in the image. , ) is the coordinate of the upper right corner sample of block E in the image. , () represents the coordinates of the lower left corner sample of block E in the image. In other words, the spatial relationship between block E and its neighboring blocks A, B, C, D, F, and G is detailed in [link to image]. Figure 10 .
[0229] for Figure 7 In this context, the existence of a neighboring block X (represented as A, B, C, D, F, or G) means that the block should be within the image to be decoded and that the block should belong to the same spatial region as block E; otherwise, the neighboring block "does not exist." Therefore, if a block "does not exist" or has not yet been decoded, then this block is "unusable"; otherwise, this block is "usable." Alternatively, if the block containing the image sample to be decoded "does not exist" or this sample has not yet been decoded, then this sample is "unusable"; otherwise, this sample is "usable."
[0230] Assume the first unidirectional motion information is represented as mvAwp0L0, mvAwp0L1, RefIdxAwp0L0, and RefIdxAwp0L1. Here, mvAwp0L0 represents the motion vector corresponding to the first reference frame list RefPicList0, and RefIdxAwp0L0 represents the reference index value of the corresponding reference frame in RefPicList0; mvAwp0L1 represents the motion vector corresponding to the second reference frame list RefPicList1, and RefIdxAwp0L1 represents the reference index value of the corresponding reference frame in RefPicList1. The second unidirectional motion information follows the same pattern.
[0231] Since the motion information here is unidirectional, one of RefIdxAwp0L0 and RefIdxAwp0L1 must be a valid value, such as 0, 1, 2, etc.; the other must be an invalid value, such as -1. If RefIdxAwp0L0 is a valid value, then RefIdxAwp0L1 is -1; in this case, the corresponding mvAwp0L0 is the required motion vector, i.e., (x, y), and mvAwp0L1 does not need to be considered. The reverse is also true.
[0232] Specifically, the steps for deriving mvAwp0L0, mvAwp0L1, RefIdxAwp0L0, RefIdxAwp0L1, mvAwp1L0, mvAwp1L1, RefIdxAwp1L0, and RefIdxAwp1L1 are as follows:
[0233] First step, such as Figure 7 As shown, F, G, C, A, B, and D are the neighboring blocks of the current block E. Determine the "availability" of F, G, C, A, B, and D:
[0234] (a) If F exists and inter-frame prediction mode is used, then F is “available”; otherwise, F is “unavailable”.
[0235] (b) If G exists and is in inter-frame prediction mode, then G is “available”; otherwise, G is “unavailable”.
[0236] (c) If C exists and is in inter-frame prediction mode, then C is “available”; otherwise, C is “unavailable”.
[0237] (d) If A exists and is in inter-frame prediction mode, then A is “available”; otherwise, A is “unavailable”.
[0238] (e) If B exists and is in inter-frame prediction mode, then B is “available”; otherwise B is “unavailable”.
[0239] (f) If D exists and inter-frame prediction mode is used, then D is “available”; otherwise, D is “unavailable”.
[0240] The second step is to add the available unidirectional motion information into the unidirectional motion information candidate list (represented by AwpUniArray) in the order of F, G, C, A, B and D, until the length of AwpUniArray is 3 (or 4) or the traversal ends.
[0241] Third step, if the length of AwpUniArray is less than 3 (or 4), split the bi-directional available motion information into uni-directional motion information pointing to reference frame list ListO and uni-directional motion information pointing to reference frame list Listl in the order of F, G, C, A, B and D, first do uni-directional motion information duplication check, if not duplicated, put into AwpUniArray, until the length is 3 (or 4) or the end of iteration.
[0242] Fourth step, split the bi-directional motion information derived from the above-mentioned method one, method two and method three into uni-directional motion information pointing to reference frame list ListO and uni-directional motion information pointing to reference frame list Listl in turn, first do uni-directional motion information duplication check, if not duplicated, put into AwpUniArray, until the length is 4 (or 5) or the end of iteration.
[0243] Fifth step, if the length of AwpUniArray is less than 4 (or 5), then repeat the last uni-directional motion information in AwpUniArray until the length of AwpUniArray is 4 (or 5).
[0244] Sixth step, assign the AwpCandIdxO+lth motion information in AwpUniArray to mvAwpOL0, mvAwpOLl, RefIdxAwpOL0 and RefIdxAwpOLl, and assign the AwpCandIdxl+lth motion information in AwpUniArray to mvAwpIL0, mvAwpILl, RefIdxAwpIL0 and RefIdxAwpILl.
[0245] Thus, for the current block, the decoder can obtain the information whether GPM or AWP is used by parsing the code stream, if it is determined that GPM or AWP is used, the decoder can parse the prediction mode parameters of GPM or AWP and two motion information index values, and the decoder constructs the motion information candidate list used by GPM or AWP of the current block, then according to the two motion information index values parsed, the two uni-directional motion information can be found in the new motion information candidate list constructed above, then the two reference blocks can be found by using the two uni-directional motion information, according to the specific prediction mode used by GPM or AWP, the weight values of the two reference blocks at each pixel position can be determined, finally the two reference blocks are weighted to obtain the prediction block of the current block.
[0246] Further, if the current mode is skip mode, the prediction block is the decoded block, meaning the decoding of the current block is finished. If the current mode is not skip mode, the quantized coefficients are entropy-decoded, followed by inverse quantization and inverse transform to obtain a residual block, and finally the decoded block is obtained by adding the residual block and the prediction block, meaning the decoding of the current block is finished.
[0247] In addition, for the decoder, if the partition mode of the current AWP is one of the 0th to 31st partition modes, the temporal motion information can be derived according to the existing method; otherwise, the temporal motion information is derived according to the aforementioned method one (or the temporal motion information is derived according to the aforementioned method one, method two, and method three in turn), and the derived temporal bidirectional motion information is split into single-directional motion information pointing to the reference frame list List0 and single-directional motion information pointing to the reference frame list List1, and a single-directional motion information duplication operation is performed first, and if not duplicated, it is put into AwpUniArray until the length is 4 (or 5) or the end of the traversal is reached. The specific text description is as follows,
[0248] Specifically, the steps of deriving mvAwp0L0, mvAwp0L1, RefIdxAwp0L0, RefIdxAwp0L1, mvAwp1L0, mvAwp1L1, RefIdxAwp1L0, and RefIdxAwp1L1 are as follows:
[0249] First, as shown in Figure 7 F, G, C, A, B, and D are adjacent blocks of the current block E, and the "availability" of F, G, C, A, B, and D is determined:
[0250] (a) If F exists and adopts an inter-prediction mode, F is "available"; otherwise, F is "unavailable".
[0251] (b) If G exists and adopts an inter-prediction mode, G is "available"; otherwise, G is "unavailable".
[0252] (c) If C exists and adopts an inter-prediction mode, C is "available"; otherwise, C is "unavailable".
[0253] (d) If A exists and adopts an inter-prediction mode, A is "available"; otherwise, A is "unavailable".
[0254] (e) If B exists and adopts an inter-prediction mode, B is "available"; otherwise, B is "unavailable".
[0255] (f) If D exists and adopts an inter-prediction mode, D is "available"; otherwise, D is "unavailable".
[0256] Second step, put the uni-directional available motion information into uni-directional motion information candidate list (denoted as AwpUniArray) in the order of F, G, C, A, B and D, until the length of AwpUniArray is 3 (or 4) or the iteration is over.
[0257] Third step, if the length of AwpUniArray is less than 3 (or 4), split the bi-directional available motion information into uni-directional motion information pointing to reference frame list List0 and uni-directional motion information pointing to reference frame list List1 in the order of F, G, C, A, B and D, and perform uni-directional motion information duplication check operation first, and if not duplicated, put into AwpUniArray, until the length is 3 (or 4) or the iteration is over.
[0258] Fourth step, if AwpIdx (in AWP prediction mode) belongs to 0~31, derive the temporal bi-directional motion information according to the existing method (i.e. the top-left corner position of the current block), otherwise derive the temporal bi-directional motion information according to the method of the embodiments of the present application (such as the aforementioned method one), split the derived bi-directional motion information into uni-directional motion information pointing to reference frame list List0 and uni-directional motion information pointing to reference frame list List1, perform uni-directional motion information duplication check operation first, and if not duplicated, put into AwpUniArray, until the length is 4 (or 5) or the iteration is over.
[0259] Fifth step, if the length of AwpUniArray is less than 4 (or 5), perform repeated padding operation on the last uni-directional motion information in AwpUniArray, until the length of AwpUniArray is 4 (or 5).
[0260] Sixth step, assign the AwpCandIdx0+1th motion information in AwpUniArray to mvAwp0L0, mvAwp0L1, RefIdxAwp0L0 and RefIdxAwp0L1, and assign the AwpCandIdx1+1th motion information in AwpUniArray to mvAwp1L0, mvAwp1L1, RefIdxAwp1L0 and RefIdxAwp1L1.
[0261] That is, in the embodiments of the present application, the temporal motion information is constructed by using the inter prediction method of the embodiments of the present application and filled into the motion information candidate list, so that the motion information candidate list is constructed by using the right lower temporal motion information. More right lower candidate positions are used for AWP or GPM, which can enhance the correlation of the right lower position. An offset can also be used to reduce the case that the selected several blocks belong to one block in the reference frame. In addition, when the selected position is not available, the right lower corner position inside the current block is used as a substitute; or the position used to derive the temporal motion information is selected according to the partition mode of GPM or AWP, or the arrangement combination of the temporal motion information derived according to different positions is selected.
[0262] In this way, the motion information candidate list of the embodiments of the present application supplements and enhances the motion information with stronger correlation with the right lower position, especially for those blocks in the right lower position in the AWP partition and with weak correlation with the upper left position in the original construction mode, the blocks can obtain motion information with stronger correlation, thereby improving the coding performance; and the correlation of the right lower position in GPM is weak, and the correlation of the right lower position can be improved by increasing the right lower temporal motion information candidate position, thereby improving the coding performance.
[0263] It should be further noted that the motion information candidate list of the embodiments of the present application generally refers to a single-direction motion information candidate list, but the construction mode of the single-direction motion information of the embodiments of the present application can be extended to the construction of bidirectional motion information, so that the construction of the single-direction motion information candidate list can also be extended to the construction of the bidirectional motion information candidate list.
[0264] The embodiments of the present application provide an inter prediction method applied to a decoder. A bitstream is parsed to obtain a prediction mode parameter of a current block; when the prediction mode parameter indicates that a preset inter prediction mode is used to determine an inter prediction value of the current block, at least one candidate position of the current block is determined; wherein the candidate position at least includes a right lower position inside the current block and a right lower position outside the current block; at least one temporal motion information of the current block is determined based on the at least one candidate position; a new motion information candidate list is constructed based on the at least one temporal motion information; and the inter prediction value of the current block is determined according to the new motion information candidate list. In this way, since the temporal motion information of the current block is determined based on the right lower position inside the current block or the right lower position outside the current block, the motion information with stronger correlation with the right lower position can be supplemented and enhanced in the motion information candidate list, thereby increasing the diversity of the motion information in the motion information candidate list; especially for GPM or AWP inter prediction mode, the correlation of the right lower position can be improved by increasing the right lower temporal motion information candidate position, thereby improving the coding and decoding performance.
[0265] The embodiment of the present application provides an inter-frame prediction method, which is applied to a video encoding device, i.e., an encoder. The function realized by the method can be realized by calling a computer program by a second processor in the encoder, and of course the computer program can be stored in a second memory. Therefore, the encoder at least includes the second processor and the second memory.
[0266] Referring to Figure 7 , a flowchart of another inter-frame prediction method is shown. As Figure 7 indicated, the method can include the following steps.
[0267] S1001: determining a prediction mode parameter of a current block;
[0268] It should be noted that the image to be encoded can be divided into a plurality of image blocks, and the current image block to be encoded can be referred to as a current block, and the image block adjacent to the current block can be referred to as a neighboring block; that is, in the image to be encoded, the current block and the neighboring block have an adjacent relationship. Here, each current block can include a first image component, a second image component and a third image component; that is, the current block is an image block in the image to be encoded, and the first image component, the second image component or the third image component prediction of the current block is performed.
[0269] Among them, assuming that the current block performs first image component prediction, and the first image component is a luminance component, that is, the image component to be predicted is a luminance component, then the current block can also be referred to as a luminance block; or, assuming that the current block performs second image component prediction, and the second image component is a chroma component, that is, the image component to be predicted is a chroma component, then the current block can also be referred to as a chroma block.
[0270] It should also be noted that the prediction mode parameter indicates the prediction mode adopted by the current block and the parameters related to the prediction mode. Here, for the determination of the prediction mode parameter, a simple decision strategy can be used, such as determining according to the size of the distortion value; or a complex decision strategy can be used, such as determining according to the result of rate distortion optimization (RDO), and the embodiment of the present application does not make any limitation. Generally, the RDO method can be used to determine the prediction mode parameter of the current block.
[0271] Specifically, in some embodiments, for S1001, the determination of the prediction mode parameter of the current block can include:
[0272] performing pre-encoding processing on the current block by using a plurality of prediction modes to obtain a rate distortion cost value corresponding to each prediction mode;
[0273] The minimum rate-distortion value is selected from the obtained multiple rate-distortion values, and a prediction mode corresponding to the minimum rate-distortion value is determined as the prediction mode parameter of the current block.
[0274] That is, at the encoder side, the current block can be pre-encoded by using multiple prediction modes respectively. Here, the multiple prediction modes usually include an inter-prediction mode, a traditional intra-prediction mode and a non-traditional intra-prediction mode; the traditional intra-prediction mode can include a direct current (DC) mode, a planar (PLANAR) mode and an angle mode, etc., the non-traditional intra-prediction mode can include a matrix-based intra prediction (MIP) mode, a cross-component linear model prediction (CCLM) mode, an intra block copy (IBC) mode and a PLT (Palette) mode, etc., and the inter-prediction mode can include a normal inter-prediction mode, a GPM prediction mode and an AWP prediction mode, etc.
[0275] In this way, after the current block is pre-encoded by using multiple prediction modes respectively, a rate-distortion value corresponding to each prediction mode can be obtained; then a minimum rate-distortion value is selected from the obtained multiple rate-distortion values, and a prediction mode corresponding to the minimum rate-distortion value is determined as the prediction mode parameter of the current block. In addition, after the current block is pre-encoded by using multiple prediction modes respectively, a distortion value corresponding to each prediction mode can be obtained; then a minimum distortion value is selected from the obtained multiple distortion values, and a prediction mode corresponding to the minimum distortion value is determined as the prediction mode parameter of the current block. In this way, the current block is finally encoded by using the determined prediction mode parameter, and in this prediction mode, the prediction residual is small, which can improve the encoding efficiency.
[0276] S1002: When the prediction mode parameter indicates that the inter-prediction value of the current block is determined by using a preset inter-prediction mode, at least one candidate position of the current block is determined; wherein the candidate position at least includes a right-down position inside the current block and a right-down position outside the current block.
[0277] It should be noted that if the prediction mode parameter indicates that the inter-prediction value of the current block is determined by using a preset inter-prediction mode, the inter-prediction method provided by the embodiments of the present application can be used. Here, the preset inter-prediction mode can be a GPM prediction mode or an AWP prediction mode, etc.
[0278] It should be noted that the motion information can include motion vector information and reference frame information. In addition, the reference frame information can be a reference frame corresponding to the reference frame list and the reference index value.
[0279] In some embodiments, for S1002, the determining the at least one candidate position of the current block can include:
[0280] obtaining a first bottom-right candidate position, a second bottom-right candidate position, a third bottom-right candidate position and a fourth bottom-right candidate position to form a candidate position set;
[0281] determining the at least one candidate position of the current block from the candidate position set;
[0282] The first bottom-right candidate position represents a bottom-right position inside the current block, and the second bottom-right candidate position, the third bottom-right candidate position and the fourth bottom-right candidate position represent bottom-right positions outside the current block.
[0283] It should be noted that, for example, Figure 7 the bottom-right position inside the current block can be the first bottom-right candidate position, as shown by the position filled with 1 in FIG. 1; Figure 7 the bottom-right position outside the current block can be the second bottom-right candidate position, as shown by the position filled with 2 in FIG. 1; Figure 11 the bottom-right position outside the current block can also be the third bottom-right candidate position, as shown by the position filled with 3 in FIG. 1; Figure 11 the bottom-right position outside the current block can also be the fourth bottom-right candidate position, as shown by the position filled with 4 in FIG. 1. Figure 11
[0284] In this way, after obtaining the first bottom-right candidate position, the second bottom-right candidate position, the third bottom-right candidate position and the fourth bottom-right candidate position, the candidate position set can be formed; and then the at least one candidate position of the current block is determined from the candidate position set.
[0285] Further, in some embodiments, the obtaining the first bottom-right candidate position, the second bottom-right candidate position, the third bottom-right candidate position and the fourth bottom-right candidate position can include:
[0286] obtaining coordinate information corresponding to a top-left pixel position of the current block, width information of the current block and height information of the current block;
[0287] performing coordinate calculation using the coordinate information corresponding to the top-left pixel position, the width information and the height information respectively to obtain first coordinate information, second coordinate information, third coordinate information and fourth coordinate information;
[0288] determining a position corresponding to the first coordinate information as the first bottom-right candidate position, determining a position corresponding to the second coordinate information as the second bottom-right candidate position, determining a position corresponding to the third coordinate information as the third bottom-right candidate position, and determining a position corresponding to the fourth coordinate information as the fourth bottom-right candidate position.
[0289] That is, the 1 position in the current block is determined according to the pixel position of the bottom-right corner of the current block. Here, still taking the example of Figure 11 assuming that the pixel position of the top-left corner of the current block is represented by (x, y), the width of the current block is represented by width, and the height of the current block is represented by height, the first coordinate information is (x+width-1, y+height-1), the second coordinate information is (x+width, y+height), the third coordinate information is (x+width-1, y+height), and the fourth coordinate information is (x+width, y+height-1); then according to the four coordinate information, the first bottom-right candidate position (1 position), the second bottom-right candidate position (2 position), the third bottom-right candidate position (3 position), and the fourth bottom-right candidate position (4 position) can be determined in turn.
[0290] Further, if it is considered that the four positions determined by the above-mentioned manner are too close, and the corresponding positions in the corresponding reference frame are more likely to belong to the same block, an offset, i.e., a first preset offset, can be set, which is represented by offset. Therefore, in some embodiments, after the first coordinate information, the second coordinate information, the third coordinate information, and the fourth coordinate information are obtained, the method can further include:
[0291] modifying the second coordinate information, the third coordinate information, and the fourth coordinate information by using the first preset offset to obtain first modified second coordinate information, first modified third coordinate information, and first modified fourth coordinate information;
[0292] determining a position corresponding to the first coordinate information as the first bottom-right candidate position, determining a position corresponding to the first modified second coordinate information as the second bottom-right candidate position, determining a position corresponding to the first modified third coordinate information as the third bottom-right candidate position, and determining a position corresponding to the first modified fourth coordinate information as the fourth bottom-right candidate position.
[0293] That is, among the four positions, the position of 1 can still be the position determined according to the first coordinate information, i.e., (x+width-1, y+height-1); the position of 2 can be the position determined according to the first modified second coordinate information, i.e., (x+width+offset, y+height+offset); the position of 3 can be the position determined according to the first modified third coordinate information, i.e., (x+width-1, y+height+offset); and the position of 4 can be the position determined according to the first modified fourth coordinate information, i.e., (x+width+offset, y+height-1).
[0294] Further, since there can be a time deviation between the current frame and the reference frame information used by the derived temporal motion information, there can also be a motion between the block on the current frame and the reference frame used by the derived temporal motion information. At this time, a second offset can also be set, denoted by (x', y'). Therefore, in some embodiments, after the first coordinate information, the second coordinate information, the third coordinate information and the fourth coordinate information are obtained, the method can further include:
[0295] When there is a motion vector between the image block on the current frame and the candidate reference frame, determining a second preset offset; wherein the candidate reference frame is a reference frame of the reference motion information used to determine the temporal motion information, and the image block at least includes a neighboring block, and the neighboring block is spatially adjacent to the current block in the current frame;
[0296] modifying the first coordinate information, the second coordinate information, the third coordinate information and the fourth coordinate information by using the second preset offset to obtain second modified first coordinate information, second modified second coordinate information, second modified third coordinate information and second modified fourth coordinate information;
[0297] determining the position corresponding to the second modified first coordinate information as the first bottom-right candidate position, determining the position corresponding to the second modified second coordinate information as the second bottom-right candidate position, determining the position corresponding to the second modified third coordinate information as the third bottom-right candidate position, and determining the position corresponding to the second modified fourth coordinate information as the fourth bottom-right candidate position.
[0298] Further, in some embodiments, the determination of the second preset offset can include:
[0299] obtaining a motion vector of a preset neighboring block on the current frame to the candidate reference frame, and determining the obtained motion vector as the second preset offset; or,
[0300] scaling motion information of the preset neighboring block on the current frame to the candidate reference frame to obtain a scaled motion vector, and determining the scaled motion vector as the second preset offset.
[0301] Here, for the determination of (x', y'), it can be that a motion vector of a certain neighboring block to the reference frame used for deriving the temporal motion information is found as (x', y'), or a motion vector of a certain neighboring block scaled to the reference frame used for deriving the temporal motion information is found as (x', y'), which is not limited here.
[0302] That is, one possible implementation is that when finding the corresponding position in the reference frame information used for deriving the temporal motion information in the above manner, it needs to consider that there is a motion vector of a block on the current frame to the reference frame used for deriving the temporal motion information. Exemplarily, if the first preset offset is considered, the 2 position is determined according to (x+width+offset, y+height+offset), and then the motion of the block (x+width+offset, y+height+offset) on the current frame to the reference frame used for deriving the temporal motion information is considered. Assuming that (x', y') is still the motion vector of the block (x+width+offset, y+height+offset) on the current frame to the reference frame used for deriving the temporal motion information, the second corrected second coordinate information is (x+width+offset+x', y+height+offset+y'), that is, the 2 position can be determined according to (x+width+offset+x', y+height+offset+y') in the reference frame used for deriving the temporal motion information.
[0303] Another possible implementation is that when finding the corresponding position in the reference frame information used for deriving the temporal motion information in the above manner, the motion vector existing between the block on the current frame and the reference frame used for deriving the temporal motion information is not considered. Exemplarily, if the first preset offset is considered, the 2 position is determined according to (x+width+offset, y+height+offset), and the motion of the block on the current frame (x+width+offset, y+height+offset) to the reference frame used for deriving the temporal motion information is not considered any more, then the second coordinate information at this time is (x+width+offset, y+height+offset), that is, the 2 position can be determined according to (x+width+offset, y+height+offset) of the reference frame used for deriving the temporal motion information. Here, the specific implementation can refer to the description on the decoder side.
[0304] In addition, in the embodiments of the present application, one possible way is to use only one of the above four positions, and one possible way is to use a combination of some positions of the above four positions. Therefore, optionally, in some embodiments, the determining of the at least one candidate position of the current block from the candidate position set can include:
[0305] selecting one candidate position from the candidate position set according to a preset manner, and determining the selected candidate position as one candidate position of the current block; or
[0306] selecting a candidate position corresponding to a high priority according to a preset priority order from the candidate position set and the selected candidate position being available, and determining the selected candidate position as one candidate position of the current block.
[0307] Optionally, in some embodiments, the determining of the at least one candidate position of the current block from the candidate position set can include:
[0308] selecting a plurality of candidate positions from the candidate position set according to a preset combination manner, and determining the selected plurality of candidate positions as a plurality of candidate positions of the current block; or
[0309] selecting a plurality of candidate positions according to a preset priority order from the candidate position set and the selected candidate positions being available, and determining the selected plurality of candidate positions as a plurality of candidate positions of the current block.
[0310] That is, one possible way is to use only one of the above four positions, and another possible way is to use a combination of some of the above four positions, such as using a combination of positions 2, 3, and 4, using a combination of positions 3 and 4, etc. Here, using all 1, 2, 3, and 4 positions is also a combination. In other words, any permutation and combination of any number (1, 2, 3, 4) of the above four positions can be used as a selection method for determining the candidate position. It should be noted that when multiple positions are needed, one possible way is to arrange 2 in the first place, i.e., position 2 has the highest priority.
[0311] Further, in some embodiments, the method can further include:
[0312] When selecting a candidate position from the set of candidate positions, if the candidate position to be selected belongs to the lower right position outside the current block and all lower right positions outside the current block are unavailable, the first lower right candidate position is determined as the candidate position to be selected.
[0313] That is, still taking Figure 12 as an example, positions 2, 3, and 4 are all outside the current block, and positions 2, 3, and 4 are sometimes unavailable. If the image boundary is encountered, the boundary that cannot be crossed by some inter-frame reference, etc., positions 2, 3, and 4 can all be unavailable. In the case where positions 2, 3, and 4 are all unavailable, one possible way is to use position 1 instead of the unavailable positions. If position 1 has already been used in the construction of the current motion information candidate list, the unavailable positions can also be skipped.
[0314] In this way, after determining at least one candidate position of the current block, the candidate position here is the position of the lower right corner of the current block, which can be the lower right position inside the current block or the lower right position outside the current block, and determining at least one temporal motion vector of the current block based on these candidate positions can increase the temporal motion information corresponding to the position of the lower right corner of the current block when constructing the motion information candidate list using temporal motion information, thereby improving the correlation of the lower right corner.
[0315] S1003: determining at least one temporal motion information of the current block based on the at least one candidate position.
[0316] It should be noted that after obtaining at least one candidate position, the temporal motion information can be determined based on the obtained candidate position. Specifically, in some embodiments, for S1003, the determining at least one temporal motion information of the current block based on the at least one candidate position can include:
[0317] determining reference frame information corresponding to each candidate position in the at least one candidate position;
[0318] For each candidate position, a time domain position associated with the candidate position is determined in the corresponding reference frame information, and motion information used by the time domain position is determined as time domain motion information corresponding to the candidate position;
[0319] Based on the at least one candidate position, at least one time domain motion information is obtained.
[0320] That is, the time domain motion information is determined according to motion information used by a corresponding position in a certain reference frame information. Moreover, different time domain motion information can be obtained for different candidate positions. In this way, after the time domain motion information is derived, the obtained time domain motion information can be filled into the motion information candidate list to obtain a new motion information candidate list.
[0321] S1004: Based on the at least one time domain motion information, a new motion information candidate list is constructed.
[0322] It should be noted that after the at least one time domain motion information is obtained, it can be filled into the motion information candidate list to obtain a new motion information candidate list. Specifically, for S304, this step can include: filling the at least one time domain motion information into the motion information candidate list to obtain the new motion information candidate list.
[0323] It should also be noted that only one filling position of the time domain motion information is reserved in the existing motion information candidate list, and in order to improve the correlation of the lower right corner, the filling position of the time domain motion information in the motion information candidate list can also be increased. Specifically, in some embodiments, the method can further include:
[0324] Adjusting a proportion value of the time domain motion information in the new motion information candidate list;
[0325] According to the adjusted proportion value, controlling at least two filling positions of the time domain motion information reserved in the new motion information candidate list.
[0326] That is, the proportion value of the time domain motion information in the motion information candidate list can be increased. If at least one position is reserved for the time domain motion information in the candidate list in the AWP prediction mode, at least two (or three) positions can be reserved for the time domain motion information in the candidate list in the AWP prediction mode, so that at least two filling positions of the time domain motion information are reserved in the new motion information candidate list.
[0327] It can be understood that when the prediction mode parameter indicates that the inter prediction value of the current block is determined using a preset inter prediction mode (such as GPM or AWP), two partitions of the current block can be determined. That is, the method can further include: when the prediction mode parameter indicates that the inter prediction value of the current block is determined using GPM or AWP, determining two partitions of the current block; wherein the two partitions include a first partition and a second partition.
[0328] Here, in the GPM or AWP prediction mode, the candidate position used for deriving the temporal motion information can also be selected according to the partition mode of GPM or AWP, or the arrangement combination of deriving the temporal motion information according to different positions can also be selected. In some embodiments, the method can further include:
[0329] Grouping a plurality of partition modes in GPM or AWP to obtain at least two sets of partition mode sets;
[0330] Determining at least one candidate position corresponding to each set of partition mode sets in the at least two sets of partition mode sets; wherein different sets of partition mode sets correspond to different at least one candidate position;
[0331] For each partition mode in each set of partition mode sets, according to the at least one candidate position determined to correspond, performing the step of determining at least one temporal motion information of the current block based on the at least one candidate position.
[0332] Further, the at least two sets of partition mode sets include a first set of partition mode sets and a second set of partition mode sets, and the method can further include:
[0333] If the current partition mode belongs to the first set of partition mode sets, the top-left pixel position inside the current block is determined as at least one candidate position of the current block, and the step of determining at least one temporal motion information of the current block based on the at least one candidate position is performed;
[0334] If the current partition mode belongs to the second set of partition mode sets, the bottom-right pixel position outside the current block is determined as at least one candidate position of the current block, and the step of determining at least one temporal motion information of the current block based on the at least one candidate position is performed.
[0335] It should be further noted that because in the GPM or AWP prediction mode, the number of closely related spatial positions that can be found by two partitions obtained by some partition modes is different, the position of which part of the current block is used to derive the temporal motion information can also be determined according to the number. Therefore, in some embodiments, the method can further include:
[0336] determining a first number corresponding to a first spatial pixel position adjacent to at least one boundary space of the first partition;
[0337] determining a second number corresponding to a second spatial pixel position adjacent to at least one boundary space of the second partition;
[0338] if the first number or the second number is less than a preset value, determining a bottom-right pixel position outside the current block as at least one candidate position of the current block, and performing the step of determining at least one temporal motion information of the current block based on the at least one candidate position;
[0339] if the first number and the second number are both greater than the preset value, determining a top-left pixel position of the current block as at least one candidate position of the current block, and performing the step of determining at least one temporal motion information of the current block based on the at least one candidate position.
[0340] It should be noted that one possible way is to select the position used to derive the temporal motion information according to the partition mode of GPM or AWP, or to select the arrangement combination of deriving the temporal motion information according to different positions. Here, the specific implementation can refer to the description of the decoder side.
[0341] In this way, after the temporal motion information is derived, a new motion information candidate list can be constructed. Subsequently, the inter prediction value of the current block is determined according to the new motion information candidate list.
[0342] S1005: determining the inter prediction value of the current block according to the new motion information candidate list.
[0343] It should be noted that when the prediction mode parameter indicates that the inter prediction value of the current block is determined using GPM or AWP, two partitions of the current block can be determined; wherein the two partitions can include a first partition and a second partition.
[0344] In this way, after the new motion information candidate list is obtained, the motion information corresponding to the first partition of the current block and the motion information of the second partition can be determined; and then the inter prediction value of the current block can be determined according to the motion information corresponding to the first partition and the motion information of the second partition.
[0345] Specifically, in some embodiments, for S1005, the determining the inter prediction value of the current block according to the new motion information candidate list can include:
[0346] determine the motion information of the first partition and the motion information of the second partition based on the new motion information candidate list, and set a first motion information index value as an index sequence value of the motion information of the first partition in the new motion information candidate list, and set a second motion information index value as an index sequence value of the motion information of the second partition in the new motion information candidate list;
[0347] calculate a first prediction value of the first partition by using the motion information of the first partition, and calculate a second prediction value of the second partition by using the motion information of the second partition;
[0348] perform weighted fusion on the first prediction value and the second prediction value to obtain an inter prediction value of the current block.
[0349] Further, in some embodiments, the method can further include:
[0350] write the first motion information index value and the second motion information index value into a bitstream.
[0351] It should be noted that GPM or AWP belongs to an inter prediction technology. At the encoder side, GPM or AWP needs to transmit a flag indicating whether GPM or AWP is used and two motion information index values (such as a first motion information index value and a second motion information index value) in a bitstream, so that the decoder side can directly obtain the flag indicating whether GPM or AWP is used and the two motion information index values by analyzing the bitstream.
[0352] That is, for the current block, GPM or AWP can be used for pre-encoding and other available prediction modes can be used for pre-encoding to determine whether to use GPM or AWP. If the pre-encoding cost of GPM or AWP is the smallest, GPM or AWP can be used. Meanwhile, when GPM or AWP is used, a motion information candidate list can also be constructed, and the construction manner is the same as that described in the decoder side embodiment.
[0353] In this way, at the encoder side, two uni-directional motion information are selected from the motion information candidate list, and then a pre-encoding is performed by selecting one mode from the division modes of GPM or AWP to determine the pre-encoding cost of GPM or AWP. One possible way is to determine the cost of all possible combinations of uni-directional motion information candidates based on all possible division modes of GPM or AWP, and then take the combination of the two uni-directional motion information and the division mode of GPM or AWP with the smallest cost as the finally determined two uni-directional motion information and the prediction mode of GPM or AWP.
[0354] Finally, information whether GPM or AWP is used is written in the bitstream. If it is determined that GPM or AWP is used, prediction mode parameters of GPM or AWP and two unidirectional motion information index values are written in the bitstream. In this way, if the current mode is the skip mode, the prediction block is the coding block, meaning that the coding of the current block ends. If the current mode is not the skip mode, quantization coefficients also need to be written in the bitstream; the quantization coefficients are composed of a residual block obtained by subtracting the inter-frame prediction value from the actual value of the current block, and the residual block is obtained by performing transform and quantization thereon, at which time the coding of the current block ends. That is, if the current mode is not the skip mode, the residual block is obtained by subtracting the inter-frame prediction block from the current block, and then the residual block is subjected to transform, quantization and entropy coding; subsequently, at the decoder side, for the case where the current mode is not the skip mode, the quantization coefficients are parsed through entropy decoding, and then the residual block is obtained by inverse quantization and inverse transform, and finally the decoded block is obtained by adding the residual block and the prediction block, meaning that the decoding of the current block ends.
[0355] The embodiment provides an inter-frame prediction method applied to an encoder. A prediction mode parameter of a current block is determined; at least one candidate position of the current block is determined when the prediction mode parameter indicates that a preset inter-frame prediction mode is used to determine an inter-frame prediction value of the current block; wherein the candidate position at least includes a right-bottom position inside the current block and a right-bottom position outside the current block; at least one time-domain motion information of the current block is determined based on the at least one candidate position; a new motion information candidate list is constructed based on the at least one time-domain motion information; and the inter-frame prediction value of the current block is determined according to the new motion information candidate list. In this way, since the time-domain motion information of the current block is determined based on the right-bottom position inside the current block or the right-bottom position outside the current block, motion information with higher correlation with the right-bottom position can be supplemented in the motion information candidate list, thereby increasing the diversity of the motion information in the motion information candidate list; especially for the GPM or AWP inter-frame prediction mode, the correlation with the right-bottom position can be improved by increasing the right-bottom time-domain motion information candidate position, thereby improving the coding and decoding performance.
[0356] Based on the same inventive concept as the foregoing embodiments, refer to Figure 12 which shows a composition structure schematic diagram of a decoder 110 provided by the embodiment of the application. As shown in Figure 13 , the decoder 110 can include an analysis unit 1101, a first determination unit 1102, a first construction unit 1103 and a first prediction unit 1104; wherein
[0357] The analysis unit 1101 is configured to analyze a bitstream and acquire a prediction mode parameter of a current block;
[0358] The first determining unit 1102 is configured to determine at least one candidate position of the current block when the prediction mode parameter indicates that a preset inter prediction mode is used to determine an inter prediction value of the current block, wherein the candidate position at least includes a lower right position inside the current block and a lower right position outside the current block.
[0359] The first determining unit 1102 is further configured to determine at least one time domain motion information of the current block based on the at least one candidate position.
[0360] The first constructing unit 1103 is configured to construct a new motion information candidate list based on the at least one time domain motion information.
[0361] The first predicting unit 1104 is configured to determine an inter prediction value of the current block according to the new motion information candidate list.
[0362] In some embodiments, the motion information includes motion vector information and reference frame information.
[0363] In some embodiments, the first determining unit 1102 is further configured to obtain a first lower right candidate position, a second lower right candidate position, a third lower right candidate position and a fourth lower right candidate position to form a candidate position set, and determine at least one candidate position of the current block from the candidate position set, wherein the first lower right candidate position represents a lower right position inside the current block, and the second lower right candidate position, the third lower right candidate position and the fourth lower right candidate position represent lower right positions outside the current block.
[0364] In some embodiments, the first determining unit 1102 is further configured to obtain coordinate information corresponding to a top left pixel position of the current block, width information of the current block and height information of the current block, perform coordinate calculation by using the coordinate information corresponding to the top left pixel position, the width information and the height information respectively to obtain first coordinate information, second coordinate information, third coordinate information and fourth coordinate information, and determine a position corresponding to the first coordinate information as the first lower right candidate position, a position corresponding to the second coordinate information as the second lower right candidate position, a position corresponding to the third coordinate information as the third lower right candidate position and a position corresponding to the fourth coordinate information as the fourth lower right candidate position.
[0365] In some embodiments, the first determining unit 1102 is further configured to correct the second coordinate information, the third coordinate information and the fourth coordinate information by using a first preset offset, to obtain first corrected second coordinate information, first corrected third coordinate information and first corrected fourth coordinate information; determine a position corresponding to the first coordinate information as the first right-down candidate position, determine a position corresponding to the first corrected second coordinate information as the second right-down candidate position, determine a position corresponding to the first corrected third coordinate information as the third right-down candidate position, and determine a position corresponding to the first corrected fourth coordinate information as the fourth right-down candidate position.
[0366] In some embodiments, the first determining unit 1102 is further configured to determine a second preset offset when there is a motion vector from an image block in a current frame to a candidate reference frame, wherein the candidate reference frame is a reference frame of reference motion information used for determining temporal motion information, and the image block includes at least a neighboring block which is spatially adjacent to the current block in the current frame; correct the first coordinate information, the second coordinate information, the third coordinate information and the fourth coordinate information by using the second preset offset, to obtain second corrected first coordinate information, second corrected second coordinate information, second corrected third coordinate information and second corrected fourth coordinate information; determine a position corresponding to the second corrected first coordinate information as the first right-down candidate position, determine a position corresponding to the second corrected second coordinate information as the second right-down candidate position, determine a position corresponding to the second corrected third coordinate information as the third right-down candidate position, and determine a position corresponding to the second corrected fourth coordinate information as the fourth right-down candidate position.
[0367] In some embodiments, the first determining unit 1102 is further configured to obtain a motion vector from a preset neighboring block in the current frame to the candidate reference frame, and determine the obtained motion vector as the second preset offset; or scale motion information of the preset neighboring block in the current frame to the candidate reference frame to obtain a scaled motion vector, and determine the scaled motion vector as the second preset offset.
[0368] In some embodiments, referring to Figure 13The decoder 110 can further include a first selection unit 1105 configured to, in a case where the number of the at least one temporal motion information is one, select one candidate position from the candidate position set according to a preset manner, and determine the selected candidate position as one candidate position of the current block; or select a candidate position corresponding to a high priority and available from the candidate position set according to a preset priority order, and determine the selected candidate position as one candidate position of the current block.
[0369] In some embodiments, the first selection unit 1105 is further configured to, in a case where the number of the at least one temporal motion information is a plurality, select a plurality of candidate positions from the candidate position set according to a preset combination manner, and determine the selected plurality of candidate positions as a plurality of candidate positions of the current block; or select a plurality of candidate positions from the candidate position set according to a preset priority order and available, and determine the selected plurality of candidate positions as a plurality of candidate positions of the current block.
[0370] In some embodiments, the first selection unit 1105 is further configured to, when selecting a candidate position from the candidate position set, if the candidate position to be selected belongs to a lower right position outside the current block and all lower right positions outside the current block are unavailable, determine the first lower right candidate position as the candidate position to be selected.
[0371] In some embodiments, referring to Figure 13 The decoder 110 can further include a first adjustment unit 1106 configured to adjust a proportion value of the temporal motion information in the new motion information candidate list; and control the filling positions reserved for at least two temporal motion information in the new motion information candidate list according to the adjusted proportion value.
[0372] In some embodiments, the first determination unit 1102 is further configured to determine reference frame information corresponding to each candidate position in the at least one candidate position; for each candidate position, determine a temporal position associated with the candidate position in the corresponding reference frame information, and determine motion information used by the temporal position as temporal motion information corresponding to the candidate position; and obtain at least one temporal motion information based on the at least one candidate position.
[0373] In some embodiments, the preset inter prediction mode includes GPM or AWP.
[0374] The first determination unit 1102 is further configured to, when the prediction mode parameter indicates that GPM or AWP is used to determine the inter prediction value of the current block, determine two partitions of the current block; and the two partitions include a first partition and a second partition.
[0375] In some embodiments, the first determining unit 1102 is further configured to group the plurality of partition modes under the GPM or the AWP to obtain at least two sets of partition mode groups; determine at least one candidate position corresponding to each of the at least two sets of partition mode groups respectively; wherein different sets of partition mode groups correspond to different at least one candidate position; and for each partition mode in each set of partition mode groups, perform the step of determining the at least one temporal motion information of the current block based on the corresponding determined at least one candidate position.
[0376] In some embodiments, the first determining unit 1102 is further configured to include a first set of partition mode groups and a second set of partition mode groups in the at least two sets of partition mode groups, if the current partition mode belongs to the first set of partition mode groups, determine the top-left pixel position inside the current block as the at least one candidate position of the current block, and perform the step of determining the at least one temporal motion information of the current block based on the at least one candidate position; and if the current partition mode belongs to the second set of partition mode groups, determine the bottom-right pixel position outside the current block as the at least one candidate position of the current block, and perform the step of determining the at least one temporal motion information of the current block based on the at least one candidate position.
[0377] In some embodiments, the first determining unit 1102 is further configured to determine a first number corresponding to a first spatial pixel position; wherein the first spatial pixel position is adjacent to at least one boundary space of the first partition; determine a second number corresponding to a second spatial pixel position; wherein the second spatial pixel position is adjacent to at least one boundary space of the second partition; and if the first number or the second number is less than a preset value, determine the bottom-right pixel position outside the current block as the at least one candidate position of the current block, and perform the step of determining the at least one temporal motion information of the current block based on the at least one candidate position; and if the first number and the second number are both greater than the preset value, determine the top-left pixel position of the current block as the at least one candidate position of the current block, and perform the step of determining the at least one temporal motion information of the current block based on the at least one candidate position.
[0378] In some embodiments, the parsing unit 1101 is further configured to parse a bitstream to determine a first motion information index value corresponding to the first partition and a second motion information index value corresponding to the second partition.
[0379] The first determining unit 1102 is further configured to determine the motion information in the new motion information candidate list indicated by the first motion information index value as the motion information of the first partition and determine the motion information in the new motion information candidate list indicated by the second motion information index value as the motion information of the second partition based on the new motion information candidate list.
[0380] The first predicting unit 1104 is further configured to calculate a first prediction value of the first partition by using the motion information of the first partition, calculate a second prediction value of the second partition by using the motion information of the second partition, and perform weighted fusion on the first prediction value and the second prediction value to obtain an inter-frame prediction value of the current block.
[0381] It can be understood that, in the embodiments of the present application, the "unit" can be a part of circuit, a part of processor, a part of program or software, etc., and of course can be a module, and can also be non-modular. Moreover, the components in the embodiments can be integrated in one processing unit, or can be physically present individually, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software function module.
[0382] When the integrated unit is realized in the form of a software function module and is not sold or used as an independent product, it can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the embodiments can be embodied in the form of a software product, and the computer software product is stored in a storage medium, and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to perform all or part of the steps of the methods described in the embodiments. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0383] Therefore, the embodiments of the present application provide a computer storage medium applied to the decoder 110, and the computer storage medium stores an inter-frame prediction program. The inter-frame prediction program is executed by a first processor to implement the method described in the decoder side in the foregoing embodiments.
[0384] Based on the components of the decoder 110 and the computer storage medium, it can be seen that Figure 13The first bus system 1204 is used to realize the connection communication between the components. The first bus system 1204 includes a data bus, a power supply bus, a control bus and a state signal bus. However, in order to clearly illustrate, all the buses are marked as the first bus system 1204 in the figure. Figure 13 The first bus system 1204 includes a data bus, a power supply bus, a control bus and a state signal bus. However, in order to clearly illustrate, all the buses are marked as the first bus system 1204 in the figure.
[0385] The first communication interface 1201 is used for receiving and sending signals in the process of transmitting information with other external network elements.
[0386] The first memory 1202 is used for storing a computer program capable of running on the first processor 1203.
[0387] The first processor 1203 is used for executing the following steps when running the computer program:
[0388] Parsing a code stream to obtain a prediction mode parameter of a current block;
[0389] When the prediction mode parameter indicates that a preset inter prediction mode is used to determine an inter prediction value of the current block, determining at least one candidate position of the current block; wherein the candidate position at least includes a right bottom position inside the current block and a right bottom position outside the current block;
[0390] Determining at least one time domain motion information of the current block based on the at least one candidate position;
[0391] Constructing a new motion information candidate list based on the at least one time domain motion information;
[0392] Determining the inter prediction value of the current block according to the new motion information candidate list.
[0393] It is to be appreciated that the first memory 1202 in the embodiments of this application can be volatile memory or nonvolatile memory, or can include both volatile and nonvolatile memory. In one embodiment, a nonvolatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically EPROM (EEPROM), or a flash memory. A volatile memory can be a Random Access Memory (RAM), which is used as the external cache. By way of example and not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DRRAM). The first memory 1202 of the system and method described herein are intended to include, without being limited to, these and any other suitable types of memory.
[0394] The first processor 1203 can be an integrated circuit chip having a processing capability of signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of hardware or the instruction in the form of software in the first processor 1203. The first processor 1203 described above can be a general processor, a digital signal processor (Digital Signal Processor, DSP), an application specific integrated circuit (Application Specific Integrated Circuit, ASIC), a ready programmable gate array (Field Programmable Gate Array, FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component. Each method, step and logic block disclosed in the embodiment of the present application can be implemented or executed. The general processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in combination with the embodiment of the present application can be directly embodied as a hardware code processor to execute, or be executed by a combination of hardware and software modules in the code processor. The software module can be located in a random access memory, a flash memory, a read only memory, a programmable read only memory or an electrically erasable programmable memory, a register or other mature storage medium in the art. The storage medium is located in the first memory 1202, and the first processor 1203 reads the information in the first memory 1202 and combines the hardware to complete the steps of the above method.
[0395] It can be understood that the embodiments described in the present application can be realized by hardware, software, firmware, middleware, microcode or combination thereof. For hardware implementation, the processing unit can be realized in one or more application specific integrated circuits (Application Specific Integrated Circuits, ASIC), digital signal processors (Digital Signal Processing, DSP), digital signal processing devices (DSP Device, DSPD), programmable logic devices (Programmable Logic Device, PLD), field programmable gate arrays (Field-Programmable Gate Array, FPGA), general processors, controllers, microcontrollers, microprocessors, other electronic units for executing functions described in the present application or combination thereof. For software implementation, the technology described in the present application can be realized by modules (such as processes, functions, etc.) for executing functions described in the present application. The software code can be stored in the memory and executed by the processor. The memory can be implemented in the processor or outside the processor.
[0396] Optionally, as another embodiment, the first processor 1203 is further configured to execute the method of any one of the preceding embodiments when running the computer program.
[0397] The embodiment provides a decoder which can include a parsing unit, a first determining unit, a first constructing unit and a first predicting unit. In the decoder, motion information with higher correlation with the lower right can be supplemented in a motion information candidate list, so that the diversity of motion information in the motion information candidate list is increased; especially for GPM or AWP inter prediction mode, the correlation of the lower right can be improved by increasing the position of the lower right temporal motion information candidate, so that the coding performance can be improved.
[0398] Based on the same inventive concept as the preceding embodiments, see Figure 14 which shows a component structure schematic diagram of an encoder 130 provided by the embodiment of the application. As shown in Figure 14 , the encoder 130 can include a second determining unit 1301, a second constructing unit 1302 and a second predicting unit 1303; wherein,
[0399] The second determining unit 1301 is configured to determine a prediction mode parameter of a current block, and when the prediction mode parameter indicates that a preset inter prediction mode is used to determine an inter prediction value of the current block, determine at least one candidate position of the current block; wherein the candidate position at least includes a lower right position inside the current block and a lower right position outside the current block;
[0400] The second determining unit 1301 is further configured to determine at least one temporal motion information of the current block based on the at least one candidate position.
[0401] The second constructing unit 1302 is configured to construct a new motion information candidate list based on the at least one temporal motion information.
[0402] The second predicting unit 1303 is configured to determine the inter prediction value of the current block according to the new motion information candidate list.
[0403] In some embodiments, the motion information includes motion vector information and reference frame information.
[0404] In some embodiments, see , the encoder 130 can further include a pre-encoding unit 1304 and a second selecting unit 1305; wherein,
[0405] The pre-encoding unit 1304 is configured to perform pre-encoding processing on the current block by using a plurality of prediction modes, to obtain a rate-distortion cost value corresponding to each prediction mode;
[0406] The second selection unit 1305 is configured to select a minimum rate-distortion value from the obtained multiple rate-distortion values, and determine a prediction mode corresponding to the minimum rate-distortion value as the prediction mode parameter of the current block.
[0407] In some embodiments, the second determination unit 1301 is further configured to obtain a first bottom-right candidate position, a second bottom-right candidate position, a third bottom-right candidate position, and a fourth bottom-right candidate position, to form a candidate position set; and determine at least one candidate position of the current block from the candidate position set; wherein the first bottom-right candidate position represents a bottom-right position inside the current block, and the second bottom-right candidate position, the third bottom-right candidate position, and the fourth bottom-right candidate position represent bottom-right positions outside the current block.
[0408] In some embodiments, the second determination unit 1301 is further configured to obtain coordinate information corresponding to a top-left pixel position of the current block, width information of the current block, and height information of the current block; perform coordinate calculation on the coordinate information corresponding to the top-left pixel position, the width information, and the height information respectively to obtain first coordinate information, second coordinate information, third coordinate information, and fourth coordinate information; determine a position corresponding to the first coordinate information as the first bottom-right candidate position, a position corresponding to the second coordinate information as the second bottom-right candidate position, a position corresponding to the third coordinate information as the third bottom-right candidate position, and a position corresponding to the fourth coordinate information as the fourth bottom-right candidate position.
[0409] In some embodiments, the second determination unit 1301 is further configured to modify the second coordinate information, the third coordinate information, and the fourth coordinate information by using a first preset offset to obtain first modified second coordinate information, first modified third coordinate information, and first modified fourth coordinate information; determine a position corresponding to the first coordinate information as the first bottom-right candidate position, a position corresponding to the first modified second coordinate information as the second bottom-right candidate position, a position corresponding to the first modified third coordinate information as the third bottom-right candidate position, and a position corresponding to the first modified fourth coordinate information as the fourth bottom-right candidate position.
[0410] In some embodiments, the second determining unit 1301 is further configured to determine a second preset offset when there is a motion vector from the image block in the current frame to the candidate reference frame; wherein the candidate reference frame is a reference frame of the reference motion information used for determining the temporal motion information, and the image block at least includes a neighboring block, and the neighboring block is spatially adjacent to the current block in the current frame; correct the first coordinate information, the second coordinate information, the third coordinate information, and the fourth coordinate information by using the second preset offset to obtain second corrected first coordinate information, second corrected second coordinate information, second corrected third coordinate information, and second corrected fourth coordinate information; determine the position corresponding to the second corrected first coordinate information as the first bottom-right candidate position, determine the position corresponding to the second corrected second coordinate information as the second bottom-right candidate position, determine the position corresponding to the second corrected third coordinate information as the third bottom-right candidate position, and determine the position corresponding to the second corrected fourth coordinate information as the fourth bottom-right candidate position.
[0411] In some embodiments, the second determining unit 1301 is further configured to obtain a motion vector from a preset neighboring block in the current frame to the candidate reference frame, and determine the obtained motion vector as the second preset offset; or scale the motion information of the preset neighboring block in the current frame to the candidate reference frame to obtain a scaled motion vector, and determine the scaled motion vector as the second preset offset.
[0412] In some embodiments, the second selecting unit 1305 is further configured to, when the number of the at least one temporal motion information is one, select one candidate position from the candidate position set in a preset manner, and determine the selected candidate position as one candidate position of the current block; or select a candidate position corresponding to a high priority from the candidate position set in a preset priority order and when the selected candidate position is available, determine the selected candidate position as one candidate position of the current block.
[0413] In some embodiments, the second selecting unit 1305 is further configured to select a plurality of candidate positions from the candidate position set in a preset combination manner, and determine the selected plurality of candidate positions as a plurality of candidate positions of the current block; or select a plurality of candidate positions from the candidate position set in a preset priority order and when the selected candidate positions are available, determine the selected plurality of candidate positions as a plurality of candidate positions of the current block.
[0414] In some embodiments, the second selecting unit 1305 is further configured to, when selecting a candidate position from the candidate position set, if the candidate position to be selected belongs to a lower right position outside the current block and all lower right positions outside the current block are unavailable, determining the first lower right candidate position as the candidate position to be selected.
[0415] In some embodiments, referring to The encoder 130 can further include a second adjusting unit 1306 configured to adjust a proportion value of the temporal motion information in the new motion information candidate list; and control filling positions reserved for at least two temporal motion information in the new motion information candidate list according to the adjusted proportion value.
[0416] In some embodiments, the second determining unit 1301 is further configured to determine reference frame information corresponding to each candidate position in the at least one candidate position; for each candidate position, determine a temporal position associated with the candidate position in the corresponding reference frame information, and determine motion information used by the temporal position as temporal motion information corresponding to the candidate position; and obtain at least one temporal motion information based on the at least one candidate position.
[0417] In some embodiments, the preset inter prediction mode includes GPM or AWP.
[0418] The second determining unit 1301 is further configured to, when the prediction mode parameter indicates that the inter prediction value of the current block is determined using GPM or AWP, determine two partitions of the current block; and the two partitions include a first partition and a second partition.
[0419] In some embodiments, the second determining unit 1301 is further configured to group a plurality of partition modes under GPM or AWP to obtain at least two sets of partition mode sets; determine at least one candidate position corresponding to each set of partition mode sets in the at least two sets of partition mode sets; different sets of partition mode sets correspond to different at least one candidate position; and for each partition mode in a set of partition mode, execute the step of determining at least one temporal motion information of the current block based on the at least one candidate position corresponding to the set of partition mode according to the at least one candidate position corresponding to the set of partition mode.
[0420] In some embodiments, the second determining unit 1301 is further configured to determine the top-left pixel position inside the current block as at least one candidate position of the current block if the partition mode to be used belongs to the first set of partition modes, and perform the step of determining at least one temporal motion information of the current block based on the at least one candidate position; and determine the bottom-right pixel position outside the current block as at least one candidate position of the current block if the partition mode to be used belongs to the second set of partition modes, and perform the step of determining at least one temporal motion information of the current block based on the at least one candidate position.
[0421] In some embodiments, the second determining unit 1301 is further configured to determine a first number corresponding to a first spatial pixel position adjacent to at least one boundary space of the first partition, determine a second number corresponding to a second spatial pixel position adjacent to at least one boundary space of the second partition, and determine the bottom-right pixel position outside the current block as at least one candidate position of the current block if the first number or the second number is less than a preset value, and perform the step of determining at least one temporal motion information of the current block based on the at least one candidate position; and determine the top-left pixel position of the current block as at least one candidate position of the current block if the first number and the second number are both greater than the preset value, and perform the step of determining at least one temporal motion information of the current block based on the at least one candidate position.
[0422] In some embodiments, the second determining unit 1301 is further configured to determine the motion information of the first partition and the motion information of the second partition based on the new motion information candidate list, set a first motion information index value as an index sequence value of the motion information of the first partition in the new motion information candidate list, and set a second motion information index value as an index sequence value of the motion information of the second partition in the new motion information candidate list.
[0423] The second prediction unit 1303 is further configured to calculate a first prediction value of the first partition by using the motion information of the first partition, calculate a second prediction value of the second partition by using the motion information of the second partition, and perform weighted fusion on the first prediction value and the second prediction value to obtain an inter-frame prediction value of the current block.
[0424] In some embodiments, referring to The encoder 130 can further include a writing unit 1307 configured to write the first motion information index value and the second motion information index value into a bitstream.
[0425] It can be understood that, in this embodiment, the "unit" can be a partial circuit, a partial processor, a partial program or software, and the like, and can also be a module, and can also be non-modular. Moreover, the components in this embodiment can be integrated in a processing unit, or can be physically present individually, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software function module.
[0426] When the integrated unit is realized in the form of a software function module and is not sold or used as an independent product, it can be stored in a computer readable storage medium. Based on this understanding, the embodiment provides a computer storage medium applied to the encoder 130, and the computer storage medium stores an inter prediction program. When the inter prediction program is executed by the second processor, the method of the encoder side in the foregoing embodiment is implemented.
[0427] Based on the components of the encoder 130 and the computer storage medium, refer to which shows a specific hardware structure example of the encoder 130 provided by the embodiment of the application, and can include a second communication interface 1401, a second memory 1402 and a second processor 1403; and various components are coupled together through a second bus system 1404. It can be understood that the second bus system 1404 is used to realize the connection and communication between the components. The second bus system 1404 includes a data bus, a power supply bus, a control bus and a state signal bus. However, for the purpose of clear illustration, all kinds of buses are marked as the second bus system 1404 in . Among them,
[0428] The second communication interface 1401 is used for receiving and sending signals in the process of transceiving information with other external network elements;
[0429] The second memory 1402 is used for storing a computer program capable of running on the second processor 1403;
[0430] The second processor 1403 is used for, when running the computer program, performing:
[0431] determining a prediction mode parameter of a current block;
[0432] determining at least one candidate position of the current block when the prediction mode parameter indicates that the inter prediction value of the current block is determined using a preset inter prediction mode; wherein the candidate position at least includes a right bottom position inside the current block and a right bottom position outside the current block;
[0433] determining at least one temporal motion information of the current block based on the at least one candidate position;
[0434] constructing a new motion information candidate list based on the at least one temporal motion information;
[0435] determining the inter prediction value of the current block according to the new motion information candidate list.
[0436] Optionally, as another embodiment, the second processor 1403 is further configured to execute the method in any one of the foregoing embodiments when the computer program is run.
[0437] It can be understood that the hardware function of the second memory 1402 is similar to that of the first memory 1202, and the hardware function of the second processor 1403 is similar to that of the first processor 1203; and details are not described herein.
[0438] The embodiment provides an encoder, which can include a second determining unit, a second constructing unit and a second predicting unit. In the encoder, motion information with higher correlation to the right bottom can be supplemented in the motion information candidate list, thereby increasing the diversity of the motion information in the motion information candidate list; and especially for the GPM or AWP inter prediction mode, the correlation to the right bottom can be improved by increasing the right bottom temporal motion information candidate position, thereby improving the coding performance.
[0439] It should be noted that in the present application, the terms "comprising", "containing" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of another identical element in the process, method, article or device including the element.
[0440] The above-mentioned sequence numbers of the embodiments of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments.
[0441] The methods disclosed in the several method embodiments provided by the present application can be combined arbitrarily without conflict to obtain new method embodiments.
[0442] The features disclosed by the several product embodiments of the present application can be arbitrarily combined, without conflict, to obtain new product embodiments.
[0443] The features disclosed by the several method or device embodiments of the present application can be arbitrarily combined, without conflict, to obtain new method embodiments or device embodiments.
[0444] The above description is merely illustrative of the application, and the scope of the application is not limited thereto. Any variations and modifications of the application, which fall within the technical scope of the application, should be covered by the scope of the application. Therefore, the scope of the application should be determined by the scope of the claims.
Claims
1. An inter prediction method, characterized by, The method is applied to a decoder, and the method comprises: parsing a code stream to obtain a prediction mode parameter of a current block; when the prediction mode parameter indicates that a preset inter-frame prediction mode is used to determine an inter-frame prediction value of the current block, determining at least one candidate position of the current block; wherein the candidate positions at least include a lower right position inside the current block and a lower right position outside the current block, and the preset inter-frame prediction mode includes an angular weighted prediction mode; based on the at least one candidate position, determining at least one time domain motion information of the current block; based on the at least one time domain motion information, constructing a motion information candidate list; determining an inter-frame prediction value of the current block according to the motion information candidate list, wherein the determining of the at least one candidate position of the current block comprises: obtaining a first lower right candidate position, a second lower right candidate position, a third lower right candidate position and a fourth lower right candidate position to form a candidate position set; determining at least one candidate position of the current block from the candidate position set; wherein the first lower right candidate position represents a lower right position inside the current block, and the second lower right candidate position, the third lower right candidate position and the fourth lower right candidate position represent lower right positions outside the current block; wherein the determining of the at least one time domain motion information of the current block based on the at least one candidate position comprises: determining reference frame information corresponding to the at least one candidate position; determining a time domain position associated with the candidate position in the corresponding reference frame information, and determining motion information used by the time domain position as time domain motion information corresponding to the candidate position, wherein the candidate position in the current frame is the same as the time domain position in the reference frame; correspondingly obtaining at least one time domain motion information based on the at least one candidate position.
2. The method of claim 1, wherein, The motion information includes motion vector information and reference frame information.
3. The method of claim 1, wherein, The obtaining of the first lower right candidate position, the second lower right candidate position, the third lower right candidate position and the fourth lower right candidate position comprises: obtaining coordinate information corresponding to a top left pixel position of the current block, width information of the current block and height information of the current block; respectively performing coordinate calculation on the coordinate information corresponding to the top left pixel position, the width information and the height information to obtain first coordinate information, second coordinate information, third coordinate information and fourth coordinate information; determining a position corresponding to the first coordinate information as the first lower right candidate position, a position corresponding to the second coordinate information as the second lower right candidate position, a position corresponding to the third coordinate information as the third lower right candidate position, and a position corresponding to the fourth coordinate information as the fourth lower right candidate position.
4. The method of claim 1, wherein, In the case that the number of the at least one time domain motion information is one, the determining of the at least one candidate position of the current block from the candidate position set comprises: selecting one candidate position from the candidate position set in a preset manner, and determining the selected candidate position as one candidate position of the current block.
5. The method of claim 1, wherein, In a case where the number of the at least one temporal motion information is one, the determining the at least one candidate position of the current block from the candidate position set comprises: selecting a candidate position corresponding to a high priority according to a preset priority order from the candidate position set, and determining the selected candidate position as one candidate position of the current block.
6. The method of claim 1, wherein, In a case where the number of the at least one temporal motion information is multiple, the determining the at least one candidate position of the current block from the candidate position set comprises: selecting multiple candidate positions according to a preset combination manner from the candidate position set, and determining the selected multiple candidate positions as multiple candidate positions of the current block; or selecting multiple candidate positions according to a preset priority order from the candidate position set, and determining the selected multiple candidate positions as multiple candidate positions of the current block.
7. The method of claim 1, wherein, The method further comprises: when selecting a candidate position from the candidate position set, if the candidate position to be selected belongs to a lower-right position outside the current block and all lower-right positions outside the current block are unavailable, determining the first lower-right candidate position as the candidate position to be selected.
8. The method of claim 1, wherein, The method further comprises: adjusting a proportion value of temporal motion information in the motion information candidate list; controlling a filling position reserved for at least two temporal motion information in the motion information candidate list according to the adjusted proportion value.
9. The method of claim 1, wherein, The method further comprises: when the prediction mode parameter indicates that the preset inter prediction mode is used to determine an inter prediction value of the current block, determining two partitions of the current block; wherein the two partitions comprise a first partition and a second partition.
10. The method of claim 9, wherein, The method further comprises: grouping multiple partition modes in the preset inter prediction mode to obtain at least two sets of partition mode sets; determining at least one candidate position corresponding to each set of partition mode sets in the at least two sets of partition mode sets; wherein different sets of partition mode sets correspond to different at least one candidate position; for each partition mode in each set of partition mode sets, performing the step of determining the at least one temporal motion information of the current block based on the at least one corresponding determined candidate position.
11. The method of claim 10, wherein, The at least two sets of partition mode sets comprise a first set of partition mode sets and a second set of partition mode sets, and the method further comprises: if a current partition mode belongs to the first set of partition mode sets, determining an upper-left pixel position inside the current block as the at least one candidate position of the current block, and performing the step of determining the at least one temporal motion information of the current block based on the at least one candidate position; if the current partition mode belongs to the second set of partition mode sets, determining a lower-right pixel position outside the current block as the at least one candidate position of the current block, and performing the step of determining the at least one temporal motion information of the current block based on the at least one candidate position.
12. The method of claim 9, wherein, The method further comprises: determining a first number corresponding to a first spatial pixel position adjacent to at least one boundary space of the first partition; determining a second number corresponding to a second spatial pixel position adjacent to at least one boundary space of the second partition; if the first number or the second number is less than a preset value, determining a bottom-right pixel position outside the current block as at least one candidate position of the current block, and performing the step of determining at least one temporal motion information of the current block based on the at least one candidate position; if the first number and the second number are both greater than the preset value, determining a top-left pixel position inside the current block as at least one candidate position of the current block, and performing the step of determining at least one temporal motion information of the current block based on the at least one candidate position.
13. The method of claim 9, wherein, The determining the inter prediction value of the current block according to the motion information candidate list comprises: determining a first motion information index value corresponding to the first partition and a second motion information index value corresponding to the second partition; determining motion information of the first partition according to motion information in the motion information candidate list indicated by the first motion information index value, and determining motion information of the second partition according to motion information in the motion information candidate list indicated by the second motion information index value based on the motion information candidate list; calculating a first prediction value of the first partition by using the motion information of the first partition, and calculating a second prediction value of the second partition by using the motion information of the second partition; performing weighted fusion on the first prediction value and the second prediction value to obtain the inter prediction value of the current block.
14. An inter prediction method, characterized by, The method applied to an encoder comprises: determining a prediction mode parameter of a current block; when the prediction mode parameter indicates that a preset inter prediction mode is used to determine an inter prediction value of the current block, determining at least one candidate position of the current block; wherein the candidate position at least comprises a bottom-right position inside the current block and a bottom-right position outside the current block, and the preset inter prediction mode comprises an angular weighted prediction mode; determining at least one temporal motion information of the current block based on the at least one candidate position; constructing a motion information candidate list based on the at least one temporal motion information; determining an inter prediction value of the current block according to the motion information candidate list, wherein the determining at least one candidate position of the current block comprises: obtaining a first bottom-right candidate position, a second bottom-right candidate position, a third bottom-right candidate position and a fourth bottom-right candidate position to form a candidate position set; determining at least one candidate position of the current block from the candidate position set; wherein the first bottom-right candidate position represents a bottom-right position inside the current block, and the second bottom-right candidate position, the third bottom-right candidate position and the fourth bottom-right candidate position represent bottom-right positions outside the current block; wherein the determining at least one temporal motion information of the current block based on the at least one candidate position comprises: Determine the reference frame information corresponding to the at least one candidate position; In the corresponding reference frame information, the temporal position associated with the candidate position is determined, and the motion information used by the temporal position is determined as the temporal motion information corresponding to the candidate position, wherein the candidate position in the current frame is the same as the temporal position in the reference frame; Based on the at least one candidate position, at least one temporal motion information is obtained.
15. The method of claim 14, wherein, The motion information includes motion vector information and reference frame information.
16. The method of claim 14, wherein, The determination of the prediction mode parameters for the current block includes: The current block is pre-encoded using multiple prediction modes to obtain the rate-distortion cost corresponding to each prediction mode. Select the minimum rate distortion value from the multiple obtained rate distortion values, and determine the prediction mode corresponding to the minimum rate distortion value as the prediction mode parameter for the current block.
17. The method of claim 14, wherein, The process of obtaining the first, second, third, and fourth lower-right candidate positions includes: Obtain the coordinate information corresponding to the top left pixel position of the current block, the width information of the current block, and the height information of the current block; Coordinates are calculated using the coordinate information corresponding to the top left pixel position, the width information, and the height information to obtain first coordinate information, second coordinate information, third coordinate information, and fourth coordinate information; The position corresponding to the first coordinate information is determined as the first lower right candidate position, the position corresponding to the second coordinate information is determined as the second lower right candidate position, the position corresponding to the third coordinate information is determined as the third lower right candidate position, and the position corresponding to the fourth coordinate information is determined as the fourth lower right candidate position.
18. The method of claim 14, wherein, When the number of at least one temporal motion information is one, determining at least one candidate position of the current block from the candidate position set includes: A candidate position is selected from the candidate position set according to a preset method, and the selected candidate position is determined as a candidate position of the current block.
19. The method of claim 14, wherein, When the number of at least one temporal motion information is one, determining at least one candidate position of the current block from the candidate position set includes: From the set of candidate positions, select the candidate positions corresponding to higher priorities according to a preset priority order. If the selected candidate positions are available, the selected candidate positions are determined as a candidate position of the current block.
20. The method of claim 14, wherein, When the number of at least one temporal motion information items is multiple, determining at least one candidate position of the current block from the candidate position set includes: Multiple candidate positions are selected from the candidate position set according to a preset combination method, and the selected multiple candidate positions are determined as multiple candidate positions of the current block; or, From the set of candidate positions, multiple candidate positions are selected in a preset priority order, and the selected candidate positions are available. The selected multiple candidate positions are then determined as multiple candidate positions of the current block.
21. The method of claim 14, wherein, The method further includes: When selecting a candidate position from the candidate position set, if the candidate position to be selected is located in the lower right position outside the current block and all lower right positions outside the current block are unavailable, then the first lower right candidate position is determined as the candidate position to be selected.
22. The method of claim 14, wherein, The method further includes: Adjust the proportion of temporal motion information in the motion information candidate list; Based on the adjusted ratio, at least two positions for filling time-domain motion information are reserved in the motion information candidate list.
23. The method of claim 14, wherein, The method further includes: When the prediction mode parameter indicates that the preset inter-frame prediction mode is used to determine the inter-frame prediction value of the current block, two partitions of the current block are determined; wherein, the two partitions include a first partition and a second partition.
24. The method of claim 23, wherein, The method further includes: The various partitioning modes under the preset inter-frame prediction mode are grouped to obtain at least two sets of partitioning modes. Determine at least one candidate position for each of the at least two sets of partitioning patterns; wherein different sets of partitioning patterns correspond to different at least one candidate position. For each partitioning pattern in a set of partitioning patterns, based on at least one candidate position, the step of determining at least one temporal motion information of the current block based on the at least one candidate position is executed.
25. The method of claim 24, wherein, The at least two sets of partitioning patterns include a first set of partitioning patterns and a second set of partitioning patterns; the method further includes: If the partitioning pattern to be used belongs to the first set of partitioning patterns, then the top left pixel position inside the current block is determined as at least one candidate position of the current block, and the step of determining at least one temporal motion information of the current block based on the at least one candidate position is executed. If the partitioning pattern to be used belongs to the second set of partitioning patterns, then the lower right pixel position outside the current block is determined as at least one candidate position of the current block, and the step of determining at least one temporal motion information of the current block based on the at least one candidate position is executed.
26. The method of claim 23, wherein, The method further includes: Determine a first number corresponding to the first spatial domain pixel position; wherein the first spatial domain pixel position is adjacent to at least one boundary space of the first partition; Determine a second quantity corresponding to the second spatial domain pixel position; wherein the second spatial domain pixel position is adjacent to at least one boundary space of the second partition; If the first quantity or the second quantity is less than a preset value, then the lower right pixel position outside the current block is determined as at least one candidate position of the current block, and the step of determining at least one temporal motion information of the current block based on the at least one candidate position is executed. If both the first quantity and the second quantity are greater than a preset value, then the top left pixel position of the current block is determined as at least one candidate position of the current block, and the step of determining at least one temporal motion information of the current block based on the at least one candidate position is executed.
27. The method of claim 23, wherein, Determining the inter-frame prediction value of the current block based on the motion information candidate list includes: Based on the motion information candidate list, the motion information of the first partition and the motion information of the second partition are determined, and a first motion information index value is set according to the index number value of the motion information of the first partition in the motion information candidate list, and a second motion information index value is set according to the index number value of the motion information of the second partition in the motion information candidate list. Calculate a first predicted value for the first partition using the motion information of the first partition, and calculate a second predicted value for the second partition using the motion information of the second partition. The first predicted value and the second predicted value are weighted and fused to obtain the inter-frame predicted value of the current block.
28. The method of claim 27, wherein, The method further includes: Write the first motion information index value and the second motion information index value into the bitstream.
29. A decoder, comprising: The decoder includes a parsing unit, a first determining unit, a first constructing unit, and a first prediction unit; wherein, The parsing unit is configured to parse the code stream and obtain the prediction mode parameters of the current block; The first determining unit is configured to determine at least one candidate position of the current block when the prediction mode parameter indicates that the inter-frame prediction value of the current block is determined using a preset inter-frame prediction mode; wherein the candidate position includes at least the lower right position inside the current block and the lower right position outside the current block, and the preset inter-frame prediction mode includes an angle-weighted prediction mode. The first determining unit is further configured to determine at least one temporal motion information of the current block based on the at least one candidate position; The first construction unit is configured to construct a motion information candidate list based on the at least one temporal motion information; The first prediction unit is configured to determine the inter-frame prediction value of the current block based on the motion information candidate list. The first determining unit is configured as follows: Obtain the first, second, third, and fourth bottom-right candidate positions, and form a candidate position set; From the set of candidate locations, determine at least one candidate location for the current block; Wherein, the first lower right candidate position represents the lower right position inside the current block, and the second lower right candidate position, the third lower right candidate position, and the fourth lower right candidate position represent the lower right position outside the current block; The first determining unit is configured as follows: Determine the reference frame information corresponding to the at least one candidate position; In the corresponding reference frame information, the temporal position associated with the candidate position is determined, and the motion information used by the temporal position is determined as the temporal motion information corresponding to the candidate position, wherein the candidate position in the current frame is the same as the temporal position in the reference frame; Based on the at least one candidate position, at least one temporal motion information is obtained.
30. A decoder, comprising: The decoder includes a first memory and a first processor; wherein, The first memory is used to store computer programs that can run on the first processor; The first processor is configured to perform the method as described in any one of claims 1 to 13 when running the computer program.
31. An encoder comprising: The encoder includes a second determining unit, a second constructing unit, and a second prediction unit; wherein... The second determining unit is configured to determine the prediction mode parameters of the current block; and when the prediction mode parameters indicate that the inter-frame prediction value of the current block is determined using a preset inter-frame prediction mode, determine at least one candidate position of the current block; wherein the candidate position includes at least the lower right position inside the current block and the lower right position outside the current block, and the preset inter-frame prediction mode includes an angle-weighted prediction mode. The second determining unit is further configured to determine at least one temporal motion information of the current block based on the at least one candidate position; The second construction unit is configured to construct a motion information candidate list based on the at least one temporal motion information; The second prediction unit is configured to determine the inter-frame prediction value of the current block based on the motion information candidate list. The second determining unit is configured as follows: Obtain the first, second, third, and fourth bottom-right candidate positions, and form a candidate position set; From the set of candidate locations, determine at least one candidate location for the current block; Wherein, the first lower right candidate position represents the lower right position inside the current block, and the second lower right candidate position, the third lower right candidate position, and the fourth lower right candidate position represent the lower right position outside the current block; The second determining unit is configured as follows: Determine the reference frame information corresponding to the at least one candidate position; In the corresponding reference frame information, the temporal position associated with the candidate position is determined, and the motion information used by the temporal position is determined as the temporal motion information corresponding to the candidate position, wherein the candidate position in the current frame is the same as the temporal position in the reference frame; Based on the at least one candidate position, at least one temporal motion information is obtained.
32. An encoder comprising: The encoder includes a second memory and a second processor; wherein... The second memory is used to store computer programs that can run on the second processor; The second processor is configured to perform the method as described in any one of claims 14 to 28 when running the computer program.
33. A computer storage medium, comprising: The computer storage medium stores a computer program that, when executed by a first processor, implements the method as described in any one of claims 1 to 13, or when executed by a second processor, implements the method as described in any one of claims 14 to 28.
Citation Information
Patent Citations
Time domain motion vector acquisition method and device, inter-frame prediction method and device, and video coding method and device
CN110213590A
Image encoding / decoding method and apparatus
WO2020004979A1
Method for processing image on basis of inter-prediction mode and device therefor
WO2020004990A1