Interpretation method and apparatus
By dividing picture blocks into smaller units and performing bidirectional optical flow prediction, the method addresses the complexity and resource issues in inter prediction, improving efficiency and reducing hardware resource consumption.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2023-11-09
- Publication Date
- 2026-05-20
AI Technical Summary
Inter prediction in video encoding and decoding becomes complex and resource-intensive when picture blocks are large, leading to high computational demands.
The method involves dividing picture blocks into smaller first picture blocks based on preset dimensions, performing bidirectional optical flow prediction on these blocks, and combining their predictors to reduce complexity and improve processing efficiency.
This approach reduces hardware resource consumption and enhances processing efficiency by limiting the size of each picture block, thereby minimizing computational complexity and resource usage.
Smart Images

Figure 0007863078000043 
Figure 0007863078000044 
Figure 0007863078000045
Abstract
Description
Technical Field
[0001] This application relates to the field of video encoding and decoding, and more particularly, to an inter prediction method and apparatus.
Background Art
[0002] Inter prediction performs picture compression by using the correlation between video picture frames, i.e., temporal correlation, and is widely applied to compression encoding or decoding in scenarios such as general televisions, video conference televisions, videophones, and high-definition televisions. A picture is processed via inter prediction on both the encoder side and the decoder side.
[0003] When inter prediction is performed on a picture, the picture is first divided into a plurality of picture blocks based on the height and width of the picture blocks corresponding to the picture, and then inter prediction is performed on each picture block obtained through the division. When the width and height of the picture blocks corresponding to the picture are relatively large, the area of each picture block obtained through the division is relatively large. As a result, when inter prediction is performed on each picture block obtained through the division, the complexity of performing inter prediction is relatively high.
Summary of the Invention
Means for Solving the Problems
[0004] Embodiments of this application provide an inter prediction method and apparatus for reducing the complexity of performing inter prediction and improving processing efficiency.
[0005] According to the first aspect, the present application provides an inter prediction method. In the method, a plurality of first picture blocks are determined within a picture block to be processed based on a preset picture division width, a preset picture division height, and the width and height of the picture block to be processed. To obtain the predictor of each first picture block, bidirectional optical flow prediction is individually performed on the plurality of first picture blocks. A predictor of the picture block to be processed is obtained using a combination of predictors of the plurality of first picture blocks. The plurality of first picture blocks are determined within the picture block to be processed based on the preset picture division width, the preset picture division height, and the width and height of the picture block to be processed. Therefore, the size of the first picture block is limited by the preset picture division width and the preset picture division height, and the area of each determined first picture block is not too large so that hardware resources such as memory resources can be consumed less, thereby reducing the complexity of performing inter prediction and improving the processing efficiency.
[0006] In a possible implementation, the width and height of the picture block to be processed are respectively the same as the width and height of the first picture block, that is, the picture block to be processed includes only one first picture block. Correspondingly, when the picture block to be processed is determined as the first picture block based on the preset picture division width, the preset picture division height, and the width and height of the picture block to be processed, bidirectional optical flow prediction is performed on the picture block to be processed used as a processing unit to obtain the predictor of the picture block to be processed.
[0007] In possible implementations, a pre-configured picture division width is compared with the width of the picture block to be processed to determine the width of the first picture block. A pre-configured picture division height is compared with the height of the picture block to be processed to determine the height of the first picture block. Multiple first picture blocks are determined within the picture block to be processed based on the width and height of the first picture blocks. In this way, the width of the first picture block is limited by the pre-configured picture division width, and the height of the first picture block is limited by the pre-configured picture division height, so that the area of each determined first picture block is not too large, thereby reducing the complexity of performing interpretation and improving processing efficiency, which in turn reduces hardware resources such as memory resources can be consumed less.
[0008] In possible implementations, the width of the first picture block is the smaller of the pre-configured picture division width and the width of the picture block to be processed, and the height of the first picture block is the smaller of the pre-configured picture division height and the height of the picture block to be processed. Therefore, the first picture is determined block The area can be reduced, the complexity of performing interpretation can be minimized, and processing efficiency can be improved.
[0009] In possible implementations, the first prediction block of the first picture block is obtained based on the motion information of the picture block to be processed. A gradient operation is performed on the first prediction block to obtain the first gradient matrix of the first picture block. Based on the first prediction block and the first gradient matrix, the motion information refinement value of each basic processing unit within the first picture block is calculated. The predictor of the first picture block is obtained based on the motion information refinement value of each basic processing unit. Since the predictor of the first picture block is obtained based on the motion information refinement value of each basic processing unit, the predictor of the first picture block can become more accurate.
[0010] In possible implementations, the first expansion is performed on the width and height of the first prediction block based on the sample values of the block edge positions of the first prediction block, such that the width and height of the first prediction block obtained after the first expansion are each two samples larger than the width and height of the first picture block, and / or the width and height of the first gradient matrix obtained after the first expansion are each two samples larger than the width and height of the first picture block, such that the width and height of the first gradient matrix obtained after the first expansion are each two samples larger than the width and height of the first picture block, such that the width and height of the first gradient matrix obtained after the first expansion are each two samples larger than the width and height of the first gradient matrix. Correspondingly, the motion information refinement value of each basic processing unit in the first picture block is calculated based on the first prediction block obtained after the first expansion and / or the first gradient matrix obtained after the first expansion. The first expansion is performed on the width and height of the first prediction block so that the width and height of the first prediction block obtained after the first expansion are each two samples larger than the width and height of the first picture block. In this way, when bidirectional prediction is performed on picture blocks within a reference picture to obtain a first prediction block, the size of the obtained first prediction block may be reduced, and correspondingly, the size of the picture blocks may also be reduced to reduce the amount of data required for bidirectional prediction, thereby consuming fewer hardware resources.
[0011] In possible implementations, interpolation filtering is performed on sample values of the block edge regions of the first prediction block, or sample values of the block edge positions of the first prediction block are duplicated, in order to perform a second expansion on the width and height of the first prediction block. Correspondingly, a gradient operation is performed on the first prediction block obtained after the second expansion. Sample values of the block edge positions of the first prediction block are duplicated to perform a second expansion on the width and height of the first prediction block. Therefore, the implementation is simple and the complexity of operation is low.
[0012] In possible implementations, the first prediction block includes a forward prediction block and a backward prediction block, and the first gradient matrix includes a forward horizontal gradient matrix, a forward vertical gradient matrix, a backward horizontal gradient matrix, and a backward vertical gradient matrix.
[0013] In possible implementations, the pre-configured picture division width is 64, 32, or 16, and the pre-configured picture division height is 64, 32, or 16. In this way, the size of the first picture block, determined under the constraints of the pre-configured picture division width and pre-configured picture division height, can be reduced.
[0014] In possible implementations, the basic processing unit is a 4x4 sample matrix.
[0015] According to a second aspect, the present application provides an interpretation device including a determination module, a prediction module, and a combination module. The determination module determines a plurality of first picture blocks within a picture block to be processed based on a preset picture division width, a preset picture division height, and the width and height of the picture block to be processed. The prediction module individually performs bidirectional optical flow prediction for the plurality of first picture blocks to obtain predictors for each first picture block. The combination module obtains predictors for the picture block to be processed using combinations of predictors for the plurality of first picture blocks. The determination module determines a plurality of first picture blocks within a picture block to be processed based on a preset picture division width, a preset picture division height, and the width and height of the picture block to be processed. Therefore, the size of the first picture block is limited by the preset picture division width and preset picture division height, and the area of each determined first picture block is not very large. As a result, less hardware resources such as memory resources can be consumed, the complexity of performing interpretation can be reduced, and processing efficiency can be improved.
[0016] In possible implementations, the decision module, prediction module, and combination module may be further configured to perform the operation of the method in any possible implementation of the first embodiment. Further details will not be described here.
[0017] According to a third aspect, an embodiment of the present application provides an interpretation device. The device includes a processor and memory, the processor being connected to the memory. The memory stores one or more programs, one or more of which are executed by the processor, and one or more of which include instructions for performing the method in the first aspect or any possible implementation of the first aspect.
[0018] According to a fourth aspect, the present application provides a non-volatile computer-readable storage medium configured to store a computer program. The computer program is loaded by a processor to execute instructions for the method in the first aspect or any possible implementation of the first aspect.
[0019] According to a fifth aspect, the present application provides a chip comprising a programmable logic circuit and / or programmable instructions. When the chip is in operation, the method of the first aspect or any possible implementation of the first aspect is performed.
[0020] According to a sixth aspect, an embodiment of the present application provides an interpretation method comprising the steps of: acquiring motion information of a picture block to be processed, wherein the picture block to be processed includes a plurality of virtual pipeline data units, and each virtual pipeline data unit includes at least one basic processing unit; acquiring a predictor matrix for each virtual pipeline data unit based on the motion information; calculating a horizontal prediction gradient matrix and a vertical prediction gradient matrix for each virtual pipeline data unit based on each predictor matrix; and calculating a motion information refinement value for each basic processing unit within each virtual pipeline data unit based on the predictor matrix, the horizontal prediction gradient matrix, and the vertical prediction gradient matrix.
[0021] In a feasible implementation of the sixth embodiment, the step of obtaining a predictor matrix for each virtual pipeline data unit based on motion information includes the step of obtaining an initial predictor matrix for each virtual pipeline data unit based on motion information, wherein the size of the initial predictor matrix is equal to the size of the virtual pipeline data unit, and the step of using the initial predictor matrix as a predictor matrix.
[0022] In a feasible implementation of the sixth embodiment, after the step of obtaining an initial prediction matrix for each virtual pipeline data unit, the method further includes the step of performing sample augmentation at the edges of the initial prediction matrix to obtain an augmented prediction matrix, wherein the size of the augmented prediction matrix is greater than the size of the initial prediction matrix, and correspondingly the step of using the initial prediction matrix as a predictor matrix includes the step of using the augmented prediction matrix as a predictor matrix.
[0023] In a feasible implementation of the sixth embodiment, the step of performing sample expansion at an edge of the initial prediction matrix includes the step of obtaining sample values of samples outside the initial prediction matrix based on interpolation of sample values of samples within the initial prediction matrix, or the step of using the sample values of samples at an edge of the initial prediction matrix as sample values of samples outside the initial prediction matrix that are adjacent to the edge.
[0024] In an executable implementation of the sixth embodiment, a virtual pipeline data unit includes a plurality of motion compensation units, and the step of obtaining a predictor matrix for each virtual pipeline data unit based on motion information includes the step of obtaining a compensation value matrix for each motion compensation unit based on motion information, and the step of combining the compensation value matrices of the plurality of motion compensation units to obtain a predictor matrix.
[0025] In a feasible implementation of the sixth embodiment, the step of calculating the horizontal and vertical gradient matrices for each virtual pipeline data unit based on each predictor matrix includes the step of performing horizontal gradient calculations and vertical gradient calculations separately on the predictor matrix in order to obtain the horizontal and vertical gradient matrices.
[0026] In an executable implementation of the sixth embodiment, prior to the step of calculating the motion information refinement value of each basic processing unit in each virtual pipeline data unit based on a predictor matrix, a horizontal predictor gradient matrix, and a vertical predictor gradient matrix, the method further includes the step of performing sample expansion at the edges of the predictor matrix to obtain a padding predictor matrix, wherein the padding predictor matrix has a predetermined size; and the step of performing gradient expansion separately at the edges of the horizontal predictor gradient matrix and the edges of the vertical predictor gradient matrix to obtain a padding horizontal gradient matrix and a padding vertical gradient matrix, wherein the padding horizontal gradient matrix and the padding vertical gradient matrix each have a predetermined size; and correspondingly, the step of calculating the motion information refinement value of each basic processing unit in each virtual pipeline data unit based on a predictor matrix, a horizontal predictor gradient matrix, and a vertical predictor gradient matrix includes the step of calculating the motion information refinement value of each basic processing unit in each virtual pipeline data unit based on a padding predictor matrix, a padding horizontal gradient matrix, and a padding vertical gradient matrix.
[0027] In a feasible implementation of the sixth embodiment, prior to the step of performing sample augmentation on the edges of the predictor matrix, the method further includes the step of determining that the size of the predictor matrix is smaller than a preset size.
[0028] In a feasible implementation of the sixth embodiment, prior to the step of performing gradient expansion on the edges of the horizontal predictive gradient matrix and the edges of the vertical predictive gradient matrix, the method further includes the step of determining that the size of the horizontal predictive gradient matrix and / or the size of the vertical predictive gradient matrix is smaller than a preset size.
[0029] In a feasible implementation of the sixth embodiment, after the step of calculating the motion information refinement value for each basic processing unit in each virtual pipeline data unit, the method further includes the step of obtaining a predictor for each basic processing unit based on the predictor matrix of the virtual pipeline data unit and the motion information refinement value for each basic processing unit in the virtual pipeline data unit.
[0030] In a feasible implementation of the sixth embodiment, the method is used for bidirectional prediction, and correspondingly, motion information includes a first reference frame list motion information and a second reference frame list motion information, the predictor matrix includes a first predictor matrix and a second predictor matrix, the first predictor matrix is obtained based on the first reference frame list motion information, the second predictor matrix is obtained based on the second reference frame list motion information, and the horizontal prediction gradient matrix includes a first horizontal prediction gradient matrix and a second horizontal prediction gradient matrix, the first horizontal prediction gradient matrix is calculated based on the first predictor matrix. The second horizontal prediction gradient matrix is calculated based on the second predictor matrix, the vertical prediction gradient matrix includes the first vertical prediction gradient matrix and the second vertical prediction gradient matrix, the first vertical prediction gradient matrix is calculated based on the first predictor matrix, the second vertical prediction gradient matrix is calculated based on the second predictor matrix, the motion information refinement value includes the first reference frame list motion information refinement value and the second reference frame list motion information refinement value, the first reference frame list motion information refinement value is calculated based on the first predictor matrix, the first horizontal prediction gradient matrix and the first vertical prediction gradient matrix, and 2 The reference frame list motion information refinement value is, 2 The predictor matrix and the 2 It is calculated based on the horizontal prediction gradient matrix and the second vertical prediction gradient matrix.
[0031] In a feasible implementation of the sixth embodiment, prior to the step of performing sample expansion on the edges of the initial prediction matrix, the method further includes the step of determining that the time domain position of the picture frame in which the picture block to be processed is located is between a first reference frame indicated by a first reference frame list motion information and a second reference frame indicated by a second reference frame list motion information.
[0032] In a feasible implementation of the sixth embodiment, after the step of obtaining the predictor matrix for each virtual pipeline data unit, the method further includes the step of determining that the difference between the first predictor matrix and the second predictor matrix is less than a first threshold.
[0033] In a feasible implementation of the sixth embodiment, the motion information refinement value of a basic processing unit corresponds to one basic predictor matrix in the predictor matrix, and prior to the step of calculating the motion information refinement value of each basic processing unit in each virtual pipeline data unit based on the predictor matrix, a horizontal predictor gradient matrix, and a vertical predictor gradient matrix, the method further includes the step of determining that the difference between a first basic predictor matrix and a second basic predictor matrix is less than a second threshold.
[0034] In the executable implementation of the sixth embodiment, the size of the basic processing unit is 4 × 4.
[0035] In the executable implementation of the sixth embodiment, the width of the virtual pipeline data unit is W, the height of the virtual pipeline data unit is H, and the size of the augmented prediction matrix is (W+n+2)×(H+n+2). Correspondingly, the size of the horizontal prediction gradient matrix is (W+n)×(H+n), and the size of the vertical prediction gradient matrix is (W+n)×(H+n), where W and H are positive integers and n is an even number.
[0036] In a feasible implementation of the sixth embodiment, n is 0, 2, or -2.
[0037] In an executable implementation of the sixth embodiment, prior to the step of obtaining motion information of the picture block to be processed, the method further includes the step of determining that the picture block to be processed includes a plurality of virtual pipeline data units.
[0038] According to a seventh aspect, an embodiment of the present application provides an interprediction device comprising: an acquisition module configured to acquire motion information of a picture block to be processed, wherein the picture block to be processed includes a plurality of virtual pipeline data units, and each virtual pipeline data unit includes at least one basic processing unit; a compensation module configured to acquire a predictor matrix for each virtual pipeline data unit based on the motion information; a calculation module configured to calculate a horizontal predictor gradient matrix and a vertical predictor gradient matrix for each virtual pipeline data unit based on each predictor matrix; and a refinement module configured to calculate a motion information refinement value for each basic processing unit within each virtual pipeline data unit based on the predictor matrix, the horizontal predictor gradient matrix, and the vertical predictor gradient matrix.
[0039] In a feasible implementation of the seventh embodiment, the compensation module is configured to perform, specifically, an operation to obtain an initial prediction matrix for each virtual pipeline data unit based on motion information, wherein the size of the initial prediction matrix is equal to the size of the virtual pipeline data unit, and an operation to use the initial prediction matrix as a predictor matrix.
[0040] In a feasible implementation of the seventh embodiment, the compensation module is configured to perform, specifically, an operation in which sample augmentation is performed at the edges of the initial prediction matrix in order to obtain an augmented prediction matrix, wherein the size of the augmented prediction matrix is greater than the size of the initial prediction matrix, and an operation in which the augmented prediction matrix is used as a predictor matrix.
[0041] In a feasible implementation of the seventh embodiment, the compensation module is configured to either obtain sample values of samples outside the initial prediction matrix based on interpolation of sample values of samples within the initial prediction matrix, or to use the sample values of samples at the edges of the initial prediction matrix as sample values of samples outside the initial prediction matrix that are adjacent to the edges.
[0042] In a feasible implementation of the seventh embodiment, the virtual pipeline data unit includes a plurality of motion compensation units, and the compensation module is configured to specifically obtain a compensation value matrix for each motion compensation unit based on motion information and to combine the compensation value matrices of the plurality of motion compensation units to obtain a predictor matrix.
[0043] In an executable implementation of the seventh embodiment, the computation module is configured to perform horizontal gradient calculations and vertical gradient calculations separately on the predictor matrix in order to obtain horizontal predicted gradient matrices and vertical predicted gradient matrices.
[0044] In a feasible implementation of the seventh embodiment, the device further includes a padding module configured to perform: an operation to perform sample expansion at the edges of a predictor matrix in order to obtain a padding prediction matrix, wherein the padding prediction matrix has a preset size; an operation to perform gradient expansion separately at the edges of the horizontal prediction gradient matrix and the edges of the vertical prediction gradient matrix in order to obtain a padding horizontal gradient matrix and a padding vertical gradient matrix, wherein the padding horizontal gradient matrix and the padding vertical gradient matrix each have a preset size; and an operation to calculate motion information refinement values for each basic processing unit in each virtual pipeline data unit based on the padding prediction matrix, the padding horizontal gradient matrix, and the padding vertical gradient matrix.
[0045] In a feasible implementation of the seventh embodiment, the device further includes a decision module configured to determine that the size of the predictor matrix is less than a preset size.
[0046] In a feasible implementation of the seventh embodiment, the decision module is further configured to determine that the size of the horizontal predictive gradient matrix and / or the size of the vertical predictive gradient matrix are less than a preset size.
[0047] In a seventh executable implementation, the refinement module is further configured to obtain predictors for each basic processing unit based on the predictor matrix of the virtual pipeline data unit and the motion information refinement values of each basic processing unit within the virtual pipeline data unit.
[0048] In a feasible implementation of the seventh embodiment, the device is used for bidirectional prediction, and accordingly, motion information includes a first reference frame list motion information and a second reference frame list motion information, the predictor matrix includes a first predictor matrix and a second predictor matrix, the first predictor matrix is obtained based on the first reference frame list motion information, the second predictor matrix is obtained based on the second reference frame list motion information, and the horizontal prediction gradient matrix includes a first horizontal prediction gradient matrix and a second horizontal prediction gradient matrix, the first horizontal prediction gradient matrix is calculated based on the first predictor matrix. The second horizontal prediction gradient matrix is calculated based on the second predictor matrix, the vertical prediction gradient matrix includes the first vertical prediction gradient matrix and the second vertical prediction gradient matrix, the first vertical prediction gradient matrix is calculated based on the first predictor matrix, the second vertical prediction gradient matrix is calculated based on the second predictor matrix, the motion information refinement value includes the first reference frame list motion information refinement value and the second reference frame list motion information refinement value, the first reference frame list motion information refinement value is calculated based on the first predictor matrix, the first horizontal prediction gradient matrix and the first vertical prediction gradient matrix, and 2 The reference frame list motion information refinement value is, 2 The predictor matrix and the 2 It is calculated based on the horizontal prediction gradient matrix and the second vertical prediction gradient matrix.
[0049] In an executable implementation of the seventh embodiment, the determination module is further configured to determine that the time domain position of the picture frame in which the picture block to be processed is located is between a first reference frame indicated by a first reference frame list motion information and a second reference frame indicated by a second reference frame list motion information.
[0050] In a feasible implementation of the seventh embodiment, the decision module is further configured to determine that the difference between the first predictor matrix and the second predictor matrix is less than a first threshold.
[0051] In a feasible implementation of the seventh embodiment, the decision module is further configured to determine that the difference between a first basic predictor matrix and a second basic predictor matrix is less than a second threshold.
[0052] In the executable implementation of the seventh embodiment, the size of the basic processing unit is 4 × 4.
[0053] In the executable implementation of the seventh embodiment, the width of the virtual pipeline data unit is W, the height of the virtual pipeline data unit is H, and the size of the augmented prediction matrix is (W+n+2)×(H+n+2). Correspondingly, the size of the horizontal prediction gradient matrix is (W+n)×(H+n), and the size of the vertical prediction gradient matrix is (W+n)×(H+n), where W and H are positive integers and n is an even number.
[0054] In a feasible implementation of the seventh embodiment, n is 0, 2, or -2.
[0055] In a seventh executable implementation, the decision module is further configured to determine that the picture block to be processed contains a plurality of virtual pipeline data units.
[0056] According to the eighth aspect, an embodiment of the present application provides an encoding device comprising coupled non-volatile memory and a processor. The processor invokes program code stored in memory to perform some or all steps of the method in the first aspect, or some or all steps of the method in the sixth aspect.
[0057] According to the ninth aspect, an embodiment of the present application provides a decoding device comprising coupled non-volatile memory and a processor. The processor invokes program code stored in memory to perform some or all steps of the method in the first aspect, or some or all steps of the method in the sixth aspect.
[0058] According to the tenth aspect, an embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores program code, which includes instructions used to perform some or all steps of the method in the first aspect, or some or all steps of the method in the sixth aspect.
[0059] According to the eleventh aspect, an embodiment of the present application provides a computer program product. When the computer program product is executed on a computer, the computer can perform some or all of the steps of the method in the first aspect, or some or all of the steps of the method in the sixth aspect.
[0060] To more clearly illustrate the technical solutions in the embodiments or background of this application, the following describes the accompanying drawings illustrating the embodiments or background of this application. [Brief explanation of the drawing]
[0061] [Figure 1A] This is a block diagram of an example of a video coding system 10 according to the embodiment of this application. [Figure 1B] This is a block diagram of an example of a video coding system 40 according to an embodiment of this application. [Figure 2] This is a block diagram of an exemplary structure of the encoder 20 according to the embodiment of this application. [Figure 3] This is a block diagram of an exemplary structure of the decoder 30 according to the embodiment of this application. [Figure 4] This is a block diagram of an example of a video coding device 400 according to the embodiment of this application. [Figure 5] This is a block diagram of another example of an encoding or decoding device according to an embodiment of the present application. [Figure 6] This is a schematic diagram of a candidate position for motion information according to the embodiment of this application. [Figure 7] This is a schematic diagram showing how motion information is used for interpretation according to an embodiment of this application. [Figure 8] This is a schematic diagram of the bidirectional weighted prediction according to the embodiment of this application. [Figure 9] This is a schematic diagram of CU boundary padding according to an embodiment of the present application. [Figure 10] This is a schematic diagram of a VPDU division according to an embodiment of this application. [Figure 11] This is a schematic diagram of an invalid VPDU division according to an embodiment of this application. [Figure 12] This is a flowchart of the interpretation prediction method according to the embodiment of this application. [Figure 13] This is another schematic diagram showing how motion information is used in interpretation according to an embodiment of this application. [Figure 14] This is a flowchart of another interpretation method according to an embodiment of this application. [Figure 15] This is another schematic diagram showing how motion information is used in interpretation according to an embodiment of this application. [Figure 16A] This is a flowchart of another interpretation method according to an embodiment of this application. [Figure 16B] This is a flowchart of another interpretation method according to an embodiment of this application. [Figure 17] This is a flowchart of the method according to the embodiment of this application. [Figure 18] This is a block diagram of the configuration of an inter-prediction device according to an embodiment of this application. [Figure 19] This is a block diagram of the configuration of another interpretation device according to an embodiment of this application. [Modes for carrying out the invention]
[0062] The embodiments of this application will be described below with reference to the accompanying drawings of embodiments of this application. In the following description, references will be made to the accompanying drawings, which form part of the present disclosure and illustrate specific aspects of embodiments of the present invention, or specific aspects in which embodiments of the present invention may be used. It should be understood that embodiments of the present invention may be used in other aspects and may include structural or logical modifications not shown in the accompanying drawings. Accordingly, the following detailed description should not be construed as limiting, and the scope of the present invention is defined by the accompanying claims. For example, it should be understood that what is disclosed with reference to a described method may also apply to a corresponding device or system configured to perform the method, and vice versa. For example, where one or more specific method steps are described, the corresponding device may include one or more units, such as functional units for performing the described method steps (e.g., one unit that performs one or more steps, or multiple units, each performing one or more of a plurality of steps), even if such one or more units are not explicitly described or illustrated in the accompanying drawings. In addition, for example, if a particular device is described based on one or more units, such as a functional unit, the corresponding method may include one step for performing the function of one or more units (e.g., one step for performing the function of one or more units, or multiple steps, each of which is used to perform the function of one or more of the units), even if such one or more steps are not explicitly described or illustrated in the accompanying drawings. Furthermore, it should be understood that the features of the various exemplary embodiments and / or aspects described herein may be combined with each other unless otherwise specified.
[0063] The technical solutions in the embodiments of this application may be applied not only to existing video coding standards (e.g., standards such as H.264 or high-efficiency video coding (HEVC)) but also to future video coding standards (e.g., the H.266 standard). The terms used in the embodiments of this application are used solely to describe the specific embodiments of this application and are not intended to limit this application. Below, some relevant concepts in the embodiments of this application are briefly described first.
[0064] Video coding typically involves processing a sequence of pictures that form a video or video sequence. In the field of video coding, the terms “picture,” “frame,” and “image” may be used synonymously. As used herein, video coding refers to video encoding or video decoding. Video encoding is performed on the source side and typically involves processing the original video picture (e.g., by compression) to reduce the amount of data required to represent the video picture (for more efficient storage and / or transmission). Video decoding is performed on the destination side and typically involves the reverse processing of the encoder to reconstruct the video picture. In embodiments, “encoding” of a video picture should be understood as “encoding” or “decoding” in relation to a video sequence. The combination of encoding and decoding is also called coding (encoding and decoding).
[0065] A video sequence contains a series of pictures, each of which is further divided into slices, and each slice into blocks. Video coding is performed by these blocks. In some newer video coding standards, the concept of a "block" is further extended. For example, the H.264 standard introduces macroblocks (MB). Macroblocks may be further divided into multiple partitions that can be used for predictive coding. In the HEVC standard, basic concepts such as coding units (CUs), prediction units (PUs), and transform units (TUs) are used, where multiple block units are functionally obtained through partitioning, and an entirely new tree-based structure is used for description. For example, a CU may be divided into smaller CUs based on a quadtree, and the smaller CUs may be further divided to generate a quadtree structure. A CU is the basic unit for partitioning and coding a coded picture. PUs and TUs also have similar tree structures. A PU may correspond to a prediction block and is the basic unit for predictive coding. A CU is further divided into multiple PUs based on the division pattern. A TU may correspond to a transformation block and is the basic unit for transforming the predicted residuals. However, CUs, PUs, and TUs are, in essence, conceptually blocks (or picture blocks).
[0066] For example, in HEVC, a coding tree unit (CTU) is divided into multiple CUs by using a quadtree structure, which is represented as the coding tree. It is determined at the CU level whether the picture region is coded by interpicture (time) prediction or intrapicture (spatial) prediction. Each CU may be further divided into one, two, or four PUs based on the PU partitioning type. Within a single PU, the same prediction process is applied, and the relevant information is sent to the decoder based on the PU. After obtaining the residual block by applying the prediction process based on the PU partitioning type, the CU may be partitioned into transform units (TUs) based on another quadtree structure similar to the coding tree used for the CU. In the latest developments of video compression technology, frames are partitioned by a quad-tree and binary tree (QTBT) to partition the coding blocks. In the QTBT block structure, the CU may be square or rectangular.
[0067] In this specification, for the sake of clarity and understanding, the picture block being encoded within the currently coded picture may be referred to as the current block. For example, during encoding, the current block is the block being encoded, and during decoding, the current block is the block being decoded. A decoded picture block within a reference picture used to predict the current block is called a reference block. In other words, a reference block is a block that provides a reference signal about the current block, and the reference signal represents a sample value within the picture block. A block within a reference picture that provides a prediction signal about the current block may be called a prediction block, and the prediction signal represents a sample value, sampling value, or sampling signal within the prediction block. For example, after traversing several reference blocks, an optimal reference block may be found that provides a prediction about the current block, and this block is called a prediction block.
[0068] In lossless video coding, the original video picture can be reconstructed, meaning that (assuming no transmission loss or other data loss occurs during storage or transmission) the reconstructed video picture will have the same quality as the original video picture. In lossy video coding, further compression is performed, such as through quantization, to reduce the amount of data required to represent the video picture, and the video picture cannot be fully reconstructed on the decoder side; that is, the quality of the reconstructed video picture is inferior to that of the original video picture.
[0069] Some H.261 video coding standards relate to "lossy hybrid video coding" (i.e., spatial and temporal predictions in the sample domain are combined with 2D transform coding to apply quantization in the transform domain). Each picture in a video sequence is typically partitioned into a set of non-overlapping blocks, and coding is typically performed at the block level. Specifically, on the encoder side, video is typically processed, i.e., encoded, at the block (video block) level. For example, spatial (intra-picture) and temporal (inter-picture) predictions generate predicted blocks, which are subtracted from the current block (the block being processed or to be processed) to obtain a residual block, which is then transformed in the transform domain and quantized to reduce the amount of data to be transmitted (compressed). On the decoder side, the reverse processing of the encoder is applied to the encoded or compressed block to reconstruct the current block for representation. In addition, the encoder replicates the decoder's processing loop so that the encoder and decoder generate the same predictions (e.g., intra-predictions and inter-predictions) and / or reconstructions to process, i.e., encode, subsequent blocks.
[0070] The following describes a system architecture to which embodiments of this application apply. Figure 1A is a schematic block diagram of a video coding system 10 to which embodiments of this application apply apply. As shown in Figure 1A, the video coding system 10 may include a source device 12 and a destination device 14. The source device 12 generates encoded video data, and therefore the source device 12 may be called a video encoder. The destination device 14 can decode the encoded video data generated by the source device 12, and therefore the destination device 14 may be called a video decoder. In various implementation solutions, the source device 12, the destination device 14, or both the source device 12 and the destination device 14 may include one or more processors and memory coupled to one or more processors. The memory may include, but is not limited to, RAM, ROM, EEPROM, flash memory, or any other medium that can be used to store desired program code in the form of computer-accessible instructions or data structures as described herein. The source device 12 and destination device 14 may include a variety of devices, including desktop computers, mobile computing devices, notebook (e.g., laptop) computers, tablet computers, set-top boxes, telephone handsets such as "smartphones," televisions, cameras, display devices, digital media players, video game consoles, in-vehicle computers, wireless communication devices, and the like.
[0071] Figure 1A depicts the source device 12 and the destination device 14 as separate devices, but the device embodiment may include both the source device 12 and the destination device 14, or both the functions of the source device 12 and the functions of the destination device 14, i.e., both the source device 12 or its corresponding function and the destination device 14 or its corresponding function. In such embodiments, the source device 12 or its corresponding function and the destination device 14 or its corresponding function may be implemented using the same hardware and / or software, separate hardware and / or software, or a combination thereof.
[0072] A communication connection may be implemented between the source device 12 and the destination device 14 via link 13, and the destination device 14 may receive encoded video data from the source device 12 via link 13. Link 13 may include one or more media or devices that can move the encoded video data from the source device 12 to the destination device 14. In one example, link 13 may include one or more communication media that enable the source device 12 to transmit the encoded video data directly to the destination device 14 in real time. In this example, the source device 12 may modulate the encoded video data according to a communication standard (e.g., a wireless communication protocol) and transmit the modulated video data to the destination device 14. One or more communication media may include wireless and / or wired communication media, such as a radio frequency (RF) spectrum or one or more physical transmission cables. One or more communication media may be part of a packet-based network, such as a local area network, a wide area network, or a global network (e.g., the Internet). One or more communication media may include a router, switch, base station, or another device that facilitates communication from source device 12 to destination device 14.
[0073] The source device 12 includes an encoder 20, and additionally or optionally, the source device 12 may further include a picture source 16, a picture preprocessor 18, and a communication interface 22. In certain embodiments, the encoder 20, picture source 16, picture preprocessor 18, and communication interface 22 may be hardware components within the source device 12, or software programs within the source device 12. Descriptions are provided separately below.
[0074] Picture Source 16 may include, or may include, any type of picture capture device configured to capture real-world pictures, etc., and / or any type of device for generating pictures or comments (for the purpose of encoding screen content, some text on the screen is also considered part of the picture or image to be encoded), such as a computer graphics processing unit configured to generate computer animated pictures, or any type of device configured to acquire and / or provide real-world pictures or computer animated pictures (e.g., screen content or virtual reality (VR) pictures) and / or any combination thereof (e.g., augmented reality (AR) pictures). Picture Source 16 may also include a camera configured to capture pictures, or a memory configured to store pictures. Picture Source 16 may further include any type of (internal or external) interface through which previously captured or generated pictures are stored and / or pictures are acquired or received. If the picture source 16 is a camera, the picture source 16 may be, for example, a local camera or an integrated camera integrated into the source device. If the picture source 16 is memory, the picture source 16 may be, for example, local memory or integrated memory integrated into the source device. If the picture source 16 includes an interface, the interface may be, for example, an external interface for receiving pictures from an external video source. The external video source may be, for example, an external picture capture device such as a camera, external memory, or an external picture generation device. The external picture generation device may be, for example, an external computer graphics processor, a computer, or a server.The interface may be any type of interface that conforms to a proprietary or standardized interface protocol, such as a wired or wireless interface, or an optical interface.
[0075] A picture may be considered a two-dimensional array or matrix of samples (picture elements). A sample in an array may also be called a sampling point. The number of sampling points in the horizontal (or axis) and vertical (or axis) directions of the array or picture defines the size and / or resolution of the picture. For color representation, three color components are typically used; that is, a picture may be represented as, or contain, three sample arrays. In the RGB format or color space, a picture contains corresponding red, green, and blue sample arrays. However, in video coding, each sample is typically represented in a luminance / chromaticity format or color space; for example, a picture in the YCbCr format contains a luminance component represented by Y (sometimes represented by L) and two chromaticity components represented by Cb and Cr. The luminance (abbreviated luma) component Y represents luminance or gray level intensity (for example, these two are the same in a grayscale picture), and the two chromaticity (abbreviated chroma) components Cb and Cr represent chromaticity or color information components. Therefore, a picture in the YCbCr format includes a luminance sample array of luminance sample values (Y) and two chromaticity sample arrays of chromaticity values (Cb and Cr). A picture in the RGB format may be converted to or transformed into a picture in the YCbCr format, and vice versa. This process is also called color conversion or transformation. If the picture is monochrome, the picture may include only the luminance sample array. In this embodiment of the present application, the picture source 16 provides the picture Puri Processor 18 The picture sent may also be called the original picture data 17.
[0076] The picture preprocessor 18 receives the original picture data 17 and may perform preprocessing on the original picture data 17 to obtain a preprocessed picture 19 or preprocessed picture data 19. For example, the preprocessing performed by the picture preprocessor 18 may include cropping, color format conversion (e.g., RGB to YUV), color correction, or noise reduction.
[0077] The encoder 20 (also called the video encoder 20) is configured to process the preprocessed picture data 19 by receiving the preprocessed picture data 19 and using the relevant prediction mode (such as the prediction mode in each embodiment herein) to provide encoded picture data 21 (structural details of the encoder 20 will be further described below based on Figures 2, 4, or 5). In some embodiments, the encoder 20 is as described in this application. Inter To implement the encoder-side application of the prediction method, the system may be configured to perform each of the embodiments described below.
[0078] The communication interface 22 may be configured to receive encoded picture data 21 and transmit the encoded picture data 21 via link 13 to destination device 14 or any other device (e.g., memory) for storage or direct reconstruction. The any other device may be any device used for decoding or storage. The communication interface 22 may be configured, for example, to encapsulate the encoded picture data 21 in an appropriate format, such as a data packet, for transmission via link 13.
[0079] The destination device 14 includes a decoder 30, and additionally or optionally, the destination device 14 may further include a communication interface 28, a picture postprocessor 32, and a display device 34. A description is provided separately below.
[0080] The communication interface 28 may be configured to receive encoded picture data 21 from the source device 12 or any other source. Any other source is, for example, a storage device, and the storage device is, for example, an encoded picture data storage device. The communication interface 28 may be configured to transmit or receive encoded picture data 21 over the link 13 between the source device 12 and the destination device 14, or over any type of network. The link 13 is, for example, a direct wired connection or a wireless connection, and any type of network is, for example, a wired network or a wireless network or any combination thereof, or any type of private network or public network or any combination thereof. The communication interface 28 may be configured, for example, to decapsulate data packets transmitted over the communication interface 22 in order to obtain encoded picture data 21.
[0081] Both communication interface 22 and communication interface 28 may be configured as unidirectional or bidirectional communication interfaces, and may be configured, for example, to send and receive messages to establish communication, and to confirm and exchange any other information related to data transmission, such as communication links and / or encoded picture data transmission.
[0082] The decoder 30 (also called decoder 30) is configured to receive encoded picture data 21 and provide decoded picture data 31 or decoded picture 31 (structural details of decoder 30 will be described in further detail below based on Figures 3, 4, or 5). In some embodiments, decoder 30 is as described in this application. Inter To implement the decoder-side application of the prediction method, the system may be configured to perform each of the embodiments described below.
[0083] The picture post-processor 32 is configured to post-process the decoded picture data 31 (also called reconstructed picture data) in order to obtain post-processed picture data 33. The post-processing performed by the picture post-processor 32 may include color format conversion (e.g., YCbCr to RGB), color correction, cropping, resampling, or any other processing. The picture post-processor 32 may be further configured to send the post-processed picture data 33 to the display device 34.
[0084] The display device 34 is configured to receive post-processed picture data 33 for displaying the picture to a user, observer, etc. The display device 34 may include or be any type of display configured to present the reconstructed picture, such as an integrated or external display or monitor. For example, the display may include a liquid crystal display (LCD), an organic light-emitting diode (OLED) display, a plasma display, a projector, a micro-LED display, a liquid crystal on silicon (LCoS) display, a digital light processor (DLP), or any other type of display.
[0085] Figure 1A depicts the source device 12 and the destination device 14 as separate devices, but the device embodiment may include both the source device 12 and the destination device 14, or both the functions of the source device 12 and the functions of the destination device 14, i.e., both the source device 12 or its corresponding function and the destination device 14 or its corresponding function. In such embodiments, the source device 12 or its corresponding function and the destination device 14 or its corresponding function may be implemented using the same hardware and / or software, separate hardware and / or software, or a combination thereof.
[0086] Based on the description, a person skilled in the art will readily understand that the functions of different units or the presence and (exact) division of multiple functions / one function of the source device 12 and / or destination device 14 shown in Figure 1A may vary depending on the actual device and application. The source device 12 and destination device 14 may each include any one of a variety of devices, including any type of handheld or stationary device, e.g., a notebook computer or laptop computer, a mobile phone, a smartphone, a tablet or tablet computer, a video camera, a desktop computer, a set-top box, a television, a camera, an in-car device, a display device, a digital media player, a video game console, a video streaming transmission device (such as a content service server or content distribution server), a broadcast receiver device, or a broadcast transmitter device, and may or may not use any type of operating system.
[0087] The encoder 20 and decoder 30 may each be implemented as any one of a variety of suitable circuits, for example, one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, or any combination thereof. If the technology is partially implemented in software, the device may store software instructions in a suitable non-temporary computer-readable storage medium and execute the instructions in hardware by using one or more processors to perform the technology in this disclosure. Any of the foregoing (including hardware, software, and combinations of hardware and software) may be considered as one or more processors.
[0088] In some cases, the video coding system 10 shown in Figure 1A is merely an example, and the technology in this application may apply to video coding configurations (e.g., video coding or video decoding) that do not require any data communication between the coding device and the decoding device. In other examples, the data may be retrieved from local memory or streamed over a network. The video coding device may code the data and store the data in memory, and / or the video decoding device may retrieve the data from memory and decode the data. In some examples, coding and decoding are performed by devices that do not communicate with each other and only code the data into memory and / or retrieve the data from memory and decode the data.
[0089] Figure 1B is an illustrative diagram of an example of a video coding system 40, including the encoder 20 in Figure 2 and / or the decoder 30 in Figure 3, according to an exemplary embodiment. The video coding system 40 can implement various combinations of technologies in the embodiments of this application. In the illustrated implementation, the video coding system 40 may include an imaging device 41, an encoder 20, a decoder 30 (and / or a video encoder / decoder implemented by the logic circuit 47 of a processing unit 46), an antenna 42, one or more processors 43, one or more memories 44, and / or a display device 45.
[0090] As shown in Figure 1B, the imaging system 41, antenna 42, processing unit 46, logic circuit 47, encoder 20, decoder 30, processor 43, memory 44, and / or display device 45 can communicate with each other. As described, the video coding system 40 is shown having both the encoder 20 and the decoder 30, but in different examples, the video coding system 40 may include only the encoder 20 or only the decoder 30.
[0091] In some examples, the antenna 42 may be configured to transmit or receive an encoded bitstream of video data. Furthermore, in some examples, the display device 45 may be configured to present the video data. In some examples, the logic circuit 47 may be implemented by a processing unit 46. The processing unit 46 may include application-specific integrated circuit (ASIC) logic, a graphics processing unit, a general-purpose processor, etc. The video coding system 40 may also include an optional processor 43. The optional processor 43 may similarly include application-specific integrated circuit (ASIC) logic, a graphics processing unit, a general-purpose processor, etc. In some examples, the logic circuit 47 may be implemented by hardware such as dedicated video coding hardware, and the processor 43 may be implemented by universal software, an operating system, etc. In addition, the memory 44 may be any type of memory, such as volatile memory (e.g., static random access memory (SRAM) or dynamic random access memory (DRAM)) or non-volatile memory (e.g., flash memory). In non-restrictive examples, memory 44 may be implemented by cache memory. In some examples, logic circuit 47 may access memory 44 (for example, to implement a picture buffer). In other examples, logic circuit 47 and / or processing unit 46 may include memory (e.g., a cache) to implement a picture buffer, etc.
[0092] In some examples, the encoder 20 implemented by the logic circuit may include a picture buffer (implemented, for example, by the processing unit 46 or memory 44) and a graphics processing unit (implemented, for example, by the processing unit 46). The graphics processing unit may be communicatively coupled to the picture buffer. The graphics processing unit may include the encoder 20 implemented by the logic circuit 47 to implement the various modules described with reference to Figure 2 and / or any other encoder systems or subsystems described herein. The logic circuit may be configured to perform the various operations described herein.
[0093] In some examples, the decoder 30 may also be implemented by logic circuits 47 to implement various modules described with reference to the decoder 30 in Figure 3, and / or any other decoder systems or subsystems described herein. In some examples, the decoder 30 implemented by logic circuits may be (e.g., a processing unit) 46 It may also include a picture buffer (implemented by memory 44) and a graphics processing unit (implemented by, for example, processing unit 46). The graphics processing unit may be communicatively coupled to the picture buffer. The graphics processing unit may include a decoder 30 implemented by logic circuit 47 to implement the various modules described with reference to Figure 3 and / or any other decoder systems or subsystems described herein.
[0094] In some examples, antenna 42 may be configured to receive an encoded bitstream of video data. As described herein, the encoded bitstream may include data related to encoding partitioning (e.g., conversion coefficients or quantization conversion coefficients, optional indicators (as described), and / or data defining the encoding partitioning), such as data related to video frame encoding as described herein, indicators, index values, and mode selection data. The video coding system 40 may further include a decoder 30 coupled to antenna 42 and configured to decode the encoded bitstream. A display device 45 is configured to present video frames.
[0095] In the examples described with reference to the encoder 20 in this embodiment of the present application, it should be understood that the decoder 30 may be configured to perform the reverse process. For signaling syntax elements, the decoder 30 may be configured to receive and parse the syntax elements and decode the associated video data accordingly. In some examples, the encoder 20 may entropy encode the syntax elements into an encoded video bitstream. In such examples, the decoder 30 may parse the syntax elements and decode the associated video data accordingly.
[0096] It should be noted that the method described in this embodiment of the present application is primarily used in the interpretation process. This process is present in both the encoder 20 and the decoder 30. The encoder 20 and decoder 30 in this embodiment of the present application may be the corresponding encoder / decoder in a video standard protocol such as H.263, H.264, HEVV, MPEG-2, MPEG-4, VP8, or VP9, or a next-generation video standard protocol (such as H.266).
[0097] Figure 2 is a schematic / conceptual block diagram of an example encoder 20 configured to implement an embodiment of the present application. In the example of Figure 2, the encoder 20 includes a residual calculation unit 204, a transformation unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transformation unit 212, a reconstruction unit 214, a buffer 216, a loop filter unit 220, a decoded picture buffer (DPB) 230, a prediction unit 260, and an entropy coding unit 270. The prediction unit 260 may include an inter-prediction unit 244, an intra-prediction unit 254, and a mode selection unit 262. The inter-prediction unit 244 may include a motion estimation unit and a motion compensation unit (not shown in this figure). The encoder 20 shown in Figure 2 may also be called a hybrid video encoder or a hybrid video codec-based video encoder.
[0098] For example, the residual calculation unit 204, the transformation processing unit 206, the quantization unit 208, the prediction processing unit 260, and the entropy coding unit 270 form the forward signal path of the encoder 20, while the inverse quantization unit 210, the inverse transformation processing unit 212, the reconstruction unit 214, the buffer 216, the loop filter 220, the decoded picture buffer (DPB) 230, the prediction processing unit 260, etc., form the reverse signal path of the encoder. The reverse signal path of the encoder corresponds to the signal path of the decoder (see decoder 30 in Figure 3).
[0099] The encoder 20 receives a picture 201 or a picture block 203 of picture 201, for example, a picture within a series of pictures forming a video or video sequence, by using an input 202, etc. The picture block 203 may also be called the current picture block or the picture block to be encoded, and picture 201 may be called the current picture or the picture to be encoded (in particular, if the current picture is distinguished from other pictures in video coding, for example, other pictures in the same video sequence also include previously encoded and / or decoded pictures in the video sequence of the current picture).
[0100] Embodiments of encoder 20 may include a partitioning unit (not shown in Figure 2) configured to partition picture 201 into multiple non-overlapping blocks, such as block 203. The partitioning unit may be configured to use the same block size for all pictures in the video sequence and to use a corresponding raster that defines the block size, or it may be configured to change the block size between pictures, subsets, or groups of pictures and partition each picture into a corresponding block.
[0101] In one example, the predictive processing unit 260 of the encoder 20 may be configured to perform any combination of the partitioning techniques described above.
[0102] The size of picture block 203 is smaller than the size of picture 201, but like picture 201, picture block 203 is or may be considered to be a two-dimensional array or matrix of samples having sample values. In other words, picture block 203 may include, for example, one sample array (e.g., a luminance array for monochrome picture 201), three sample arrays (e.g., one luminance array and two chromaticity arrays for color picture), or any other quantity and / or type of array based on the color format used. The quantity of samples in the horizontal and vertical (or axis) directions of picture block 203 defines the size of picture block 203.
[0103] The encoder 20 shown in Figure 2 is configured to encode the picture 201 block by block, and is configured, for example, to perform encoding and prediction for each picture block 203.
[0104] The residual calculation unit 204 is configured to calculate the residual block 205 based on the picture block 203 and the prediction block 265 (further details regarding the prediction block 265 are provided below), for example, by calculating the sample values of the prediction block 265 from the sample values of the picture block 203 sample by sample. Reduce The system is configured to obtain residual block 205 in the sample domain by performing calculations.
[0105] The transformation processing unit 206 is configured to apply a transformation such as a discrete cosine transform (DCT) or a discrete sine transform (DST) to the sample values of the residual block 205 in order to obtain transformation coefficients 207 in the transformation domain. The transformation coefficients 207 may also be called residual transformation coefficients and represent the residual block 205 in the transformation domain.
[0106] The conversion processing unit 206 may be configured to apply an integer approximation of the DCT / DST, for example, the conversion specified in HEVC / H.265. This integer approximation is typically scaled proportionally by a certain factor compared to the orthogonal DCT conversion. An additional scale factor is applied as part of the conversion process to maintain the norm of the residual blocks obtained by the forward and inverse conversions. The scale factor is typically selected based on several constraints, such as powers of 2, the bit depth of the conversion coefficients, or a trade-off between the precision used in the shift operation and the implementation cost. For example, a specific scale factor may be specified for the inverse conversion on the decoder 30 side using the inverse conversion processing unit 212 (and correspondingly, for the inverse conversion on the encoder 20 side using the inverse conversion processing unit 212, etc.), and correspondingly, a corresponding scale factor may be specified for the forward conversion on the encoder 20 side using the conversion processing unit 206.
[0107] The quantization unit 208 is configured to quantize the transformation coefficients 207 by applying scalar quantization, vector quantization, etc., in order to obtain quantization transformation coefficients 209. The quantization transformation coefficients 209 may also be called quantization residual coefficients 209. The quantization process may reduce the bit depth associated with some or all of the transformation coefficients 207. For example, n-bit transformation coefficients may be truncated to m-bit transformation coefficients during quantization, where n is greater than m. The degree of quantization may be changed by adjusting the quantization parameter (QP). For example, for scale quantization, different scales may be applied to achieve finer or coarser quantization. Smaller quantization steps correspond to finer quantization, and larger quantization steps correspond to coarser quantization. An appropriate quantization step may be indicated by the quantization parameter (QP). For example, the quantization parameter may be an index to a predefined set of appropriate quantization steps. For example, smaller quantization parameters may correspond to finer quantization (smaller quantization steps), larger quantization parameters may correspond to coarser quantization (larger quantization steps), or vice versa. Quantization may include division by the quantization step and the corresponding quantization, or inverse quantization performed by an inverse quantization unit 210, etc., or it may include multiplication by the quantization step. In embodiments of some standards such as HEVC, the quantization parameter may be used to determine the quantization step. Generally, the quantization step may be calculated based on the quantization parameter by a fixed-point approximation of the equations involving division. Additional scale factors may be introduced for quantization and dequantization to restore the norm of the residual block, which may be modified for the scale used in the fixed-point approximation of the equations used for the quantization step and quantization parameter. In exemplary implementations, the scale of the inverse transform may be combined with the scale of the dequantization.Alternatively, a customized quantization table may be used, for example, to signal from the encoder to the decoder in the bitstream. Quantization is a lossy operation, and larger quantization steps result in greater losses.
[0108] The inverse quantization unit 210 uses the quantization unit 208 to obtain the dequantized coefficient 211. Quantization applied by The inverse quantization of is configured to be applied to the quantization coefficients, for example, by applying the inverse quantization of the quantization scheme applied by the quantization unit 208 based on or using the same quantization steps as the quantization unit 208. The dequantized coefficients 211 may also be called the dequantized residual coefficients 211 and may correspond to the transformation coefficients 207, but the losses resulting from quantization are usually different from the transformation coefficients.
[0109] The inverse transform processing unit 212 is configured to apply the inverse transform of the transform applied by the transform processing unit 206, for example, the inverse discrete cosine transform (DCT) or the inverse discrete sine transform (DST), in order to obtain the inverse transform block 213 in the sample domain. The inverse transform block 213 may also be called the inverse transform dequantized block 213 or the inverse transform residual block 213.
[0110] The reconstruction unit 214 (e.g., adder 214) is configured to add the inverse transform block 213 (i.e., the reconstructed residual block 213) to the prediction block 265 in order to obtain the reconstructed block 215 in the sample domain by adding the sample values of the reconstructed residual block 213 to the sample values of the prediction block 265.
[0111] Optionally, a buffer unit 216 (or simply "buffer" 216), such as a line buffer 216, is configured to buffer or store the reconstructed block 215 and its corresponding sample values for purposes such as intra-prediction. In other embodiments, the encoder may be configured to use the unfiltered reconstructed block and / or its corresponding sample values stored in the buffer unit 216 for any type of estimation and / or prediction, such as intra-prediction.
[0112] For example, in one embodiment of encoder 20, the buffer unit 216 is intra Measurement In addition to being configured to store the reconfigured block 215, the loop filter unit 220 (not shown in Figure 2) may also be configured to store the filtered block 221, and / or the buffer unit 216 and the decoded picture buffer unit 230 may be configured to form a single buffer. In other embodiments, blocks or samples from the filtered block 221 and / or the decoded picture buffer 230 (not shown in Figure 2) may be used intrapre Measurement It may be used as input or basis for...
[0113] The loop filter unit 220 (or simply "loop filter" 220) is configured to perform filtering on the reconstructed block 215 to obtain a filtered block 221 in order to perform sample conversion smoothly or to improve video quality. The loop filter unit 220 is intended to represent one or more loop filters, such as a deblocking filter, a sample-adaptive offset (SAO) filter, or another filter such as a bilateral filter, an adaptive loop filter (ALF), a sharpening or smoothing filter, or a co-filter. Although the loop filter unit 220 is shown as an in-loop filter in Figure 2, the loop filter unit 220 may be implemented as a post-loop filter in other configurations. The filtered block 221 may also be called a filtered reconstructed block 221. The decoded picture buffer 230 may store the reconstructed coding block after the loop filter unit 220 has performed filtering operations on the reconstructed coding block.
[0114] Embodiments of encoder 20 (correspondingly to loop filter unit 220) may be used to output loop filter parameters (e.g., sample adaptive offset information) so that decoder 30 can receive and apply the same loop filter parameters for decoding, etc., for example, by directly outputting the loop filter parameters or by having entropy coding unit 270 or any other entropy coding unit output the loop filter parameters after performing entropy coding.
[0115] The decoded picture buffer (DPB) 230 may be a reference picture memory that stores reference picture data for the video encoder 20 to encode video data. The DPB 230 may be any one of several memories, for example, a dynamic random access memory (DRAM) (including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), or resistive RAM (RRAM)), or another type of memory. The DPB 230 and buffer 216 may be provided by the same memory or separate memories. In one example, the decoded picture buffer (DPB) 230 is configured to store filtered blocks 221. The decoded picture buffer 230 may be further configured to store other previously filtered blocks, such as the previously reconfigured and filtered block 221 of the same current picture or a different picture, such as the previously reconfigured picture, and may provide a fully previously reconfigured, i.e., decoded picture (and the corresponding reference block and corresponding sample) and / or a partially reconfigured current picture (and the corresponding reference block and corresponding sample) for interpretation, etc. In one example, if the reconfigured block 215 is reconfigured without intra-loop filtering, the decoded picture buffer (DPB) 230 is configured to store the reconfigured block 215.
[0116] The prediction processing unit 260, also known as the block prediction processing unit 260, is configured to receive or acquire block 203 (the current block 203 of the current picture) and reconstructed picture data, for example, reference samples from the same (current) picture in buffer 216, and / or reference picture data 231 from one or more previous decoded pictures in decoded picture buffer 230, and to process such data for prediction, i.e., to provide a prediction block 265 which may be an inter-prediction block 245 or an intra-prediction block 255.
[0117] The mode selection unit 262 may be configured to select a prediction mode (e.g., intra or inter-prediction mode) and / or select the corresponding prediction block 245 or 255 as the prediction block 265 in order to calculate the residual block 205 and reconstruct the reconstructed block 215.
[0118] Embodiments of the mode selection unit 262 may be used to select a prediction mode (for example, from prediction modes supported by the prediction processing unit 260). A prediction mode provides either the best match or the smallest residual (smallest residual means better compression in transmission or storage), or the smallest signaling overhead (smallest signaling overhead means better compression in transmission or storage), or considers or balances both. The mode selection unit 262 may be configured to determine the prediction mode based on rate distortion optimization (RDO), i.e., to select a prediction mode that provides the smallest rate distortion optimization, or to select a prediction mode in which the associated rate distortion satisfies at least the prediction mode selection criteria.
[0119] The prediction processing (e.g., by using the prediction processing unit 260) and mode selection (e.g., by using the mode selection unit 262) performed by an example of the encoder 20 are described in detail below.
[0120] As described above, the encoder 20 is configured to determine or select the best or most optimal prediction mode from a (predetermined) set of prediction modes. The set of prediction modes may include, for example, an intra-prediction mode and / or an inter-prediction mode.
[0121] The intra-predictive mode set may include 35 different intra-predictive modes, e.g., omnidirectional modes such as DC (or average) mode and planar mode, or directional modes as defined in H.265, or it may include 67 different intra-predictive modes, e.g., omnidirectional modes such as DC (or average) mode and planar mode, or developing directional modes as defined in H.266.
[0122] In possible implementations, the interprediction mode set depends on available reference pictures (e.g., at least a portion of the decoded pictures stored in DBP230) and other interprediction parameters, for example, whether the entire reference picture is used or only a portion of the reference picture is used, whether a search window area surrounding the current block region is searched for the best-matching reference block, and / or whether sample interpolation such as half-sample interpolation and / or quarter-sample interpolation is applied. The interprediction mode set may include, for example, an Advanced Motion Vector Prediction (AMVP) mode and a merge mode. In a particular implementation, the interprediction mode set may include an improved control point-based AMVP mode and a control point-based merge mode as described in this embodiment of the application. In one example, the intraprediction unit 254 may be configured to perform any combination of the interprediction techniques described below.
[0123] In addition to the prediction mode described above, a skip mode and / or a direct mode may also be applied in this embodiment of the present application.
[0124] The prediction processing unit 260 may be further configured to partition the picture block 203 into smaller block partitions or subblocks by iteratively using, for example, quad-tree (QT) partitioning, binary-tree (BT) partitioning, triple-tree (TT) partitioning, or any combination thereof, and to perform predictions for each of the block partitions or subblocks. Mode selection includes selecting the tree structure of the partitioned picture block 203 and selecting a prediction mode to be applied to each of the block partitions or subblocks.
[0125] The interpretation unit 244 may include a motion estimation (ME) unit (not shown in Figure 2) and a motion compensation (MC) unit (not shown in Figure 2). The motion estimation unit performs motion estimation using picture block 203 (the current picture block 203 of the current picture 201) and decoded picture 231, or at least one or more previously reconstructed blocks, e.g., a previous decoded picture. 31 It is configured to receive or acquire one or more other reconfigured blocks that are different from the current one. For example, a video sequence may include the current picture and the previous decoded picture 31. In other words, the current picture and the previous decoded picture 31 may be part of a sequence of pictures that form a video sequence or a picture sequence.
[0126] For example, the encoder 20 may be configured to select a reference block from multiple reference blocks of the same picture or different pictures within multiple other pictures, and to provide the motion estimation unit (not shown in Figure 2) as an interpretation parameter an offset (spatial offset) between the position (XY coordinates) of the reference picture and / or reference block and the position of the current block. This offset is also called a motion vector (MV).
[0127] The motion compensation unit is configured to obtain inter-prediction parameters and perform inter-prediction based on or using the inter-prediction parameters to obtain inter-prediction blocks 245. Motion compensation performed by the motion compensation unit (not shown in Figure 2) may include fetching or generating (or performing interpolation with subsample precision) predicted blocks based on motion / block vectors determined by motion estimation. During interpolation filtering, additional samples may be generated from known samples, thereby potentially increasing the amount of candidate predicted blocks that may be used to encode picture blocks. Once the motion vector to be used for the current picture block's PU is received, the motion compensation unit 246 may locate the predicted block pointed to by the motion vector in the reference picture list. The motion compensation unit 246 may further generate syntax elements associated with blocks and video slices so that the video decoder 30 uses the syntax elements when decoding picture blocks in video slices.
[0128] Specifically, the interprediction unit 244 may transmit a syntax element to the entropy coding unit 270, the syntax element containing interprediction parameters (such as instruction information for selecting the interprediction mode to be used for predicting the current block after multiple interprediction modes have been traversed). In possible application scenarios where there is only one interprediction mode, the interprediction parameters may alternatively not be carried within the syntax element. In this case, the decoder 30 may perform decoding directly in the default prediction mode. It can be understood that the interprediction unit 244 may be configured to perform any combination of interprediction techniques.
[0129] The intra-prediction unit 254 is configured to receive, for example, the picture block 203 (the current picture block) of the same picture and one or more previously reconstructed blocks, such as reconstructed adjacent blocks, in order to perform intra-prediction. For example, the encoder 20 may be configured to select an intra-prediction mode from a plurality of (predetermined) intra-prediction modes.
[0130] Embodiments of the encoder 20 may be configured to select an intra-prediction mode based on optimization criteria, for example, based on minimum residual (e.g., an intra-prediction mode that provides a prediction block 255 most similar to the current picture block 203) or minimum rate distortion.
[0131] The intra-prediction unit 254 is further configured to determine an intra-prediction block 255 based on the intra-prediction parameters of the selected intra-prediction mode. In any case, after selecting the intra-prediction mode to be used for the block, the intra-prediction unit 254 is further configured to provide the intra-prediction parameters to the entropy coding unit 270, i.e., to provide information indicating the selected intra-prediction mode to be used for the block. In one example, the intra-prediction unit 254 may be configured to perform any combination of the following intra-prediction techniques.
[0132] Specifically, the intra-prediction unit 254 may transmit a syntax element to the entropy coding unit 270, the syntax element including intra-prediction parameters (such as instruction information for selecting the intra-prediction mode to be used for predicting the current block after multiple intra-prediction modes have been traversed). In possible application scenarios where there is only one intra-prediction mode, the intra-prediction parameters may alternatively not be carried within the syntax element. In this case, the decoder 30 may perform decoding directly in the default prediction mode.
[0133] The entropy coding unit 270, for example, encodes the bitstory MuTo obtain encoded picture data 21 that can be output by using output 272 in form, an entropy coding algorithm or scheme (e.g., variable length coding, VLC, context adaptive VLC, CAVLC, arithmetic coding, context adaptive binary arithmetic coding, CABAC, syntax-based context-adaptive binary arithmetic coding, SBAC, probability interval partitioning entropy, PIPE, or another entropy coding method or technique) is configured to be applied to one or more (or none) of the quantization residual coefficients 209, inter-prediction parameters, intra-prediction parameters, and / or loop filter parameters. The encoded bitstream may be transmitted to the video decoder 30 or archived by the video decoder 30 for later transmission or retrieval. The entropy coding unit 270 may be further configured to perform entropy coding on another syntax element of the current video slice being coded.
[0134] Another structural variation of the video encoder 20 may be configured to encode a video stream. For example, a non-conversion-based encoder 20 may directly quantize the residual signal for some blocks or frames without a conversion processing unit 206. In another implementation, the encoder 20 may have a quantization unit 208 and an inverse quantization unit 210 coupled into a single unit.
[0135] Specifically, in this embodiment of the present application, the encoder 20 may be configured to implement the interpretation method described in the following embodiments.
[0136] It should be understood that another structural modification of the video encoder 20 may be configured to encode a video stream. For example, for some picture blocks or picture frames, the video encoder 20 may directly quantize the residual signal without requiring processing by the conversion processing unit 206, and correspondingly without requiring processing by the inverse conversion processing unit 212. Alternatively, for some picture blocks or picture frames, the video encoder 20 may not generate residual data, and correspondingly, the conversion processing unit 206, quantization unit 208, inverse quantization unit 210, and inverse conversion processing unit 212 may not perform any processing. Alternatively, the video encoder 20 may directly store the reconstructed picture blocks as reference blocks without requiring processing by the filter 220. Alternatively, the quantization unit 208 and inverse quantization unit 210 within the video encoder 20 may be combined with each other. The loop filter 220 is optional, and in the case of lossless compression coding, the transform processing unit 206, quantization unit 208, inverse quantization unit 210, and inverse transform processing unit 212 are optional. It should be understood that in different application scenarios, the inter-prediction unit 244 and intra-prediction unit 254 may be used selectively.
[0137] Figure 3 is a schematic / conceptual block diagram of an example of a decoder 30 configured to implement an embodiment of the present application. The video decoder 30 is configured to receive encoded picture data (e.g., encoded bitstream) 21 encoded by an encoder 20 or the like in order to obtain a decoded picture 231. In the decoding process, the video decoder 30 receives video data from the video encoder 20, and receives, for example, an encoded video bitstream and associated syntax elements that represent picture blocks of an encoded video slice.
[0138] In the example in Figure 3, the decoder 30 includes an entropy decoding unit 304, an inverse quantization unit 310, an inverse transformation unit 312, a reconstruction unit 314 (e.g., an adder 314), a buffer 316, a loop filter 320, a decoded picture buffer 330, and a prediction unit 360. The prediction unit 360 may include an inter-prediction unit 344, an intra-prediction unit 354, and a mode selection unit 362. In some examples, the video decoder 30 may perform a decoding traversal that is approximately the inverse of the coding traversal described with reference to the video encoder 20 in Figure 2.
[0139] The entropy decoding unit 304 is configured to perform entropy decoding on the encoded picture data 21 to obtain one or all of the following: quantization coefficients 309, decoding coding parameters (not shown in Figure 3), such as inter-prediction parameters, intra-prediction parameters, loop filter parameters, and / or other syntax elements (decoded). The entropy decoding unit 304 is further configured to transfer the inter-prediction parameters, intra-prediction parameters, and / or other syntax elements to the prediction processing unit 360. The video decoder 30 may receive syntax elements at the video slice level and / or at the video block level.
[0140] The inverse quantization unit 310 may have the same function as the inverse quantization unit 110, the inverse transformation processing unit 312 may have the same function as the inverse transformation processing unit 212, the reconstruction unit 314 may have the same function as the reconstruction unit 214, the buffer 316 may have the same function as the buffer 216, the loop filter 320 may have the same function as the loop filter 220, and the decoded picture buffer 330 may have the same function as the decoded picture buffer 230.
[0141] The prediction processing unit 360 may include an inter-prediction unit 344 and an intra-prediction unit 354. The inter-prediction unit 344 may have functions similar to those of the inter-prediction unit 244, and the intra-prediction unit 354 may have functions similar to those of the intra-prediction unit 254. The prediction processing unit 360 is typically configured to perform block prediction and / or obtain prediction blocks 365 from encoded data 21, and to receive or (explicitly or implicitly) information regarding prediction-related parameters and / or selected prediction modes from, for example, the entropy decoding unit 304.
[0142] If a video slice is encoded as an intra-coded (I) slice, the intra-prediction unit 354 of the prediction processing unit 360 is configured to generate a prediction block 365 to be used for the picture block of the current video slice, based on the signaled intra-prediction mode and data from a previous decoded block of the current frame or picture. If a video frame is encoded as an intercoded (i.e., B or P) slice, the inter-prediction unit 344 (e.g., a motion compensation unit) of the prediction processing unit 360 is configured to generate a prediction block 365 to be used for the video block of the current video slice, based on a motion vector received from the entropy decoding unit 304 and another syntax element. For inter-prediction, the prediction block may be generated from one of the reference pictures in a single reference picture list. The video decoder 30 may construct the reference frame lists of List 0 and List 1 by using default construction techniques based on the reference pictures stored in the DPB 330.
[0143] The prediction processing unit 360 is configured to determine prediction information to be used for the video block of the current video slice by analyzing motion vectors and other syntax elements, and to use the prediction information to generate prediction blocks to be used for the current video block being decoded. For example, the prediction processing unit 360 uses several received syntax elements to determine, in order to decode the video block of the current video slice, the prediction mode used to encode the video block of the video slice (e.g., intra or interpredict), the interpredict slice type (e.g., B slice, P slice, or GPB slice), configuration information of one or more pictures in the reference picture list used for the slice, the motion vector of each intercoded video block used for the slice, the interpredict state of each intercoded video block used for the slice, and other information. In another example of this disclosure, the syntax elements received by the video decoder 30 from the bitstream include syntax elements in one or more of the following: an adaptive parameter set (APS), a sequence parameter set (SPS), a picture parameter set (PPS), or a slice header.
[0144] The inverse quantization unit 310 may be provided in the bitstream and configured to perform inverse quantization (i.e., dequantization) on the quantization transformation coefficients decoded by the entropy decoding unit 304. The inverse quantization process may include using quantization parameters calculated by the video encoder 20 for each video block in the video slice to determine the degree of quantization to be applied and the degree of inverse quantization to be applied.
[0145] The inverse transformation processing unit 312 is configured to apply an inverse transformation (e.g., inverse DCT, inverse integer transformation, or a conceptually similar inverse transformation process) to the transformation coefficients in order to generate residual blocks in the sample domain.
[0146] The reconstruction unit 314 (e.g., adder 314) is configured to add the inverse transform block 313 (i.e., the reconstructed residual block 313) to the prediction block 365 in order to obtain the reconstructed block 315 in the sample domain by, for example, adding the sample values of the reconstructed residual block 313 to the sample values of the prediction block 365.
[0147] The loop filter unit 320 (either within or after the encoding loop) is configured to filter the reconstructed block 315 to obtain a filtered block 321 in order to perform a smooth sample transformation or to improve video quality. In one example, the loop filter unit 320 may be configured to perform any combination of the following filtering techniques. The loop filter unit 320 is intended to represent one or more loop filters, such as a deblocking filter, a sample-adaptive offset (SAO) filter, or another filter such as a bilateral filter, an adaptive loop filter (ALF), a sharpening or smoothing filter, or a co-filter. Although the loop filter unit 320 is shown as an in-loop filter in Figure 3, the loop filter unit 320 may be implemented as a post-loop filter in other configurations.
[0148] A decoded video block 321 within a given frame or picture is then stored in a decoded picture buffer 330 that stores a reference picture to be used for subsequent motion compensation.
[0149] The decoder 30 is configured to output the decoded picture 31 by using output 332, etc., in order to present the decoded picture 31 to the user or to provide the decoded picture 31 for the user to view.
[0150] Another variation of the video decoder 30 may be configured to decode a compressed bitstream. For example, the decoder 30 may generate an output video stream without the loop filter unit 320. For example, the non-conversion-based decoder 30 may directly dequantize the residual signal for some blocks or frames without the inverse conversion processing unit 312. In another implementation, the video decoder 30 may have an inverse quantization unit 310 and an inverse conversion processing unit 312 coupled into a single unit.
[0151] Specifically, in this embodiment of the present invention, the decoder 30 is configured to implement the interpretation method described in the following embodiments.
[0152] It should be understood that another structural modification of the video decoder 30 may be configured to decode an encoded video bitstream. For example, the video decoder 30 may generate an output video stream without the filter 320 needing to perform any processing. Alternatively, for some picture blocks or picture frames, the entropy decoding unit 304 of the video decoder 30 does not obtain quantization coefficients by decoding, and correspondingly, the inverse quantization unit 310 and the inverse transformation processing unit 312 do not need to perform any processing. The loop filter 320 is optional, and in the case of lossless compression, the inverse quantization unit 310 and the inverse transformation processing unit 312 are optional. It should be understood that in different application scenarios, the inter-prediction unit and the intra-prediction unit may be used selectively.
[0153] In the encoder 20 and decoder 30 of this application, it should be understood that the processing result of a procedure may be further processed and then output to the next procedure. For example, after a procedure such as interpolation filtering, motion vector derivation, or loop filtering, further operations such as clipping or shifting may be performed on the processing result of the corresponding procedure.
[0154] For example, the motion vectors of the control points of the current picture block, or the motion vectors of the subblocks of the current picture block derived from the motion vectors of adjacent affine coding blocks, may be further processed. This is not limited to the present application. For example, the values of the motion vectors are restricted to be within a certain bit width range. Assuming that the allowable bit width of the motion vector is bitDepth, the values of the motion vectors are in the range of -2^(bitDepth-1) to 2^(bitDepth-1)-1, where the symbol "^" represents exponentiation. If bitDepth is 16, the values are in the range of -32768 to 32767. If bitDepth is 18, the values are in the range of -131072 to 131071. In another example, the values of the motion vectors (e.g., the motion vectors MV of four 4x4 subblocks in one 8x8 picture block) are restricted so that the maximum difference between the integer parts of the MVs of the four 4x4 subblocks does not exceed N samples, for example, not exceeding 1 sample.
[0155] The following two methods may be used to restrict the values of the motion vector to be within a specific bit width range.
[0156] Method 1: The most significant bit of the motion vector overflow is removed. ux = (vx + 2 bitDepth )%2 bitDepth vx=(ux≧2 bitDepth-1 )?(ux-2 bitDepth ):ux uy=(vy+2 bitDepth )%2 bitDepth vy=(uy≧2bitDepth-1 )?(uy - 2 bitDepth ): uy
[0157] Here, vx is the horizontal component of the motion vector of a picture block or a sub - block of a picture block, vy is the vertical component of the motion vector of a picture block or a sub - block of a picture block, ux and uy are intermediate values, and bitDepth represents the bit width.
[0158] For example, the value of vx is - 32769, and is 32767 obtained according to the above formula. The value is stored in the computer in the form of two's complement code. The two's complement code of - 32769 is 1,0111,1111,1111,1111 (17 bits). When an overflow occurs, the computer discards the most significant bit. Therefore, the value of vx is 0111,1111,1111,1111, that is, 32767, which is consistent with the result obtained according to the formula.
[0159] Method 2: As shown in the following formula, Clipping is performed on the motion vector. vx = Clip3(-2 bitDepth-1 , 2 bitDepth-1 , - 1, vx) vy = Clip3(-2 bitDepth-1 , 2 bitDepth-1 , - 1, vy)
[0160] Here, vx is the horizontal component of the motion vector of a picture block or a sub - block of a picture block, vy is the vertical component of the motion vector of a picture block or a sub - block of a picture block, x, y, and z correspond to the three input values of the MV clipping process Clip3, and Clip3 represents clipping the value of z to the range [x, y]. [[ID=三十三]] [[ID=三十四]]
Number
[0161] [[ID=四十]] Figure 4 is a schematic diagram of a video coding device 400 (e.g., a video encoding device 400 or a video decoding device 400) according to an embodiment of this application. The video coding device 400 is suitable for implementing the embodiments described herein. In one embodiment, the video coding device 400 may be a video decoder (e.g., the decoder 30 in Figure 1A) or a video encoder (e.g., the encoder 20 in Figure 1A). In another embodiment, the video coding device 400 may be one or more components of the decoder 30 in Figure 1A or the encoder 20 in Figure 1A.
[0162] The video coding device 400 includes an input port 410 and a receiver unit (Rx) 420 configured to receive data, a processor, logic unit, or central processing unit (CPU) 430 configured to process data, a transmitter unit (Tx) 440 and an output port 450 configured to transmit data, and a memory 460 configured to store data. The video coding device 400 may further include optical-electrical components and electrical-optical (EO) components coupled to the input port 410, the receiver unit 420, the transmitter unit 440, and the output port 450 for outputting or inputting optical or electrical signals.
[0163] The processor 430 is implemented by hardware and software. The processor 430 may be implemented as one or more CPU chips, cores (e.g., a multi-core processor), FPGAs, ASICs, and DSPs. The processor 430 communicates with input ports 410, a receiver unit 420, a transmitter unit 440, output ports 450, and memory 460. The processor 430 includes a coding module 470 (e.g., an encoding module 470 or a decoding module 470). The encoding / decoding module 470 implements embodiments disclosed herein to implement the saturation block prediction method provided in embodiments of this application. For example, the encoding / decoding module 470 performs, processes, or provides various coding operations. Thus, the encoding / decoding module 470 substantially extends the functionality of the video coding device 400 and affects the switching of the video coding device 400 between different states. Alternatively, the encoding / decoding module 470 is implemented as instructions stored in memory 460 and executed by the processor 430.
[0164] Memory 460 may include one or more disks, tape devices, and solid-state drives, and may be used as overflow data storage devices to store programs when these programs are selected for execution, and to store instructions and data read during program execution. Memory 460 may be volatile and / or non-volatile, and may be read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random access memory (SRAM).
[0165] Figure 5 is a simplified block diagram of a device 500, which may be used as either or both of the source device 12 and destination device 14 in Figure 1A, according to an exemplary embodiment. The device 500 may implement the technology of the present application. In other words, Figure 5 is a schematic block diagram of an implementation of an encoding device or decoding device (abbreviated as coding device 500) according to one embodiment of the present application. The coding device 500 may include a processor 510, a memory 530, and a bus system 550. The processor and memory are connected via the bus system. The memory is configured to store instructions. The processor is configured to execute instructions stored in the memory. The memory of the coding device stores program code, and the processor of the coding device may invoke the program code stored in the memory to perform the video encoding or decoding method described in the present application. For the sake of avoiding repetition, further details are not described here.
[0166] In this embodiment of the present application, the processor 510 may be a Central Processing Unit (CPU). Alternatively, the processor 510 may be another general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or another programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, etc. The general-purpose processor may be a microprocessor, any conventional processor, etc.
[0167] Memory 530 may include read-only memory (ROM) or random access memory (RAM). Any other suitable type of storage device may be used as memory 530. Memory 530 may include code and data 531 accessed by the processor 510 via the bus 550. Memory 530 may further include an operating system 533 and an application program 535. The application program 535 includes at least one program that enables the processor 510 to perform the video coding or decoding method described in this application (in particular, the interpretation method described in this application). For example, the application program 535 may include applications 1 to N, further including a video coding or decoding application (abbreviated as a video coding application) for performing the video coding or decoding method described in this application.
[0168] In addition to the data bus, the bus system 550 may further include a power bus, a control bus, a status signal bus, and so on. However, for clarity, the various types of buses shown in the figure are all represented as the bus system 550.
[0169] Optionally, the coding device 500 may further include one or more output devices, such as a display 570. In one example, the display 570 may be a touch-sensitive display combining a display with a touch-sensing element that can operate to sense touch input. The display 570 may be connected to the processor 510 via the bus 550.
[0170] The following describes in detail the solution in the embodiment of this application.
[0171] Video coding primarily involves processes such as intra-prediction, inter-prediction, transformation, quantization, entropy encoding, and in-loop filtering (mainly de-blocking filtering). After the picture is partitioned into coding blocks, intra-prediction or inter-prediction is performed. Then, after the residuals are obtained, transformation and quantization are performed. Finally, entropy encoding is performed, and the bitstream is output. Here, the coding block is an M×N array containing samples (M may or may not be equal to N). In addition, the sample value at each sample position is known.
[0172] Intra prediction is the process of predicting the sample values of samples in the current coding block by using the sample values of samples in the reconstructed region of the current picture.
[0173] Interpretation involves searching for a reconstructed picture for a matching reference block within the current coding block, obtaining motion information for the current coding block, and then calculating predictive information or predictors (information and values are not distinguished below) for the sample values of the samples within the current coding block based on the motion information. The process of calculating motion information is called motion estimation (ME), and the process of calculating predictors for the sample values of the samples within the current coding block is called motion compensation (MC).
[0174] It should be noted that the motion information of a current coding block includes predicted direction information (generally forward, backward, or bidirectional), two motion vectors (MV) pointing to a reference block, and information indicating the picture in which the reference block is located (generally indicated as a reference index).
[0175] Forward prediction is the process of selecting a reference picture from a forward reference picture set for the current coding block in order to obtain a reference block. Backward prediction is the process of selecting a reference picture from a backward reference picture set for the current coding block in order to obtain a reference block. Bidirectional prediction is the process of selecting a reference picture from a forward reference picture set and then selecting a reference picture from a backward reference picture set in order to obtain a reference block. When bidirectional prediction is performed, there are two reference blocks for the current coding block. Each reference block must be indicated by using a motion vector and a reference frame index. Then, the predictor of the sample value of a sample in the current block is determined based on the sample values of the sample in the two reference blocks.
[0176] During motion estimation, multiple reference blocks must be tried for the current coding block in the reference picture, and the specific reference block ultimately used for prediction is determined by rate-distortion optimization (RDO) or other methods.
[0177] After prediction information is obtained by intra-prediction or inter-prediction, residual information is obtained by subtracting the corresponding prediction information from the sample value of the sample in the current coding block. The residual information is then transformed using a method such as the discrete cosine transform (DCT), and the bitstream is obtained by quantization and entropy coding. After the prediction signal is combined with the reconstructed residual signal, filtering must be performed to obtain the reconstructed signal. The reconstructed signal is used as a reference signal for subsequent coding.
[0178] Decoding is the reverse process of encoding. For example, residual information is first obtained through entropy decoding, inverse quantization, and inverse transform, and the bitstream is decoded to determine whether intra-prediction or inter-prediction is performed for the current coding block. If intra-prediction is performed, prediction information is constructed based on the sample values of samples in the reconstructed region around the current coding block by using the intra-prediction method. If inter-prediction is performed, motion information must be obtained by analysis, and based on the motion information obtained by the analysis, a reference block is determined in the reconstructed picture, and the sample values of samples in the block are used as prediction information. This process is called motion compensation (MC). By combining the prediction information and residual information and performing a filtering operation, the reconstructed information can be obtained.
[0179] In HEVC, two interpretation modes are used: Advanced Motion Vector Prediction (AMVP) mode and Merge mode.
[0180] In AMVP mode, a list of candidate motion vectors is constructed based on the motion information of spatially or temporally adjacent coding blocks to the current coding block. The optimal motion vector is then determined from this list and used as the Motion Vector Predictor (MVP) for the current coding block. The rate distortion cost is calculated according to the equation J = SAD + λ, where J is the rate distortion cost (RD Cost), SAD is the sum of absolute differences between the original sample value and the predicted sample value obtained by motion estimation using the candidate motion vector predictor, R is the bit rate, and λ is the Lagrange multiplier. The encoder transfers the index value of the selected motion vector predictor from the candidate motion vector list and the index value of the reference frame to the decoder. Furthermore, a motion search is performed in the neighborhood around the MVP to obtain the actual motion vector of the current coding block. The encoder then transfers the motion vector difference between the MVP and the actual motion vector to the decoder.
[0181] In Merge mode, a list of candidate motion information is constructed based on the motion information of spatially or temporally adjacent coding blocks to the current coding block. Then, the optimal motion information within the candidate motion information list is determined based on the rate distortion cost and used as the motion information for the current coding block. Next, the index value of the optimal motion information position within the candidate motion information list (hereinafter referred to as the merge index) is transferred to the decoder. Figure 6 shows the spatial and temporal candidate motion information for the current coding block. The spatial candidate motion information comes from five spatially adjacent blocks (A0, A1, B0, B1, and B2). If adjacent blocks are unavailable, or if inter-prediction mode is used, the motion information of adjacent blocks is not added to the candidate motion information list. The temporal candidate motion information for the current coding block is obtained by scaling the MV of the block at the corresponding position in the reference frame based on the picture order count (POC) of the reference frame and the current frame. It is first determined whether the block at position T in the reference frame is available. If the block is unavailable, the block at position C is selected.
[0182] In interpretation in HEVC, all samples within a coding block have the same motion information, and then motion compensation is performed based on the motion information to obtain predictors for the samples within the coding block.
[0183] A video sequence typically contains a certain number of pictures, usually called frames. Adjacent pictures are usually similar, i.e., they have a lot of redundancy. Motion compensation is performed to increase the compression ratio by eliminating redundancy between adjacent frames. Motion compensation is a method for describing the difference between adjacent frames (where "adjacent" means that two frames are adjacent in terms of coding relations, but not necessarily adjacent in terms of playback order) and is part of the interpretation process. Before motion compensation is performed, the motion information of the coding blocks is obtained by motion estimation or bitstream decoding. The motion information includes (1) the prediction direction of a coding block, including forward prediction, backward prediction, and bidirectional prediction, where forward prediction indicates that the coding block is predicted by using the previous coded frame, backward prediction indicates that the coding block is predicted by using the subsequent coded frame, and bidirectional prediction indicates that the coding block is predicted by using both the forward coded frame and the backward coded frame; (2) the reference frame index of the coding block, indicating the frame in which the reference block of the current coding block is located; and (3) the motion vector MV of the coding block, indicating the motion displacement of the coding block relative to the reference block, where MV is a horizontal component (MV) that indicates the motion displacement of the coding block relative to the reference block in the horizontal direction and the motion displacement of the coding block relative to the reference block in the vertical direction, respectively. x (shown as) and vertical component (MV yIt mainly includes motion vectors (MVs), including (shown as). When forward or backward prediction is performed on a coding block, there is only one MV. When bidirectional prediction is performed on a coding block, there are two MVs. Figure 7 illustrates the above explanation of motion information. In Figure 7, and in the following explanation of motion and prediction information, 0 indicates "forward" and 1 indicates "backward". For example, Ref0 represents a forward reference frame, Ref1 represents a backward reference frame, MV0 represents a forward motion vector, and MV1 represents a backward motion vector. A, B, and C represent the forward reference block, the current coding block, and the backward reference block, respectively. Cur shows the current coding frame, and the dashed line shows the movement trajectory of B. Motion compensation is the process of finding reference blocks based on motion information and processing the reference blocks to obtain predicted blocks for the coding blocks.
[0184] The basic motion compensation process for forward prediction is as follows: As shown in Figure 7, the current coding block is block B, and the width and height of B are W and H, respectively. In this case, based on the motion information, the forward reference frame of the current coding block B is frame Ref0, and the forward motion vector of the current coding block B is MV0 = (MV0 x MV0 yIt can be seen that ). When coding block B in the Cur frame is coded, first, based on the coordinates (i,j) of the upper left corner of B in the Cur frame, the same coordinate point may be found in the Ref0 frame, and based on the width and height of block B, block B' in Ref0 may be obtained. Block B' is then moved to block A based on the MV0 of B'. Finally, interpolation is performed on block A to obtain the predicted block for the current coding block B. The sample value of each sample in the predicted block for the current coding block B is called the predictor of the corresponding sample in block B. The motion compensation process for backward prediction is the same as the motion compensation process for forward prediction, the only difference being the direction of reference. It should be noted that the predicted blocks obtained by backward prediction motion compensation and forward prediction motion compensation are called forward prediction blocks and backward prediction blocks, respectively. If bidirectional prediction is not performed on the coding block, the forward prediction block and backward prediction block obtained are the predicted blocks for the current coding block.
[0185] For bidirectional prediction, based on motion information, forward prediction blocks and backward prediction blocks are acquired during forward prediction motion compensation and backward prediction motion compensation, respectively. Then, the prediction block of coding block B is acquired by weighted prediction and bidirectional optical flow (BIO or BDOF) on the sample values at the same location within the forward prediction block and backward prediction block.
[0186] In the weighted prediction method, when the predictor for the current coding block is calculated, only weighted addition must be performed sequentially on the sample values of the forward prediction block and the isotropic sample values of the backward prediction block; that is, PredB(i,j)=ω0PredA(i,j)+ω1PredC(i,j) (1) That is the case.
[0187] In equation (1), PredB(i,j), PredA(i,j), and PredC(i,j) are the predictors for the prediction block, the forward prediction block, and the backward prediction block of the current coding block at coordinate (i,j), respectively, and ω0 and ω1 (0≦ω0≦1, 0≦ω1≦1, and ω0+ω1=1) are weighting coefficients, the values of ω0 and ω1 may differ depending on the encoder. In general, both ω0 and ω1 are 1 / 2.
[0188] Figure 8 shows an example of obtaining the predicted block of the current coding block using weighted addition. In Figure 8, PredB, PredA, and PredC are the predicted block, forward predicted block, and backward predicted block of the current coding block, respectively, and have a size of 4x4. The values of the small blocks within the predicted block are point predictors, and a coordinate system with the upper left corner as the origin is established for PredB, PredA, and PredC. For example, the predictor of PredB at coordinate (0,0) is: PredB(0,0)=ω0PredA(0,0)+ω1PredC(0,0) =ω0a 0,0 +ω1c 0,0 That is the case.
[0189] The PredB predictor at coordinate (0,1) is: PredB(0,1)=ω0PredA(0,1)+ω1PredC(0,1) =ω0a 0,1 +ω1c 0,1 That is the case.
[0190] Other points will be calculated sequentially, and I will not explain the details.
[0191] While bidirectional weighted prediction techniques are computationally simple, such block-based prediction and compensation methods can be found to be very coarse, achieving only insufficient prediction effects, especially for pictures with complex textures, and resulting in low compression efficiency.
[0192] In BIO, bidirectional predictive motion compensation is performed on the current CU to obtain forward and backward predicted blocks. Then, refined motion vectors for each 4x4 subblock within the current CU are derived based on the forward and backward predictors. Finally, compensation is performed again on each sample within the current coding block to ultimately obtain the predicted blocks of the current CU.
[0193] Refinement motion vector of each 4x4 subblock (v x ,v y ) is obtained by applying BIO to a 6x6 window Ω around the subblock to minimize the L0 and L1 predictors. Specifically, (v x ,v y ) is derived according to the formula.
number
[0194] Here,
number
number
[0195] S1, S2, S3, S5, and S6 are calculated according to the following formulas.
number
[0196] In the formula
number
[0197] Here, I (k) (i,j) is the predictor of the current CU sample position (i,j) (where k is equal to 0 or 1, 0 indicates "forward", 1 indicates "backward", and so on),
number
number
number
[0198] After the refined motion vector is obtained according to equation (2), the final predictor for each sample in the current block is determined according to the following equation.
number
[0199] shift and o offset 15-BD and 1<<(14-BD)+2·(1<<13), and rnd(.) is the rounding function (rounding).
[0200] The refined motion vectors of the 4x4 subblock are forward and backward predictors I (k)(x,y) and the forward, backward, horizontal, and vertical slopes of the 6x6 region where the 4x4 subblock is located.
number
number
[0201] To reduce the complexity of BIO, conventional techniques perform special processing at the CU boundary.
[0202] First, predictors for the W*H region are obtained by using an 8-tap filter, and then an extension is performed around them by only one row and one column. Predictors for the extended region are obtained by using a bilinear filter to obtain predicted sample values for the (W+2)*(H+2) region.
[0203] Next, the gradient of the W*H region may be calculated based on the predicted sample values of the (W+2)*(H+2) region and equation (5).
[0204] Finally, according to the padding method, expansion is performed on the gradient of the surrounding W*H region to obtain the gradient of the (W+2)*(H+2) region, and expansion is performed on the predictor of the surrounding W*H region to obtain the predictor of the (W+2)*(H+2) region. The padding is shown in Figure 9, i.e., the sample values of the edges are assigned to the expanded region.
[0205] The specific implementation process for BIO is as follows:
[0206] Step 1: Determine the current CU's movement information.
[0207] The current CU motion information may be determined by using merge mode, AMVP mode (see background explanation), or another mode, which is not limited herein.
[0208] It should be noted that other methods for determining motion information may also be applicable to this application. Details are not described herein.
[0209] Step 2: Determine whether the current CU meets the BIO usage requirements.
[0210] If bidirectional prediction is performed on the current CU, and the relationship between the forward-referenced frame number POC_L0, the backward-referenced frame number POC_L1, and the current frame number POC_Cur satisfies the following equation, then the current CU satisfies the BIO usage conditions. (POC_L0-POC_Cur)*(POC_L1-POC_Cur)<0
[0211] It should be noted that whether BIO is used may also be determined by determining whether the current CU size is greater than a preset threshold. For example, the current CU height. H The current width of CU is 8 or more. W BIO can only be used if the value is 8 or greater.
[0212] It should be noted that other conditions of use for BIO may also apply to this application. Details are not described herein.
[0213] If the current CU meets the BIO usage requirements, step 3 is performed; otherwise, motion compensation is performed in another manner.
[0214] Step 3: Calculate the forward and backward predictors for the current CU.
[0215] Forward and backward predictor I (k) Motion compensation is performed by using motion information to obtain (i,j), where i=-1..cuW and j=-1..cuH (a prediction matrix of (cuW+2)*(cuH+2) is obtained).
[0216] I obtained by performing interpolation using an 8-tap interpolation filter (k) At (i,j), i=0..cuW-1, and j=0..cuH-1, predictors at other locations (where an extension of one row and one column is performed) are obtained by performing interpolation using a bilinear interpolation filter.
[0217] It should be noted that the predictor for the extended region may also be obtained by using other methods, for example, by using an 8-tap interpolation filter, or by directly using a reference sample at an integer sample position. This is not limited herein.
[0218] It should be noted that in order to determine whether SAD is less than the threshold TH_CU, the SAD between the forward and backward predictors is calculated, and if SAD is less than the threshold TH_CU, BIO is not performed; otherwise, BIO is performed. Other determination methods may also be applied to this application, and their details are not described herein.
[0219] The formula for calculating SAD is as follows:
number
[0220] The threshold TH_CU may be set to (1<<(BD-8+shift))*cuW*cuH, and shift may be set to Max(2,14-BD).
[0221] Step 4: Calculate the horizontal gradient and the vertical gradient based on the forward predictor and the backward predictor of the current CU.
[0222] Horizontal gradient and vertical gradient
Number
Number
[0223] Step 5: Perform padding on the forward predictor and the backward predictor of the current CU, and the horizontal gradient and the vertical gradient.
[0224] I (k) (i, j), and
Number
Number
[0225] Step 6: Derive the refined motion vector for each 4×4 sub-block, and then perform weighting.
[0226] For each 4x4 subblock, vx and vy are obtained according to equation (2). Finally, weighting is performed according to equation (6) to obtain the predictor for each 4x4 subblock.
[0227] It should be noted that the SAD between the forward and backward predictors of each 4x4 subblock may be calculated to determine whether the SAD is less than the threshold TH_SCU. If the SAD is less than the threshold TH_SCU, a weighted average is performed directly; otherwise, vx and vy are obtained according to equation (2), and then weighting is performed according to equation (6). Other determination methods may also be applied to this application and are not described in detail herein. TU_SCU may be set to 1 << (BD - 3 + shift).
[0228] Virtual pipeline data units (VPDUs) are non-overlapping M×M luminance / N×M chromaticity processing units. For hardware decoders, consecutive VPDUs are processed simultaneously at different pipeline levels. Different VPDUs are processed simultaneously at different pipeline levels.
[0229] The VPDU partitioning principle is as follows:
[0230] (1) If a VPDU contains one or more CUs, the CUs are entirely contained within the VPDU.
[0231] (2) If a CU contains one or more VPDUs, the VPDUs are entirely contained within the CU.
[0232] In conventional technology, the size of the VPDU is 64 × 64. As shown in Figure 10, the dashed lines represent the boundaries of the VPDU, and the solid lines represent the boundaries of the CU. Figure 11 shows an invalid CU division.
[0233] If a CU contains multiple VPDUs, the hardware decoder divides the VPDUs into contiguous VPDUs for processing. For example, if the CU is 128x128 and the VPDUs are 64x64, four contiguous VPDUs will be processed.
[0234] The technical problem to be solved in this application is that when motion compensation is performed on a CU via BIO, the method for processing boundary samples of the CU differs from the method for processing internal samples of the CU. If there are VPDU partition boundaries within the CU, during BIO prediction, in order to ensure that the results of VPDU processing match the results of CU processing, the boundaries need to be processed in the same way as internal samples of the CU, which increases implementation complexity.
[0235] Referring to Figure 12, one embodiment of the present application provides an interpretation method. The method may be applied to an interpretation unit 244 in an encoder shown in Figure 2 or an interpretation unit 344 in a decoder shown in Figure 3. The method may be a bidirectional interpretation method and includes the following steps.
[0236] Step S101: Select a smaller width from the preset picture division width Width and the width of the picture block to be processed cuW, where the smaller width is shown as blkW and is used as the width of the first picture block and the preset picture division height Height Then, a smaller height is selected from the height cuH of the picture block to be processed, where the smaller height is denoted as blkH and used as the height of the first picture block.
[0237] When the method in this embodiment is applied to an encoder, when encoding a picture, the encoder divides the picture into picture blocks to be processed. In this step, the picture blocks to be processed are obtained, and then a smaller width blkW = min(Width, cuW) is selected, the smaller width blkW is used as the width of the first picture block, and a smaller height blkH = min( Height , cuH) is selected, and the smaller height blkH is used as the height of the first picture block.
[0238] When the method in this embodiment is applied to a decoder, the decoder receives a video bitstream from an encoder, and the video bitstream includes a picture block to be processed. In this step, the picture block to be processed is extracted from the video bitstream, and then a smaller width blkW = min(Width, cuW) is selected, and the smaller width blkW is used as the width of the first picture block, and the smaller height blkH = min( Height , cuH) is selected, and the smaller height blkH is used as the height of the first picture block.
[0239] The preset picture division width Width and the preset picture division height Height may each be equal to the width and height of the VPDU. Alternatively, the preset picture division width Width is a value such as 64, 32, or 16, and the preset picture division height Height is a value such as 64, 32, or 16. For example, Width = 64 and Height = 64, or Width = 32 and Height = 32, or Width = 16 and Height = 16, or Width = 64 and Height = 32, or Width = 32 and Height = 64, or Width = 64 and Height = 16, or Width = 16 and Height = 64, or Width = 32 and Height = 16, or Width = 16 and Height = 32.
[0240] Step 102: Based on the width blkW and height blkH of the first picture block, determine a plurality of first picture blocks within the picture block to be processed.
[0241] In a feasible implementation, the width and height of the picture block to be processed are the same as those of the first picture block, respectively; that is, the picture block to be processed contains only one first picture block. Obtaining the predictor of the first picture block is equivalent to obtaining the predictor of the picture block to be processed.
[0242] The predictor for any first picture block is obtained according to the operations in steps 103 to 107 below.
[0243] Step 103: Based on the motion information of the picture block to be processed, obtain the first predicted block of the first picture block, where the width of the first predicted block is greater than the width of the first picture block, and the height of the first predicted block is greater than the height of the first picture block.
[0244] The motion information of the picture block to be processed includes the motion information of a first picture block, and the motion information of the first picture block includes information such as a reference picture and motion vectors. In this embodiment, an optical flow-based bidirectional prediction method (i.e., the aforementioned BIO or BDOF-related techniques) is used for interpretation. Therefore, the motion information of the first picture block includes information such as a forward reference picture, a backward reference picture, a forward motion vector, and a backward motion vector.
[0245] When the method in this embodiment is applied to an encoder, the encoder may determine motion information of the picture block to be processed in merge mode, AMVP mode, or another mode, and the motion information of the picture block to be processed includes motion information of each first picture block within the picture block to be processed. In this step, the motion information of the picture block to be processed determined by the encoder is obtained, and the motion information of the first picture block is obtained from the motion information of the picture block to be processed.
[0246] When the method in this embodiment is applied to a decoder, the video bitstream received by the decoder from the encoder includes motion information of the picture block to be processed, and the motion information of the picture block to be processed includes motion information of each first picture block within the picture block to be processed. In this step, the motion information of the picture block to be processed is extracted from the video bitstream, and the motion information of the first picture block is obtained from the motion information of the picture block to be processed.
[0247] The first prediction block of the first picture block includes a first forward prediction block and a first backward prediction block. In this step, the first forward prediction block and the first backward prediction block of the first picture block may be obtained in the following steps (1) to (8). Steps (1) to (8) may be as follows.
[0248] (1) Based on the first position of the first picture block within the picture block to be processed and the motion information of the first picture block, the first forward region within the forward reference picture is determined, where the width of the first forward region is blkW+2 and the height of the first forward region is blkH+2.
[0249] For example, referring to Figure 13, the motion information of the first picture block B includes a forward reference picture Ref0, a backward reference picture Ref1, a forward motion vector MV0, and a backward motion vector MV1. The second forward region B11 is determined within the forward reference picture Ref0 based on the first position of the first picture block B, where the width of the second forward region B11 is blkW and the height of the second forward region B11 is blkH. The third forward region B12 is determined based on the forward motion vector MV0 and the position of the second forward region B11, where the width of the third forward region B12 is blkW and the height of the third forward region B12 is blkH. A first forward region A1 is determined which includes a third forward region B12, where the width of the first forward region A1 is blkW+2, the height of the first forward region A1 is blkH+2, and the center of the third forward region B12 coincides with the center of the first forward region A1.
[0250] (2) Determine whether the corner position of the first forward region matches the sample position in the forward reference picture. If the corner position of the first forward region matches the sample position in the forward reference picture, obtain the picture block in the first forward region from the forward reference picture so that it functions as the first forward prediction block of the first picture block. If the corner position of the first forward region does not match the sample position in the forward reference picture, perform step (3).
[0251] For example, referring to Figure 13, the upper left corner of the first front region A1 is used as an example. Assuming that the corner position of the upper left corner of the first front region A1 is (15,16) in the front reference picture Ref0, the corner position of the upper left corner coincides with the sample position in the front reference picture Ref0, and the sample position in the front reference picture Ref0 is (15,16). In another example, assuming that the corner position of the upper left corner of the first front region A1 is (15.3,16.2) in the front reference picture Ref0, the corner position of the upper left corner does not coincide with the sample position in the front reference picture Ref0, that is, there is no sample at the position (15.3,16.2) in the front reference picture Ref0.
[0252] (3) Determine the sample closest to the corner position of the first forward region in the forward reference picture, and determine the fourth forward region by using the sample as the corner, where the width of the fourth forward region is blkW+2 and the height of the fourth forward region is blkH+2.
[0253] For any corner position of the first forward region, assume that the position of the upper-left corner of the first forward region is used as an example. The sample closest to the upper-left corner position in the forward reference picture is determined, and by using this sample as the upper-left corner, the fourth forward region is determined. The width of the fourth forward region is blkW+2, and the height of the fourth forward region is blkH+2.
[0254] For example, referring to Figure 13, the corner position of the upper left corner of the first forward region A1 is (15.3, 16.2), and the position of the sample closest to the corner position (15.3, 16.2) is determined to be (15, 16) in the forward reference picture Ref0. By using the sample at position (15, 16) as the upper left corner, the fourth forward region A2 is determined. The width of the fourth forward region A2 is blkW+2, and the height of the fourth forward region A2 is blkH+2.
[0255] (4) Determine a fifth forward region that includes a fourth forward region, where the center of the fourth forward region coincides with the center of the fifth forward region, the width of the fifth forward region is blkW+n+1, and the height of the fifth forward region is blkH+n+1. Obtain a picture block in the fifth forward region from the forward reference picture, and perform interpolation filling on the picture block by using an interpolation filter to obtain the first forward prediction block of the first picture block, where the width of the first forward prediction block is blkW+2, the height of the first forward prediction block is blkH+2, and n is the number of taps in the interpolation filter.
[0256] For example, an 8-tap interpolation filter is used as an example. Referring to Figure 13, a fifth forward region A3 is determined, which includes the fourth forward region A2. The center of the fourth forward region A2 coincides with the center of the fifth forward region A3, the width of the fifth forward region A3 is blkW+9, and the height of the fifth forward region A3 is blkH+9. The picture block in the fifth forward region A3 is obtained from the forward reference picture Ref0, and interpolation filtering is performed on the picture block by using an interpolation filter to obtain the first forward prediction block of the first picture block B. The width of the first forward prediction block is blkW+2, and the height of the first forward prediction block is blkH+2.
[0257] (5) Based on the first position and motion information of the first picture block, a first rear region in the rear reference picture is determined, where the width of the first rear region is blkW+2 and the height of the first rear region is blkH+2.
[0258] For example, referring to Figure 13, a second rear region C11 is determined within the back reference picture Ref1 based on the first position of the first picture block B, where the width of the second rear region C11 is blkW and the height of the second rear region C11 is blkH. A third rear region C12 is determined based on the back motion vector MV1 and the position of the second rear region C11, where the width of the third rear region C12 is blkW and the height of the third rear region C12 is blkH. A first rear region D1 containing the third rear region C12 is determined, where the width of the first rear region D1 is blkW+2 and the height of the first rear region D1 is blkH+2, and the center of the third rear region C12 may coincide with the center of the first rear region D1.
[0259] (6) Determine whether the corner position of the first rear region matches the sample position in the rear reference picture. If the corner position of the first rear region matches the sample position in the rear reference picture, retrieve the picture block in the first rear region from the rear reference picture so that it functions as the first rear prediction block of the first picture block. If the corner position of the first rear region does not match the sample position in the rear reference picture, perform step (7).
[0260] For example, referring to Figure 13, the upper left corner of the first rear region A1 is used as an example. The corner position of the upper left corner of the first rear region D1 is the back reference picture Ref 1 Assuming the position is (5,6) within the background, the upper left corner position coincides with the sample position in the backreference picture Ref0, and the sample position in the backreference picture Ref0 is (5,6). In another example, assuming the upper left corner position of the first back region D1 is (5.3,6.2) within the backreference picture Ref0, the upper left corner position does not coincide with the sample position in the backreference picture Ref0, i.e., there is no sample at position (5.3,6.2) in the backreference picture Ref0.
[0261] (7) Determine the sample closest to the corner position of the first rear region in the back reference picture, and use the sample as a corner to determine the fourth rear region, where the width of the fourth rear region is blkW+2 and the height of the fourth rear region is blkH+2.
[0262] For any corner position of the first rear region, assume that the position of the upper-left corner of the first rear region is used as an example. The sample closest to the upper-left corner position in the rear reference picture is determined, and by using this sample as the upper-left corner, the fourth rear region is determined. The width of the fourth rear region is blkW+2, and the height of the fourth rear region is blkH+2.
[0263] For example, referring to Figure 13, the corner position of the upper left corner of the first rear region D1 is (5.3, 6.2), and the position of the sample closest to the corner position (5.3, 6.2) is determined as (5, 6) in the rear reference picture Ref1. By using the sample at position (5, 6) as the upper left corner, the fourth rear region D2 is determined. The width of the fourth rear region D2 is blkW+2, and the height of the fourth rear region D2 is blkH+2.
[0264] (8) Determine a fifth rear region that includes a fourth rear region, where the center of the fourth rear region coincides with the center of the fifth rear region, the width of the fifth rear region is blkW+n+1, and the height of the fifth rear region is blkH+n+1. Perform interpolation filtering on the picture block by using an interpolation filter to obtain a picture block in the fifth rear region from the rear reference picture and obtain a first rear prediction block of the first picture block, where the width of the first rear prediction block is blkW+2 and the height of the first rear prediction block is blkH+2.
[0265] For example, an 8-tap interpolation filter is used as an example. Referring to Figure 13, a fifth back region D3 containing a fourth back region D2 is determined. The center of the fourth back region D2 coincides with the center of the fifth back region D3, the width of the fifth back region D3 is blkW+9, and the height of the fifth back region D3 is blkH+9. The picture block in the fifth back region D3 is obtained from the back reference picture Ref1, and interpolation filtering is performed on the picture block by using an interpolation filter to obtain the first back prediction block of the first picture block B. The width of the first back prediction block is blkW+2, and the height of the first back prediction block is blkH+2.
[0266] The number of taps n in the interpolation filter may be a value such as 6, 8, or 10.
[0267] When this step is performed, it may be further determined whether interpretation should be performed via BIO based on the motion information of the picture block to be processed, and if it is determined that interpretation should be performed via BIO, this step is performed. The decision process may be as follows:
[0268] It is determined whether the frame numbers of the picture block to be processed, the frame numbers of the forward-referenced picture, and the frame numbers of the backward-referenced picture satisfy the pre-configured BIO usage conditions. If the pre-configured BIO usage conditions are satisfied and it is determined that interpretation will be performed via BIO, this step is executed. If the pre-configured BIO usage conditions are not satisfied, it is determined that interpretation will be performed by a method other than BIO. The implementation process of the alternative method is not described in detail herein.
[0269] The pre-configured BIO usage conditions may also be those shown in the first equation below.
[0270] The first equation is (POC_L0-POC_Cur)*(POC_L1-POC_Cur)<0.
[0271] In the first equation, POC_L0 is the frame number of the forward-referenced picture, POC_Cur is the frame number of the picture block to be processed, POC_L1 is the frame number of the backward-referenced picture, and * is a multiplication operation.
[0272] In this step, it may be further determined whether interpretation is performed via BIO based on the first forward prediction block and the first backward prediction block of the first picture block, and if it is determined that interpretation is performed via BIO, step 104 is performed. The decision process may be as follows:
[0273] SAD is calculated based on the first forward prediction block and the first backward prediction block of the first picture block according to the second equation below. If SAD exceeds the preset threshold TH_CU, it is determined that interpretation will be performed via BIO, and step 104 is performed. If SAD does not exceed the preset threshold TH_CU, it is determined that interpretation will be performed by a method other than BIO. The implementation process of the alternative method is not described in detail herein.
[0274] The second equation is,
number
[0275] In the second equation, (1) (i,j) is the predictor of the sample in the i-th row and j-th column of the first back prediction block, (0) (i,j) is the predictor for the sample in the i-th row and j-th column of the first forward prediction block.
[0276] TH_CU = (1 << (BD - 8 + shift)) * blkW * blkH, where shift = Max(2, 14 - BD), where BD is the current sample bit width, abs() is the operation to obtain the absolute value, and << is the left shift operation.
[0277] Step 104: To obtain the first gradient matrix of the first picture block, perform a gradient operation on the first prediction block of the first picture block, where the width of the first gradient matrix is blkW and the height of the first gradient matrix is blkH.
[0278] The first gradient matrix includes a first forward horizontal gradient matrix, a first forward vertical gradient matrix, a first backward horizontal gradient matrix, and a first backward vertical gradient matrix.
[0279] In this step, horizontal and vertical gradients are calculated based on the predictor of each sample contained within the first prediction block, according to the following third equation. Each calculated horizontal gradient corresponds to one row number and one column number, and each calculated vertical gradient corresponds to one row number and one column number. Based on the row and column numbers corresponding to the calculated horizontal gradients, the first horizontal gradient matrix of the first picture block is formed by the calculated horizontal gradients, and based on the row and column numbers corresponding to the calculated vertical gradients, the first vertical gradient matrix of the first picture block is formed by the calculated vertical gradients.
[0280] When the gradient of a row or column in the gradient matrix is calculated, two sample predictors are obtained from the first prediction block based on the row and column numbers, and the horizontal or vertical gradient is calculated based on the two sample predictors according to the following third equation. The horizontal gradient corresponds to the row and column numbers individually, and the vertical gradient corresponds to the row and column numbers individually.
[0281] The first prediction block includes a first forward prediction block and a first backward prediction block. Based on the first forward prediction block, the forward horizontal slope and forward vertical slope are calculated according to the following third equation. Each calculated forward horizontal slope corresponds to one row number and one column number, and each calculated forward vertical slope corresponds to one row number and one column number. Based on the row and column numbers corresponding to the calculated forward horizontal slopes, the first forward horizontal slope matrix of the first picture block is formed by the calculated forward horizontal slopes, and based on the row and column numbers corresponding to the calculated forward vertical slopes, the first forward vertical slope matrix of the first picture block is formed by the calculated forward vertical slopes.
[0282] Based on the first backward prediction block, the backward horizontal and backward vertical slopes are calculated according to the following third equation. Each calculated backward horizontal slope corresponds to one row number and one column number, and each calculated backward vertical slope corresponds to one row number and one column number. Based on the row and column numbers corresponding to the calculated backward horizontal slopes, the first backward horizontal slope matrix of the first picture block is formed by the calculated backward horizontal slopes, and based on the row and column numbers corresponding to the calculated backward vertical slopes, the first backward vertical slope matrix of the first picture block is formed by the calculated backward vertical slopes.
[0283] The third equation is,
number
[0284] In the third equation, the value of k may be 0 or 1, where 0 indicates "forward" and 1 indicates "backward".
number
number
number
[0285] I (k) (i+1,j) is the predictor of the sample in the (i+1)th row and jth column of the first prediction block, where if k=0, I (k) (i+1,j) is the predictor of the sample in the (i+1)th row and jth column of the first forward prediction block, where if k=1, I (k) (i+1,j) is the predictor of the sample in the (i+1)th row and jth column of the first back prediction block, I (k) (i-1,j) is the predictor of the sample in the (i-1)th row and jth column of the first prediction block, where if k=0, I (k) (i-1,j) is the predictor of the sample in the (i-1)th row and jth column of the first forward prediction block, where if k=1, I (k) (i-1,j) is the predictor for the sample in the (i-1)th row and jth column of the first back-prediction block.
[0286] I (k) (i,j+1) is the predictor of the sample in the i-th row and (j+1)-th column of the first prediction block, where if k=0, I (k) (i,j+1) is the predictor of the sample in the i-th row and (j+1)-th column of the first forward prediction block, where if k=1, I (k) (i,j+1) is the predictor of the sample in the i-th row and (j+1)-th column of the first back prediction block, I (k) (i,j-1) is the predictor of the sample in the i-th row and (j-1)-th column of the first prediction block, where if k=0, I (k) (i,j-1) is the predictor of the sample in the i-th row and (j-1)-th column of the first forward prediction block, where if k=1, I (k)(i,j-1) is the predictor value for the sample in the i-th row and (j-1)-th column of the first back-prediction block.
[0287] It should be noted that for a first prediction block having a width of blkW+2 and a height of blkH+2, a first gradient matrix having the width of blkW and the height of blkH may be obtained based on the first prediction block according to the third equation above. The first gradient matrix includes a first horizontal gradient matrix having the width of blkW and the height of blkH, and a first vertical gradient matrix having the width of blkW and the height of blkH. That is, for a first forward prediction block having a width of blkW+2 and a height of blkH+2, a first forward horizontal gradient matrix having the width of blkW and the height of blkH may be obtained based on the first forward prediction block according to the third equation above. For a first backward prediction block having a width of blkW+2 and a height of blkH+2, a first backward horizontal gradient matrix having a width of blkW and a height of blkH, and a first backward vertical gradient matrix having a width of blkW and a height of blkH may be obtained based on the first backward prediction block according to the third equation above.
[0288] Step 105: Perform the first expansion on the width and height of the first gradient matrix based on the gradients at the matrix edge positions of the first gradient matrix, such that the width and height of the first gradient matrix obtained after the first expansion are each 2 samples larger than the width and height of the first picture block.
[0289] The width and height of the first gradient matrix obtained after the first extension are equal to the width and height of the first prediction block, respectively. The width of the first prediction block is blkW+2, and the height of the first prediction block is blkH+2. The width of the first gradient matrix is also blkW+2, and the height of the first gradient matrix is also blkH+2.
[0290] In this step, the first extension is performed individually on the width and height of the first forward horizontal gradient matrix, the width and height of the first forward vertical gradient matrix, the width and height of the first backward horizontal gradient matrix, and the width and height of the first backward vertical gradient matrix, so that the width of the first forward horizontal gradient matrix, the width of the first forward vertical gradient matrix, the width of the first backward horizontal gradient matrix, and the width and height of the first backward vertical gradient matrix obtained after the first extension all become blkW+2, and the height of the first forward horizontal gradient matrix, the height of the first forward vertical gradient matrix, the height of the first backward horizontal gradient matrix, and the height of the first backward vertical gradient matrix obtained after the first extension all become blkH+2.
[0291] In this step, the first gradient matrix contains four edges. For the gradient at the left edge of the first gradient matrix, a single column gradient is obtained by performing an extension on the left side of the first gradient matrix based on the gradient at the left edge. For the gradient at the right edge of the first gradient matrix, a single column gradient is obtained by performing an extension on the right side of the first gradient matrix based on the gradient at the right edge. For the gradient at the top edge of the first gradient matrix, a single row gradient is obtained by performing an extension on the upper side of the first gradient matrix based on the gradient at the top edge. For the gradient at the bottom edge of the first gradient matrix, a single row gradient is obtained by performing an extension on the lower side of the first gradient matrix based on the gradient at the bottom edge. Therefore, the width and height of the first gradient matrix obtained after the first extension are 2 samples larger than the width and height of the first picture block, respectively.
[0292] Step 106: Based on the first prediction block and the first gradient matrix, calculate the motion information refinement value for each basic processing unit in the first picture block.
[0293] The width of the basic processing unit may be M, and the height of the basic processing unit may also be M; that is, the basic processing unit is a picture block containing M*M samples. The value of M may be 2, 3, or 4, for example.
[0294] The refined motion information value of the basic processing unit includes the refined horizontal motion information value and the refined vertical motion information value.
[0295] This step can be implemented by 1061 to 1064, which may be as follows:
[0296] 1061: In order to obtain each basic processing unit contained within the first picture block, the first picture block is divided, where each basic processing unit is a picture block having a size of M*M.
[0297] 1062: Based on the position of the basic processing unit, the basic prediction block of any basic processing unit in the first prediction block is determined, where the width of the basic prediction block is M+2 and the height of the basic prediction block is M+2.
[0298] Assuming that the basic processing unit covers the first to M rows and the first to M columns of the first picture block, the picture block covering the 0th to (M+1) rows and the 0th to (M+1) columns of the first prediction block is used as the basic prediction block of the basic processing unit.
[0299] The basic prediction block of the basic processing unit includes a forward basic prediction block and a backward basic prediction block. Specifically, a picture block covering rows 0 to (M+1) and columns 0 to (M+1) within the first forward prediction block is used as the forward prediction block of the basic processing unit, and a picture block covering rows 0 to (M+1) and columns 0 to (M+1) within the first backward prediction block is used as the backward basic prediction block of the basic processing unit.
[0300] 1063: Based on the position of the basic processing unit, the basic gradient matrix of the basic processing unit in the first gradient matrix is determined, where the width of the basic gradient matrix is M+2 and the height of the basic gradient matrix is M+2.
[0301] Assuming that the basic processing unit covers the first to M rows and the first to M columns of the first picture block, the matrix covering the 0th to (M+1) rows and the 0th to (M+1) columns in the first gradient matrix is used as the basic gradient matrix of the basic processing unit.
[0302] The basic gradient matrix of the basic processing unit includes a forward horizontal basic gradient matrix, a forward vertical basic gradient matrix, a backward horizontal basic gradient matrix, and a backward vertical basic gradient matrix. Specifically, the matrix covering the 0th to (M+1)th rows and the 0th to (M+1)th columns of the first forward horizontal gradient matrix is used as the forward horizontal basic gradient matrix of the basic processing unit, the matrix covering the 0th to (M+1)th rows and the 0th to (M+1)th columns of the first forward vertical gradient matrix is used as the forward vertical basic gradient matrix of the basic processing unit, the matrix covering the 0th to (M+1)th rows and the 0th to (M+1)th columns of the first backward horizontal gradient matrix is used as the backward horizontal basic gradient matrix of the basic processing unit, and the matrix covering the 0th to (M+1)th rows and the 0th to (M+1)th columns of the first backward vertical gradient matrix is used as the backward vertical basic gradient matrix of the basic processing unit.
[0303] 1064. Based on the basic prediction block and basic gradient matrix of the basic processing unit, the refined motion information value of the basic processing unit is calculated.
[0304] In 1064, based on the forward basic prediction block, backward basic prediction block, forward horizontal basic gradient matrix, forward vertical basic gradient matrix, backward horizontal basic gradient matrix, and backward vertical basic gradient matrix of the basic processing unit, the horizontal motion information refinement value and vertical motion information refinement value of the basic processing unit are calculated according to the following fourth and fifth equations.
[0305] The fourth equation is,
number
[0306] The fifth equation is,
number
[0307] In the fourth equation above, (i,j)∈Ω indicates that i=0,1,..., and M+1, and j=0,1,..., and M+1. In the fifth equation above, v x This is the refined value of the horizontal motion information of the basic processing unit, and v y This is the refined value of the vertical motion information of the basic processing unit, and th' BIO =2 13-BD And,
number
[0308] By repeatedly executing steps 1062 through 1064, the refined motion information values for each basic processing unit contained within the first picture block may be obtained.
[0309] Step 107: Obtain the predictor for the first picture block based on the motion information refinement value of each basic processing unit contained within the first picture block.
[0310] The predictor for the first picture block includes the predictor for each sample within each basic processing unit in the first picture block.
[0311] Based on the forward basic prediction block, backward basic prediction block, forward horizontal basic gradient matrix, forward vertical basic gradient matrix, backward horizontal basic gradient matrix, and backward vertical basic gradient matrix of the basic processing unit, the predictor for each sample contained within any basic processing unit contained within the first picture block is calculated according to the following sixth equation.
[0312] The sixth equation is,
number
[0313] In the sixth equation, pred BIO (i,j) is the predictor of the sample in the i-th row and j-th column within the basic processing unit, with shift=15-BD, and O offset = 1 << (14 - BD) + 2 · (1 << 13), and rnd() rounds to the nearest integer.
[0314] By repeatedly executing steps 103 through 107, the predictor for each first picture block within the picture block being processed is obtained.
[0315] Step 108: Obtain the predictor for the picture block to be processed using a combination of predictors from multiple first picture blocks contained within the picture block to be processed.
[0316] The inter prediction method shown in Figure 12 may be summarized as steps 1 to 6, and steps 1 to 6 may be as follows.
[0317] Step 1: Determine the current CU's movement information.
[0318] The current CU motion information may be determined by using merge mode, AMVP mode (see background explanation), or another mode, which is not limited herein.
[0319] It should be noted that other methods for determining motion information may also be applicable to this application. Details are not described herein.
[0320] Step 2: Determine whether the current CU meets the BIO usage requirements.
[0321] If bidirectional prediction is performed on the current CU, and the relationship between the forward-referenced frame number POC_L0, the backward-referenced frame number POC_L1, and the current frame number POC_Cur satisfies the following equation, then the current CU satisfies the BIO usage conditions. (POC_L0-POC_Cur)*(POC_L1-POC_Cur)<0
[0322] It should be noted that whether BIO is used may also be determined by determining whether the current CU size is greater than a preset threshold. For example, the current CU height. H The current width of CU is 8 or more. W BIO can only be used if the value is 8 or greater.
[0323] It should be noted that other conditions of use for BIO may also apply to this application. Details are not described herein.
[0324] If the current CU meets the BIO usage requirements, step 3 is performed; otherwise, motion compensation is performed in another manner.
[0325] The VPDU size is retrieved, and VPDU_X and VPDU_Y, as well as the parameters blkW and blkH, are set. blkW=Min(cuW,VPDU_X) blkH=Min(cuH,VPDU_Y)
[0326] The Min function indicates that the minimum value will be selected.
[0327] For example, if the size of the CU is 128x128 and the size of the VPDU is 64x64, then BlkW is 64 and blkH is 64.
[0328] For example, if the size of the CU is 128 x 128 and the size of the VPDU is 128 x 32, then BlkW is 128 and blkH is 32.
[0329] For example, if the size of CU is 128 x 128 and the size of VPDU is 32 x 128, then BlkW is 32 and blkH is 128.
[0330] Optionally, if the size of the maximum interpredictive processing unit is smaller than the VPDU size, blkW and blkH may be set according to the following formulas. blkW = Min(cuW, MAX_MC_X) blkH=Min(cuH,MAX_MC_Y)
[0331] For example, if the size of the CU is 128 x 128 and the size of the maximum interpredictive processing unit is 32 x 32, then BlkW is 32 and blkH is 32.
[0332] Each CU is divided based on blkW and blkH in order to perform BIO.
[0333] Step 3: Calculate the forward and backward predictors for the current CU.
[0334] Forward and backward predictor I (k) Motion compensation is performed by using motion information to obtain (i,j), where i=-1..blkW and j=-1..blkH (a prediction matrix of (blkW+2)*(blkH+2) is obtained).
[0335] I obtained by performing interpolation using an 8-tap interpolation filter (k)At (i,j), i=0..blkW-1, and j=0..blkH, predictors at other locations (where an extension of one row and one column is performed) are obtained by performing interpolation using a bilinear interpolation filter.
[0336] Predictors may be obtained by using VPDU as the minimum predictor acquisition unit, or by using a block smaller than VPDU as the minimum predictor acquisition unit. This is not limited to these methods.
[0337] It should be noted that the predictor for the extended region may also be obtained by using other methods, for example, by using an 8-tap interpolation filter, or by directly using a reference sample at an integer sample position. This is not limited herein.
[0338] It should be noted that in order to determine whether SAD is less than the threshold TH_CU, the SAD between the forward and backward predictors is calculated, and if SAD is less than the threshold TH_CU, BIO is not performed; otherwise, BIO is performed. Other determination methods may also be applied to this application, and their details are not described herein.
[0339] The formula for calculating SAD is as follows:
number
[0340] The threshold TH_CU may be set to (1<<(BD-8+shift))*blkW*blkH, and shift may be set to Max(2,14-BD).
[0341] Step 4: Calculate the horizontal and vertical gradients based on the current forward and backward predictors of the CU.
[0342] Horizontal slope and vertical slope
number
number
[0343] Step 5: Perform padding on the current CU's forward and backward predictors, as well as the horizontal and vertical gradients.
[0344] I (k) (i,j) and,
number
number
[0345] Step 6: Derive the refined motion vectors for each 4x4 subblock, and then perform weighting.
[0346] For each 4x4 subblock, vx and vy are obtained according to equation (2). Finally, weighting is performed according to equation (6) to obtain the predictor for each 4x4 subblock.
[0347] It should be noted that the SAD between the forward and backward predictors of each 4x4 subblock may be calculated to determine whether the SAD is less than the threshold TH_SCU. If the SAD is less than the threshold TH_SCU, a weighted average is performed directly; otherwise, vx and vy are obtained according to equation (2), and then weighting is performed according to equation (6). Other determination methods may also be applied to this application and are not described in detail herein. TU_SCU may be set to 1 << (BD - 3 + shift).
[0348] In this embodiment of the present application, a smaller width is selected in the pre-configured picture division width Width and the width of the picture block to be processed cuW, and is shown as blkW, and a smaller height is selected in the pre-configured picture division height Height The first picture block to be processed is selected at height cuH of the picture block to be processed and is shown as blkH, and is determined based on blkW and blkH. Therefore, when interpretation processing is performed for each first picture block, the area of each determined first picture block is not too large so that hardware resources such as memory space resources are consumed less, thereby reducing the complexity of the implementation and improving the efficiency of interpretation processing.
[0349] Referring to Figure 14, an embodiment of the present application provides an interpretation method. The method may be applied to an interpretation unit 244 in an encoder shown in Figure 2, or an interpretation unit 344 in a decoder shown in Figure 3. The method may also be a bidirectional interpretation method and includes the following steps.
[0350] Steps 201 and 202 are the same as steps 101 and 102, respectively, and will not be described in detail again herein.
[0351] Step 203: Based on the motion information of the picture block to be processed, obtain the first predicted block of the first picture block, where the width of the first predicted block is equal to the width of the first picture block, and the height of the first predicted block is equal to the height of the first picture block.
[0352] The motion information of the first picture block includes information such as a reference picture and motion vectors. In this embodiment, an optical flow-based bidirectional prediction method is used for interpretation. Therefore, the motion information of the first picture block includes information such as a forward reference picture, a backward reference picture, a forward motion vector, and a backward motion vector.
[0353] When the method in this embodiment is applied to an encoder, the encoder may determine motion information of the picture block to be processed in merge mode, AMVP mode, or another mode, and the motion information of the picture block to be processed includes motion information of each first picture block within the picture block to be processed. In this step, the motion information of the picture block to be processed determined by the encoder is obtained, and the motion information of the first picture block is obtained from the motion information of the picture block to be processed.
[0354] When the method in this embodiment is applied to a decoder, the video bitstream received by the decoder from the encoder includes motion information of the picture block to be processed, and the motion information of the picture block to be processed includes motion information of each first picture block within the picture block to be processed. In this step, the motion information of the picture block to be processed is extracted from the video bitstream, and the motion information of the first picture block is obtained from the motion information of the picture block to be processed.
[0355] The first prediction block of the first picture block includes a first forward prediction block and a first backward prediction block. In this step, the first forward prediction block and the first backward prediction block of the first picture block may be obtained in the following steps (1) to (8). Steps (1) to (8) may be as follows.
[0356] (1) Based on the first position of the first picture block and the motion information of the first picture block, a first forward region in the forward reference picture is determined, where the width of the first forward region is blkW and the height of the first forward region is blkH.
[0357] For example, referring to Figure 15, the motion information of the first picture block B includes a forward reference picture Ref0, a backward reference picture Ref1, a forward motion vector MV0, and a backward motion vector MV1. Based on the first position of the first picture block B, a second forward region B11 is determined within the forward reference picture Ref0, where the width of the second forward region B11 is blkW and the height of the second forward region B11 is blkH. Based on the forward motion vector MV0 and the position of the second forward region B11, a first forward region B12 is determined, where the width of the first forward region B12 is blkW and the height of the first forward region B12 is blkH.
[0358] (2) Determine whether the corner position of the first forward region matches the sample position in the forward reference picture. If the corner position of the first forward region matches the sample position in the forward reference picture, obtain the picture block in the first forward region from the forward reference picture so that it functions as the first forward prediction block of the first picture block. If the corner position of the first forward region does not match the sample position in the forward reference picture, perform step (3).
[0359] For example, referring to Figure 15, the upper left corner of the first front region B12 is used as an example. Assuming that the corner position of the upper left corner of the first front region B12 is (15,16) in the front reference picture Ref0, the corner position of the upper left corner coincides with the sample position in the front reference picture Ref0, and the sample position in the front reference picture Ref0 is (15,16). In another example, assuming that the corner position of the upper left corner of the first front region B12 is (15.3,16.2) in the front reference picture Ref0, the corner position of the upper left corner does not coincide with the sample position in the front reference picture Ref0, that is, there is no sample at the position (15.3,16.2) in the front reference picture Ref0.
[0360] (3) Determine the sample closest to the corner position of the first forward region in the forward reference picture, and determine the third forward region by using the sample as the corner, where the width of the third forward region is blkW and the height of the third forward region is blkH.
[0361] For any corner position of the first front region, assume that the position of the upper-left corner of the first front region is used as an example. The sample closest to the upper-left corner position in the front reference picture is determined, and by using this sample as the upper-left corner, the third front region is determined. The width of the third front region is blkW, and the height of the third front region is blkH.
[0362] For example, referring to Figure 15, the corner position of the upper left corner of the first forward region B12 is (15.3, 16.2), and the position of the sample closest to the corner position (15.3, 16.2) is determined as (15, 16) in the forward reference picture Ref0. By using the sample at position (15, 16) as the upper left corner, the third forward region A1 is determined. The width of the third forward region A1 is blkW, and the height of the third forward region A1 is blkH.
[0363] (4) Determine a fourth forward region that includes a third forward region, where the center of the third forward region coincides with the center of the fourth forward region, the width of the fourth forward region is blkW+n-1, and the height of the fourth forward region is blkH+n-1. Perform interpolation filtering on the picture block by using an interpolation filter to obtain a picture block in the fourth forward region from the forward reference picture, and to obtain a first forward prediction block of the first picture block, where the width of the first forward prediction block is blkW, the height of the first forward prediction block is blkH, and n is the number of taps in the interpolation filter.
[0364] For example, an 8-tap interpolation filter is used as an example. Referring to Figure 15, a fourth forward region A2 containing a third forward region A1 is determined. The center of the third forward region A1 coincides with the center of the fourth forward region A2, the width of the fourth forward region A2 is blkW+7, and the height of the fourth forward region A2 is blkH+7. The picture block in the fourth forward region A2 is obtained from the forward reference picture Ref0, and interpolation filtering is performed on the picture block by using an interpolation filter to obtain the first forward prediction block of the first picture block B. The width of the first forward prediction block is blkW, and the height of the first forward prediction block is blkH.
[0365] (5) Based on the first position and motion information of the first picture block, a first rear region in the rear reference picture is determined, where the width of the first rear region is blkW and the height of the first rear region is blkH.
[0366] For example, referring to Figure 15, based on the first position of the first picture block B, a second rear region C11 is determined within the back reference picture Ref1, where the width of the second rear region C11 is blkW and the height of the second rear region C11 is blkH. Based on the back motion vector MV1 and the position of the second rear region C11, a first rear region C12 is determined, where the width of the first rear region C12 is blkW and the height of the first rear region C12 is blkH.
[0367] (6) Determine whether the corner position of the first rear region matches the sample position in the rear reference picture. If the corner position of the first rear region matches the sample position in the rear reference picture, retrieve the picture block in the first rear region from the rear reference picture so that it functions as the first rear prediction block of the first picture block. If the corner position of the first rear region does not match the sample position in the rear reference picture, perform step (7).
[0368] For example, referring to Figure 15, the upper left corner of the first rear region C12 is used as an example. Assuming that the corner position of the upper left corner of the first rear region C12 is (5,6) in the rear reference picture Ref1, the corner position of the upper left corner coincides with the sample position in the rear reference picture Ref1, and the sample position in the rear reference picture Ref1 is (5,6). In another example, assuming that the corner position of the upper left corner of the first rear region C12 is (5.3,6.2) in the rear reference picture Ref1, the corner position of the upper left corner does not coincide with the sample position in the rear reference picture Ref0, that is, there is no sample at position (5.3,6.2) in the rear reference picture Ref0.
[0369] (7) Determine the sample closest to the corner position of the first rear region in the back reference picture, and determine the third rear region by using the sample as the corner, where the width of the third rear region is blkW and the height of the third rear region is blkH.
[0370] For any corner position of the first rear region, assume that the position of the upper-left corner of the first rear region is used as an example. The sample closest to the upper-left corner position in the rear reference picture is determined, and by using this sample as the upper-left corner, the third rear region is determined. The width of the third rear region is blkW, and the height of the third rear region is blkH.
[0371] For example, referring to Figure 15, the corner position of the upper left corner of the first rear region C12 is (5.3, 6.2), and the position of the sample closest to the corner position (5.3, 6.2) is determined as (5, 6) in the rear reference picture Ref1. By using the sample at position (5, 6) as the upper left corner, the third rear region D1 is determined. The width of the third rear region D1 is blkW, and the height of the third rear region D1 is blkH.
[0372] (8) Determine a fourth rear region that includes a third rear region, where the center of the third rear region coincides with the center of the fourth rear region, the width of the fourth rear region is blkW+n-1, and the height of the fourth rear region is blkH+n-1. Perform interpolation filtering on the picture block by using an interpolation filter to obtain a picture block in the fourth rear region from the rear reference picture and obtain a first rear prediction block of the first picture block, where the width of the first rear prediction block is blkW and the height of the first rear prediction block is blkH.
[0373] For example, an 8-tap interpolation filter is used as an example. Referring to Figure 15, a fourth back region D2 containing a third back region D1 is determined. The center of the third back region D1 coincides with the center of the fourth back region D2, the width of the fourth back region D2 is blkW+7, and the height of the fourth back region D2 is blkH+7. The picture block in the fourth back region D2 is obtained from the back reference picture Ref1, and interpolation filtering is performed on the picture block by using an interpolation filter to obtain the first back prediction block of the first picture block B. The width of the first back prediction block is blkW, and the height of the first back prediction block is blkH.
[0374] When this step is performed, it may be further determined whether interpretation is performed via BIO based on the motion information of the picture block to be processed. If it is determined that interpretation is performed via BIO, this step is performed. For details of the determination process, please refer to the relevant content in step 103 in the embodiment shown in Figure 12. Further details will not be described again in this specification.
[0375] In this step, it may be further determined whether interpretation is performed via BIO based on the first forward prediction block and the first backward prediction block of the first picture block. If it is determined that interpretation is performed via BIO, step 204 is performed. For details of the determination process, please refer to the relevant content in step 103 in the embodiment shown in Figure 12. Further details will not be described again in this specification.
[0376] Step 204: To obtain the first gradient matrix of the first picture block, perform a gradient operation on the first prediction block of the first picture block, where the width of the first gradient matrix is blkW-2 and the height of the first gradient matrix is blkH-2.
[0377] The first gradient matrix includes a first forward horizontal gradient matrix, a first forward vertical gradient matrix, a first backward horizontal gradient matrix, and a first backward vertical gradient matrix.
[0378] The width of the first forward horizontal gradient matrix, the width of the first forward vertical gradient matrix, the width of the first backward horizontal gradient matrix, and the width of the first backward vertical gradient matrix may all be blkW-2, and the height of the first forward horizontal gradient matrix, the height of the first forward vertical gradient matrix, the height of the first backward horizontal gradient matrix, and the height of the first backward vertical gradient matrix may all be blkH-2.
[0379] For a detailed implementation process of performing gradient calculations on the first prediction block of the first picture block in this step, please refer to the relevant content in step 104 of the embodiment shown in Figure 12. Further details will not be described again in this specification.
[0380] Step 205: Perform the first expansion on the width and height of the first gradient matrix based on the gradients at the matrix edge positions of the first gradient matrix, such that the width and height of the first gradient matrix obtained after the first expansion are each 2 samples larger than the width and height of the first picture block.
[0381] The width and height of the first gradient matrix obtained after the first extension are equal to the width blkW+2 and height blkH+2 of the first prediction block, respectively.
[0382] In this step, the first extension is performed individually on the width and height of the first forward horizontal gradient matrix, the width and height of the first forward vertical gradient matrix, the width and height of the first backward horizontal gradient matrix, and the width and height of the first backward vertical gradient matrix, so that the width of the first extended forward horizontal gradient matrix, the width of the first forward vertical gradient matrix, the width of the first backward horizontal gradient matrix, and the width and height of the first backward vertical gradient matrix obtained after the first extension all become blkW+2, and the height of the first forward horizontal gradient matrix, the height of the first forward vertical gradient matrix, the height of the first backward horizontal gradient matrix, and the height of the first backward vertical gradient matrix, all become blkH+2.
[0383] For a method of performing the first extension on the first gradient matrix, please refer to the relevant content in step 205 of the embodiment shown in Figure 12. Further details will not be described again in this specification.
[0384] Step 206: To perform a second expansion on the width and height of the first prediction block, the sample values at the block edge positions of the first prediction block are duplicated, where the width and height of the first prediction block obtained after the second expansion are blkW+2 and blkH+2.
[0385] In this step, sample values at the block edge positions of the first forward prediction block and sample values at the block edge positions of the first backward prediction block are duplicated in order to perform a second expansion on the width and height of the first forward prediction block and the width and height of the first backward prediction block. That is, in this step, the width and height of the first forward prediction block obtained after the second expansion are blkW+2 and blkH+2, respectively, and the width and height of the first backward prediction block obtained after the second expansion are blkW+2 and blkH+2, respectively.
[0386] In this step, interpolation filtering may be further performed on the sample area of the block edge region of the first prediction block in order to perform a second expansion on the width and height of the first prediction block.
[0387] Optionally, in step 203, a picture block in the reference picture having width blkW and height blkH is directly used as the first prediction block of the first picture block; that is, referring to Figure 15, a picture block in the first forward region B12 is used as the first forward reference block in the forward reference figure Ref0, and a picture block in the first backward region C12 is used as the first backward prediction block in the backward reference figure Ref1. In this case, the first prediction block is the picture block in the reference picture. In this case, a circle of samples surrounding the first prediction block and closest to the first prediction block is selected from the reference picture, and the selected circle of samples and the first prediction block form a first prediction block having width blkW+2 and height blkH+2 obtained after the second expansion.
[0388] Optionally, in step 203, the first prediction block of the first picture block is obtained by using an interpolation filter. In this case, the first prediction block is not a picture block in the reference picture. For any sample on any edge of the first prediction block (for ease of explanation, the edge is referred to as the first edge), the second position of each sample contained within the second edge is obtained based on the first position of each sample on the first edge in the reference picture. The second edge is outside the first prediction block, and the distance between the second edge and the first edge is 1 sample. The second edge contains blkW+2 samples or blkH+2 samples. For each sample on the second edge, the second position of the sample in the reference picture is located between two adjacent samples or between four adjacent samples, and interpolation filtering is performed on two adjacent samples or four adjacent samples by using an interpolation filter to obtain the samples. A second edge corresponding to each edge of the first prediction block is obtained in the manner described above, and each obtained second edge and the first prediction block form a first prediction block having a width of blkW+2 and a height of blkH+2 obtained after the second extension.
[0389] Step 206 may be performed alternatively before step 204. In this way, once the first prediction block obtained after the second expansion is acquired, a gradient operation may be performed on the first prediction block obtained after the second expansion to obtain the first gradient matrix of the first picture block. Since the width of the first prediction block obtained after the second expansion is blkW+2 and the height of the first prediction block obtained after the second expansion is blkH+2, the width of the obtained first gradient matrix is blkW and the height of the obtained first gradient matrix is blkH. Then, the first expansion is performed on the width and height of the first gradient matrix based on the gradient at the matrix edge positions of the first gradient matrix, such that the width and height of the obtained first gradient matrix after the first expansion are 2 samples larger than the width and height of the first picture block, respectively.
[0390] Steps 207 through 209 are the same as steps 106 through 108, respectively, and will not be described in detail again in this specification.
[0391] The inter prediction method shown in Figures 16A and 16B may be summarized as steps 1 to 6, which may be as follows:
[0392] Step 1: Determine the current CU's movement information.
[0393] The current CU motion information may be determined by using merge mode, AMVP mode (see background explanation), or another mode, which is not limited herein.
[0394] It should be noted that other methods for determining motion information may also be applicable to this application. Details are not described herein.
[0395] Step 2: Determine whether the current CU meets the BIO usage requirements.
[0396] If bidirectional prediction is performed on the current CU, and the relationship between the forward-referenced frame number POC_L0, the backward-referenced frame number POC_L1, and the current frame number POC_Cur satisfies the following equation, then the current CU satisfies the BIO usage conditions. (POC_L0-POC_Cur)*(POC_L1-POC_Cur)<0
[0397] It should be noted that whether BIO is used may also be determined by determining whether the current CU size is greater than a preset threshold. For example, the current CU height. H The current width of CU is 8 or more. W BIO can only be used if the value is 8 or greater.
[0398] It should be noted that other conditions of use for BIO may also apply to this application. Details are not described herein.
[0399] If the current CU meets the BIO usage requirements, step 3 is performed; otherwise, motion compensation is performed in another manner.
[0400] The VPDU size is retrieved, and VPDU_X and VPDU_Y, as well as the parameters blkW and blkH, are set. blkW=Min(cuW,VPDU_X) blkH=Min(cuH,VPDU_Y)
[0401] For example, if the size of the CU is 128 x 128 and the size of the VPDU is 64 x 64, then blkW is 64 and blkH is 64.
[0402] For example, if the size of the CU is 128 x 128 and the size of the VPDU is 128 x 32, then blkW is 128 and blkH is 32.
[0403] For example, if the size of CU is 128 x 128 and the size of VPDU is 32 x 128, then blkW is 32 and blkH is 128.
[0404] Optionally, if the size of the maximum interpredictive processing unit is smaller than the size of the VPDU, blkW and blkH may be set according to the following formulas. blkW = Min(cuW, MAX_MC_X) blkH=Min(cuH,MAX_MC_Y)
[0405] For example, if the size of the CU is 128 x 128 and the size of the maximum interpredictive processing unit is 32 x 32, then blkW is 32 and blkH is 32.
[0406] Each CU is divided based on blkW and blkH in order to perform BIO.
[0407] Step 3: Calculate the forward and backward predictors for the current CU.
[0408] Forward predictor and backward predictor I (k) Motion compensation is performed by using motion information to obtain (i,j), where i=0..blkW-1 and j=0..blkH-1 (a prediction matrix of blkW*blkH is obtained).
[0409] It should be understood that predictors may be obtained by using VPDU as the minimum predictor acquisition unit, or by using a block smaller than VPDU as the minimum predictor acquisition unit. This is not limited to this.
[0410] Step 4: Calculate the horizontal and vertical gradients based on the current forward and backward predictors of the CU.
[0411] Horizontal slope and vertical slope
number
number
[0412] Step 5: Perform padding on the current CU's forward and backward predictors, as well as the horizontal and vertical gradients.
[0413] I (k) (i,j) and,
number
number
[0414] Step 6: Derive the refined motion vectors for each 4x4 subblock, and then perform weighting.
[0415] For each 4x4 subblock, vx and vy are obtained according to equation (2). Finally, weighting is performed according to equation (6) to obtain the predictor for each 4x4 subblock.
[0416] In this embodiment of the present application, a smaller width is selected in the pre-configured picture division width Width and the width of the picture block to be processed cuW, and is shown as blkW, and a smaller height is selected in the pre-configured picture division height HeightThe first picture block to be processed is selected at height cuH of the picture block to be processed and shown as blkH, and is determined based on blkW and blkH. Therefore, the area of each determined first picture block is not very large so that less memory space is consumed when interpretation processing is performed for each first picture block. In addition, the first prediction block of the first picture block is obtained based on the motion information of the first picture block. The width of the first prediction block is equal to the width of the first picture block, and the height of the first prediction block is equal to the height of the first picture block. Therefore, the first prediction block may be relatively small so that less hardware resources such as CPU resources and memory resources are consumed to obtain the first prediction block, thereby reducing implementation complexity and improving processing efficiency.
[0417] Referring to Figures 16A and 16B, embodiments of the present application provide an interpretation method. The method may be applied to an interpretation unit 244 in an encoder shown in Figure 2, or an interpretation unit 344 in a decoder shown in Figure 3. The method may be a bidirectional interpretation method and includes the following steps.
[0418] Step 301: Compare the width cuW of the first picture block with the preset picture division width Width, and compare the height cuH of the first picture block with the preset picture division height Height Compared to, if cuW is greater than or equal to Width, and / or cuH is Height If the above is true, perform step 302, or if cuW is less than Width and cuH is Height If it is less than, perform step 305.
[0419] When the method in this embodiment is applied to an encoder, when encoding a picture, the encoder divides the picture into first picture blocks. Prior to this step, the first picture blocks are obtained from the encoder.
[0420] When the method in this embodiment is applied to a decoder, the decoder receives a video bitstream from an encoder, and the video bitstream includes a first picture block. Prior to this step, the first picture block is extracted from the video bitstream.
[0421] When this step is performed, it may be further determined, based on the motion information of the first picture block, whether interpretation should be performed via BIO. If it is determined that interpretation should be performed via BIO, this step is performed. For a detailed implementation process, please refer to the relevant content in step 103 in the embodiment shown in Figure 12. Further details will not be described again in this specification.
[0422] Step 302: Based on the motion information of the first picture block, obtain the second predicted block of the first picture block, where the width of the second predicted block is cuW+4 and the height of the second predicted block is cuH+4.
[0423] The motion information of the first picture block includes information such as a reference picture and motion vectors. In this embodiment, an optical flow-based bidirectional prediction method is used for interpretation. Therefore, the motion information of the first picture block includes information such as a forward reference picture, a backward reference picture, a forward motion vector, and a backward motion vector.
[0424] When the method in this embodiment is applied to an encoder, the encoder may determine the motion information of the first picture block in merge mode, AMVP mode, or another mode. In this step, the motion information of the first picture block determined by the encoder is obtained.
[0425] When the method in this embodiment is applied to a decoder, the video bitstream received by the decoder from the encoder includes motion information of a first picture block. In this step, the motion information of the first picture block is extracted from the video bitstream.
[0426] The second prediction block of the first picture block includes a second forward prediction block and a second backward prediction block. In this step, the second forward prediction block and the second backward prediction block of the first picture block may be obtained in the following steps (1) to (8). Steps (1) to (8) may be as follows.
[0427] (1) Based on the first position of the first picture block within the picture block to be processed and the motion information of the first picture block, the first forward region within the forward reference picture is determined, where the width of the first forward region is blkW+4 and the height of the first forward region is blkH+4.
[0428] For example, referring to Figure 13, the motion information of the first picture block B includes a forward reference picture Ref0, a backward reference picture Ref1, a forward motion vector MV0, and a backward motion vector MV1. Based on the first position of the first picture block B, a second forward region B11 is determined within the forward reference picture Ref0, where the width of the second forward region B11 is blkW and the height of the second forward region B11 is blkH. Based on the forward motion vector MV0 and the position of the second forward region B11, a third forward region B12 is determined, where the width of the third forward region B12 is blkW and the height of the third forward region B12 is blkH. A first forward region A1 is determined which includes a third forward region B12, where the width of the first forward region A1 is blkW+4, the height of the first forward region A1 is blkH+4, and the center of the third forward region B12 coincides with the center of the first forward region A1.
[0429] (2) Determine whether the corner position of the first forward region matches the sample position in the forward reference picture. If the corner position of the first forward region matches the sample position in the forward reference picture, obtain the picture block in the first forward region from the forward reference picture so that it functions as the second forward prediction block of the first picture block. If the corner position of the first forward region does not match the sample position in the forward reference picture, perform step (3).
[0430] For example, referring to Figure 13, the upper left corner of the first front region A1 is used as an example. Assuming that the corner position of the upper left corner of the first front region A1 is (15,16) in the front reference picture Ref0, the corner position of the upper left corner coincides with the sample position in the front reference picture Ref0, and the sample position in the front reference picture Ref0 is (15,16). In another example, assuming that the corner position of the upper left corner of the first front region A1 is (15.3,16.2) in the front reference picture Ref0, the corner position of the upper left corner does not coincide with the sample position in the front reference picture Ref0, that is, there is no sample at the position (15.3,16.2) in the front reference picture Ref0.
[0431] (3) Determine the sample closest to the corner position of the first forward region in the forward reference picture, and determine the fourth forward region by using the sample as the corner, where the width of the fourth forward region is blkW+4 and the height of the fourth forward region is blkH+4.
[0432] For any corner position of the first front region, assume that the position of the upper-left corner of the first front region is used as an example. The sample closest to the upper-left corner position in the front reference picture is determined, and by using this sample as the upper-left corner, the fourth front region is determined. The width of the fourth front region is blkW+4, and the height of the fourth front region is blkH+4.
[0433] For example, referring to Figure 13, the corner position of the upper left corner of the first forward region A1 is (15.3, 16.2), and the position of the sample closest to the corner position (15.3, 16.2) is determined as (15, 16) in the forward reference picture Ref0. By using the sample at position (15, 16) as the upper left corner, the fourth forward region A2 is determined. The width of the fourth forward region A2 is blkW+4, and the height of the fourth forward region A2 is blkH+4.
[0434] (4) Determine a fifth forward region that includes a fourth forward region, where the center of the fourth forward region coincides with the center of the fifth forward region, the width of the fifth forward region is blkW+n+3, and the height of the fifth forward region is blkH+n+3. Perform interpolation filtering on the picture block by using an interpolation filter to obtain a picture block in the fifth forward region from the forward reference picture and to obtain a second forward prediction block of the first picture block, where the width of the second forward prediction block is blkW+4, the height of the second forward prediction block is blkH+4, and n is the number of taps in the interpolation filter.
[0435] For example, an 8-tap interpolation filter is used as an example. Referring to Figure 13, a fifth forward region A3 is determined, which includes the fourth forward region A2. The center of the fourth forward region A2 coincides with the center of the fifth forward region A3, the width of the fifth forward region A3 is blkW+11, and the height of the fifth forward region A3 is blkH+11. The picture block in the fifth forward region A3 is obtained from the forward reference picture Ref0, and interpolation filtering is performed on the picture block by using an interpolation filter to obtain the second forward prediction block of the first picture block B. The width of the second forward prediction block is blkW+4, and the height of the second forward prediction block is blkH+4.
[0436] (5) Based on the first position and motion information of the first picture block, the first rear region in the rear reference picture is determined, where the width of the first rear region is blkW+4 and the height of the first rear region is blkH+4.
[0437] For example, referring to Figure 13, a second rear region C11 is determined within the back reference picture Ref1 based on the first position of the first picture block B, where the width of the second rear region C11 is blkW and the height of the second rear region C11 is blkH. A third rear region C12 is determined based on the back motion vector MV1 and the position of the second rear region C11, where the width of the third rear region C12 is blkW and the height of the third rear region C12 is blkH. A first rear region D1 containing the third rear region C12 is determined, where the width of the first rear region D1 is blkW+4 and the height of the first rear region D1 is blkH+4, and the center of the third rear region C12 may coincide with the center of the first rear region D1.
[0438] (6) Determine whether the corner position of the first rear region matches the sample position in the rear-reference picture. If the corner position of the first rear region matches the sample position in the rear-reference picture, retrieve the picture block in the first region from the rear-rear-reference picture so that it functions as a second rear-predicted block of the first picture block. If the corner position of the first rear region does not match the sample position in the rear-reference picture, perform step (7).
[0439] For example, referring to Figure 13, the upper left corner of the first rear region A1 is used as an example. Assuming that the corner position of the upper left corner of the first rear region A1 is (5,6) in the rear reference picture Ref0, the corner position of the upper left corner coincides with the sample position in the rear reference picture Ref0, and the sample position in the rear reference picture Ref0 is (5,6). In another example, assuming that the corner position of the upper left corner of the first rear region D1 is (5.3,6.2) in the rear reference picture Ref0, the corner position of the upper left corner does not coincide with the sample position in the rear reference picture Ref0, that is, there is no sample at position (5.3,6.2) in the rear reference picture Ref0.
[0440] (7) Determine the sample closest to the corner position of the first rear region in the back reference picture, and use the sample as a corner to determine the fourth rear region, where the width of the fourth rear region is blkW+4 and the height of the fourth rear region is blkH+4.
[0441] For any corner position of the first rear region, assume that the position of the upper-left corner of the first rear region is used as an example. The sample closest to the upper-left corner position in the rear reference picture is determined, and by using this sample as the upper-left corner, the fourth rear region is determined. The width of the fourth rear region is blkW+4, and the height of the fourth rear region is blkH+4.
[0442] For example, referring to Figure 13, the corner position of the upper left corner of the first rear region D1 is (5.3, 6.2), and the position of the sample closest to the corner position (5.3, 6.2) is determined as (5, 6) in the rear reference picture Ref1. By using the sample at position (5, 6) as the upper left corner, the fourth rear region D2 is determined. The width of the fourth rear region D2 is blkW+4, and the height of the fourth rear region D2 is blkH+4.
[0443] (8) Determine a fifth rear region that includes a fourth rear region, where the center of the fourth rear region coincides with the center of the fifth rear region, the width of the fifth rear region is blkW+n+3, and the height of the fifth rear region is blkH+n+3. Obtain a picture block in the fifth rear region from the rear reference picture, and perform interpolation filtering on the picture block by using an interpolation filter to obtain a second rear prediction block of the first picture block, where the width of the second rear prediction block is blkW+4 and the height of the second rear prediction block is blkH+4.
[0444] For example, an 8-tap interpolation filter is used as an example. Referring to Figure 13, a fifth back region D3 containing a fourth back region D2 is determined. The center of the fourth back region D2 coincides with the center of the fifth back region D3, the width of the fifth back region D3 is blkW+11, and the height of the fifth back region D3 is blkH+11. The picture block in the fifth back region D3 is obtained from the back reference picture Ref1, and interpolation filtering is performed on the picture block by using an interpolation filter to obtain the second back prediction block of the first picture block B. The width of the second back prediction block is blkW+4, the height of the second back prediction block is blkH+4, and n is the number of taps in the interpolation filter.
[0445] Step 303: To obtain the first gradient matrix of the first picture block, perform a gradient operation on the second prediction block of the first picture block, where the width of the first gradient matrix is cuW+2 and the height of the first gradient matrix is cuH+2.
[0446] The first gradient matrix includes a first forward horizontal gradient matrix, a first forward vertical gradient matrix, a first backward horizontal gradient matrix, and a first backward vertical gradient matrix.
[0447] In this step, to obtain the first gradient matrix, perform a gradient operation on the second prediction block of the first picture block. For a detailed implementation process of this step, please refer to the detailed process of obtaining the first gradient matrix in step 104 of the embodiment shown in Figure 12. Further details will not be described again in this specification.
[0448] The first prediction block includes a second forward prediction block and a second backward prediction block. Based on the second forward prediction block, a second forward horizontal gradient matrix having a width of cuW+2 and a height of cuH+2, and a second forward vertical gradient matrix having a width of cuW+2 and a height of cuH+2 may be obtained. Based on the second backward prediction block, a second backward horizontal gradient matrix having a width of cuW+2 and a height of cuH+2, and a second backward vertical gradient matrix having a width of cuW+2 and a height of cuH+2 may be obtained.
[0449] Step 304: Determine the first prediction block of the first picture block within the second prediction block, where the width of the first prediction block is cuW+2 and the height of the first prediction block is cuH+2. be .
[0450] The center of the first prediction block coincides with the center of the second prediction block.
[0451] The first prediction block includes a first forward prediction block and a first backward prediction block.
[0452] In this step, a first forward prediction block having the width cuW+2 and height cuH+2 of the first picture block is determined within the second forward prediction block, and a first backward prediction block having the width cuW+2 and height cuH+2 of the first picture block is determined within the second backward prediction block.
[0453] Step 305: Based on the motion information of the first picture block, obtain the first predicted block of the first picture block, where the width of the first predicted block is cuW+2 and the height of the first predicted block is cuH+2.
[0454] For a detailed process of determining the first prediction block in this step, please refer to the relevant content in step 103 of the embodiment shown in Figure 12. Further details will not be described again in this specification.
[0455] Step 306: To obtain the first gradient matrix of the first picture block, perform a gradient operation on the first prediction block of the first picture block, where the width of the first gradient matrix is cuW and the height of the first gradient matrix is cuH.
[0456] The first gradient matrix includes a first forward horizontal gradient matrix, a first forward vertical gradient matrix, a first backward horizontal gradient matrix, and a first backward vertical gradient matrix.
[0457] For a detailed implementation process of this step, please refer to the relevant content in step 104 of the embodiment shown in Figure 12. Further details will not be described again in this specification.
[0458] Step 307: Perform the first expansion on the width and height of the first gradient matrix based on the gradients at the matrix edge positions of the first gradient matrix, such that the width and height of the first gradient matrix obtained after the first expansion are each 2 samples larger than the width and height of the first picture block.
[0459] For a detailed implementation process of this step, please refer to the relevant content in step 105 of the embodiment shown in Figure 12. Further details will not be described again in this specification.
[0460] Steps 308 to 310 are the same as steps 106 to 108, respectively, and will not be described in detail again herein.
[0461] The inter prediction method shown in Figure 15 may be summarized as steps 1 to 6, and steps 1 to 6 may be as follows.
[0462] Step 1: Determine the current CU's movement information.
[0463] The current CU motion information may be determined by using merge mode, AMVP mode (see background explanation), or another mode, which is not limited herein.
[0464] It should be noted that other methods for determining motion information may also be applicable to this application. Details are not described herein.
[0465] Step 2: Determine whether the current CU meets the BIO usage requirements.
[0466] If bidirectional prediction is performed on the current CU, and the relationship between the forward-referenced frame number POC_L0, the backward-referenced frame number POC_L1, and the current frame number POC_Cur satisfies the following equation, then the current CU satisfies the BIO usage conditions. (POC_L0-POC_Cur)*(POC_L1-POC_Cur)<0
[0467] It should be noted that other conditions of use for BIO may also apply to this application. Details are not described herein.
[0468] If the current CU meets the BIO usage requirements, step 3 is performed; otherwise, motion compensation is performed in another manner.
[0469] Step 3: Calculate the forward and backward predictors for the current CU.
[0470] If cuW is greater than or equal to VPDU_X, or if cuH is greater than or equal to VPDU_Y, then the forward predictor and backward predictor I (k) Motion compensation is performed by using motion information to obtain (i,j), where i = -2..cuW+1 and j = -2..cuH+1 (the prediction matrix (cuW+4)*(cuH+4) is obtained by using the same interpolation filter).
[0471] If cuW is less than VPDU_X or cuH is less than VPDU_Y, then the forward predictor and backward predictor I (k) Motion compensation is performed by using motion information to obtain (i,j), where i=-1..cuW and j=-1..cuH (a prediction matrix of (cuW+2)*(cuH+2) is obtained).
[0472] I obtained by performing interpolation using an 8-tap interpolation filter (k) At (i,j), i=0..cuW-1, and j=0..cuH-1, predictors at other locations (where an extension of one row and one column is performed) are obtained by performing interpolation using a bilinear interpolation filter.
[0473] It should be understood that predictors may be obtained by using VPDU as the minimum predictor acquisition unit, or by using a block smaller than VPDU as the minimum predictor acquisition unit. This is not limited to this.
[0474] It should be noted that the predictor for the extended region may also be obtained by using other methods, for example, by using an 8-tap interpolation filter, or by directly using a reference sample at an integer sample position. This is not limited herein.
[0475] It should be noted that in order to determine whether SAD is less than the threshold TH_CU, the SAD between the forward and backward predictors is calculated, and if SAD is less than the threshold TH_CU, BIO is not performed; otherwise, BIO is performed. Other determination methods may also be applied to this application, and their details are not described herein.
[0476] The formula for calculating SAD is as follows:
number
[0477] The threshold TH_CU may be set to (1<<(BD-8+shift))*cuW*cuH, and shift may be set to Max(2,14-BD).
[0478] Step 4: Calculate the horizontal and vertical gradients based on the current forward and backward predictors of the CU.
[0479] If cuW is greater than or equal to VPDU_X, or if cuH is greater than or equal to VPDU_Y, then the horizontal and vertical gradients are...
number
number
[0480] If cuW is less than VPDU_X or cuH is less than VPDU_Y, then the horizontal and vertical gradients are...
number
number
[0481] Step 5: If cuW is less than VPDU_X and cuH is less than VPDU_Y, perform padding on the current CU's forward and backward predictors, as well as the horizontal and vertical gradients.
[0482] I (k) (i,j) and,
number
number
[0483] Step 6: Derive the refined motion vectors for each 4x4 subblock, and then perform weighting.
[0484] For each 4x4 subblock, vx and vy are obtained according to equation (2). Finally, weighting is performed according to equation (6) to obtain the predictor for each 4x4 subblock.
[0485] It should be noted that the SAD between the forward and backward predictors of each 4x4 subblock may be calculated to determine whether the SAD is less than the threshold TH_SCU. If the SAD is less than the threshold TH_SCU, a weighted average is performed directly; otherwise, vx and vy are obtained according to equation (2), and then weighting is performed according to equation (6). Other determination methods may also be applied to this application and are not described in detail herein. TU_SCU may be set to 1 << (BD - 3 + shift).
[0486] In this embodiment of the present application, BIO prediction is performed on the boundaries of the VPDU and the boundaries of the CU in the same manner. When the CU contains multiple VPDUs, the complexity of implementing motion-compensated prediction is reduced.
[0487] In this embodiment of the present application, if cuW is greater than or equal to Width, and / or cuH is Height If the above conditions are met, the second prediction block of the first picture block is obtained based on the motion information of the first picture block. Since the width of the second prediction block is cuW+4 and the height of the second prediction block is cuH+4, a gradient operation is performed on the second prediction block of the first picture block to obtain a first gradient matrix having a width of cuW+2 and a height of cuH+2, so that the extension process from the edges of the first gradient matrix can be omitted, thereby improving interpretation efficiency.
[0488] Figure 17 is a schematic flowchart of the method according to the embodiment of this application. As shown in the figure, an interpretation method is provided, which includes the following steps.
[0489] S1201: Obtain motion information of the picture block to be processed, where the picture block to be processed includes multiple virtual pipeline data units, and each virtual pipeline data unit includes at least one basic processing unit.
[0490] S1202: Based on motion information, obtain the predictor matrix for each virtual pipeline data unit.
[0491] S1203: Based on each predictor matrix, calculate the horizontal and vertical predictor gradient matrices for each virtual pipeline data unit.
[0492] S1204: Based on the predictor matrix, the horizontal prediction gradient matrix, and the vertical prediction gradient matrix, the motion information refinement value for each basic processing unit within each virtual pipeline data unit is calculated.
[0493] In a feasible implementation, the step of obtaining a predictor matrix for each virtual pipeline data unit based on motion information includes the step of obtaining an initial predictor matrix for each virtual pipeline data unit based on motion information, wherein the size of the initial predictor matrix is equal to the size of the virtual pipeline data unit, and the step of using the initial predictor matrix as a predictor matrix.
[0494] In a feasible implementation, after the step of obtaining an initial prediction matrix for each virtual pipeline data unit, the method further includes the step of performing sample augmentation on the edges of the initial prediction matrix to obtain an augmented prediction matrix, wherein the size of the augmented prediction matrix is greater than the size of the initial prediction matrix, and correspondingly the step of using the initial prediction matrix as the predictor matrix includes the step of using the augmented prediction matrix as the predictor matrix.
[0495] In a feasible implementation, the step of performing sample extension on the edges of the initial prediction matrix includes the step of obtaining sample values of samples outside the initial prediction matrix based on interpolation of sample values of samples within the initial prediction matrix, or the step of using the sample values of samples at the edges of the initial prediction matrix as sample values of samples outside the initial prediction matrix that are adjacent to the edges.
[0496] In a feasible implementation, a virtual pipeline data unit includes a plurality of motion compensation units, and the step of obtaining the predictor matrix of each virtual pipeline data unit based on motion information includes the step of obtaining the compensation value matrix of each motion compensation unit based on motion information, and the step of combining the compensation value matrices of the plurality of motion compensation units to obtain the predictor matrix.
[0497] In a feasible implementation, the step of calculating the horizontal and vertical gradient matrices for each virtual pipeline data unit based on each predictor matrix includes the step of performing horizontal gradient calculations and vertical gradient calculations separately on the predictor matrix in order to obtain the horizontal and vertical gradient matrices.
[0498] In a feasible implementation, prior to the step of calculating the motion information refinement value for each basic processing unit in each virtual pipeline data unit based on a predictor matrix, a horizontal predictor gradient matrix, and a vertical predictor gradient matrix, the method further includes the step of performing sample expansion at the edges of the predictor matrix to obtain a padding predictor matrix, wherein the padding predictor matrix has a predetermined size; and the step of performing gradient expansion separately at the edges of the horizontal predictor gradient matrix and the edges of the vertical predictor gradient matrix to obtain a padding horizontal gradient matrix and a padding vertical gradient matrix, wherein the padding horizontal gradient matrix and the padding vertical gradient matrix each have a predetermined size; and correspondingly, the step of calculating the motion information refinement value for each basic processing unit in each virtual pipeline data unit based on a predictor matrix, a horizontal predictor gradient matrix, and a vertical predictor gradient matrix includes the step of calculating the motion information refinement value for each basic processing unit in each virtual pipeline data unit based on a padding predictor matrix, a padding horizontal gradient matrix, and a padding vertical gradient matrix.
[0499] In a feasible implementation, prior to the step of performing sample augmentation on the edges of the predictor matrix, the method further includes the step of determining that the size of the predictor matrix is smaller than a pre-set size.
[0500] In a feasible implementation, prior to the step of performing gradient augmentation on the edges of the horizontal predictive gradient matrix and the edges of the vertical predictive gradient matrix, the method further includes the step of determining that the size of the horizontal predictive gradient matrix and / or the size of the vertical predictive gradient matrix is smaller than a predetermined size.
[0501] In a feasible implementation, after the step of calculating the motion information refinement value for each basic processing unit within each virtual pipeline data unit, the method further includes the step of obtaining the predictor for each basic processing unit based on the predictor matrix of the virtual pipeline data unit and the motion information refinement value for each basic processing unit within the virtual pipeline data unit.
[0502] In a feasible implementation, the method is used for bidirectional prediction, and accordingly, motion information includes a first reference frame list motion information and a second reference frame list motion information, the predictor matrix includes a first predictor matrix and a second predictor matrix, the first predictor matrix is obtained based on the first reference frame list motion information, the second predictor matrix is obtained based on the second reference frame list motion information, and the horizontal prediction gradient matrix includes a first horizontal prediction gradient matrix and a second horizontal prediction gradient matrix, the first horizontal prediction gradient matrix is calculated based on the first predictor matrix, and the second The horizontal prediction gradient matrix is calculated based on the second predictor matrix, the vertical prediction gradient matrix includes the first vertical prediction gradient matrix and the second vertical prediction gradient matrix, the first vertical prediction gradient matrix is calculated based on the first predictor matrix, the second vertical prediction gradient matrix is calculated based on the second predictor matrix, the motion information refinement value includes the first reference frame list motion information refinement value and the second reference frame list motion information refinement value, the first reference frame list motion information refinement value is calculated based on the first predictor matrix, the first horizontal prediction gradient matrix and the first vertical prediction gradient matrix, and2 The reference frame list motion information refinement value is, 2 The predictor matrix and the 2 It is calculated based on the horizontal prediction gradient matrix and the second vertical prediction gradient matrix.
[0503] In a feasible implementation, prior to the step of performing sample augmentation on the edges of the initial prediction matrix, the method further includes the step of determining that the time domain position of the picture frame in which the picture block to be processed is located is between a first reference frame indicated by a first reference frame list motion information and a second reference frame indicated by a second reference frame list motion information.
[0504] In a feasible implementation, after the step of obtaining the predictor matrix for each virtual pipeline data unit, the method further includes the step of determining that the difference between the first predictor matrix and the second predictor matrix is less than a first threshold.
[0505] In a feasible implementation, the motion information refinement value of a basic processing unit corresponds to one basic predictor matrix in the predictor matrix, and prior to the step of calculating the motion information refinement value of each basic processing unit in each virtual pipeline data unit based on the predictor matrix, the horizontal predictor gradient matrix, and the vertical predictor gradient matrix, the method further includes the step of determining that the difference between a first basic predictor matrix and a second basic predictor matrix is less than a second threshold.
[0506] In a feasible implementation, the size of the basic processing unit is 4x4.
[0507] In a feasible implementation, the width of the virtual pipeline data unit is W, the height of the virtual pipeline data unit is H, and the size of the augmented prediction matrix is (W+n+2)×(H+n+2). Correspondingly, the size of the horizontal prediction gradient matrix is (W+n)×(H+n), and the size of the vertical prediction gradient matrix is (W+n)×(H+n), where W and H are positive integers and n is an even number.
[0508] In a feasible implementation, n is 0, 2, or -2.
[0509] In a feasible implementation, prior to the step of obtaining motion information of the picture block to be processed, the method further includes the step of determining that the picture block to be processed contains multiple virtual pipeline data units.
[0510] Figure 18 is a schematic flowchart of the method according to the embodiment of this application. As shown in the figure, an interpretation device is provided. An acquisition module 1301 configured to acquire motion information of a picture block to be processed, wherein the picture block to be processed includes a plurality of virtual pipeline data units, and each virtual pipeline data unit includes at least one basic processing unit, A compensation module 1302 is configured to obtain the predictor matrix for each virtual pipeline data unit based on motion information, A computing module 1303 is configured to calculate the horizontal and vertical gradient matrices for each virtual pipeline data unit based on each predictor matrix, A refinement module 1304 is configured to calculate the motion information refinement value for each basic processing unit within each virtual pipeline data unit based on the predictor matrix, the horizontal prediction gradient matrix, and the vertical prediction gradient matrix. Includes.
[0511] In a feasible implementation, the compensation module 1302 is configured to perform, specifically, an operation to obtain an initial prediction matrix for each virtual pipeline data unit based on motion information, wherein the size of the initial prediction matrix is equal to the size of the virtual pipeline data unit, and an operation to use the initial prediction matrix as a predictor matrix.
[0512] In a feasible implementation, the compensation module 1302 is configured to perform, specifically, an operation in which sample augmentation is performed at the edges of the initial prediction matrix in order to obtain an augmented prediction matrix, wherein the size of the augmented prediction matrix is greater than the size of the initial prediction matrix, and an operation in which the augmented prediction matrix is used as the predictor matrix.
[0513] In a feasible implementation, the compensation module 1302 is configured to either obtain sample values of samples outside the initial prediction matrix based on interpolation of sample values of samples within the initial prediction matrix, or to use the sample values of samples at the edges of the initial prediction matrix as sample values of samples outside the initial prediction matrix that are adjacent to the edges.
[0514] In a feasible implementation, a virtual pipeline data unit includes multiple motion compensation units, and the compensation module is configured to obtain the compensation value matrix of each motion compensation unit based on motion information, and to combine the compensation value matrices of the multiple motion compensation units to obtain a predictor matrix.
[0515] In a feasible implementation, the computation module 1303 is configured to perform horizontal gradient calculations and vertical gradient calculations separately on the predictor matrix in order to obtain the horizontal predicted gradient matrix and the vertical predicted gradient matrix.
[0516] In a feasible implementation, the device further includes a padding module 1305 configured to perform the following operations: an operation to perform sample expansion at the edges of the predictor matrix in order to obtain a padding prediction matrix, wherein the padding prediction matrix has a preset size; an operation to perform gradient expansion separately at the edges of the horizontal prediction gradient matrix and the edges of the vertical prediction gradient matrix in order to obtain a padding horizontal gradient matrix and a padding vertical gradient matrix, wherein the padding horizontal gradient matrix and the padding vertical gradient matrix each have a preset size; and an operation to calculate motion information refinement values for each basic processing unit in each virtual pipeline data unit based on the padding prediction matrix, the padding horizontal gradient matrix, and the padding vertical gradient matrix.
[0517] In a feasible implementation, the device further includes a decision module 1306 configured to determine that the size of the predictor matrix is less than a preset size.
[0518] In a feasible implementation, the decision module 1306 is further configured to determine that the size of the horizontal prediction gradient matrix and / or the size of the vertical prediction gradient matrix are less than a preset size.
[0519] In a seventh executable implementation, the refinement module 1304 is further configured to obtain predictors for each basic processing unit based on the predictor matrix of the virtual pipeline data unit and the motion information refinement values of each basic processing unit within the virtual pipeline data unit.
[0520] In a feasible implementation, the device is used for bidirectional prediction, and accordingly, motion information includes a first reference frame list motion information and a second reference frame list motion information, the predictor matrix includes a first predictor matrix and a second predictor matrix, the first predictor matrix is obtained based on the first reference frame list motion information, the second predictor matrix is obtained based on the second reference frame list motion information, the horizontal prediction gradient matrix includes a first horizontal prediction gradient matrix and a second horizontal prediction gradient matrix, the first horizontal prediction gradient matrix is calculated based on the first predictor matrix, and the second The horizontal prediction gradient matrix is calculated based on the second predictor matrix, the vertical prediction gradient matrix includes the first vertical prediction gradient matrix and the second vertical prediction gradient matrix, the first vertical prediction gradient matrix is calculated based on the first predictor matrix, the second vertical prediction gradient matrix is calculated based on the second predictor matrix, the motion information refinement value includes the first reference frame list motion information refinement value and the second reference frame list motion information refinement value, the first reference frame list motion information refinement value is calculated based on the first predictor matrix, the first horizontal prediction gradient matrix and the first vertical prediction gradient matrix, and 2 The reference frame list motion information refinement value is the first predictor matrix and the second 2 It is calculated based on the horizontal prediction gradient matrix and the second vertical prediction gradient matrix.
[0521] In an executable implementation, the determination module 1306 is further configured to determine that the time domain position of the picture frame in which the picture block to be processed is located is between a first reference frame indicated by a first reference frame list motion information and a second reference frame indicated by a second reference frame list motion information.
[0522] In a feasible implementation of the seventh embodiment, the decision module 1306 is further configured to determine that the difference between the first predictor matrix and the second predictor matrix is less than a first threshold.
[0523] In a feasible implementation, the decision module 1306 is further configured to determine that the difference between the first basic predictor matrix and the second basic predictor matrix is less than a second threshold.
[0524] In a feasible implementation, the size of the basic processing unit is 4x4.
[0525] In a feasible implementation, the width of the virtual pipeline data unit is W, the height of the virtual pipeline data unit is H, and the size of the augmented prediction matrix is (W+n+2)×(H+n+2). Correspondingly, the size of the horizontal prediction gradient matrix is (W+n)×(H+n), and the size of the vertical prediction gradient matrix is (W+n)×(H+n), where W and H are positive integers and n is an even number.
[0526] In a feasible implementation, n is 0, 2, or -2.
[0527] In a feasible implementation, the decision module 1306 is further configured to determine that the picture block to be processed contains multiple virtual pipeline data units.
[0528] Figure 19 is a schematic flowchart of the method according to the embodiment of this application. As shown in the figure, an interpretation device 1400 is provided. A determination module 1401 is configured to determine multiple first picture blocks within a picture block to be processed based on a pre-configured picture division width, a pre-configured picture division height, and the width and height of the picture block to be processed. A prediction module 1402 is configured to individually perform bidirectional optical flow prediction for multiple first picture blocks in order to obtain a predictor for each first picture block, A combination module 1403 configured to obtain a predictor for the picture block to be processed using a combination of multiple predictors for the first picture block, Includes.
[0529] In a feasible implementation, the decision module 1401 is: To determine the width of the first picture block, compare the pre-configured picture division width with the width of the picture block to be processed. To determine the height of the first picture block, compare the pre-configured picture division height with the height of the picture block to be processed. Based on the width and height of the first picture block, determine multiple first picture blocks within the picture block to be processed. It is configured in this way.
[0530] In a feasible implementation, the width of the first picture block is the smaller of the pre-configured picture division width and the width of the picture block to be processed, and the height of the first picture block is the smaller of the pre-configured picture division height and the height of the picture block to be processed.
[0531] In an executable implementation, the prediction module 1402 obtains the first predicted block of the first picture block based on the motion information of the picture block to be processed, To obtain the first gradient matrix of the first picture block, perform a gradient operation on the first prediction block, Based on the first prediction block and the first gradient matrix, the motion information refinement value for each basic processing unit in the first picture block is calculated. Based on the refined motion information values of each basic processing unit, obtain the predictor for the first picture block. It is configured in this way.
[0532] In a feasible implementation, the device 1400 further includes a first expansion module 1404.
[0533] The first extension module is configured to perform a first extension on the width and height of the first prediction block based on the sample values of the block edge positions of the first prediction block, such that the width and height of the first prediction block obtained after the first extension are each two samples greater than the width and height of the first picture block, and / or to perform a first extension on the width and height of the first gradient matrix based on the gradients of the matrix edge positions of the first gradient matrix, such that the width and height of the first gradient matrix obtained after the first extension are each two samples greater than the width and height of the first picture block.
[0534] Accordingly, the prediction module 1402 is configured to calculate motion information refinement values for each basic processing unit in the first picture block based on the first prediction block and / or the first gradient matrix obtained after the first extension.
[0535] In a feasible implementation, the device further includes a second extension module 1405.
[0536] The second extension module is configured to perform interpolation filtering on the sample values of the block edge regions of the first prediction block, or to duplicate the sample values of the block edge locations of the first prediction block, in order to perform the second extension on the width and height of the first prediction block.
[0537] Accordingly, the prediction module 1402 is configured to perform gradient calculations on the first prediction block acquired after the second extension.
[0538] In a feasible implementation, the first prediction block includes a forward prediction block and a backward prediction block, and the first gradient matrix includes a forward horizontal gradient matrix, a forward vertical gradient matrix, a backward horizontal gradient matrix, and a backward vertical gradient matrix.
[0539] In a feasible implementation, the pre-configured picture division width is 64, 32, or 16, and the pre-configured picture division height is 64, 32, or 16.
[0540] In a feasible implementation, the basic processing unit is a 4x4 sample matrix.
[0541] In this embodiment of the present application, the determination module determines a plurality of first picture blocks within a picture block to be processed based on a pre-configured picture division width, a pre-configured picture division height, and the width and height of the picture block to be processed. Thus, the size of the first picture blocks is limited by the pre-configured picture division width and pre-configured picture division height, and the area of each determined first picture block is not too large, so that hardware resources such as memory resources can be consumed less, the complexity of performing interpretation can be reduced, and processing efficiency can be improved.
[0542] Those skilled in the art will understand that the functions described with reference to the various exemplary logic blocks, modules, and algorithmic steps disclosed and described herein may be implemented by hardware, software, firmware, or any combination thereof. When implemented by software, the functions described with reference to the exemplary logic blocks, modules, and steps may be recorded as one or more instructions or codes in a computer-readable medium, or transmitted through a computer-readable medium, and executed by a hardware-based processing unit. The computer-readable medium may include computer-readable storage media corresponding to tangible media such as data storage media, or any communication media that facilitates the transmission of computer programs from one place to another (e.g., by communication protocols). Thus, the computer-readable medium may generally correspond to (1) non-temporary tangible computer-readable storage media, or (2) communication media such as signals or carriers. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to extract instructions, codes, and / or data structures for implementing the technology described herein. Computer program products may include computer-readable media.
[0543] As an example, and not an limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other compact disk storage devices, magnetic disk storage devices or other magnetic storage devices, flash memory, or any other media that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. In addition, any connection is appropriately called a computer-readable medium. For example, if instructions are transmitted from a website, server, or another remote source via coaxial cable, optical fiber, twisted pair, digital subscriber line (DSL), or wireless technology such as infrared, radio, or microwave, then coaxial cable, optical fiber, twisted pair, DSL, or wireless technology such as infrared, radio, or microwave are included within the definition of a medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carriers, signals, or any other temporary media, and are actually directed toward non-temporary, tangible storage media. As used herein, the terms "disk" and "disc" include compact discs (CDs), laser discs, optical discs, digital multipurpose discs (DVDs), and Blu-ray® discs. A disk typically reproduces data magnetically, while a disc reproduces data optically using a laser. Any combination of the aforementioned items should also fall within the scope of computer-readable media.
[0544] Instructions may be executed by one or more processors, such as digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Thus, the term “processor” as used herein may refer to any of the aforementioned structures or any other structure suitable for implementing the techniques described herein. In addition, in some embodiments, the functions described with reference to the exemplary logic blocks, modules, and steps described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into a composite codec. Furthermore, all techniques may be implemented within one or more circuits or logic elements.
[0545] The technology described herein may be implemented in a variety of devices or apparatus, including wireless handsets, integrated circuits (ICs), or IC sets (e.g., chipsets). Various components, modules, or units are described herein to highlight the functional aspects of apparatus configured to implement the disclosed technology, but these are not necessarily implemented by different hardware units. In practice, as described above, the various units may be combined with a codec hardware unit in combination with appropriate software and / or firmware, or provided by an interoperable hardware unit (including one or more processors as described above).
[0546] In the embodiments described above, each embodiment has its own focus. For aspects not described in detail in one embodiment, please refer to the relevant descriptions in other embodiments.
[0547] The foregoing description is merely an example of a particular implementation of this application and is not intended to limit the scope of protection of this application. Any modifications or substitutions that are readily conceivable by a person skilled in the art within the scope of the art disclosed herein shall fall within the scope of protection of this application. Accordingly, the scope of protection of this application shall be subject to the scope of protection of the claims. [Explanation of Symbols]
[0548] 10 Video Coding Systems 12 Source Devices 13 links 14 Destination device 16 Picture Sources 17 Original picture data 18 Picture Preprocessor 19 Pre-processed pictures, pre-processed picture data 20 encoders, video encoders 21 Encode bitstream, encode picture data 22 Communication Interfaces 28 Communication Interfaces 30 decoders, video decoders 31 Decrypted picture data, decrypted picture 32 Picture Post-Processors 33 Post-processed picture data 34 Display Devices 40 Video Coding Systems 41 Imaging devices 42 Antennas 43 processors 44 memory 45 Display Devices 46 Processing Units 201 Pictures 202 Input Signal 203 Picture Block, Block 204 Residual Calculation Unit 205 Residual Block 206 Conversion Processing Unit 207 Conversion coefficient 208 Quantization Units 209 Quantization conversion coefficients, Quantization residual coefficients 210 Inverse Quantization Unit 211 Inverse quantization coefficient, inverse quantization residual coefficient 212 Inverse Transform Processing Unit 213 Inverse transform block, inverse transform inverse quantization block, inverse transform residual block, reconstructed residual block 214 Reconfiguration Unit, Adder 215 Reconstructed Blocks 216 buffers, buffer units 220 Loop Filter Units, Loop Filters, Filters 221 Filtered block, filtered reconfigured block, reconfigured filtered block 230 Decoded picture buffer, Decoded picture buffer unit, DPB 231 Reference picture data, decoded picture 244 Interpretation Units 245 Interpretation Block, Prediction Block 254 Intra Prediction Units, Intra Prediction 255 Intra Prediction Blocks, Prediction Blocks 260 Prediction Processing Units 262 Mode Selection Unit 265 Prediction Blocks 270 Entropy Coding Units 272 Output 302 Compensation Module 304 Entropy Decoding Unit 309 Quantization coefficient 310 Inverse Quantization Unit 312 Inverse Transform Processing Unit 313 Reconstructed residual blocks, inverse transform blocks 314 Reconfiguration Unit, Adder 315 Reconstructed Blocks 316 buffers 320 Loop Filters, Loop Filter Units, Filters 321 Filtered blocks, decoded video blocks 330 Decode picture buffer, DPB 332 output 344 Interpretation Units 354 Intra Prediction Units 360 Predictive Processing Unit 362 Mode Selection Unit 365 Prediction Block 400 video coding devices, video encoding devices, video decoding devices 410 Input Ports 420 Receiver Unit (Rx), Receiver Unit 430 processors 440 Transmitter Unit (Tx), Transmitter Unit 450 output ports 460 memory 470 coding modules, encoding modules, decoding modules, encoding / decoding modules 500 devices, coding devices 510 Processor 530 memory 531 Data 533 Operating Systems 535 Application Programs 550 bus system, bus 570 displays 1301 Acquisition Module 1302 Compensation Module 1303 Computation Module 1304 Refinement Module 1305 Padding Module 1306 Decision Module 1400 Interpretation device, device 1401 Decision Module 1402 Prediction Module 1403 Combination Module 1404 First expansion module 1405 Second expansion module
Claims
1. Interpretation method applicable to an encoding device or decoding device, A step of obtaining a picture block to be processed, wherein the picture block to be processed is obtained by dividing an image; A step of determining a plurality of first picture blocks within a picture block to be processed, based on a pre-set picture division width, a pre-set picture division height, and the width and height of the picture block to be processed, wherein the width of the first picture block among the plurality of first picture blocks is equal to the smaller of the pre-set picture division width and the width of the picture block to be processed, and the height of the first picture block is equal to the smaller of the pre-set picture division height and the height of the picture block to be processed, A step of obtaining predictors for each of the first picture blocks, comprising individually performing bidirectional optical flow prediction for a plurality of first picture blocks, wherein the first picture block comprises a plurality of basic processing units, the predictors of the first picture block comprise predictors for each sample within each basic processing unit in the first picture block, and the predictors for each sample within a basic processing unit are calculated based on a forward basic prediction block, a backward basic prediction block, a forward horizontal basic gradient matrix, a forward vertical basic gradient matrix, a backward horizontal basic gradient matrix, and a backward vertical basic gradient matrix of the basic processing unit. The steps include obtaining a predictor for the picture block to be processed using the combination of predictors for the plurality of first picture blocks, and Methods that include...
2. The method according to claim 1, wherein the preset picture division width is 64, 32, or 16, and the preset picture division height is 64, 32, or 16.
3. The method according to claim 1 or 2, wherein the basic processing unit is a 4x4 sample matrix.
4. An interpretermination device, wherein the interpretermination device is configured for encoding or decoding, A module configured to acquire a picture block to be processed, wherein the picture block to be processed is acquired by dividing an image; A determination module configured to determine a plurality of first picture blocks within a picture block to be processed based on a pre-configured picture division width, a pre-configured picture division height, and the width and height of the picture block to be processed, wherein the width of the first picture block among the plurality of first picture blocks is equal to the smaller of the pre-configured picture division width and the width of the picture block to be processed, and the height of the first picture block is equal to the smaller of the pre-configured picture division height and the height of the picture block to be processed, A prediction module configured to individually perform bidirectional optical flow prediction on a plurality of first picture blocks in order to obtain predictors for each first picture block, wherein the first picture block comprises a plurality of basic processing units, the predictors of the first picture block comprise predictors for each sample within each basic processing unit in the first picture block, and the predictors for each sample within the basic processing unit are calculated based on the forward basic prediction block, the backward basic prediction block, the forward horizontal basic gradient matrix, the forward vertical basic gradient matrix, the backward horizontal basic gradient matrix, and the backward vertical basic gradient matrix of the basic processing unit, the prediction module, A combination module configured to obtain the predictor of the picture block to be processed using the combination of the predictors of the plurality of first picture blocks, A device equipped with the following features.
5. The apparatus according to claim 4, wherein the preset picture division width is 64, 32, or 16, and the preset picture division height is 64, 32, or 16.
6. The apparatus according to claim 4 or 5, wherein the basic processing unit is a 4x4 sample matrix.
7. An interpretation device comprising coupled memory and a processor, wherein the processor invokes program code stored in the memory to perform the method according to any one of claims 1 to 3.
8. A computer program for generating a bitstream by causing a processor to execute an interpretation method, wherein the interpretation method is A step of obtaining a picture block to be processed, wherein the picture block to be processed is obtained by dividing an image; A step of determining a plurality of first picture blocks within a picture block to be processed, based on a pre-set picture division width, a pre-set picture division height, and the width and height of the picture block to be processed, wherein the width of the first picture block among the plurality of first picture blocks is equal to the smaller of the pre-set picture division width and the width of the picture block to be processed, and the height of the first picture block is equal to the smaller of the pre-set picture division height and the height of the picture block to be processed, A step of obtaining predictors for each of the first picture blocks, comprising individually performing bidirectional optical flow prediction for a plurality of first picture blocks, wherein the first picture block comprises a plurality of basic processing units, the predictors of the first picture block comprise predictors for each sample within each basic processing unit in the first picture block, and the predictors for each sample within a basic processing unit are calculated based on a forward basic prediction block, a backward basic prediction block, a forward horizontal basic gradient matrix, a forward vertical basic gradient matrix, a backward horizontal basic gradient matrix, and a backward vertical basic gradient matrix of the basic processing unit. The steps include obtaining a predictor for the picture block to be processed using the combination of predictors for the plurality of first picture blocks, and A computer program that includes [this].