A video processing method and related devices
By performing composite prediction of the current block in video encoding, and determining the target weighting based on the importance of the reference prediction value for weighting, the problem of low composite prediction accuracy is solved and the encoding performance is improved.
Patent Information
- Application Number
- CN202211230593.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-30
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-09-30
AI Technical Summary
In the existing video encoding technology, the accuracy of composite prediction in weighted prediction is not high, which affects the encoding performance.
By performing composite predictions on the current block in the video code stream, N reference prediction values are obtained, and target weight reorganization is determined based on the importance of these reference prediction values, and weighted prediction processing is performed to improve prediction accuracy.
Improve the accuracy of weighted prediction during the composite prediction process, thereby improving coding performance.
Smart Images

Figure CN115633168B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of audio - video technologies, specifically to the field of video coding technologies, and particularly to a video processing method, a video processing device, a computer device, a computer - readable storage medium, and a computer program product. Background Art
[0002] In existing video coding technologies, a block - based hybrid coding framework is adopted. The original video data can be divided into a series of coding blocks (CodingUnit, CU), and video coding methods such as prediction, transformation, and entropy coding are combined to achieve compression of video data. In order to obtain better prediction effects, current mainstream video coding standards such as the AV1 (the first - generation video coding standard developed by the Alliance for Open Media, Alliance for Open Media Video 1) standard and the AV2 (the second - generation video coding standard developed by the Alliance for Open Media, Alliance for Open Media Video 2) which is under research and development in the Alliance for Open Media (AOM) include a prediction mode called compound prediction. This compound - prediction mode allows the use of multiple reference video signals for weighted prediction. However, it is found in practice that the current compound prediction has low prediction accuracy when performing weighted prediction. Summary of the Invention
[0003] Embodiments of the present application provide a video processing method and related devices, which can improve the accuracy of weighted prediction in the compound - prediction process, thereby improving the coding performance.
[0004] On the one hand, embodiments of the present application provide a video processing method, which includes:
[0005] Performing compound prediction on a current block in a video bitstream to obtain N reference prediction values of the current block, where N is an integer greater than 1;
[0006] Determining a target weight group for weighted prediction for the current block according to the importance of the N reference prediction values, where the target weight group includes one or more weight values;
[0007] Performing weighted - prediction processing on the N reference prediction values based on the target weight group to obtain a prediction value of the current block, and the prediction value of the current block is used to reconstruct a decoded image corresponding to the current block.
[0008] On the one hand, embodiments of the present application provide a video processing method, which includes:
[0009] Performing partitioning processing on a current frame in a video to obtain a current block;
[0010] Perform a composite prediction on the current block to obtain N reference prediction values for the current block;
[0011] Determine a target weight group for weighted prediction for the current block according to the importance of the N reference prediction values, where the target weight group includes one or more weight values;
[0012] Perform a weighted prediction process on the N reference prediction values based on the target weight group to obtain a prediction value for the current block, and the prediction value for the current block is used to reconstruct the decoded image corresponding to the current block;
[0013] Encode the video based on the prediction value of the current block to generate a video bitstream.
[0014] On the one hand, an embodiment of the present application provides a video processing device, and the device includes:
[0015] A processing unit, configured to perform a composite prediction on the current block in the video bitstream to obtain N reference prediction values for the current block, where N is an integer greater than 1;
[0016] A determination unit, configured to determine a target weight group for weighted prediction for the current block according to the importance of the N reference prediction values, where the target weight group includes one or more weight values;
[0017] The processing unit is further configured to perform a weighted prediction process on the N reference prediction values based on the target weight group to obtain a prediction value for the current block, and the prediction value for the current block is used to reconstruct the decoded image corresponding to the current block.
[0018] In one embodiment, the N reference prediction values are derived from N reference blocks; one reference prediction value corresponds to one reference block; the video frame where the reference block is located is a reference frame, and the video frame where the current block is located is the current frame;
[0019] The positional relationship between the N reference blocks and the current block includes any one of the following:
[0020] The N reference blocks are respectively located in N reference frames, and the N reference frames and the current frame belong to different video frames in the video bitstream;
[0021] The N reference blocks are located in the same reference frame, and the same reference frame and the current frame belong to different video frames in the video bitstream;
[0022] One or more of the N reference blocks are located in the current frame, the remaining reference blocks among the N reference blocks are located in one or more reference frames, and the one or more reference frames and the current frame belong to different video frames in the video bitstream;
[0023] The N reference blocks and the current block are both located in the current frame.
[0024] In one embodiment, the processing unit is further configured to:
[0025] Determine whether the current block meets the conditions for adaptive weighted prediction;
[0026] If the current block meets the conditions for adaptive weighted prediction, determine a target weight group for weighted prediction for the current block according to the importance of N reference prediction values.
[0027] In one embodiment, the current block meeting the conditions for adaptive weighted prediction includes at least one of the following:
[0028] The sequence header of the frame sequence to which the current block belongs includes a first indication field, and the first indication field indicates that the coded blocks in the frame sequence are allowed to use adaptive weighted prediction;
[0029] The header of the current slice to which the current block belongs includes a second indication field, and the second indication field indicates that the coded blocks in the current slice are allowed to use adaptive weighted prediction;
[0030] The frame header of the current frame in which the current block is located includes a third indication field, and the third indication field indicates that the coded blocks in the current frame are allowed to use adaptive weighted prediction;
[0031] During the composite prediction process, the current block uses at least two reference frames for inter-frame prediction;
[0032] During the composite prediction process, the current block uses at least one reference frame for inter-frame prediction and uses the current frame for intra-frame prediction;
[0033] The motion type of the current block is a specified motion type;
[0034] The current block uses a preset motion vector prediction mode;
[0035] The current block uses a preset interpolation filter;
[0036] The current block does not use a specific coding tool;
[0037] The reference frames used by the current block during the composite prediction process meet specific conditions, and the specific conditions include one or more of the following: the orientation relationship between the reference frames used and the current frame in the video bitstream meets a preset relationship; the absolute value of the importance difference between the reference prediction values corresponding to the reference frames used is greater than or equal to a preset threshold;
[0038] Among them, the orientation relationship meeting the preset relationship includes any one of the following: all the reference frames used are before the current frame; all the reference frames used are after the current frame; some of the reference frames used are before the current frame, and the remaining reference frames are after the current frame.
[0039] In one embodiment, the video bitstream includes one or more weight lists, each weight list includes one or more weight groups, each weight group includes one or more weight values, the number of weight values included in each weight group is allowed to be the same or different, and the values of the weight values included in each weight group are allowed to be the same or different; the determining unit is specifically configured to:
[0040] Determine a target weight list from one or more weight lists according to the importance of N reference prediction values; and,
[0041] Select a target weight group for weighted prediction from the target weight list.
[0042] In one embodiment, the number of weight lists in the video bitstream is M + 1, one weight list corresponds to one threshold interval, and M is a positive integer greater than or equal to 1; the determining unit is specifically configured to:
[0043] Obtain the importance measurement values of N reference prediction values, and calculate the importance difference between the N reference prediction values. The importance difference between any two reference prediction values is measured by the difference between the importance measurement values of any two reference prediction values;
[0044] Determine the threshold interval in which the absolute value of the importance difference between the N reference prediction values is located;
[0045] Determine the weight list corresponding to the threshold interval in which the absolute value of the importance difference between the N reference prediction values is located as the target weight list.
[0046] In one embodiment, the N reference prediction values of the current block include a first reference prediction value and a second reference prediction value; the video bitstream includes a first weight list and a second weight list; the determining unit is specifically configured to:
[0047] Compare the magnitude of the importance measurement value of the first reference prediction value and the importance measurement value of the second reference prediction value;
[0048] If the importance measurement value of the first reference prediction value is greater than the importance measurement value of the second reference prediction value, determine the first weight list as the target weight list;
[0049] If the importance measurement value of the first reference prediction value is less than or equal to the importance measurement value of the second reference prediction value, determine the second weight list as the target weight list;
[0050] Wherein, the weight values in the first weight list are opposite to the weight values in the second weight list.
[0051] In one embodiment, among the N reference prediction values of the current block, there are a first reference prediction value and a second reference prediction value; the video bitstream includes a first weight list, a second weight list, and a third weight list; the determining unit is specifically configured to:
[0052] Call a mathematical symbol function to process the difference between the importance measure value of the first reference prediction value and the importance measure value of the second reference prediction value to obtain a symbol value;
[0053] If the symbol value is a first preset value, determine the first weight list as the target weight list;
[0054] If the symbol value is a second preset value, determine the second weight list as the target weight list;
[0055] If the symbol value is a third preset value, determine the third weight list as the target weight list;
[0056] Wherein, the first weight list, the second weight list, and the third weight list are different weight lists respectively; or, two of the first weight list, the second weight list, and the third weight list are allowed to be the same weight list.
[0057] In one embodiment, any one of the N reference prediction values is denoted as reference prediction value i, the reference prediction value i is derived from reference block i, and the video frame where the reference block i is located is reference frame i; i is an integer and i is less than or equal to N; the video frame where the current block is located is the current frame;
[0058] The importance measure value of the reference prediction value i is determined by any of the following methods:
[0059] Method 1: Calculated according to the frame display order of the current frame in the video bitstream and the frame display order of the reference frame i in the video bitstream;
[0060] Method 2: Calculated according to the frame display order of the current frame in the video bitstream, the frame display order of the reference frame i in the video bitstream, and the quality metric Q;
[0061] Method 3: Calculate the importance measure score of the reference frame i according to the calculation results of Method 1 and Method 2, sort the importance measure scores of the reference frames corresponding to the N reference prediction values in ascending order, and determine the index of the reference frame i in the sorting as the importance measure value of the reference prediction value i;
[0062] Method 4: Adjust the calculation results of Method 1, Method 2, or Method 3 based on the prediction mode of the reference prediction value i to obtain the importance measure value of the reference prediction value i;
[0063] Among them, the prediction mode of the reference prediction value i includes any one of the following: inter-frame prediction mode, intra-frame prediction mode.
[0064] In one embodiment, the number of weight groups included in the target weight list is greater than 1; the determining unit is specifically configured to:
[0065] Decode the index of the target weight group for weighted prediction from the video bitstream;
[0066] Select the target weight group from the target weight list according to the index of the target weight group;
[0067] Among them, the index of the target weight group uses a binary coding method of truncated unary code or an entropy coding method with multiple symbols.
[0068] In one embodiment, the processing unit is specifically configured to:
[0069] Perform weighted summation processing on N reference prediction values respectively using the weight values in the target weight group to obtain the prediction value of the current block; or,
[0070] Perform weighted processing on N reference prediction values respectively using the weight values in the target weight group in the form of integer calculation to obtain the prediction value of the current block.
[0071] On the one hand, an embodiment of the present application provides a video processing device, and the device includes:
[0072] A processing unit, configured to perform partitioning processing on the current frame in the video to obtain a current block, where N is an integer greater than 1;
[0073] The processing unit is further configured to perform composite prediction on the current block to obtain N reference prediction values of the current block;
[0074] A determining unit, configured to determine a target weight group for weighted prediction for the current block according to the importance of N reference prediction values, and the target weight group includes one or more weight values;
[0075] The processing unit is further configured to perform weighted prediction processing on N reference prediction values based on the target weight group to obtain the prediction value of the current block, and the prediction value of the current block is used to reconstruct the decoded image corresponding to the current block;
[0076] The processing unit is further configured to encode the video based on the prediction value of the current block to generate a video bitstream.
[0077] In one embodiment, the determining unit is specifically configured to:
[0078] Determine a target weight list from one or more weight lists according to the importance of N reference prediction values; each weight list contains one or more weight groups, each weight group contains one or more weight values, the number of weight values contained in each weight group is allowed to be the same or different, and the values of the weight values contained in each weight group are allowed to be the same or different;
[0079] Select a target weight group for weighted prediction from the target weight list.
[0080] In one embodiment, the bitrate consumed when performing weighted prediction processing on the N reference prediction values using the weight values in the target weight group is less than a preset bitrate threshold; or,
[0081] Perform weighted prediction processing on the N reference prediction values based on the weight values in the target weight group, such that the quality loss of the current block during the encoding process is less than a preset loss threshold; or,
[0082] The bitrate consumed when performing weighted prediction processing on the N reference prediction values using the weight values in the target weight group is less than a preset bitrate threshold, and perform weighted prediction processing on the N reference prediction values based on the weight values in the target weight group, such that the quality loss of the current block during the encoding process is less than a preset loss threshold.
[0083] In another embodiment, the target weight group is the weight group with the optimal encoding performance in the target weight list;
[0084] Among them, the optimal encoding performance includes: the bitrate consumed when performing weighted prediction processing on the N reference prediction values using the weight values in the target weight group is the minimum of the consumed bitrates corresponding to all weight groups in the target weight list; or, the quality loss of the current block corresponding to performing weighted prediction processing on the N reference prediction values based on the weight values in the target weight group is the minimum of the quality losses corresponding to all weight groups in the target weight list; or, the bitrate consumed when performing weighted prediction processing on the N reference prediction values using the weight values in the target weight group is the minimum of the consumed bitrates corresponding to all weight groups in the target weight list, and the quality loss of the current block corresponding to performing weighted prediction processing on the N reference prediction values based on the weight values in the target weight group is the minimum of the quality losses corresponding to all weight groups in the target weight list.
[0085] On the one hand, an embodiment of the present application provides a computer device, which includes:
[0086] A processor, adapted to execute a computer program;
[0087] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the above-described video processing method is implemented.
[0088] On the one hand, an embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and the computer program is loaded and executed by a processor to implement the above video processing method.
[0089] On the one hand, an embodiment of the present application provides a computer program product. The computer program product includes a computer program, and the computer program is stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the above video processing method.
[0090] In an embodiment of the present application, composite prediction is performed on a current block in a video bitstream to obtain N reference prediction values of the current block. The current block refers to an encoded block being decoded in the video bitstream. According to the importance of the N reference prediction values, a target weight group for weighted prediction is determined for the current block; weighted prediction processing is performed on the N reference prediction values based on the target weight group to obtain a prediction value of the current block, and the prediction value of the current block is used to reconstruct a decoded image corresponding to the current block. Composite prediction is used in the decoding process of the video, and the importance of the reference prediction values is fully considered in the composite prediction, so as to determine appropriate weights for weighted prediction for the current block based on the importance of each reference prediction value in the composite prediction, which can improve the prediction accuracy of the current block and thus improve the codec performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0091] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0092] Figure 1a is a basic workflow diagram of a video encoding provided by an exemplary embodiment of the present application;
[0093] Figure 1b is a schematic diagram of an inter-frame prediction provided by an exemplary embodiment of the present application;
[0094] Figure 2 is a schematic structural diagram of a video processing system provided by an exemplary embodiment of the present application;
[0095] Figure 3It is a schematic flowchart of a video processing method provided by an exemplary embodiment of the present application;
[0096] Figure 4 It is a schematic flowchart of a video processing method provided by another exemplary embodiment of the present application;
[0097] Figure 5 It is a schematic structural diagram of a video processing device provided by an exemplary embodiment of the present application;
[0098] Figure 6 It is a schematic structural diagram of a video processing device provided by another exemplary embodiment of the present application;
[0099] Figure 7 It is a schematic structural diagram of a computer device provided by an exemplary embodiment of the present application. Detailed implementation manners
[0100] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without making creative efforts shall fall within the protection scope of the present application.
[0101] Next, the technical terms involved in the present application will be introduced:
[0102] I. Video coding
[0103] A video can be composed of one or more video frames, and each video frame contains part of the video signal of the video. The acquisition methods of video signals can be divided into two types: captured by a camera or generated by a computer. Since the statistical characteristics corresponding to different acquisition methods are different, the video compression coding methods may also be different.
[0104] In mainstream video coding technologies, taking HEVC (High Efficiency Video Coding, the international video coding standard HEVC / H.265), VVC (versatile video coding, the international video coding standard VVC / H.266), AV1 (Alliance for Open Media Video 1, the first-generation video coding standard developed by the Alliance for Open Media), AV2 (Alliance for Open Media Video 2, the second-generation video coding standard developed by the Alliance for Open Media), and AVS3 (audio and video source coding standard) as examples, a hybrid coding framework is adopted. Under this hybrid coding framework, a series of operations and processes on the video are allowed as follows:
[0105] 1) Block partition structure: According to the size of the current frame (i.e., the video frame being encoded or decoded) input, the current frame can be divided into several non-overlapping processing units, and each processing unit will perform similar compression operations. This processing unit is called CTU (Coding Tree Unit) or LCU (Largest Coding Unit). The CTU can be further divided more finely to obtain one or more basic coding units, called CUs (Coding Units or Coding Blocks). Each CU is the most basic element in a coding process. The various encoding and decoding processing flows that may be adopted for each CU are described in the subsequent embodiments of this application.
[0106] 2) Prediction coding: It includes modes such as intra-frame prediction and inter-frame prediction. After the original video signal contained in the current CU (i.e., the CU being encoded or decoded in the current frame) in the current frame is predicted by the reconstructed video signal in the selected reference CU, a residual video signal is obtained. Here, the current CU can also be called the current block, the video frame where the current block is located is called the current frame, the reference CU used for predicting the current block can also be called the reference block of the current block, and the video frame where the reference block is located is called the reference frame. Among them, the encoding end needs to determine the most suitable prediction coding mode from among many possible prediction coding modes for the current CU and tell the decoding end. Among them, the prediction coding modes can include:
[0107] a. Intra(picture)Prediction: The reconstructed video signal used for prediction comes from the already encoded and reconstructed area within the same video frame, that is, the current block and the reference block are in the same video frame. Among them, the basic idea of intra-frame prediction is to use the correlation of adjacent pixels within the same video frame to remove spatial redundancy. In video coding, adjacent pixels refer to the reconstructed pixels of the already encoded CUs around the current CU within the same video frame.
[0108] b. Inter(picture)Prediction: The reconstructed video signal used for prediction comes from other video frames that have been encoded and are different from the current frame, that is, the current block and the reference block are in different video frames respectively.
[0109] 3) Transform & Quantization: The residual video signal is transformed through operations such as DFT (Discrete Fourier Transform) and DCT (Discrete Cosine Transform), and can be converted into the transform domain, which is called transform coefficients. The residual video signal in the transform domain is further subjected to a lossy quantization operation, losing certain information, making the quantized signal conducive to compressed representation.
[0110] In some video coding standards, there may be more than one transform method to choose from. Therefore, the encoder also needs to select one of the transforms for the current CU and inform the decoder. The fineness of quantization is usually determined by QP (Quantization Parameters). When the QP value is large, it means that a larger range of transform coefficients will be quantized to the same output, so usually a larger distortion and a lower bitrate will be brought; on the contrary, when the QP value is small, it means that a smaller range of transform coefficients will be quantized to the same output, so usually a smaller distortion will be brought, and at the same time, a higher bitrate will be corresponding.
[0111] 4) Entropy Coding or Statistical Coding: The quantized transform domain signal will be statistically compressed encoded according to the frequency of each value, and finally a binary (0 or 1) video bitstream will be output. At the same time, the encoding will generate other information, such as the selected predictive coding mode, motion vectors, etc. These other information also need to be entropy encoded to reduce the bitrate. Among them, statistical coding is a lossless coding method, which can effectively reduce the bitrate required to represent the same signal. Common statistical coding methods include Variable Length Coding (VLC) or Context Adaptive Binary Arithmetic Coding (CABAC).
[0112] 5) Loop Filtering: For the already encoded CU, after inverse quantization, inverse transformation, and prediction compensation operations (the reverse operations of the above 2) - 4)), the decoded image corresponding to this CU can be reconstructed. Compared with the original image, due to the influence of quantization, some information is different from the original image, resulting in distortion. Therefore, a filter can be used to filter the reconstructed decoded image to effectively reduce the degree of distortion caused by quantization. The filter can be, for example, deblocking, SAO (Sample Adaptive Offset), or ALF (Adaptive Loop Filter), etc. Since these reconstructed decoded images after filtering will be used as reference CUs for other CUs that need to be encoded subsequently and applied in the prediction process of other CUs, the above filtering operation is also called loop filtering, that is, the filtering operation within the encoding loop.
[0113] Based on the relevant descriptions of the above steps 1) - 5), an embodiment of the present application provides a basic flowchart of a video encoder. Please refer to Figure 1a , Figure 1a which is an exemplary basic flowchart of a video encoder. Among them, Figure 1a taking the current block as the kth CU in the current frame (current image) (such as Figure 1a the s shown k [x, y]) as an example for illustration, where k is a positive integer and k is less than or equal to the total number of CUs included in the current frame. s k [x, y] represents the pixel at coordinates [x, y] in the kth CU (abbreviated as pixel), x represents the abscissa of the pixel, and y represents the ordinate of the pixel; s k [x, y] can obtain a prediction signal after motion compensation or intra-frame prediction and other processes. Subtracting the prediction signal from the original signal s k [x, y] gives the residual video signal u k [x, y]; then the residual video signal u k [x, y] is subjected to transformation and quantization processing.
[0114] Among them, the data output by the quantization processing has two different destinations:
[0115] A: The data output by the quantization processing can be sent to the entropy encoder for entropy encoding to obtain the encoded bitstream (i.e., the video bitstream), and this bitstream is output to a buffer for storage and waiting to be transmitted.
[0116] B. The data output by the quantization process can be dequantized and inversely transformed to obtain the residual video signal u′ after inverse transformation. k Then, the inverse transformed residual video signal u′ k [x, y] and prediction signal Add together to get a new prediction signal And the new prediction signal The new prediction signal can then be sent to the current image buffer for storage. After intra-frame prediction processing, we get The new prediction signal After loop filtering, the reconstructed signal s′ can be obtained k [x, y], and the reconstructed signal s′ k [x, y] is sent to the decoded image buffer for storage to generate the reconstructed video. Reconstructed signal s′ k [x, y] is obtained through motion compensation prediction processing in Can represent the reference block, m x and m y Represent the horizontal and vertical components of the motion vector of the reference block, respectively.
[0117] In some embodiments, mainstream video coding standards such as HEVC, VVC, AVS3, AV1, and AV2 all adopt a block-based hybrid coding framework, which can divide the original video data into a series of coding blocks, and combine video coding methods such as prediction, transformation, and entropy coding to achieve video data compression. Among them, motion compensation is a commonly used prediction method for video coding. Motion compensation is based on the redundant characteristics of video content in the time domain or spatial domain, and derives the prediction value of the current block from the encoded reference block. This type of prediction method includes: inter-frame prediction, intra-frame block copy prediction, intra-frame string copy prediction, etc. In specific coding implementations, these prediction methods may be used alone or in combination. For coding blocks using these prediction methods, it is usually necessary to explicitly or implicitly encode one or more two-dimensional displacement vectors in the video bitstream. The displacement vector is used to indicate the displacement of the current block (or the same-position block of the current block) relative to one or more reference blocks.
[0118] It should be noted that, under different prediction modes and different implementations, the displacement vectors may have different names. In this article, they are uniformly described as follows: 1) The displacement vector in inter-frame prediction is called the motion vector (Motion Vector, abbreviated as MV); 2) The displacement vector in intra-block copy prediction is called the block vector (Block Vector, abbreviated as BV); 3) The displacement vector in intra-string copy prediction is called the string vector (String Vector, abbreviated as SV). Hereinafter, taking inter-frame prediction as an example, the technologies related to inter-frame prediction will be introduced.
[0119] Inter-frame prediction: Inter-frame prediction utilizes the correlation in the video time domain and uses the pixels of adjacent encoded images to predict the pixels of the current image, so as to effectively remove the redundancy in the video time domain and can effectively save the bits of the encoded residual data. As Figure 1b shown, Figure 1b This is a schematic diagram of inter-frame prediction provided by an embodiment of the present application. Among them, in Figure 1b , P is the current frame, Pr is the reference frame, B is the current block, and Br is the reference block of B. The coordinate positions of B' and B in the image are the same (that is, B' is the co-located block of B), the coordinates of Br are (x r , y r ), and the coordinates of B' are (x, y). The displacement between the current block and its reference block is called the motion vector (MV), that is: MV = (x r - x, y r - y).
[0120] Among them, considering that adjacent blocks in the time domain or spatial domain have strong correlations, the MV prediction technology can be used to further reduce the bits required for encoding the MV. In H.265 / HEVC, inter-frame prediction includes two MV prediction technologies: Merge (merge) and AMVP (Advanced Motion Vector Prediction). The Merge mode will establish an MV candidate list for the current PU (Prediction Unit), and there are 5 candidate MVs (and their corresponding reference images) in it. Traverse these 5 candidate MVs and select the one with the minimum rate-distortion cost as the optimal MV. If the codec establishes the candidate list in the same way, the encoder only needs to transmit the index of the optimal MV in the candidate list. In AV1 and AV2, a technology called Dynamic motion vector prediction (DMVP) is used to predict the MV.
[0121] To achieve better prediction performance, multi-reference frames are allowed to be used for inter-frame prediction in current mainstream video coding standards. In the AV1 standard and the upcoming AOM next-generation standard AV2, there is a prediction mode called compound prediction. This compound prediction mode allows the current block to perform inter-frame prediction using two reference frames, and the predicted inter-frame values are weighted and combined to derive the predicted value of the current block, or the predicted inter-frame value derived from one reference frame and the predicted intra-frame value derived from the current frame are weighted and combined to derive the predicted value of the current block. Here, the current block refers to the coding block being encoded (or decoded). In the embodiments of this application, the predicted inter-frame value and the predicted intra-frame value are both referred to as reference predicted values hereinafter. The following is the formula for deriving the predicted value of the current block during the compound prediction process:
[0122] P(x, y) = (w(x, y) · P 0 (x, y) + (1 - w(x, y)) · P 1 (x, y)) / 2
[0123] Wherein, P(x, y) is the predicted value of the current block; P 0 (x, y) and P 1 (x, y) are the two reference predicted values corresponding to the current block (x, y), and w(x, y) is the weight applied to the first reference predicted value P 0 (x, y).
[0124] Optionally, in video coding, to reduce the complexity of weighted prediction, integer calculations are usually used instead of floating-point calculations. The following is the formula for deriving the predicted value of the current block using integer calculation methods:
[0125] P(x, y) = (w(x, y) · P 0 (x, y) + (64 - w(x, y)) · P 1 (x, y) + 32) >> 6
[0126] Wherein, the weight value w(x, y) and the reference predicted values P 0 (x, y), P 1 (x, y) are all integers, and the right shift operation is used to replace division. ">> 6" means shifting 6 bits to the right. By shifting 6 bits to the right, it can be divided by 64, and 32 is the bias added for rounding.
[0127] According to the current video coding standard, compound prediction adopts a special weighting mode, that is, P0 and P1 have equal weight values, and the weights corresponding to the reference predicted values at different positions are all set to fixed values. The specific formula is as follows:
[0128] P(x, y) = (32 × P0 (x, y) + 32 × P 1 ((x, y) + 32) >> 6
[0129] II. Video Decoding
[0130] On the decoding side, for each CU, after obtaining the video bitstream, on the one hand, first perform entropy decoding on the video bitstream to obtain information on various prediction coding modes and quantized transform coefficients, and then perform inverse quantization and inverse transformation on each transform coefficient to obtain a residual video signal. On the other hand, according to the known information on the prediction coding mode, the prediction signal corresponding to this CU (subsequently referred to as the prediction value) can be obtained, and the residual video signal and the prediction signal are added together to obtain a reconstructed video signal, which can be used to reconstruct the decoded image corresponding to this CU. Finally, this reconstructed video signal needs to undergo loop filtering operations to generate the final output signal.
[0131] Based on the above related descriptions, an embodiment of the present application provides a video processing solution, which can be applied to video encoders or video compression products using composite prediction (or weighted prediction based on multiple reference frames). The general principle of this video processing solution is as follows:
[0132] At the encoding end: Composite prediction can be used for the CUs included in the video frame to obtain N reference prediction values of the CU, where N is an integer greater than 1; and appropriate weights are adaptively selected for the CU according to the importance of the N reference prediction values for weighted prediction to obtain the prediction value of the CU. The so-called weighted prediction means performing weighted prediction processing on the N reference prediction values of the CU using the adaptively selected weights. Then, the video is encoded based on the prediction value of the CU to generate a video bitstream, and the video bitstream is sent to the decoding end.
[0133] At the decoding end: When decoding the CU in the video bitstream, it is possible to determine to perform composite prediction on the CU in the video bitstream according to the information on the prediction coding mode, and adaptively select appropriate weights for the CU according to the importance of the N reference prediction values for weighted prediction; then perform weighted prediction on the N reference prediction values based on the adaptively selected weights to obtain the prediction value of this CU, and use the prediction value of the CU to reconstruct the decoded image corresponding to this CU.
[0134] As described above, the current video coding standard stipulates that equal weight values are used in composite prediction to perform weighted prediction on the reference prediction values derived from different reference blocks. However, in practical applications, the reference prediction values derived from different reference blocks may have unequal importance. Using equal weight values cannot reflect the importance difference of the prediction reference values. In this case, using the existing standard will affect the prediction accuracy. The embodiments of the present application improve the current video coding standard, fully consider the importance of the reference prediction values derived from different reference blocks in composite prediction, and allow to adaptively select appropriate weights for weighted prediction of the CU based on the importance of each reference prediction value in composite prediction, expanding the weighted prediction method in the video coding standard, which can improve the prediction accuracy of the CU and thus improve the coding performance.
[0135] Next, the video processing system provided by the embodiments of the present application will be described in detail. Please refer to Figure 2 , Figure 2 which is a schematic structural diagram of a video processing system provided by an embodiment of the present application. The video processing system 20 may include an encoding device 201 and a decoding device 202. The encoding device 201 is located at the encoding end, and the decoding device is located at the decoding end. The encoding device 201 may be a terminal or a server, and the decoding device 202 may be a terminal or a server. A communication connection may be established between the encoding device 201 and the decoding device 202. Among them, the terminal may be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, a smart TV, etc., but is not limited thereto. The server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0136] (1) For the encoding device 201:
[0137] The encoding device 201 can obtain the video to be encoded, and the video can be obtained by shooting with a camera device or generated by a computer. The camera device may be a hardware component set in the encoding device 201. For example, the camera device may be an ordinary camera, a stereo camera, and a light field camera, etc. set in the terminal. The camera device may also refer to a hardware device connected to the encoding device 201, such as a camera connected to the server.
[0138] Among them, a video includes one or more video frames. The encoding device 201 can divide each video frame into one or more CUs and encode each CU. When encoding any CU, composite prediction can be performed on the CU being encoded (subsequently referred to as the current block) to obtain N reference prediction values of the current block, and factors such as the bitrate consumed during weighted prediction processing and the quality loss of the current block during the encoding process are comprehensively considered to determine the importance of each reference prediction value. Then, a suitable target weight group is adaptively selected for the current block based on the importance of each reference prediction value. The target weight group may include one or more weight values. Then, the weight values in the target weight group are used to perform weighted prediction processing on the N reference prediction values to obtain the prediction value of the current block. Among them, the prediction value of the current block can be understood as the prediction signal corresponding to the current block, and the prediction value of the current block can be used to reconstruct the decoded image corresponding to the current block.
[0139] Among them, the N reference prediction values of the current block are derived from N reference blocks of the current block; one reference prediction value corresponds to one reference block. The video frame where the reference block is located is the reference frame, and the video frame where the current block is located is the current frame. The positional relationship between the N reference blocks and the current block may include but is not limited to any of the following: ① The N reference blocks are respectively located in N reference frames, and the N reference frames and the current frame belong to different video frames in the video stream; ② The N reference blocks are located in the same reference frame, and the same reference frame and the current frame belong to different video frames in the video stream; ③ One or more of the N reference blocks are located in the current frame, and the remaining reference blocks among the N reference blocks are located in one or more reference frames, and the one or more reference frames and the current frame belong to different video frames in the video stream; ④ The N reference blocks and the current block are all located in the current frame. It can be seen that the prediction mode of the composite prediction in the embodiments of the present application includes both an inter-frame prediction mode, that is, allowing inter-frame prediction using at least two reference frames; and a combined prediction mode, that is, allowing inter-frame prediction using at least one reference frame, and at the same time allowing intra-frame prediction using the current frame; and also includes an intra-frame prediction mode, that is, allowing intra-frame prediction using the current frame.
[0140] Corresponding to different prediction modes of the composite prediction, the derivation methods of the N reference prediction values of the current block may include any of the following: ① The N reference prediction values of the current block are derived after performing inter-frame prediction using the N reference blocks of the current block respectively. At this time, the N reference prediction values can all be referred to as inter-frame prediction values. ② Among the N reference prediction values of the current block, at least one reference prediction value is derived after performing inter-frame prediction using at least one of the N reference blocks of the current block. This part of the reference prediction values can be referred to as inter-frame prediction values; the remaining reference prediction values are derived after performing intra-frame prediction using the remaining reference blocks among the N reference blocks. This part of the reference prediction values can be referred to as intra-frame prediction values.
[0141] Then, the encoding device 201 performs operations such as transform coding, quantization, and entropy coding on the video based on the predicted values of the CUs included in the video frames, obtains a video bitstream, and sends the video bitstream to the decoding device 202 so that the decoding device 202 can perform decoding processing on the video bitstream.
[0142] (2) For the decoding device 202:
[0143] After receiving the video bitstream sent by the encoding device 201, the decoding device 202 can perform decoding processing on the video bitstream and reconstruct the video corresponding to the video bitstream. Specifically, on the one hand, the decoding device 202 can perform entropy decoding on the video bitstream to obtain the prediction modes of each CU in the video bitstream and the quantized transform coefficients, and perform composite prediction on the current block (i.e., the CU being decoded) according to the prediction mode of the current block to obtain N reference prediction values of the current block; and determine whether the current block allows the use of adaptive weighted prediction.
[0144] If it is determined that the current block allows the use of adaptive weighted prediction, the target weight list can be determined from one or more weight lists according to the importance of the N reference prediction values, and the target weight group for weighted prediction can be determined for the current block from the target weight list, and the target weight group contains one or more weight values. Then, the weight values in the target weight group are directly used to perform weighted prediction processing on the N reference prediction values to obtain the predicted value of the current block. If it is determined that the current block does not allow the use of adaptive weighted prediction, the weighted prediction processing can be performed on the N reference prediction values according to the existing video coding standard. For example, equal weight values are used to perform weighted prediction processing on each reference prediction value to obtain the predicted value of the current block.
[0145] On the other hand, the decoding device 202 performs inverse quantization and inverse transformation on the quantized transform coefficients to obtain the residual signal value of the current block, and superimposes the predicted value and the residual signal value of the current block to obtain the reconstructed value of the current block, and then reconstructs the decoded image corresponding to the current block according to the reconstructed value. Among them, the decoded decoded image can be used as a reference image for decoding other CUs and can also be used to reconstruct the video.
[0146] In the embodiments of the present application, composite prediction is used in both the video encoding and decoding processes, and the importance of the reference prediction values derived from different reference blocks is fully considered in the composite prediction. It is allowed to adaptively select appropriate weights for weighted prediction for the CU based on the importance of each reference prediction value in the composite prediction, which expands the weighted prediction method in the video coding standard, can improve the prediction accuracy of the CU, and thus improve the encoding and decoding performance.
[0147] Next, the video processing method provided by the embodiments of the present application will be elaborated. Please refer to Figure 3 , Figure 3Schematic flowchart of a video processing method provided by an embodiment of this application. This video processing method can be executed by a decoding device in the above video processing system. The video processing method described in this embodiment may include the following steps S301 - S303:
[0148] S301. Perform composite prediction on a current block in a video bitstream to obtain N reference prediction values for the current block. The current block refers to an encoded block being decoded in the video bitstream, and N is an integer greater than 1.
[0149] The video bitstream contains one or more video frames, and each video frame may contain one or more encoded blocks. The above N reference prediction values can be derived from N reference blocks, and one reference prediction value corresponds to one reference block. In the embodiment of this application, the video frame where the reference block is located may be a reference frame, and the video frame where the current block is located is the current frame. Among them, the positional relationship between the N reference blocks and the current block includes any of the following: ① The N reference blocks are respectively located in N reference frames, and the N reference frames and the current frame may belong to different video frames in the video bitstream. For example, N = 2, one of the two reference blocks is located in reference frame one, and the other reference block is located in reference frame two. Reference frame one, reference frame two, and the current frame belong to different video frames in the video bitstream. ② The N reference blocks are located in the same reference frame, and the same reference frame and the current frame may belong to different video frames in the video bitstream. For example, N = 2, both reference blocks are located in reference frame one, and reference frame one and the current frame belong to different video frames in the video bitstream. ③ One or more of the N reference blocks are located in the current frame, and the remaining reference blocks among the N reference blocks are located in one or more reference frames, and the one or more reference frames and the current frame belong to different video frames in the video bitstream. For example, N = 4, the four reference blocks are respectively: reference block 1, reference block 2, reference block 3, reference block 4; reference block 1 is located in the current frame, and the remaining reference blocks 2, 3, and 4 are all located in reference frame one, and this reference frame one and the current frame belong to different video frames in the video bitstream; another example is that reference block 1 and reference block 2 are located in the current frame, the remaining reference block 3 may be located in reference frame one, and reference block 4 is located in a reference frame two, and reference frame one, reference frame two, and the current frame belong to different video frames in the video bitstream. ④ The N reference blocks and the current block are all located in the current frame. For example, N = 2, both reference blocks are located in the current frame.
[0150] According to the positional relationship between the N reference blocks and the current block shown in the above ① - ④, it can be known that the prediction mode of the composite prediction in the embodiment of this application includes both an inter - frame prediction mode, that is, allowing inter - frame prediction using at least two reference frames; and may also include a combined prediction mode, that is, allowing inter - frame prediction using at least one reference frame, and at the same time allowing intra - frame prediction using the current frame; and also includes an intra - frame prediction mode, that is, allowing intra - frame prediction using the current frame.
[0151] Corresponding to different prediction modes of the composite prediction, the derivation method of the N reference prediction values of the current block may include any of the following: ① The N reference prediction values of the current block are obtained by performing inter-frame prediction on the N reference blocks of the current block respectively. In this case, these N reference prediction values can all be referred to as inter-frame prediction values. ② Among the N reference prediction values of the current block, at least one reference prediction value is obtained by performing inter-frame prediction on at least one reference block among the N reference blocks of the current block. This part of the reference prediction values can be referred to as inter-frame prediction values; the remaining reference prediction values are obtained by performing intra-frame prediction on the remaining reference blocks among the N reference blocks. This part of the reference prediction values can be referred to as intra-frame prediction values.
[0152] In one embodiment, before executing step S302, it can be first determined whether the current block meets the conditions for adaptive weighted prediction. If the current block meets the conditions for adaptive weighted prediction, then step S302 is executed; by determining whether the current block meets the conditions for adaptive weighted prediction, it is possible to adaptively select weights for weighted prediction, further improving the prediction accuracy and thus the coding performance.
[0153] Among them, the condition that the current block meets the adaptive weighted prediction includes at least one of the following:
[0154] a) The sequence header of the frame sequence to which the current block belongs includes a first indication field, and the first indication field indicates that all coding blocks in the frame sequence are allowed to use adaptive weighted prediction; where the frame sequence refers to a sequence composed of multiple video frames in sequence. It should be understood that when the sequence header of the frame sequence includes the first indication field, the first indication field is used to indicate that all coding blocks included in the entire frame sequence are allowed to use adaptive weighted prediction.
[0155] As an implementation manner, the first indication field can be represented as seq_acp_flag. This first indication field can indicate whether the coding blocks in the frame sequence are allowed to use adaptive weighted prediction according to its value. If the above first indication field is a first preset value (such as 1), it indicates that all coding blocks in the frame sequence are allowed to use adaptive weighted prediction, and it can be determined that the current block meets the conditions for adaptive weighted prediction; if the above first indication field is a second preset value (such as 0), it indicates that all coding blocks in the frame sequence are not allowed to use adaptive weighted prediction, and it can be determined that the current block does not meet the conditions for adaptive weighted prediction.
[0156] b) The start header of the current slice to which the current block belongs contains a second indication field, and the second indication field indicates that the coded blocks in the current slice are allowed to use adaptive weighted prediction. Here, a video frame can be divided into multiple slices, and each slice contains one or more coded blocks; the current slice refers to the slice to which the current block belongs, that is, the slice being decoded. It should be understood that when the start header of the current slice contains the second indication field, the second indication field can be used to indicate that all the coded blocks included in the current slice are allowed to use adaptive weighted prediction.
[0157] As an implementation manner, the second indication field can be represented as slice_acp_flag. The second indication field can indicate whether the coded blocks in the current slice are allowed to use adaptive weighted prediction according to its value. If the second indication field is a first preset value (such as 1), it indicates that all the coded blocks in the current slice are allowed to use adaptive weighted prediction, and thus it can be determined that the current block meets the adaptive weighted prediction condition; if the second indication field is a second preset value (such as 0), it indicates that all the coded blocks in the current slice are not allowed to use adaptive weighted prediction, and thus it can be determined that the current block does not meet the adaptive weighted prediction condition.
[0158] c) The frame header of the current frame in which the current block is located contains a third indication field, and the third indication field indicates that the coded blocks in the current frame are allowed to use adaptive weighted prediction. It should be understood that when the frame header of the current frame contains the third indication field, the third indication field can be used to indicate that all the coded blocks included in the current frame are allowed to use adaptive weighted prediction.
[0159] As an implementation manner, the third indication field can be represented as pic_acp_flag. The third indication field indicates whether the current frame is allowed to use adaptive weighted prediction according to its value. If the third indication field is a first preset value (such as 1), it indicates that the coded blocks in the current frame are allowed to use adaptive weighted prediction, and thus it can be determined that the current block meets the adaptive weighted prediction condition; if the third indication field is a second preset value (such as 0), it can indicate that the coded blocks in the current frame are not allowed to use adaptive weighted prediction, and thus it can be determined that the current block does not meet the adaptive weighted prediction condition.
[0160] d) During the composite prediction process, the current block uses at least two reference frames for inter-frame prediction.
[0161] e) During the composite prediction process, the current block uses at least one reference frame for inter-frame prediction and uses the current frame for intra-frame prediction. For example, if the current block uses reference frame one for inter-frame prediction and uses the current frame for intra-frame prediction, then it can be determined that the current block meets the conditions for adaptive weighted prediction.
[0162] f) The motion type of the current block is the specified motion type. For example, if the motion type of the current block is simple_translation (simple balance), then it is determined that the current block meets the conditions for adaptive weighted prediction.
[0163] g) The current block uses a preset motion vector prediction mode. For example, if the preset motion vector prediction mode used by the current block is NEAR_NEARMV, then it can be determined that the current block meets the conditions for adaptive weighted prediction.
[0164] In the AV1 and AV2 standards, a technique called Dynamic motion vector prediction is used to predict the MV. The MV can be predicted by spatially adjacent blocks in the current frame or temporally adjacent blocks in the reference frame. For single-reference frame inter prediction, each reference frame has a separate predicted MV list; for composite frame inter prediction, the predicted mv lists corresponding to different reference frames form a predicted MV group list, and multiple predicted MV modes are allowed, such as NEAR_NEARMV, NEAR_NEWMV, NEW_NEAR_MV, NEW_NEWMV, GLOBAL_GLOBALMV, JOINT_NEWMV, etc. Among them:
[0165] NEAR_NEARMV: It means that the MVs corresponding to the two reference frames are the MVs in the predicted MV group.
[0166] NEAR_NEWMV: The first MV is the first predicted mv in the predicted MV group, and the second MV is derived from the MVD decoded from the video bitstream and the second predicted MV in the predicted MV group. Here, MVD (Motion Vector Difference) refers to the difference between the current MV and the predicted MV (Motion Vector Prediction, MVP).
[0167] NEW_NEARMV: The first MV is derived from the MVD decoded from the video bitstream and the first predicted MV in the predicted MV group, and the second MV is the second predicted MV in the predicted MV group.
[0168] NEW_NEWMV: The first MV is derived from the MVD1 decoded from the video bitstream and the first predicted MV in the predicted MV group, and the second MV is derived from the MVD2 decoded from the video bitstream and the second predicted MV in the predicted MV group.
[0169] GLOBAL_GLOBALMV: The MV is derived according to the per-frame global motion information.
[0170] NEW_NEWMV: Similar to NEW_NEWMV, but only one MVD is included in the video bitstream, and the other MVD is derived based on the information between reference frames.
[0171] In the embodiments of the present application, the preset motion vector prediction mode can be one or more of the above: NEAR_NEARMV, NEAR_NEWMV, NEW_NEARMV, NEW_NEWMV, GLOBAL_GLOBALMV, JOINT_MV. It should be noted that the preset motion vector prediction mode is not limited to the above MV prediction modes. For other standards, such as H.265 and H.266, the motion vector prediction mode can also be determined by combining merge and AMVP.
[0172] h) The current block uses a preset interpolation filter. The video encoding process usually includes many encoding tools, including various types of interpolation filters, such as: linear interpolation filter, cic (integrated-comb filter) interpolation filter, and so on. The preset interpolation filter in the embodiments of the present application can be any one of these various types of interpolation filters. For example, the preset interpolation filter can be a linear interpolation filter. When the current block uses a linear interpolation filter, it can be determined that the current block meets the conditions for adaptive weighted prediction.
[0173] i) The current block does not use a specific encoding tool. For example, if the current block does not use an optical flow-based motion vector optimization method, then it can be determined that the current block meets the conditions for adaptive weighted prediction. Among them, video coding standards such as AV2 and H.266 allow the use of an optical flow-based motion vector optimization method, which is a method for refining motion vectors based on the derivation of the optical flow equation.
[0174] j) If the reference frames used by the current block in the composite prediction process meet specific conditions, then it can be determined that the current block meets the conditions for adaptive weighted prediction. Among them, the specific conditions include one or more of the following (that is, the specific conditions can include one or more of ① and ②): ① The azimuth relationship between the reference frames used in the composite prediction and the current frame in the video bitstream meets a preset relationship; among them, the azimuth relationship meeting the preset relationship includes any one of the following: all the reference frames used are before the current frame; all the reference frames used are after the current frame; some of the reference frames used are before the current frame, and the remaining reference frames are after the current frame. The video bitstream contains multiple video frames, and each video frame corresponds to a frame display time in the video. This azimuth relationship can actually be understood as the order of frame display times. For example, all the reference frames used being before the current frame can be understood as: the frame display time of the reference frame is earlier than the frame display time of the current frame; all the reference frames used being after the current frame can be understood as: the frame display time of the reference frame is later than the frame display time of the current frame.
[0175] ② The absolute value of the importance difference between the reference prediction values corresponding to the reference frames used in the composite prediction is greater than or equal to a preset threshold. At this time, it can be determined that the current block meets the conditions for adaptive weighted prediction; among them, the preset threshold can be set according to requirements. As an implementation, the importance of the reference prediction values corresponding to the reference frames can be measured by importance metric values. At this time, the importance difference between the reference prediction values corresponding to the reference frames used in the composite prediction can be determined according to the importance metric values of the reference prediction values corresponding to the used reference frames. For example, N = 2. In the composite prediction, the two reference prediction values are: the reference prediction value 1 corresponding to the first reference frame and the reference prediction value 2 corresponding to the second reference frame; let the importance metric value of the reference prediction value 1 be D0, and the importance metric value of the reference prediction value 2 be D1. Then, the importance difference between the reference prediction values corresponding to the two reference frames is: the difference D0 - D1 between the importance metric value D0 of the reference prediction value 1 and the importance metric value D1 of the reference prediction value 2, that is, the absolute value of the importance difference between the two reference prediction values is: ΔD = abs(D0 - D1), and abs() represents taking the absolute value.
[0176] It should be understood that the above conditions for the current block to meet the adaptive weighted prediction can be used alone or in combination. For example, the conditions for the current block to meet the adaptive weighted prediction may include: the sequence header of the frame sequence to which the current block belongs contains a first indication field, and the first indication field indicates that the coded blocks in the frame sequence are allowed to use adaptive weighted prediction, and the motion type of the current block is a specified motion type. Another example is that the frame header of the current frame where the current block is located contains a third indication field, and the third indication field indicates that the coded blocks in the current frame are allowed to use adaptive weighted prediction and the current block uses a preset motion vector prediction mode; and so on; the present application does not make any limitation on this.
[0177] S302. Determine a target weight group for weighted prediction for the current block according to the importance of the N reference prediction values.
[0178] The importance of the reference prediction values can be comprehensively determined based on factors such as the bitrate consumed during weighted prediction processing and the quality loss of the current block during the encoding process. For example, when a certain reference prediction value is used for weighted prediction of the current block and the bitrate consumption increases significantly, it indicates that the help of this reference prediction value in reducing bitrate consumption is not obvious enough, and its importance is relatively low. Another example: when a certain reference prediction value is used for weighted prediction of the current block and the quality loss increases significantly, it indicates that the help of this reference prediction value in reducing quality loss is relatively small, and its importance is relatively low. The target weight group contains one or more weight values, and these weight values will act on each reference prediction value during the weighted prediction process. If the importance of a certain reference prediction value is relatively low, then the weight value corresponding to this reference prediction value in the target weight group is relatively small, while if the importance of a certain reference prediction value is relatively high, then the weight value corresponding to this reference prediction value in the target weight group is relatively large. That is to say, the target weight group is selected based on considering the influence of each reference prediction value on factors such as bitrate consumption and quality loss, so that the comprehensive cost of the weighted prediction process (i.e., the cost of bitrate consumption, the cost of quality loss, the cost of bitrate consumption and quality loss) is relatively small and the encoding and decoding performance is relatively good.
[0179] In one embodiment, the video bitstream contains one or more weight lists, each weight list contains one or more weight groups, and one weight group contains one or more weight values; the number of weight values contained in each weight group is allowed to be the same or different, the values of the weight values contained in each weight group are allowed to be the same or different, and the order of the weight values contained in each weight group is allowed to be the same or different. The following is an example of a weight list:
[0180] i. The weight list 1 is expressed as: {2, 4, 6, 8, 10, 12, 14} / 16. This indicates that the weight list 1 contains 7 weight groups, which are: weight group 1: {2} / 16; weight group 2: {4} / 16; weight group 3: {6} / 16, and so on, weight group 7: {14} / 16.
[0181] ii. The weight list 2 is expressed as: {14, 8, 4, 12, 2} / 16. This indicates that the weight list 2 contains 5 weight groups, which are: weight group 1: {14} / 16; weight group 2: {8} / 16; and so on, weight group 5: {2} / 16.
[0182] iii. The weight list 3 is expressed as: {4, {8, 8}, 12} / 16. This indicates that the weight list 3 contains 3 weight groups, which are: weight group 1: {4} / 16; weight group 2: {8, 8} / 16; weight group 3: {12} / 16.
[0183] As can be seen from the above examples: (1) The number of weight groups included in each weight list is allowed to be different. For example, the number of weight groups in weight list 1 is 7, the number of weight groups in weight list 2 is 5, and the number of weight groups in weight list 1 is 3. (2) The number of weight values included in each weight group is allowed to be the same. For example, the number of weight values included in each weight group in weight list 1 is 1. (3) The data of the weight values included in each weight group is also allowed to be different. For example, weight group 1 in weight list 3 contains 1 weight value, while weight group 2 in weight list 3 contains 2 weight values. (4) The weight values in each weight group are allowed to be the same. For example, weight group 1 in weight list 1 contains the weight value 2 / 16, and weight group 1 in weight list 2 also contains the weight value 2 / 16. (5) The weight values in each weight group are also allowed to be different. For example, weight group 1 in weight list 1 contains the weight value 2 / 16, but weight group 2 in weight list 1 contains the weight value 4 / 16. It can be understood that the sum of the weight values provided by a weight group should be equal to 1. In the above examples, although each weight group only contains one weight value, each weight value is less than 1. Therefore, each weight group implicitly contains another weight value. That is, in specific applications, each weight group actually provides two weight values, and the sum of these two weight values is 1. For example, weight group 1 in weight list 1 only contains the weight value 2 / 16, but in specific applications, this weight group 1 actually provides two weight values, namely the weight value 2 / 16 and the weight value 14 / 16. Another example: weight group 3 in weight list 3 only contains the weight value 12 / 16, but in specific applications, this weight group 3 actually provides two weight values, namely the weight value 12 / 16 and the weight value 4 / 16. It can be seen that in the embodiments of the present application, when the sum of the weight values included in the weight groups in the weight list is less than 1, the implicitly included weight value of the weight group can be obtained by calculation.
[0184] The following is another example of a weight list:
[0185] i. Weight list 4 contains 4 weight groups, namely weight group 1: {2,14} / 16; weight group 2: {4,12} / 16; weight group 3: {6,10} / 16; weight group 4: {8,8}} / 16.
[0186] ii. Weight list 5 contains 2 weight groups, namely weight group 1: {4,12} / 16; weight group 2: {10,6} / 16.
[0187] As can be seen from the above examples: (1) The sum of all the weight values included in the weight groups in the weight list is equal to 1. (2) The order of the weight values included in each weight group is allowed to be different. For example, weight group 3 in weight list 4 and weight group 2 in weight list 5 contain the same weight values, but the order of the weight values is different.
[0188] In one embodiment, step S302 may include steps s31 - s32:
[0189] s31. Determine a target weight list from one or more weight lists according to the importance of N reference prediction values.
[0190] Among them, the specific implementation manner of determining the target weight list from one or more weight lists according to the importance of N reference prediction values may include the following methods:
[0191] Method 1: When the number of weight lists in the video bitstream is only one, this only weight list in the video bitstream can be directly determined as the target weight list.
[0192] Method 2: When there are multiple weight lists in the video bitstream, an importance metric value of the reference prediction values can be introduced, and the target weight list can be determined from the multiple weight lists according to the importance metric values of the N reference prediction values.
[0193] (1) Determine the target weight list according to the absolute value of the difference between the importance metric values of the N reference prediction values.
[0194] When the number of weight lists in the video bitstream is M + 1, where M is a positive integer greater than or equal to 1, the M + 1 groups of weight lists can be respectively denoted as {w_list1, w_list2... w_listM + 1}. One weight list corresponds to one threshold interval, that is, the number of threshold intervals is also M + 1. As an implementation manner, M + 1 threshold intervals can be directly set according to requirements. As another implementation manner, M thresholds can be obtained, and M + 1 threshold intervals can be divided according to the M thresholds. For example, M thresholds are obtained, which are T1, T2,..., TM respectively; then they are divided into M + 1 threshold intervals, which are [0, T1], (T1, T2], (T2, T3],..., (TM, +∞) respectively; among them, [0, T1] can correspond to w_list1, (T1, T2] can correspond to w_list2, and so on. T is an integer greater than or equal to 0.
[0195] The decoding device can obtain the importance metric values of the N reference prediction values, calculate the importance differences between the N reference prediction values, and the importance difference between any two reference prediction values is measured by the difference between the importance metric values of any two reference prediction values. Then, the decoding device can determine the threshold interval where the absolute value of the importance difference between the N reference prediction values is located; and determine the weight list corresponding to the threshold interval where the absolute value of the importance difference between the N reference prediction values is located as the target weight list.
[0196] Among them, let the importance metric values of any two reference prediction values be D0 and D1 respectively, and the two references
[0197] If the importance difference between the reference prediction values is represented as D0 - D1, then the absolute value of the importance difference between any two reference prediction values is represented as ΔD = abs(D0 - D1). It should be noted that the importance metric value is an indicator used to measure the degree of importance, but there can be multiple cases for the measurement benchmark. For example, the measurement benchmark can be that the larger the importance metric value, the higher the degree of importance; another example is that the measurement benchmark can also be that the smaller the importance metric value, the higher the degree of importance. For this measurement benchmark, the present application does not make any limitations.
[0198] In one implementation, if N = 2, that is, only two reference prediction values are included, and the importance metric values of these two reference prediction values are D0 and D1 respectively, then the weight list corresponding to the threshold interval where ΔD = abs(D0 - D1) is located is directly determined as the target weight list. For example, when the number of weight lists in the video bitstream is 2, these two weight lists are respectively denoted as {w_list1, w_list2}, where w_list1 is {8, 12, 14} / 16 and w_list2 is {12, 8, 4} / 16. Suppose the threshold interval corresponding to w_list1 is [0, 1], and the threshold interval corresponding to w_list2 is (1, +∞); then, when ΔD = abs(D0 - D1) is less than or equal to 1, the threshold interval where it is located is [0, 1], and w_list1 is determined as the target weight list; otherwise, when ΔD = abs(D0 - D1) is greater than 1, the threshold interval where it is located is (1, +∞), and w_list2 is determined as the target weight list.
[0199] In another implementation, if N > 2, the absolute values of the importance differences between any two reference prediction values can be calculated respectively, and the weight lists corresponding to the threshold intervals where each absolute value is located are found respectively, and the weight list with the largest corresponding quantity is determined as the target weight list. For example, when N = 3, the importance metric values of the three reference prediction values are D0, D1, and D2 respectively. Then the absolute values of the importance differences between any two reference prediction values can be calculated respectively, that is, calculate ΔD = abs(D0 - D1), ΔD' = abs(D1 - D2), and ΔD″ = abs(D0 - D1) respectively, and then judge the weight lists corresponding to the threshold intervals where ΔD, ΔD', and ΔD″ are located respectively. If there are two or more of them corresponding to the same weight list, then the same weight list is determined as the target weight list.
[0200]
[0201] In another embodiment, if N>2, the absolute value of the importance difference between any two reference prediction values can be calculated respectively, and then the maximum value among the absolute values can be found, and the weight list corresponding to the threshold interval where the maximum value is located is determined as the target weight list. For example, in the above example where N = 3, after calculating ΔD, ΔD', and ΔD″, find the maximum value among ΔD, ΔD', and ΔD″. Suppose the maximum value is ΔD', then the weight list corresponding to the threshold interval where ΔD' is located is determined as the target weight list.
[0202] In another embodiment, if N>2, the absolute value of the importance difference between any two reference prediction values can be calculated respectively, and then the minimum value among the absolute values can be found, and the weight list corresponding to the threshold interval where the minimum value is located is determined as the target weight list. For example, in the above example where N = 3, after calculating ΔD, ΔD', and ΔD″, find the minimum value among ΔD, ΔD', and ΔD″. Suppose the minimum value is ΔD, then the weight list corresponding to the threshold interval where ΔD is located is determined as the target weight list.
[0203] In another embodiment, if N>2, the absolute value of the importance difference between any two reference prediction values can be calculated respectively, and then the average value of the absolute values can be calculated, and the weight list corresponding to the threshold interval where the average value is located is determined as the target weight list. For example, in the above example where N = 3, after calculating ΔD, ΔD', and ΔD″, then calculate the average value among ΔD, ΔD', and ΔD″=(ΔD + ΔD'+ ΔD″) / 3, and the weight list corresponding to the threshold interval where the average value is located is determined as the target weight list.
[0204] It should be noted that the embodiments of the present application can also use other numerical characteristics of the absolute values of the importance differences between N reference prediction values, such as the maximum value, minimum value, average value, etc. after squaring each absolute value; to determine the target weight list, and the present application does not limit this.
[0205] (2) Determine the target weight list by comparing the magnitudes of the importance measurement values of the reference prediction values.
[0206] The N reference prediction values of the current block may include a first reference prediction value and a second reference prediction value, and the video bitstream may include a first weight list and a second weight list. The decoding device can compare the magnitude of the importance measurement value of the first reference prediction value with the magnitude of the importance measurement value of the second reference prediction value; if it is determined that the importance measurement value of the first reference prediction value is greater than the importance measurement value of the second reference prediction value, then the first weight list is determined as the target weight list; if it is determined that the importance measurement value of the first reference prediction value is less than or equal to the importance measurement value of the second reference prediction value, then the second weight list is determined as the target weight list.
[0207] For example, the importance measurement value of the above first reference prediction value is D0, the importance measurement value of the second reference prediction value is D1, the above first weight list is w_list1, and the second weight list is w_list2; compare the magnitudes of D0 and D1; if D0 > D1, the first weight list w_list1 can be determined as the target weight list; if D0 ≤ D1, the first weight list w_list2 can be determined as the target weight list.
[0208] Optionally, the weight values in the first weight list are opposite to those in the second weight list, that is, w_list2[x] = 1 - w_list1[x], where x represents the weight value in the weight list. For example, if w_list1 = {0.2, 0.4}, then w_list2[x] = 1 - w_list1[x], that is, w_list2[x] = {0.8, 0.6}. Additionally, the weight values in the first weight list and the second weight list can also be set separately.
[0209] (3) Use the mathematical sign function and the importance measurement value of the reference prediction value to determine the target weight list.
[0210] Among the N reference prediction values of the current block, there are a first reference prediction value and a second reference prediction value; the video bitstream includes a first weight list, a second weight list, and a third weight list. The decoding device can call the mathematical sign function to process the difference between the importance measurement value of the first reference prediction value and the importance measurement value of the second reference prediction value to obtain a sign value. If the above sign value is a first preset value (such as -1), the first weight list is determined as the target weight list; if the above sign value is a second preset value (such as 0), the second weight list is determined as the target weight list; if the above sign value is a third preset value (such as 1), the third weight list is determined as the target weight list. Among them, the above first weight list, second weight list, and third weight list are different weight lists respectively; or, two of the first weight list, second weight list, and third weight list are allowed to be the same weight list.
[0211] Among them, the importance measurement value of the above first reference prediction value is D0, the importance measurement value of the second reference prediction value is D1, and the data sign function is called to process the difference between D0 and D1 to obtain a sign value, that is, sign value = sign(D0 - D1), where sign() represents the data sign function.
[0212] (4) The above methods (1), (2), and (3) can be used alone or in combination to determine the target weight list. As an implementation method, there are M + 1 weight lists. For any two reference prediction values, first, method (1) can be used to determine the candidate weight list according to the threshold interval where the absolute value of the importance difference between the two reference prediction values is located. Then, method (2) can be used to compare the importance measurement values of the first reference prediction value and the second reference prediction value. If it is determined that the importance measurement value of the first reference prediction value is greater than that of the second reference prediction value, the candidate weight list is directly determined as the target weight list; if it is determined that the importance measurement value of the first reference prediction value is less than or equal to that of the second reference prediction value, the weight list corresponding to the weight value opposite to the weight value in the candidate weight list is determined as the target weight list.
[0213] For example, the importance measurement value of the first reference prediction value is D0, and the importance measurement threshold of the second reference prediction value is D1; method (1) can be used to determine the candidate weight list w_list1. If D0 > D1, the weight list w_list1 is determined as the target weight list; if D0 ≤ D1, the weight list w_list2 is determined as the target weight list; where the weight values in w_list2 are opposite to those in w_list1, that is, w_list2[x] = 1 - w_list1[x].
[0214] As another implementation method, assume that there are 3*(M + 1) weight lists in the video bitstream, that is, one threshold interval can correspond to 3 weight lists; for any two reference prediction values, first, method (1) can be used to determine the threshold interval where the absolute value of the importance difference between the first reference prediction value and the second reference prediction value is located, and the 3 weight lists corresponding to this threshold interval are all determined as candidate weight lists, that is, the candidate weight list can include the first weight list, the second weight list, and the third weight list. Then, method (2) can be used to call the mathematical symbol function to process the difference between the importance measurement values of the first reference prediction value and the second reference prediction value to obtain the symbol value. If the symbol value is the first preset value, the first weight list is determined as the target weight list; if the symbol value is the second preset value, the second weight list is determined as the target weight list; if the symbol value is the third preset value, the third weight list is determined as the target weight list.
[0215] For example, three weight lists {w_list1, w_list2, w_list3} are determined as candidate weight lists by using method (1), and then a mathematical symbol function is called to process the difference between the importance metric value of the first reference prediction value and the importance metric value of the second reference prediction value to obtain a symbol value. When the symbol value is -1, w_list1 is determined as the target weight list; when the symbol value is 0, w_list2 is determined as the target weight list; when the symbol value is 1, w_list3 is determined as the target weight list.
[0216] Any one of the N reference prediction values is denoted as reference prediction value i, which is derived from reference block i, and the video frame where reference block i is located is reference frame i; i is an integer and i is less than or equal to N; the video frame where the current block is located is the current frame; the importance metric value of reference prediction value i can be determined by any of the following methods:
[0217] Method 1: It is calculated according to the frame display order (Picture Order Count, POC) of the current frame in the video bitstream and the frame display order of reference frame i in the video bitstream. Specifically, the difference between the frame display order of the current frame in the video bitstream and the frame display order of reference frame i in the video bitstream can be calculated, and the absolute value of this difference is used as the importance metric value of reference prediction value i. For example, let the frame display serial number of the current frame in the video bitstream be denoted as cur_poc, and the frame display order of reference frame i in the video bitstream be ref_poc; the importance metric value D of reference prediction value i = abs(cur_poc - ref_poc), where abs() represents taking the absolute value.
[0218] Method 2: It is calculated according to the frame display order of the current frame in the video bitstream, the frame display order of reference frame i in the video bitstream, and the quality metric Q, where the quality metric Q can be determined according to various situations, and this application does not limit this. Specifically, the quality metric Q of reference frame i can be derived from the quantization information of the current block. For example: the quality metric Q can be set to the base_qindex (basic quantization index) of reference frame i, and the base_qindex of any reference frame can be different or the same. In another implementation, the quality metric Q of reference frame i can also be derived from other coding information. For example: the quality metric Q of reference frame i can be derived based on the coding information difference between the encoded CUs in reference frame i and the encoded CUs in the current frame.
[0219] As an implementation, the decoding device can calculate the difference between the frame display order of the current frame in the video bitstream and the frame display order of reference frame i in the video bitstream; then, using an objective function, the importance metric value of reference prediction value i is determined according to the above difference, the quality metric Q, and the importance metric value list.
[0220] Among them, the objective function is: D = f(cur_poc - ref_poc) + Q, where D represents the importance metric value of the reference prediction value i, f(x) is an increasing function, x = cur_poc - ref_poc, cur_poc represents the frame display sequence number of the current frame in the video bitstream, and ref_poc represents the frame display order of the reference frame i in the video bitstream. The above list of importance metric values contains the correspondence between f(x) and the reference importance metric value, as shown in Table 1:
[0221] Table 1
[0222] x 0 1 2 3 4 5 6 7 8 9 f(x) 0 64 96 112 120 124 126 127 128 129
[0223] It should be noted that the function expression of f(x) above is only an example, and the embodiments of the present application allow the function expression of f(x) to change, that is, the embodiments of the present application do not limit the specific expression form of f(x).
[0224] In one embodiment, the importance metric value of the above reference frame i can be calculated according to the orientation relationship between the reference frame i and the current frame and the quality metric Q. As a way of implementation, a correspondence between the reference orientation relationship and the reference importance metric value can be established. For example, if the reference orientation relationship is that the reference frame is before the current frame, the reference importance metric value can correspond to a first value; if the reference orientation relationship is that the reference frame is after the current frame, the reference importance metric value can correspond to a second value. Then, the decoding device can calculate the importance metric value of the reference prediction value i according to the reference importance metric value corresponding to the reference frame i and the quality metric Q.
[0225] Method 3: Calculate the importance metric score of the reference frame i according to the calculation results of Method 1 and Method 2, and sort the importance metric scores of the reference frames corresponding to the N reference prediction values in ascending order, and determine the index of the reference frame i in the sorting as the importance metric value of the reference prediction value i.
[0226] As a way of implementation, the first importance metric value of the reference prediction value i can be calculated through the above Method 1, and the second importance metric value of the reference prediction value i can be calculated through the above Method 2. Then, calculate the importance metric score (such as score) of the reference frame i according to the first importance metric value and the second importance metric value. Then, sort the importance metric scores of the reference frames corresponding to the N reference prediction values in ascending order, and determine the index of the reference frame i in the sorting as the importance metric value of the reference prediction value i.
[0227] For example, the importance metric score of reference frame one corresponding to reference prediction value 1 is 20, the importance metric score of reference frame two corresponding to reference prediction value 2 is 30, and the importance metric score of reference frame three corresponding to reference prediction value 3 is 40; then, sort the importance metric scores of the reference frames corresponding to the 3 reference prediction values in ascending order, and the sorting result is: reference frame one, reference frame two, reference frame three; among them, the index of reference frame one in the sorting is 1, so the importance metric value of reference prediction value 1 is 1; the index of reference frame two in the sorting is 2, so the importance metric value of reference prediction value 2 is 2; the index of reference frame three in the sorting is 3, so the importance metric value of reference prediction value 3 is 3.
[0228] Among them, there are the following ways to calculate the importance metric score of reference frame i according to the first importance metric value and the second importance metric value: ① Perform weighted summation on the first importance metric value and the second importance metric value to obtain the importance metric value of reference prediction value i. ② Perform averaging on the first importance metric value and the second importance metric value to obtain the importance metric value of reference prediction value i.
[0229] Method 4: In order to obtain a more accurate importance metric value, the calculation result of Method 1, Method 2 or Method 3 can be adjusted based on the prediction mode of reference prediction value i to obtain the importance metric value of reference prediction value i; among them, the prediction mode of reference prediction value i includes any one of the following: inter-frame prediction mode, intra-frame prediction mode. As an implementation method, the calculation result of Method 1, Method 2 or Method 3 can be adjusted according to the adjustment function to obtain the importance metric value of reference prediction value i. The adjustment function can be, for example, D’ = g(D) = a*D + b; where D’ represents the importance metric value of reference prediction value i; D is the calculation result of the above Method 1, Method 2 or Method 3, and a and b can be adaptively set according to the prediction mode. For example, the following are 3 specific examples:
[0230] a) If the prediction mode of reference prediction value i is inter-frame prediction, then a = 1 and b = 0;
[0231] b) If the prediction mode of reference prediction value i is intra-frame prediction, then a = 2 and b = 0;
[0232] c) If the prediction mode of reference prediction value i is intra-frame prediction, then a = 0 and b = 160.
[0233] It should be understood that in the actual process, any one of the above Methods 1-4 can be used according to the requirements to determine the importance metric value of the reference prediction value, and the present application does not make any limitations in this regard.
[0234] s32. Select the target weight group for weighted prediction from the target weight list.
[0235] Among them, the decoding device selects a target weight group from the target weight list, which can be divided into the following two cases:
[0236] (1) When the number of weight groups included in the target weight list is equal to 1, there is no need to decode the index of the target weight group from the video bitstream at this time, and the weight group in the target weight list is directly used as the target weight group for weighted prediction.
[0237] (2) When the number of weight groups included in the target weight list is greater than 1, that is, there are multiple weight groups in the target weight list. For example, the target weight list is expressed as {{2,14},{4,12},{6,10},{8,8}} / 16, and there are 4 weight groups in this target weight list. At this time, the index of the target weight group for weighted prediction of the current block can be indicated in the video bitstream. The index of the target weight group uses a coding method of binarization encoding with truncated unary code or an entropy coding method with multiple symbols. Among them, the truncated unary code is in the case where the maximum value Max of the syntax element to be encoded is known. Assume the symbol to be encoded is x: If 0 < x < Max, x is binarized in the way of unary code; if x = Max, the binary string of x is all composed of 1s and the length is Max. Then, the decoding device needs to decode the index of the target weight group for weighted prediction from the video bitstream, and select the target weight group from the target weight list according to the index of the target weight group. Among them, the index of the target weight group can indicate the position in the target weight list; for example, the above-mentioned target weight list contains 4 weight groups; the decoding device decodes the index of the target weight group from the video bitstream as 2, then determines the position in the target weight list according to the index of the target weight group, that is, it is in the second position in the target weight list (i.e., weight group 2), then the determined target weight group is {4,12} / 16.
[0238] S303. Perform weighted prediction processing on N reference prediction values based on the target weight group to obtain a prediction value of the current block, and the prediction value of the current block is used to reconstruct the decoded image corresponding to the current block.
[0239] The number of weight values actually provided by the target weight set should correspond to the number of reference prediction values. For example, if the number of reference prediction values is N, correspondingly, the number of weight values actually provided by the target weight set is also N. For example, N = 3, and the three reference prediction values are reference prediction value 1, reference prediction value 2, and reference prediction value 3 respectively. The target weight set may include three weight values, namely weight value 1, weight value 2, and weight value 3; among them, reference prediction value 1 corresponds to weight value 1, reference prediction value 2 corresponds to weight value 2, and reference prediction value 3 corresponds to weight value 3. In the above example, the target weight set may also only include two weight values, namely weight value 1 and weight value 2, and the sum of weight value 1 and weight value 2 is less than 1. Then, the implicit weight value 3 in the target weight set can be obtained by calculation: weight value 3 = 1 - weight value 1 - weight value 2.
[0240] In one embodiment, there are the following two ways (1) and (2) to perform weighted prediction processing on N reference prediction values based on the target weights to obtain the prediction value of the current block:
[0241] (1) The decoding device can perform weighted summation processing on the N reference prediction values respectively using the weight values in the target weight set to obtain the prediction value of the current block. Among them, the prediction value P(x, y) of the current block can be:
[0242] P(x, y) = (w1·P 0 (x, y) + w2·P 1 (x, y) + … + wn·P N-1 (x, y)) / N
[0243] Among them, P(x, y) is the prediction value of the current block, and P 0 (x, y), P 1 (x, y) …… P N-1 (x, y) respectively represent N reference prediction values; w1 represents the weight value corresponding to the first reference prediction value corresponding to the current block (x, y), w2 represents the weight value corresponding to the second reference prediction value corresponding to the current block (x, y), and so on, wn represents the weight value corresponding to the Nth reference prediction value corresponding to the current block (x, y).
[0244] In one embodiment, when N = 2, that is, the current block uses two reference prediction values for weighted prediction. At this time, the number of weight values included in the target weight set is one, that is, it includes an implicit weight value. The prediction value P(x, y) of the current block can be:
[0245] P(x, y) = (w(x, y)·P 0 (x, y) + (1 - w(x, y))·P 1 (x, y)) / 2.
[0246] Among them, P(x, y) is the predicted value of the current block, P 0 (x, y) and P 1 (x, y) are two reference predicted values corresponding to the current block (x, y), and w(x, y) is the weight value in the target weight group applied to the first reference predicted value P 0 (x, y). 1 - w(x, y) is the implicit weight value and is applied to the second reference predicted value P 1 (x, y).
[0247] (2) In video coding, considering the complexity of predicted value calculation, a right shift operation can be used to replace division to achieve integer calculation, and the weight values in the target weight group are respectively used to weight N reference predicted values in the form of integer calculation to obtain the predicted value of the current block. The complexity of predicted value calculation can be reduced to a certain extent through the form of integer calculation.
[0248] For example, taking the number of reference predicted values as 2 (i.e., N = 2) and the number of weight values in the target weight group as 1 as an example, the decoding device can respectively weight the 2 reference predicted values in the form of integer calculation using the weight values in the target weight group to obtain the predicted value of the current block. At this time, the predicted value P(x, y) of the current block can be:
[0249] P(x, y) = (w(x, y) × P 0 (x, y) + (16 - w(x, y)) × P 1 (x, y) + 8) >> 4
[0250] Among them, ">> 4" means shifting right by 4 bits, that is, shifting right by 4 bits can make the data in the weighting process divisible by 16, and 8 is the bias added for rounding. The above P 0 (x, y), P 1 (x, y), w(x, y), and P(x, y) are all integers; P(x, y) is the predicted value of the current block, P 0 (x, y) and P 1 (x, y) are two reference predicted values corresponding to the current block (x, y), and w(x, y) is the weight value (i.e., the weight value in the target weight group) applied to the first predicted value P 0 (x, y).
[0251] Also for example, ">> 6" means shifting right by 6 bits, that is, shifting right by 6 bits can make the data in the weighting process divisible by 64, and 32 is the bias added for rounding. At this time, the predicted value P(x, y) of the current block can be:
[0252] P(x, y) = (w · P 0 (x, y) + (64 - w) · P1 ((x, y) + 32) >> 6
[0253] In one embodiment, after obtaining the predicted value of the current block, the residual video signal of the current block and the predicted value may be subjected to superposition processing to obtain the reconstructed value of the current block, and the decoded image corresponding to the current block is reconstructed based on the reconstructed value of the current block. On the one hand, the decoded image corresponding to the current block can be used as a reference image for weighted prediction of other coding blocks. On the other hand, the decoded image corresponding to the current block can also be used to reconstruct the current frame of the current block. Finally, the video can be reconstructed based on multiple reconstructed video frames.
[0254] In the embodiment of the present application, composite prediction is performed on the current block in the video bitstream to obtain N reference predicted values of the current block. The current block refers to the coding block being decoded in the video bitstream, and N is an integer greater than 1. According to the importance of the N reference predicted values, a target weight group for weighted prediction is determined for the current block. The N reference predicted values are subjected to weighted prediction processing based on the target weight group to obtain the predicted value of the current block. The predicted value of the current block is used to reconstruct the decoded image corresponding to the current block. Composite prediction is used in the decoding process of the video, and the importance of the reference predicted values is fully considered in the composite prediction, so as to adaptively select appropriate weights for weighted prediction of the current block based on the importance of each reference predicted value in the composite prediction, which can improve the prediction accuracy of the current block and thus improve the coding and decoding performance.
[0255] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of another video processing method provided by the embodiment of the present application. This video processing method can be executed by an encoding device in a video processing system. The video processing method described in this embodiment may include the following steps S401 - S405:
[0256] S401. Perform partitioning processing on the current frame in the video to obtain the current block. The current block refers to the coding block being encoded in the video. Among them, the video may include one or more video frames, and the current frame refers to the video frame being encoded. The encoding device can partition the current frame in the video to obtain one or more coding blocks, and the current block refers to any one of the one or more coding blocks in the current frame that is being encoded.
[0257] S402. Perform composite prediction on the current block to obtain N reference predicted values of the current block, where N is an integer greater than 1.
[0258] The above N reference prediction values can be derived from N reference blocks, with one reference prediction value corresponding to one reference block. In the embodiments of the present application, the video frame where the reference block is located can be a reference frame, and the video frame where the current block is located is the current frame. Among them, the positional relationship between the N reference blocks and the current block includes any one of the following: ① The N reference blocks are respectively located in N reference frames, and the N reference frames and the current frame can belong to different video frames in the video. ② The N reference blocks are located in the same reference frame, and the same reference frame and the current frame can belong to different video frames in the video. ③ One or more of the N reference blocks are located in the current frame, and the remaining reference blocks among the N reference blocks are located in one or more reference frames, and the one or more reference frames and the current frame belong to different video frames in the video. ④ The N reference blocks and the current block are both located in the current frame.
[0259] According to the positional relationship between the N reference blocks and the current block shown in the above ① - ④, it can be known that the prediction mode of the composite prediction in the embodiments of the present application includes both an inter-frame prediction mode, that is, allowing inter-frame prediction using at least two reference frames; and can also include a combined prediction mode, that is, allowing inter-frame prediction using at least one reference frame and allowing intra-frame prediction using the current frame at the same time; and also includes an intra-frame prediction mode, that is, allowing intra-frame prediction using the current frame.
[0260] Corresponding to different modes of the composite prediction, obtaining the N reference prediction values of the current block through composite prediction of the current block can include any one of the following: ① During the process of performing composite prediction on the current block, the current block uses the N reference blocks for inter-frame prediction to obtain the N reference prediction values of the current block. At this time, the N reference prediction values of the current block are derived after performing inter-frame prediction using the N reference blocks of the current block respectively. ② During the process of performing composite prediction on the current block, the current block uses at least one of the N reference prediction values for inter-frame prediction and uses the remaining reference blocks among the N reference blocks for intra-frame prediction, thereby obtaining the N reference prediction values. At this time, some of the N reference prediction values are derived after performing inter-frame prediction using at least one of the N reference blocks of the current block; the remaining reference prediction values are derived after performing intra-frame prediction using the remaining reference blocks among the N reference blocks.
[0261] S403. Determine a target weight group for weighted prediction for the current block according to the importance of the N reference prediction values.
[0262] The importance of the reference prediction values can be comprehensively determined based on factors such as the bitrate consumed during weighted prediction processing and the quality loss of the current block during the encoding process. For example, when the current block uses a certain reference prediction value for weighted prediction and the bitrate consumption increases significantly, it indicates that the help of this reference prediction value in reducing bitrate consumption is not obvious enough, and its importance is relatively low. Another example is that when the current block uses a certain reference prediction value for weighted prediction and the quality loss increases significantly, it means that the help of this reference prediction value in reducing quality loss is relatively small, and its importance is relatively low. The target weight group contains one or more weight values, and these weight values will act on each reference prediction value during the weighted prediction process. If the importance of a certain reference prediction value is relatively low, then the weight value corresponding to this reference prediction value in the target weight group is relatively small, while if the importance of a certain reference prediction value is relatively high, then the weight value corresponding to this reference prediction value in the target weight group is relatively large. That is to say, the target weight group is selected based on considering the influence of each reference prediction value on factors such as bitrate consumption and quality loss, so that the comprehensive cost of the weighted prediction process (i.e., the cost of bitrate consumption, the cost of quality loss, the cost of bitrate consumption and quality loss) is relatively small and the encoding and decoding performance is relatively good. In one implementation, when using the weight values in the target weight group to perform weighted prediction processing on N reference prediction values, the consumed bitrate is less than a preset bitrate threshold; or, based on the weight values in the target weight group, performing weighted prediction processing on N reference prediction values, so that the quality loss of the current block during the encoding process is less than a preset loss threshold; or, when using the weight values in the target weight group to perform weighted prediction processing on N reference prediction values, the consumed bitrate is less than a preset bitrate threshold, and based on the weight values in the target weight group, performing weighted prediction processing on N reference prediction values, so that the quality loss of the current block during the encoding process is less than a preset loss threshold.
[0263] In one embodiment, there are one or more weight lists during the encoding process, which can be understood as: one or more weight lists can all be used during the encoding process. Each weight list contains one or more weight groups, each weight group contains one or more weight values, the number of weight values contained in each weight group is allowed to be the same or different, and the values of the weight values contained in each weight group are allowed to be the same or different. At this time, according to the importance of N reference prediction values, determining the target weight group for weighted prediction of the current block may include steps S41 - S42:
[0264] S41. Determine the target weight list from one or more weight lists according to the importance of N reference prediction values.
[0265] Among them, the specific implementation methods for determining the target weight list from one or more weight lists according to the importance of N reference prediction values may include the following several methods:
[0266] Method 1: When there is a weight list during the encoding process, the weight list can be directly used as the target weight list.
[0267] Method 2: When there are multiple weight lists during the encoding process, an importance metric value of the reference prediction value can be introduced, and the target weight list can be determined from the multiple weight lists according to the importance metric values of the N reference prediction values.
[0268] (1) Determine the target weight list according to the absolute value of the difference between the importance metric values of the N reference prediction values.
[0269] When the number of weight lists existing during the encoding process is M + 1, where M is a positive integer greater than or equal to 1, one weight list corresponds to one threshold interval, that is, the number of threshold intervals is also M + 1. Then, the encoding device can obtain the importance metric values of the N reference prediction values, and calculate the importance difference between the N reference prediction values. The importance difference between any two reference prediction values is measured by the difference between the importance metric values of any two reference prediction values; specifically, the difference between the importance metrics of any two reference prediction values can be calculated, and the difference between the importance metric values of any two reference prediction values is used as the importance difference between any two reference prediction values; then, determine the threshold interval where the absolute value of the importance difference between the N reference prediction values is located; and determine the weight list corresponding to the threshold interval where the absolute value of the importance difference between the N reference prediction values is located as the target weight list.
[0270] (2) Determine the target weight list by comparing the magnitudes of the importance metric values of the reference prediction values.
[0271] The N reference prediction values of the current block may include a first reference prediction value and a second reference prediction value, and the weight lists existing during the encoding process include a first weight list and a second weight list. The encoding device can compare the magnitude of the importance metric value of the first reference prediction value with the magnitude of the importance metric value of the second reference prediction value; if it is determined that the importance metric value of the first reference prediction value is greater than the importance metric value of the second reference prediction value, the first weight list is determined as the target weight list; if it is determined that the importance metric value of the first reference prediction value is less than or equal to the importance metric value of the second reference prediction value, the second weight list is determined as the target weight list.
[0272] Optionally, the weight values in the first weight list are opposite to the weight values in the second weight list.
[0273] (3) Determine the target weight list by using the mathematical sign function and the importance metric value of the reference prediction value.
[0274] Among the N reference prediction values of the current block, there are a first reference prediction value and a second reference prediction value; among the weight lists existing during the encoding process, there are a first weight list, a second weight list, and a third weight list. The encoding device can call a mathematical symbol function to process the difference between the importance metric value of the first reference prediction value and the importance metric value of the second reference prediction value to obtain a symbol value. If the above symbol value is a first preset value (such as -1), the first weight list is determined as the target weight list; if the above symbol value is a second preset value (such as 0), the second weight list is determined as the target weight list; if the above symbol value is a third preset value (such as 1), the third weight list is determined as the target weight list.
[0275] (4) The above methods (1), (2), and (3) can be used alone, or methods (1), (2), and (3) can be combined to determine the target weight list. As an implementation method, there are M + 1 weight lists. Using method (1), the weight list corresponding to the threshold interval where the absolute value of the importance difference is located can be determined as the candidate weight list. Then, among the N reference prediction values, there are a first reference prediction value and a second prediction value. Using method (2), compare the magnitudes of the importance metric value of the first reference prediction value and the importance metric value of the second reference prediction value. If it is determined that the importance metric value of the first reference prediction value is greater than the importance metric value of the second reference prediction value, the candidate weight list is directly determined as the target weight list; if it is determined that the importance metric value of the first reference prediction value is less than or equal to the importance metric value of the second reference prediction value, the weight list corresponding to the weight value opposite to the weight value in the candidate weight list is determined as the target weight list.
[0276] Among them, any one of the N reference prediction values is denoted as reference prediction value i. The reference prediction value i is derived from reference block i, and the video frame where reference block i is located is reference frame i; i is an integer and i is less than or equal to N; the video frame where the current block is located is the current frame; the importance metric value of the reference prediction value i can be determined by any of the following methods:
[0277] Method 1: It is calculated according to the frame display order (POC) of the current frame in the video and the frame display order of reference frame i in the video. Specifically, the difference between the frame display order of the current frame in the video and the frame display order of reference frame i in the video can be calculated, and the absolute value of this difference is used as the importance metric value of the reference prediction value i. For example, let the frame display serial number of the current frame in the video be denoted as cur_poc, and the frame display order of reference frame i in the video be ref_poc; the importance metric value D of the reference prediction value i = abs(cur_poc - ref_poc), where abs() represents taking the absolute value.
[0278] Method 2: It is calculated based on the frame display order of the current frame in the video, the frame display order of reference frame i in the video, and the quality metric Q. Among them, the quality metric Q can be determined according to various situations, and this application does not limit this. Specifically, the quality metric Q of reference frame i can be derived from the quantization information of the current block. For example, the quality metric Q can be set to the base_qindex (basic quantization index) of reference frame i. The base_qindex of any reference frame can be different or the same. In another implementation, the quality metric Q of reference frame i can also be derived from other coding information. For example, the quality metric Q of reference frame i can be derived based on the coding information difference between the coded CUs in reference frame i and the coded CUs in the current frame.
[0279] As an implementation, the encoding device can calculate the difference between the frame display order of the current frame in the video and the frame display order of reference frame i in the video; then, using the objective function, determine the importance metric value of reference prediction value i based on the above difference, quality metric Q, and the list of importance metric values.
[0280] In one embodiment, the importance metric value of the above reference frame i can be calculated based on the orientation relationship between reference frame i and the current frame and the quality metric Q. As an implementation, a correspondence relationship can be established between the reference orientation relationship and the reference importance metric value. For example, if the reference orientation relationship is that the reference frame is before the current frame, the reference importance metric value can correspond to the first value; if the reference orientation relationship is that the reference frame is after the current frame, the reference importance metric value can correspond to the second value. Then, the encoding device can calculate the importance metric value of reference prediction value i based on the reference importance metric value corresponding to reference frame i and the quality metric Q.
[0281] Method 3: Calculate the importance metric score of reference frame i based on the calculation results of Method 1 and Method 2, and sort the importance metric scores of the reference frames corresponding to the N reference prediction values in ascending order, and determine the index of reference frame i in the sorting as the importance metric value of reference prediction value i.
[0282] As an implementation, the first importance metric value of reference prediction value i can be calculated through the above Method 1, and the second importance metric value of reference prediction value i can be calculated through the above Method 2. Then, calculate the importance metric score (such as score) of reference frame i based on the first importance metric value and the second importance metric value. Then, sort the importance metric scores of the reference frames corresponding to the N reference prediction values in ascending order, and determine the index of reference frame i in the sorting as the importance metric value of reference prediction value i.
[0283] Among them, there are the following ways to calculate the importance metric score of reference frame i based on the first importance metric value and the second importance metric value: ① Perform a weighted sum on the first importance metric value and the second importance metric value to obtain the importance metric value of reference prediction value i. ② Perform an averaging process on the first importance metric value and the second importance metric value to obtain the importance metric value of reference prediction value i.
[0284] Method 4: To obtain a relatively accurate importance metric value, the calculation result of Method 1, Method 2, or Method 3 can be adjusted based on the prediction mode of reference prediction value i to obtain the importance metric value of reference prediction value i; among them, the prediction mode of reference prediction value i includes any one of the following: inter-frame prediction mode, intra-frame prediction mode.
[0285] It should be understood that in the actual process, any one of the above Methods 1 - 4 can be used to determine the importance metric value of the reference prediction value according to requirements, and this application does not make any limitations in this regard.
[0286] S42. Select the target weight group for weighted prediction from the target weight list.
[0287] Among them, the target weight list contains one or more weight groups.
[0288] (1) When the number of weight groups contained in the target weight list is equal to 1, directly use the weight group in the target weight list as the target weight group for weighted prediction.
[0289] (2) The target weight list contains multiple weight groups, and select the target weight group for weighted prediction from the target weight list.
[0290] When performing weighted prediction processing, there may be cases of bitrate consumption, and there may also be cases of quality loss in the encoding process of encoding blocks. Therefore, the encoding device can first obtain the encoding performance when using each weight group in the target weight list to perform weighted prediction processing on N reference prediction values, and then determine the weight group with the optimal encoding performance in the target weight list as the target weight group. As an implementation method, the bitrate consumed when performing weighted prediction processing on N reference prediction values using the weight values in the target weight group is the minimum of the bitrate consumptions corresponding to all weight groups in the target weight list; or, the quality loss of the current block during the encoding process corresponding to performing weighted prediction processing on N reference prediction values based on the weight values in the target weight group is the minimum of the quality losses corresponding to all weight groups in the target weight list; or, the bitrate consumed when performing weighted prediction processing on N reference prediction values using the weight values in the target weight group is the minimum of the bitrate consumptions corresponding to all weight groups in the target weight list, and the quality loss of the current block during the encoding process corresponding to performing weighted prediction processing on N reference prediction values based on the weight values in the target weight group is the minimum of the quality losses corresponding to all weight groups in the target weight list.
[0291] It should be understood that when selecting the target weight group for weighted prediction on the encoding side, it is necessary to continuously try each weight group from the target weight list to obtain the target weight group for weighted prediction. On the decoding side, there is no need for continuous trial. Instead, the index of the target weight group to be used will be indicated in the video bitstream. The decoding side only needs to decode the index of the target weight group from the video bitstream and find the target weight group from the target weight list according to the index of the target weight group.
[0292] S404. Perform weighted prediction processing on N reference prediction values based on the target weight group to obtain the prediction value of the current block.
[0293] Among them, the specific implementation method of step S404 can refer to the specific implementation method of step S303 above and will not be elaborated here.
[0294] S405. Encode the video based on the prediction value of the current block to generate a video bitstream. The encoding device can encode the video based on the prediction value of the current block and the index of the target weight group to generate a video bitstream. Among them, encoding the video based on the prediction value of the current block and the index of the target weight group to generate a video bitstream can refer to the corresponding encoding description above and will not be elaborated here.
[0295] In the embodiments of the present application, the current frame in the video is partitioned to obtain a current block, where the current block refers to the coding block being encoded in the video; the current block is subjected to composite prediction to obtain N reference prediction values of the current block; according to the importance of the N reference prediction values, a target weight group for weighted prediction is determined for the current block; the N reference prediction values are subjected to weighted prediction processing based on the target weight group to obtain the prediction value of the current block; and the video is encoded based on the prediction value of the current block to generate a video bitstream. By using composite prediction during the encoding process of the video and the importance of the reference prediction values derived in the composite prediction, it is possible to adaptively select appropriate weights for the current block based on the importance of each reference prediction value in the composite prediction for weighted prediction, which can improve the prediction accuracy of the current block and thus improve the codec performance.
[0296] Please refer to Figure 5 , Figure 5 FIG. is a schematic structural diagram of a video processing device provided by an embodiment of the present application. The video processing device can be disposed in the computer device provided by the embodiment of the present application, and the computer device can be the decoding device mentioned in the above method embodiment. Figure 5 The video processing device shown can be a computer program (including program code) running in a computer device, and the video processing device can be used to execute Figure 3 part or all of the steps in the method embodiment shown. Please refer to Figure 5 , and the video processing device can include the following units:
[0297] A processing unit 501, configured to perform composite prediction on a current block in a video bitstream to obtain N reference prediction values of the current block, where N is an integer greater than 1;
[0298] A determination unit 502, configured to determine a target weight group for weighted prediction for the current block according to the importance of the N reference prediction values, where the target weight group includes one or more weight values;
[0299] The processing unit 501 is further configured to perform weighted prediction processing on the N reference prediction values based on the target weight group to obtain the prediction value of the current block, and the prediction value of the current block is used to reconstruct the decoded image corresponding to the current block.
[0300] In one embodiment, the N reference prediction values are derived from N reference blocks; one reference prediction value corresponds to one reference block; the video frame where the reference block is located is a reference frame, and the video frame where the current block is located is the current frame;
[0301] The positional relationship between the N reference blocks and the current block includes any one of the following:
[0302] The N reference blocks are respectively located in N reference frames, and the N reference frames and the current frame belong to different video frames in the video bitstream;
[0303] N reference blocks are located in the same reference frame, and the same reference frame and the current frame belong to different video frames in the video bitstream;
[0304] One or more of the N reference blocks are located in the current frame, the remaining reference blocks among the N reference blocks are located in one or more reference frames, and the one or more reference frames and the current frame belong to different video frames in the video bitstream;
[0305] The N reference blocks and the current block are both located in the current frame.
[0306] In one embodiment, the processing unit 501 is further configured to:
[0307] Determine whether the current block meets the conditions for adaptive weighted prediction;
[0308] If the current block meets the conditions for adaptive weighted prediction, determine a target weight group for weighted prediction for the current block according to the importance of the N reference prediction values.
[0309] In one embodiment, the condition that the current block meets the adaptive weighted prediction includes at least one of the following:
[0310] The sequence header of the frame sequence to which the current block belongs includes a first indication field, and the first indication field indicates that the coding blocks in the frame sequence are allowed to use adaptive weighted prediction;
[0311] The header of the current slice to which the current block belongs includes a second indication field, and the second indication field indicates that the coding blocks in the current slice are allowed to use adaptive weighted prediction;
[0312] The frame header of the current frame where the current block is located includes a third indication field, and the third indication field indicates that the coding blocks in the current frame are allowed to use adaptive weighted prediction;
[0313] During the composite prediction process, the current block uses at least two reference frames for inter-frame prediction;
[0314] During the composite prediction process, the current block uses at least one reference frame for inter-frame prediction and uses the current frame for intra-frame prediction;
[0315] The motion type of the current block is a specified motion type;
[0316] The current block uses a preset motion vector prediction mode;
[0317] The current block uses a preset interpolation filter;
[0318] The current block does not use a specific coding tool;
[0319] The reference frames used by the current block in the composite prediction process satisfy specific conditions, and the specific conditions include one or more of the following: the orientation relationship between the reference frames used and the current frame in the video bitstream satisfies a preset relationship; the absolute value of the importance difference between the reference prediction values corresponding to the reference frames used is greater than or equal to a preset threshold;
[0320] Among them, the orientation relationship satisfying the preset relationship includes any one of the following: all the reference frames used are before the current frame; all the reference frames used are after the current frame; some of the reference frames used are before the current frame, and the remaining reference frames are after the current frame.
[0321] In one embodiment, the video bitstream includes one or more weight lists, each weight list includes one or more weight groups, each weight group includes one or more weight values, the number of weight values in each weight group is allowed to be the same or different, and the values of the weight values included in each weight group are allowed to be the same or different;
[0322] The determining unit 502 may be specifically configured to:
[0323] Determine a target weight list from one or more weight lists according to the importance of N reference prediction values;
[0324] Select a target weight group for weighted prediction from the target weight list.
[0325] In one embodiment, the number of weight lists in the video bitstream is M + 1, one weight list corresponds to one threshold interval, and M is a positive integer greater than or equal to 1; the determining unit 502 may be specifically configured to:
[0326] Obtain the importance measurement values of N reference prediction values, and calculate the importance difference between the N reference prediction values. The importance difference between any two reference prediction values is measured by the difference between the importance measurement values of any two reference prediction values;
[0327] Determine the threshold interval where the absolute value of the importance difference between the N reference prediction values is located;
[0328] Determine the weight list corresponding to the threshold interval where the absolute value of the importance difference between the N reference prediction values is located as the target weight list.
[0329] In one embodiment, the N reference prediction values of the current block include a first reference prediction value and a second reference prediction value; the video bitstream includes a first weight list and a second weight list; the determining unit 502 may be specifically configured to:
[0330] Compare the magnitude of the importance measurement value of the first reference prediction value and the importance measurement value of the second reference prediction value;
[0331] If the importance metric value of the first reference prediction value is greater than the importance metric value of the second reference prediction value, then determine the first weight list as the target weight list;
[0332] If the importance metric value of the first reference prediction value is less than or equal to the importance metric value of the second reference prediction value, then determine the second weight list as the target weight list;
[0333] Wherein, the weight values in the first weight list are opposite to the weight values in the second weight list.
[0334] In one embodiment, among the N reference prediction values of the current block, the first reference prediction value and the second reference prediction value are included; in the video bitstream, the first weight list, the second weight list, and the third weight list are included; the determining unit 502 can be specifically configured to:
[0335] Call a mathematical symbol function to process the difference between the importance metric value of the first reference prediction value and the importance metric value of the second reference prediction value to obtain a symbol value;
[0336] If the symbol value is the first preset value, then determine the first weight list as the target weight list;
[0337] If the symbol value is the second preset value, then determine the second weight list as the target weight list;
[0338] If the symbol value is the third preset value, then determine the third weight list as the target weight list;
[0339] Wherein, the first weight list, the second weight list, and the third weight list are different weight lists respectively; or, among the first weight list, the second weight list, and the third weight list, two weight lists are allowed to be the same weight list.
[0340] In one embodiment, any one of the N reference prediction values is denoted as reference prediction value i, the reference prediction value i is derived from reference block i, and the video frame where the reference block i is located is reference frame i; i is an integer and i is less than or equal to N; the video frame where the current block is located is the current frame;
[0341] The importance metric value of the reference prediction value i is determined by any one of the following methods:
[0342] Method 1: Calculate according to the frame display order of the current frame in the video bitstream and the frame display order of the reference frame i in the video bitstream;
[0343] Method 2: Calculate according to the frame display order of the current frame in the video bitstream, the frame display order of the reference frame i in the video bitstream, and the quality metric Q;
[0344] Method 3: Calculate the importance metric score of reference frame i based on the calculation results of Method 1 and Method 2. Sort the importance metric scores of the reference frames corresponding to the N reference prediction values in ascending order, and determine the index of reference frame i in the sorting as the importance metric value of reference prediction value i;
[0345] Method 4: Adjust the calculation results of Method 1, Method 2, or Method 3 based on the prediction mode of reference prediction value i to obtain the importance metric value of reference prediction value i;
[0346] Among them, the prediction mode of reference prediction value i includes any one of the following: inter-frame prediction mode, intra-frame prediction mode.
[0347] In one embodiment, the number of weight groups included in the target weight list is greater than 1; the determining unit 502 may be specifically configured to:
[0348] Decode the index of the target weight group for weighted prediction from the video bitstream;
[0349] Select the target weight group from the target weight list according to the index of the target weight group;
[0350] Among them, the index of the target weight group uses a binary coding method of truncated unary code or a multi-symbol entropy coding method.
[0351] In one embodiment, the processing unit 501 may be specifically configured to:
[0352] Perform weighted summation processing on the N reference prediction values respectively using the weight values in the target weight group to obtain the prediction value of the current block; or,
[0353] Perform weighted processing on the N reference prediction values respectively using the weight values in the target weight group in the form of integer calculation to obtain the prediction value of the current block.
[0354] In the embodiments of the present application, perform composite prediction on the current block in the video bitstream to obtain N reference prediction values of the current block. The current block refers to the coding block being decoded in the video bitstream. According to the importance of the N reference prediction values, adaptively select the target weight group for weighted prediction for the current block; perform weighted prediction processing on the N reference prediction values based on the target weight group to obtain the prediction value of the current block. The prediction value of the current block is used to reconstruct the decoded image corresponding to the current block. Using composite prediction during the decoding process of the video and fully considering the importance of the reference prediction values in the composite prediction can realize determining appropriate weights for weighted prediction for the current block based on the importance of each reference prediction value in the composite prediction, which can improve the prediction accuracy of the current block and thus improve the codec performance.
[0355] Please refer to Figure 6 ,Figure 6 FIG. Figure 6 is a schematic structural diagram of a video processing apparatus provided by an embodiment of the present application. The video processing apparatus may be disposed in a computer device provided by an embodiment of the present application, and the computer device may be an encoding device mentioned in the above method embodiment. Figure 6 The video processing apparatus shown may be a computer program (including program code) running on a computer device, and the video processing apparatus may be used to execute Figure 4 part or all of the steps in the method embodiment shown. Please refer to Figure 6 . The video processing apparatus may include the following units:
[0356] A processing unit 601, configured to perform a partitioning process on a current frame in a video to obtain a current block;
[0357] The processing unit 601 is further configured to perform composite prediction on the current block to obtain N reference prediction values of the current block, where N is an integer greater than 1;
[0358] A determining unit 602, configured to determine a target weight group for weighted prediction for the current block according to the importance of the N reference prediction values, where the target weight group includes one or more weight values;
[0359] The processing unit 601 is further configured to perform weighted prediction processing on the N reference prediction values based on the target weight group to obtain a prediction value of the current block, and the prediction value of the current block is used to reconstruct a decoded image corresponding to the current block;
[0360] The processing unit 601 is further configured to encode the video based on the prediction value of the current block to generate a video bitstream.
[0361] In one embodiment, the determining unit 602 may be specifically configured to:
[0362] Determine a target weight list from one or more weight lists according to the importance of the N reference prediction values; each weight list includes one or more weight groups, each weight group includes one or more weight values, the number of weight values included in each weight group is allowed to be the same or different, and the values of the weight values included in each weight group are allowed to be the same or different;
[0363] Select a target weight group for weighted prediction from the target weight list.
[0364] In one embodiment, the bit rate consumed when performing weighted prediction processing on the N reference prediction values using the weight values in the target weight group is less than a preset bit rate threshold; or,
[0365] Based on the weight values in the target weight group, perform weighted prediction processing on the N reference prediction values, so that the quality loss of the current block during the encoding process is less than a preset loss threshold; or,
[0366] When weighted prediction processing is performed on N reference prediction values using the weight values in the target weight combination, the code rate consumed is less than a preset code rate threshold, and weighted prediction processing is performed on the N reference prediction values based on the weight values in the target weight combination, such that the quality loss of the current block during the encoding process is less than a preset loss threshold.
[0367] In one embodiment, the target weight combination is the weight combination with the optimal encoding performance in the target weight list;
[0368] Among them, the optimal encoding performance includes: the code rate consumed when weighted prediction processing is performed on N reference prediction values using the weight values in the target weight combination is the minimum of the code rates consumed corresponding to all weight combinations in the target weight list; or, the quality loss of the current block corresponding to when weighted prediction processing is performed on the N reference prediction values based on the weight values in the target weight combination is the minimum of the quality losses corresponding to all weight combinations in the target weight list; or, the code rate consumed when weighted prediction processing is performed on N reference prediction values using the weight values in the target weight combination is the minimum of the code rates consumed corresponding to all weight combinations in the target weight list, and the quality loss of the current block corresponding to when weighted prediction processing is performed on the N reference prediction values based on the weight values in the target weight combination is the minimum of the quality losses corresponding to all weight combinations in the target weight list.
[0369] In the embodiments of the present application, the current frame in the video is partitioned to obtain a current block, where the current block refers to the encoding block being encoded in the video; composite prediction is performed on the current block to obtain N reference prediction values of the current block; according to the importance of the N reference prediction values, a target weight combination for weighted prediction is determined for the current block; weighted prediction processing is performed on the N reference prediction values based on the target weight combination to obtain a prediction value of the current block; and the video is encoded based on the prediction value of the current block to generate a video bitstream. Composite prediction is used during the encoding process of the video, and by fully considering the importance of the reference prediction values and implementing weighted prediction by determining appropriate weights for the current block based on the importance of each reference prediction value in the composite prediction, the prediction accuracy of the current block can be improved, thereby improving the encoding and decoding performance.
[0370] Further, the embodiments of the present application also provide a schematic structural diagram of a computer device, and the schematic structural diagram of the computer device can be referred to Figure 7; The computer device may be the above-mentioned encoding device or decoding device; the computer device may include: a processor 701, an input device 702, an output device 703, and a memory 704. The above-mentioned processor 701, input device 702, output device 703, and memory 704 are connected through a bus. The memory 704 is used to store a computer program, and the computer program includes program instructions. The processor 701 is used to execute the program instructions stored in the memory 704.
[0371] When the computer device is the above-mentioned decoding device, in an embodiment of the present application, the processor 701 performs the following operations by running the executable program code in the memory 704:
[0372] Perform composite prediction on the current block in the video bitstream to obtain N reference prediction values of the current block, where N is an integer greater than 1;
[0373] Determine a target weight group for weighted prediction for the current block according to the importance of the N reference prediction values, and the target weight group includes one or more weight values;
[0374] Perform weighted prediction processing on the N reference prediction values based on the target weight group to obtain a prediction value of the current block, and the prediction value of the current block is used to reconstruct the decoded image corresponding to the current block.
[0375] In one embodiment, the N reference prediction values are derived from N reference blocks; one reference prediction value corresponds to one reference block; the video frame where the reference block is located is a reference frame, and the video frame where the current block is located is the current frame;
[0376] The positional relationship between the N reference blocks and the current block includes any one of the following:
[0377] The N reference blocks are respectively located in N reference frames, and the N reference frames and the current frame belong to different video frames in the video bitstream;
[0378] The N reference blocks are located in the same reference frame, and the same reference frame and the current frame belong to different video frames in the video bitstream;
[0379] One or more of the N reference blocks are located in the current frame, the remaining reference blocks among the N reference blocks are located in one or more reference frames, and the one or more reference frames and the current frame belong to different video frames in the video bitstream;
[0380] The N reference blocks and the current block are both located in the current frame.
[0381] In one embodiment, the processor 701 is further used to:
[0382] Determine whether the current block meets the conditions for adaptive weighted prediction;
[0383] If the current block meets the conditions for adaptive weighted prediction, a target weight set for weighted prediction is determined for the current block according to the importance of N reference prediction values.
[0384] In one embodiment, the conditions for the current block to meet the adaptive weighted prediction include at least one of the following:
[0385] A first indication field is included in the sequence header of the frame sequence to which the current block belongs, and the first indication field indicates that the coded blocks in the frame sequence are allowed to use adaptive weighted prediction;
[0386] A second indication field is included in the header of the current slice to which the current block belongs, and the second indication field indicates that the coded blocks in the current slice are allowed to use adaptive weighted prediction;
[0387] A third indication field is included in the frame header of the current frame in which the current block is located, and the third indication field indicates that the coded blocks in the current frame are allowed to use adaptive weighted prediction;
[0388] During the composite prediction process, the current block uses at least two reference frames for inter-frame prediction;
[0389] During the composite prediction process, the current block uses at least one reference frame for inter-frame prediction and uses the current frame for intra-frame prediction;
[0390] The motion type of the current block is a specified motion type;
[0391] The current block uses a preset motion vector prediction mode;
[0392] The current block uses a preset interpolation filter;
[0393] The current block does not use a specific coding tool;
[0394] The reference frames used by the current block during the composite prediction process meet specific conditions, and the specific conditions include one or more of the following: the orientation relationship between the reference frames used and the current frame in the video bitstream meets a preset relationship; the absolute value of the importance difference between the reference prediction values corresponding to the reference frames used is greater than or equal to a preset threshold;
[0395] Among them, the orientation relationship meeting the preset relationship includes any one of the following: all the reference frames used are before the current frame; all the reference frames used are after the current frame; some of the reference frames used are before the current frame, and the remaining reference frames are after the current frame.
[0396] In one embodiment, the video bitstream includes one or more weight lists, each weight list includes one or more weight groups, each weight group includes one or more weight values, the number of weight values included in each weight group is allowed to be the same or different, and the values of the weight values included in each weight group are allowed to be the same or different; the processor 701 may be specifically configured to:
[0397] Determine a target weight list from one or more weight lists according to the importance of N reference prediction values;
[0398] Select a target weight group for weighted prediction from the target weight list.
[0399] In one embodiment, the number of weight lists in the video bitstream is M + 1, one weight list corresponds to one threshold interval, and M is a positive integer greater than or equal to 1; the processor 701 may be specifically configured to:
[0400] Obtain the importance measurement values of N reference prediction values, and calculate the importance difference between the N reference prediction values. The importance difference between any two reference prediction values is measured by the difference between the importance measurement values of any two reference prediction values;
[0401] Determine the threshold interval in which the absolute value of the importance difference between the N reference prediction values is located;
[0402] Determine the weight list corresponding to the threshold interval in which the absolute value of the importance difference between the N reference prediction values is located as the target weight list.
[0403] In one embodiment, the N reference prediction values of the current block include a first reference prediction value and a second reference prediction value; the video bitstream includes a first weight list and a second weight list; the processor 701 may be specifically configured to:
[0404] Compare the magnitude of the importance measurement value of the first reference prediction value and the importance measurement value of the second reference prediction value;
[0405] If the importance measurement value of the first reference prediction value is greater than the importance measurement value of the second reference prediction value, determine the first weight list as the target weight list;
[0406] If the importance measurement value of the first reference prediction value is less than or equal to the importance measurement value of the second reference prediction value, determine the second weight list as the target weight list;
[0407] Wherein, the weight values in the first weight list are opposite to the weight values in the second weight list.
[0408] In one embodiment, among the N reference prediction values of the current block, there are a first reference prediction value and a second reference prediction value; the video bitstream includes a first weight list, a second weight list, and a third weight list; when the processor 701 determines the target weight list from one or more weight lists according to the importance of the N reference prediction values, it may specifically be used for:
[0409] Invoking a mathematical symbol function to process the difference between the importance metric value of the first reference prediction value and the importance metric value of the second reference prediction value to obtain a symbol value;
[0410] If the symbol value is a first preset value, determining the first weight list as the target weight list;
[0411] If the symbol value is a second preset value, determining the second weight list as the target weight list;
[0412] If the symbol value is a third preset value, determining the third weight list as the target weight list;
[0413] Wherein, the first weight list, the second weight list, and the third weight list are different weight lists respectively; or, two of the first weight list, the second weight list, and the third weight list are allowed to be the same weight list.
[0414] In one embodiment, any one of the N reference prediction values is represented as reference prediction value i, reference prediction value i is derived from reference block i, and the video frame where reference block i is located is reference frame i; i is an integer and i is less than or equal to N; the video frame where the current block is located is the current frame;
[0415] The importance metric value of reference prediction value i is determined by any of the following methods:
[0416] Method 1: Calculated according to the frame display order of the current frame in the video bitstream and the frame display order of reference frame i in the video bitstream;
[0417] Method 2: Calculated according to the frame display order of the current frame in the video bitstream, the frame display order of reference frame i in the video bitstream, and the quality metric Q;
[0418] Method 3: Calculating the importance metric score of reference frame i according to the calculation results of Method 1 and Method 2, sorting the importance metric scores of the reference frames corresponding to the N reference prediction values in ascending order, and determining the index of reference frame i in the sorting as the importance metric value of reference prediction value i;
[0419] Method 4: Adjusting the calculation results of Method 1, Method 2, or Method 3 based on the prediction mode of the reference prediction value i to obtain the importance metric value of the reference prediction value i;
[0420] Among them, the prediction mode of the reference prediction value i includes any one of the following: an inter-frame prediction mode, an intra-frame prediction mode.
[0421] In one embodiment, the number of weight groups included in the target weight list is greater than 1; the processor 701 may specifically be configured to:
[0422] Decode the index of the target weight group for weighted prediction from the video bitstream;
[0423] Select the target weight group from the target weight list according to the index of the target weight group;
[0424] Among them, the index of the target weight group uses a binary coding method of truncated unary code or an entropy coding method of multiple symbols.
[0425] In one embodiment, the processor 701 may specifically be configured to:
[0426] Perform weighted summation processing on N reference prediction values respectively using the weight values in the target weight group to obtain the prediction value of the current block; or,
[0427] Perform weighted processing on N reference prediction values respectively using the weight values in the target weight group in the form of integer calculation to obtain the prediction value of the current block.
[0428] In the embodiment of the present application, perform composite prediction on the current block in the video bitstream to obtain N reference prediction values of the current block. The current block refers to the encoded block being decoded in the video bitstream. According to the importance of the N reference prediction values, determine the target weight group for weighted prediction for the current block; perform weighted prediction processing on the N reference prediction values based on the target weight group to obtain the prediction value of the current block. The prediction value of the current block is used to reconstruct the decoded image corresponding to the current block. Using composite prediction during the decoding process of the video and fully considering the importance of the reference prediction values in the composite prediction, it is possible to determine appropriate weights for weighted prediction for the current block based on the importance of each reference prediction value in the composite prediction, which can improve the prediction accuracy of the current block and thus improve the encoding and decoding performance.
[0429] Optionally, when the computer device is the above-mentioned encoding device, in the embodiment of the present application, the processor 701 executes the following operations by running the executable program code in the memory 704:
[0430] Perform partitioning processing on the current frame in the video to obtain the current block;
[0431] Perform composite prediction on the current block to obtain N reference prediction values of the current block, where N is an integer greater than 1;
[0432] Determine a target weight group for weighted prediction for the current block according to the importance of N reference prediction values, where the target weight group includes one or more weight values;
[0433] Perform weighted prediction processing on the N reference prediction values based on the target weight group to obtain a prediction value for the current block, and the prediction value of the current block is used to reconstruct the decoded image corresponding to the current block;
[0434] Encode the video based on the prediction value of the current block to generate a video bitstream.
[0435] In one embodiment, the processor 701 may specifically be used for:
[0436] Determine a target weight list from one or more weight lists according to the importance of the N reference prediction values; each weight list includes one or more weight groups, each weight group includes one or more weight values, the number of weight values included in each weight group is allowed to be the same or different, and the values of the weight values included in each weight group are allowed to be the same or different;
[0437] Select a target weight group for weighted prediction from the target weight list.
[0438] In one embodiment, the bit rate consumed when performing weighted prediction processing on the N reference prediction values using the weight values in the target weight group is less than a preset bit rate threshold; or,
[0439] Perform weighted prediction processing on the N reference prediction values based on the weight values in the target weight group, such that the quality loss of the current block during the encoding process is less than a preset loss threshold; or,
[0440] The bit rate consumed when performing weighted prediction processing on the N reference prediction values using the weight values in the target weight group is less than a preset bit rate threshold, and perform weighted prediction processing on the N reference prediction values based on the weight values in the target weight group, such that the quality loss of the current block during the encoding process is less than a preset loss threshold.
[0441] In one embodiment, the target weight group is the weight group with the optimal encoding performance in the target weight list;
[0442] Among them, the optimal coding performance includes: when performing weighted prediction processing on N reference prediction values using the weight values in the target weight combination, the bitrate consumed is the minimum among the bitrates consumed by all weight combinations in the target weight list; or, when performing weighted prediction processing on N reference prediction values based on the weight values in the target weight combination, the quality loss of the current block during the coding process is the minimum among the quality losses of all weight combinations in the target weight list; or, when performing weighted prediction processing on N reference prediction values using the weight values in the target weight combination, the bitrate consumed is the minimum among the bitrates consumed by all weight combinations in the target weight list, and when performing weighted prediction processing on N reference prediction values based on the weight values in the target weight combination, the quality loss of the current block during the coding process is the minimum among the quality losses of all weight combinations in the target weight list.
[0443] In the embodiments of the present application, the current frame in the video is divided to obtain a current block, where the current block refers to the coding block being encoded in the video; a composite prediction is performed on the current block to obtain N reference prediction values of the current block; according to the importance of the N reference prediction values, a target weight combination for weighted prediction is determined for the current block; a weighted prediction process is performed on the N reference prediction values based on the target weight combination to obtain a prediction value of the current block; and the video is encoded based on the prediction value of the current block to generate a video bitstream. Using composite prediction during the video encoding process and fully considering the importance of the reference prediction values, and realizing determining appropriate weights for weighted prediction for the current block based on the importance of each reference prediction value in the composite prediction can improve the prediction accuracy of the current block, thereby improving the encoding and decoding performance.
[0444] In addition, it should be pointed out here that: the embodiments of the present application also provide a computer-readable storage medium, and a computer program is stored in the computer-readable storage medium, and the computer program includes program instructions. When the processor executes the above program instructions, it can execute the methods in the corresponding embodiments described above. Figure 3 and Figure 4 Therefore, it will not be elaborated here. For the technical details not disclosed in the embodiments of the computer-readable storage medium involved in the present application, please refer to the description of the method embodiments of the present application. As an example, the program instructions can be deployed on a computer device, or executed on multiple computer devices located at one place, or executed on multiple computer devices distributed at multiple places and interconnected through a communication network.
[0445] According to one aspect of the present application, a computer program product is provided. The computer program product includes a computer program, and the computer program is stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device can execute the aboveFigure 3 and Figure 4 the methods in the corresponding embodiments, and therefore, will not be elaborated here.
[0446] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0447] The above-disclosed is only a preferred embodiment of the present application. Of course, the scope of the rights of the present application cannot be limited thereby. Those of ordinary skill in the art can understand all or part of the processes of the above embodiments, and the equivalent changes made according to the claims of the present application still fall within the scope covered by the invention.
Claims
1. A video processing method, characterized in that, it includes: Performing composite prediction on a current block in a video bitstream to obtain N reference prediction values of the current block, where N is an integer greater than 1; Determining a target weight group for weighted prediction for the current block according to the importance of the N reference prediction values, the target weight group including one or more weight values; Performing weighted prediction processing on the N reference prediction values based on the target weight group to obtain a prediction value of the current block, and the prediction value of the current block is used to reconstruct a decoded image corresponding to the current block.
2. The method according to claim 1, characterized in that, the N reference prediction values are derived from N reference blocks; one reference prediction value corresponds to one reference block; the video frame where the reference block is located is a reference frame, and the video frame where the current block is located is a current frame; the positional relationship between the N reference blocks and the current block includes any one of the following: the N reference blocks are respectively located in N reference frames, and the N reference frames and the current frame belong to different video frames in the video bitstream; the N reference blocks are located in the same reference frame, and the same reference frame and the current frame belong to different video frames in the video bitstream; one or more of the N reference blocks are located in the current frame, the remaining reference blocks among the N reference blocks are located in one or more reference frames, and the one or more reference frames and the current frame belong to different video frames in the video bitstream; the N reference blocks and the current block are both located in the current frame.
3. The method according to claim 1, characterized in that, the method further includes: Determining whether the current block meets the conditions for adaptive weighted prediction; If the current block meets the conditions for adaptive weighted prediction, then perform the step of determining a target weight group for weighted prediction for the current block according to the importance of the N reference prediction values.
4. The method according to claim 3, characterized in that, the current block meeting the conditions for adaptive weighted prediction includes at least one of the following: a first indication field is included in the sequence header of the frame sequence to which the current block belongs, and the first indication field indicates that the coding blocks in the frame sequence are allowed to use adaptive weighted prediction; a second indication field is included in the header of the current slice to which the current block belongs, and the second indication field indicates that the coding blocks in the current slice are allowed to use adaptive weighted prediction; a third indication field is included in the frame header of the current frame where the current block is located, and the third indication field indicates that the coding blocks in the current frame are allowed to use adaptive weighted prediction; during the composite prediction process, the current block uses at least two reference frames for inter-frame prediction; during the composite prediction process, the current block uses at least one reference frame for inter-frame prediction and uses the current frame for intra-frame prediction; the motion type of the current block is a specified motion type; the current block uses a preset motion vector prediction mode; the current block uses a preset interpolation filter; the current block does not use a specific coding tool; The reference frames used by the current block in the composite prediction process satisfy specific conditions, and the specific conditions include one or more of the following: the orientation relationship between the reference frames used and the current frame in the video bitstream satisfies a preset relationship; the absolute value of the importance difference between the reference prediction values corresponding to the reference frames used is greater than or equal to a preset threshold; wherein, the orientation relationship satisfying the preset relationship includes any one of the following: all the reference frames used are before the current frame; all the reference frames used are after the current frame; some of the reference frames used are before the current frame, and the remaining reference frames are after the current frame.
5. The method according to claim 1, wherein, the video bitstream includes one or more weight lists, each weight list includes one or more weight groups, each weight group includes one or more weight values, the number of weight values included in each weight group is allowed to be the same or different, and the values of the weight values included in each weight group are allowed to be the same or different; determining a target weight group for weighted prediction for the current block according to the importance of the N reference prediction values includes: determining a target weight list from the one or more weight lists according to the importance of the N reference prediction values; selecting a target weight group for weighted prediction from the target weight list.
6. The method according to claim 5, wherein, the number of weight lists in the video bitstream is M + 1, one weight list corresponds to one threshold interval, and M is a positive integer greater than or equal to 1; determining a target weight list from the one or more weight lists according to the importance of the N reference prediction values includes: obtaining an importance measurement value of the N reference prediction values, and calculating the importance difference between the N reference prediction values, and the importance difference between any two reference prediction values is measured by the difference between the importance measurement values of the any two reference prediction values; determining the threshold interval where the absolute value of the importance difference between the N reference prediction values is located; determining the weight list corresponding to the threshold interval where the absolute value of the importance difference between the N reference prediction values is located as the target weight list.
7. The method according to claim 5, wherein, the N reference prediction values of the current block include a first reference prediction value and a second reference prediction value; the video bitstream includes a first weight list and a second weight list; determining a target weight list from the one or more weight lists according to the importance of the N reference prediction values includes: comparing the magnitude of the importance measurement value of the first reference prediction value and the importance measurement value of the second reference prediction value; if the importance measurement value of the first reference prediction value is greater than the importance measurement value of the second reference prediction value, then determining the first weight list as the target weight list; if the importance measurement value of the first reference prediction value is less than or equal to the importance measurement value of the second reference prediction value, then determining the second weight list as the target weight list; Among them, the weight values in the first weight list are opposite to the weight values in the second weight list.
8. The method according to claim 5, wherein, the N reference prediction values of the current block include a first reference prediction value and a second reference prediction value; the video bitstream includes a first weight list, a second weight list, and a third weight list; determining a target weight list from the one or more weight lists according to the importance of the N reference prediction values includes: invoking a mathematical symbol function to process the difference between the importance measurement value of the first reference prediction value and the importance measurement value of the second reference prediction value to obtain a symbol value; if the symbol value is a first preset value, determining the first weight list as the target weight list; if the symbol value is a second preset value, determining the second weight list as the target weight list; if the symbol value is a third preset value, determining the third weight list as the target weight list; wherein, the first weight list, the second weight list, and the third weight list are respectively different weight lists; or, two of the first weight list, the second weight list, and the third weight list are allowed to be the same weight list.
9. The method according to any one of claims 6-8, wherein, any one of the N reference prediction values is represented as reference prediction value i, the reference prediction value i is derived from reference block i, and the video frame where the reference block i is located is reference frame i; i is an integer and i is less than or equal to N; the video frame where the current block is located is the current frame; the importance measurement value of the reference prediction value i is determined by any one of the following methods: Method 1: Calculated according to the frame display order of the current frame in the video bitstream and the frame display order of the reference frame i in the video bitstream; Method 2: Calculated according to the frame display order of the current frame in the video bitstream, the frame display order of the reference frame i in the video bitstream, and the quality metric Q; Method 3: Calculating the importance metric score of the reference frame i according to the calculation results of Method 1 and Method 2, sorting the importance metric scores of the reference frames corresponding to the N reference prediction values in ascending order, and determining the index of the reference frame i in the sorting as the importance measurement value of the reference prediction value i; Method 4: Adjusting the calculation results of Method 1, Method 2, or Method 3 based on the prediction mode of the reference prediction value i to obtain the importance measurement value of the reference prediction value i; wherein, the prediction mode of the reference prediction value i includes any one of the following: inter-frame prediction mode, intra-frame prediction mode.
10. The method according to claim 5, wherein, the number of weight groups included in the target weight list is greater than 1; selecting a target weight group for weighted prediction from the target weight list includes: decoding an index of the target weight group for weighted prediction from the video bitstream; selecting the target weight group from the target weight list according to the index of the target weight group; Among them, the index of the target weight group is binarized and encoded using a truncated unary code, or using an entropy encoding method with multiple symbols.
11. The method according to claim 1, wherein, the weighted prediction process of the N reference prediction values based on the target weight group to obtain the prediction value of the current block includes: performing weighted summation processing on the N reference prediction values respectively using the weight values in the target weight group to obtain the prediction value of the current block; or, performing weighted processing on the N reference prediction values respectively using the weight values in the target weight group in the form of integer calculation to obtain the prediction value of the current block.
12. A video processing method, wherein, including: performing partitioning processing on the current frame in the video to obtain a current block; performing composite prediction on the current block to obtain N reference prediction values of the current block, where N is an integer greater than 1; determining a target weight group for weighted prediction for the current block according to the importance of the N reference prediction values, the target weight group including one or more weight values; performing weighted prediction processing on the N reference prediction values based on the target weight group to obtain the prediction value of the current block, and the prediction value of the current block is used to reconstruct the decoded image corresponding to the current block; encoding the video based on the prediction value of the current block to generate a video bitstream.
13. The method according to claim 12, wherein, the determining a target weight group for weighted prediction for the current block according to the importance of the N reference prediction values includes: determining a target weight list from one or more weight lists according to the importance of the N reference prediction values; each weight list includes one or more weight groups, each weight group includes one or more weight values, the number of weight values included in each weight group is allowed to be the same or different, and the values of the weight values included in each weight group are allowed to be the same or different; selecting a target weight group for weighted prediction from the target weight list.
14. The method according to claim 13, wherein, the bit rate consumed when performing weighted prediction processing on the N reference prediction values using the weight values in the target weight group is less than a preset bit rate threshold; or, performing weighted prediction processing on the N reference prediction values based on the weight values in the target weight group, such that the quality loss of the current block during the encoding process is less than a preset loss threshold; or, the bit rate consumed when performing weighted prediction processing on the N reference prediction values using the weight values in the target weight group is less than a preset bit rate threshold, and performing weighted prediction processing on the N reference prediction values based on the weight values in the target weight group, such that the quality loss of the current block during the encoding process is less than a preset loss threshold.
15. The method according to claim 13, wherein, the target weight group is the weight group with the optimal encoding performance in the target weight list; Among them, the optimal coding performance includes: when weighted prediction processing is performed on N reference prediction values using the weight values in the target weight combination, the code rate consumed is the minimum among the code rates consumed by all weight combinations in the target weight list; or, when weighted prediction processing is performed on N reference prediction values based on the weight values in the target weight combination, the quality loss of the current block during the coding process is the minimum among the quality losses of all weight combinations in the target weight list; or, when weighted prediction processing is performed on N reference prediction values using the weight values in the target weight combination, the code rate consumed is the minimum among the code rates consumed by all weight combinations in the target weight list, and when weighted prediction processing is performed on N reference prediction values based on the weight values in the target weight combination, the quality loss of the current block during the coding process is the minimum among the quality losses of all weight combinations in the target weight list.
16. A video processing device, characterized in that, it includes: a processing unit for performing composite prediction on the current block in the video bitstream to obtain N reference prediction values of the current block, where N is an integer greater than 1; a determination unit for determining a target weight combination for weighted prediction for the current block according to the importance of the N reference prediction values, the target weight combination including one or more weight values; the processing unit is further configured to perform weighted prediction processing on the N reference prediction values based on the target weight combination to obtain a prediction value of the current block, and the prediction value of the current block is used to reconstruct the decoded image corresponding to the current block.
17. A video processing device, characterized in that, it includes: a processing unit for dividing the current frame in the video to obtain a current block; the processing unit is further configured to perform composite prediction on the current block to obtain N reference prediction values of the current block, where N is an integer greater than 1; a determination unit for determining a target weight combination for weighted prediction for the current block according to the importance of the N reference prediction values, the target weight combination including one or more weight values; the processing unit is further configured to perform weighted prediction processing on the N reference prediction values based on the target weight combination to obtain a prediction value of the current block, and the prediction value of the current block is used to reconstruct the decoded image corresponding to the current block; the processing unit is further configured to encode the video based on the prediction value of the current block to generate a video bitstream.
18. A computer device, characterized in that, it includes: a processor adapted to execute a computer program; a computer-readable storage medium storing a computer program, and when the computer program is executed by the processor, it executes the video processing method according to any one of claims 1-15.
19. A computer-readable storage medium, characterized in that, the computer storage medium stores a computer program, and when the computer program is executed by a processor, it executes the video processing method according to any one of claims 1-15.
20. A computer program product, characterized in that, The computer program product includes a computer program which, when executed by a processor, implements the video processing method according to any one of claims 1-15.
Citation Information
Patent Citations
Encoding method and device therefor, and decoding method and device therefor
CN112438045A
Video coding and decoding method and device, computer readable medium and electronic equipment
CN114885160A