VIDEO PROCESSING METHOD, VIDEO PROCESSING APPARATUS, COMPUTER DEVICE, AND COMPUTER PROGRAM

By adopting a composite prediction method in video encoding, using the importance of the reference prediction value to determine the weight and perform weighted prediction, the problem of low composite prediction accuracy in the prior art is solved, and the performance of video encoding and decoding is improved.

JP2025515342AActive Publication Date: 2025-05-14TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
JP2024563425
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-30
Filing Date
2023-07-07
Publication Date
2025-05-14
Estimated Expiration
2043-07-07

AI Technical Summary

Technical Problem

In the existing video encoding technology, composite prediction has low prediction accuracy during weighted prediction, which affects video encoding and decoding performance.

Method used

By using a composite prediction method in video encoding, the current block is predicted, and N reference prediction values ​​are used to determine an appropriate weight combination based on the importance of these reference prediction values, and weighted prediction is performed to improve the prediction accuracy.

Benefits of technology

By considering the importance of the reference predicted value, the prediction accuracy of video encoding and decoding is improved and the encoding performance is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025515342000001_ABST
    Figure 2025515342000001_ABST
Patent Text Reader

Abstract

A video processing method executed by a computer device, comprising: a step (S301) of obtaining N reference prediction values ​​of a current block by performing composite prediction of the current block in a video bitstream, the current block being a coding block to be decoded in the video bitstream, the N reference prediction values ​​being derived from N reference blocks of the current block, the N reference blocks being coding blocks in the video bitstream that are referenced when the current block is decoded, and the reference prediction values ​​and the reference blocks having a one-to-one correspondence; and The method includes a step (S302) of determining a target weight group for weighted prediction for a current block, where the target weight group includes one or more weight values, and the importance is for indicating the degree of influence of each reference prediction value on the decoding performance of the current block; and a step (S303) of performing weighted prediction processing for the N reference prediction values ​​based on the weight values ​​in the target weight group to obtain a prediction value for the current block, where the prediction value for the current block is for reconstructing a decoded image corresponding to the current block.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] This application claims priority to a Chinese patent application filed with the China Patent Office on September 30, 2022, bearing application number 2022112305931 and entitled "Video Processing Method and Related Apparatus," the entire contents of which are incorporated herein by reference. The present application relates to the technical field of audio-video, in particular to the technical field of video encoding, and in particular to a video processing method, a video processing device, a computer apparatus, a computer-readable storage medium, and a computer program product. [Background technology]

[0002] In conventional video coding technology, a block-based hybrid coding framework is adopted to divide the original video data into a series of coding blocks (CUs) and combine video coding methods such as prediction, transformation, and entropy coding to realize video data compression. To achieve better prediction effect, current mainstream video coding standards, such as AV1 (the first generation video coding standard formulated by the Alliance for Open Media Video 1) standard and the Alliance for Open Media (AOM) next-generation standard AV2 (the next generation video coding standard formulated by the Alliance for Open Media Video 2), which is under development, include a prediction mode called compound prediction. In this compound prediction mode, multiple types of reference video signals are allowed to be used for weighted prediction. However, it can be seen from practice that the current compound prediction does not have high prediction accuracy during weighted prediction, which affects the performance of video encoding and decoding. Summary of the Invention [Problem to be solved by the invention]

[0003] According to various embodiments provided herein, a video processing method and related apparatus are provided. [Means for solving the problem]

[0004] According to one aspect, an embodiment of the present application provides a computing device implemented video processing method, the method comprising: The method includes the steps of: obtaining N (N is an integer greater than 1) reference prediction values ​​of a current block by performing composite prediction of the current block in a video bitstream, where the current block is a coding block to be decoded in the video bitstream, the N reference prediction values ​​are derived from N reference blocks of the current block, the N reference blocks are coding blocks in the video bitstream that are referenced when decoding the current block, and there is a one-to-one correspondence between the reference prediction values ​​and the reference blocks; determining a target weight group for weighted prediction for the current block based on importance of the N reference prediction values, where the target weight group includes one or more weight values, and the importance is for representing the degree of influence of each reference prediction value on the decoding performance of the current block; and obtaining a prediction value of the current block by performing weighted prediction processing of the N reference prediction values ​​based on the weight values ​​in the target weight group, where the prediction value of the current block is for reconstructing a decoded image corresponding to the current block.

[0005] According to one aspect, an embodiment of the present application provides a computing device implemented video processing method, the method comprising: The method includes the steps of: obtaining a current block by performing a division process of a current frame in a video; obtaining N (N is an integer greater than 1) reference prediction values ​​of the current block by performing composite prediction of the current block, where the N reference prediction values ​​are derived from N reference blocks of the current block, the N reference blocks being coding blocks in the video that are referenced when coding the current block, and the reference prediction values ​​and the reference blocks correspond one-to-one; determining a target weight group for weighted prediction for the current block based on importance of the N reference prediction values, where the target weight group includes one or more weight values, and the importance is for representing a degree of influence of each reference prediction value on coding performance of the current block; obtaining a prediction value of the current block by performing a weighted prediction process of the N reference prediction values ​​based on the weight values ​​in the target weight group, where the prediction value of the current block is for reconstructing a decoded image corresponding to the current block; and generating a video bitstream by encoding the video based on the prediction value of the current block.

[0006] According to one aspect, an embodiment of the present application provides a video processing apparatus, the apparatus comprising: a processing unit for obtaining N (N is an integer greater than 1) reference prediction values ​​of a current block in a video bitstream by performing composite prediction of the current block, where the current block is a coding block to be decoded in the video bitstream, the N reference prediction values ​​being derived from N reference blocks of the current block, the N reference blocks being coding blocks in the video bitstream that are referenced when decoding the current block, and where the reference prediction values ​​and the reference blocks correspond one-to-one; and a determination unit for determining a target weight group for weighted prediction for the current block based on importance of the N reference prediction values, where the target weight group includes one or more weight values, and the importance is for representing the degree of influence of each reference prediction value on the decoding performance of the current block, where the processing unit further obtains a prediction value of the current block by performing weighted prediction processing of the N reference prediction values ​​based on the weight values ​​in the target weight group, and the prediction value of the current block is for reconstructing a decoded image corresponding to the current block.

[0007] According to one aspect, an embodiment of the present application provides a video processing apparatus, the apparatus comprising: the N reference prediction values ​​are derived from N reference blocks of the current block, the N reference blocks being coding blocks in the video that are referenced when the current block is coded, and the reference prediction values ​​and the reference blocks correspond one-to-one to each other; and a determination unit that determines a target weight group for weighted prediction for the current block based on importance of the N reference prediction values, the target weight group including one or more weight values, and the importance is for representing a degree of influence of each reference prediction value on coding performance of the current block. The processing unit further performs weighted prediction processing of the N reference prediction values ​​based on the weight values ​​in the target weight group to obtain a prediction value of the current block, the prediction value of the current block being for reconstructing a decoded image corresponding to the current block, and the processing unit further encodes the video based on the prediction value of the current block to generate a video bitstream.

[0008] According to one aspect, an embodiment of the present application provides a computing device, the computing device comprising: The system comprises a processor suitable for executing computer-readable instructions, and a computer-readable storage medium having stored thereon computer-readable instructions which, when executed by the processor, cause the system to implement the video processing method as described above.

[0009] According to one aspect, an embodiment of the present application provides a computer readable storage medium having stored thereon computer readable instructions which, when loaded by a processor, cause the processor to execute the video processing method as described above.

[0010] According to one aspect, an embodiment of the present application further provides a computer program product including computer readable instructions stored in a computer readable storage medium, the computer readable instructions being read by a processor of a computing device from the computer readable storage medium and, when executed by the processor, causing the computing device to perform the video processing method described above.

[0011] The details of one or more embodiments of the application are set forth in the drawings and description below. Other features, objects, and advantages of the application will become apparent from the description, drawings, and claims. [Brief description of the drawings]

[0012] In order to more clearly describe the configuration of the embodiments of the present application or the prior art, the following briefly introduces drawings necessary for the description of the embodiments or the prior art. Obviously, the drawings in the following description only show some embodiments of the present application, and those skilled in the art can obtain other drawings from these drawings without creative labor. [Figure 1a] 1 is a basic operational flowchart of video encoding provided in one exemplary embodiment of the present application; [Figure 1b] FIG. 2 is a schematic diagram of inter prediction provided in one exemplary embodiment of the present application; [Diagram 2] FIG. 1 is a schematic diagram of a configuration of a video processing system provided in one exemplary embodiment of the present application. [Diagram 3] FIG. 2 is a schematic diagram of the flow of a video processing method provided in one exemplary embodiment of the present application. [Figure 4] FIG. 2 is a schematic diagram of the flow of a video processing method provided in another exemplary embodiment of the present application. [Diagram 5] 1 is a schematic diagram of a configuration of a video processing device provided in one exemplary embodiment of the present application; [Figure 6] FIG. 2 is a schematic diagram of a configuration of a video processing device provided in another exemplary embodiment of the present application. [Figure 7]FIG. 2 is a schematic diagram of a configuration of a computer device provided in one exemplary embodiment of the present application. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0013] Hereinafter, the configuration of the embodiment of the present application will be described clearly and completely with reference to the drawings of the embodiment of the present application. It is obvious that the described embodiment is only a part of the embodiment of the present application, not all of the embodiments. All other embodiments that a person skilled in the art can obtain from the embodiment of the present application without creative labor belong to the scope of protection of the present application.

[0014] The following introduces technical terms related to this application.

[0015] I. Video Coding

[0016] A video may be composed of one or more video frames, each of which contains a part of the video signal of the video. The acquisition method of the video signal can be divided into two types: a camera-captured method and a computer-generated method. Different acquisition methods have different statistical characteristics, so the video compression encoding methods may also be different.

[0017] In mainstream video coding technologies, such as High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), AV1 (Alliance for Open Media Video 1), AV2 (Alliance for Open Media Video 2), and AVS3 (Audio Video coding Standard 3), a hybrid coding framework is used. This hybrid coding framework allows a series of operations and processes to be performed on video, such as:

[0018] 1) Block partition structure: Based on the size of the input current frame (i.e., the video frame being encoded or decoded), the current frame can be divided into a few non-overlapping processing units, and similar compression operations are performed on each processing unit. This processing unit is called a coding tree unit (CTU) or a largest coding unit (LCU). The CTU can be further divided downward to obtain one or more basic coding units called coding units or coding blocks (CUs). Each CU is the most basic element of the encoding process. Various encoding and decoding process flows that can be used for each CU are described in the following embodiments of the present application.

[0019] 2) Predictive coding: includes modes such as intra prediction and inter prediction. A residual video signal is obtained after performing prediction using a reconstructed video signal in a selected reference CU for an original video signal included in a current CU in a current frame (i.e., a CU being coded or decoded in a current frame). Here, the current CU is also called a current block, the video frame in which the current block is located is called a current frame, the reference CU used to predict the current block is also called a reference block of the current block, and the video frame in which the reference block is located is called a reference frame. Here, the coding side needs to decide to select the most appropriate one from among many possible predictive coding modes for the current CU and inform the decoding side. Here, the predictive coding modes may include the following:

[0020] a. Intra (picture) Prediction: The reconstructed video signal used for prediction is from an already coded and reconstructed region in the same video frame, i.e., the current block and the reference block are located in the same video frame. Here, the basic idea of ​​intra prediction is to remove spatial redundancy by utilizing the correlation of adjacent pixels in the same video frame. In video coding, the adjacent pixels refer to the reconstructed pixels of the coded CUs around the current CU in the same video frame.

[0021] b. Inter(picture) Prediction: The reconstructed video signal used for prediction is from another coded video frame different from the current frame, i.e. the current block and the reference block are located in different video frames.

[0022] 3) Transform & Quantization: The residual video signal can be transformed into a transform domain through a transform operation called transform coefficients, such as Discrete Fourier Transform (DFT), Discrete Cosine Transform (DCT), etc. The residual video signal in the transform domain is further subjected to a lossy quantization operation, in which certain information is lost, making the quantized signal more favorable for compressed representation.

[0023] In some video coding standards, there may be multiple transform methods to choose from. Therefore, the encoding side also needs to select one of the transforms for the current CU and inform the decoding side. The quantization granularity is usually determined by Quantization Parameters (QP). A large QP value indicates that a larger range of transform coefficients are quantized to the same output, and therefore usually results in larger distortion and lower bitrate. Conversely, a small QP value indicates that a smaller range of transform coefficients are quantized to the same output, and therefore usually results in smaller distortion and corresponds to a higher bitrate.

[0024] 4) Entropy Coding or Statistical Coding: Statistical compression coding is performed on the quantized transform domain signal according to the frequency of occurrence of each value, and a binarized (0 or 1) video bitstream is finally output. In addition, other information, such as a selected predictive coding mode and a motion vector, is generated by the coding. Entropy coding is also required for these other information to reduce the bit rate. Here, statistical coding is one of the lossless coding methods, and can effectively reduce the bit rate required to represent the same signal. Common statistical coding methods include variable length coding (VLC) and content adaptive binary arithmetic coding (CABAC).

[0025] 5) Loop Filtering: By performing operations of inverse quantization, inverse transformation, and prediction compensation (the inverse operations of 2) to 4) above) on an encoded CU, a decoded image corresponding to the CU can be reconstructed. Compared to the original image, the reconstructed decoded image is affected by quantization, and some information differs from the original image, resulting in distortion. Therefore, by performing a filtering operation on the reconstructed decoded image with a filter, the degree of distortion caused by quantization can be effectively reduced. The filter may be, for example, deblocking, sample adaptive offset (SAO), or adaptive loop filtering (ALF), etc. These filtered reconstructed decoded images are used for predicting other CUs as reference CUs for other CUs that need to be coded subsequently, and therefore the above filtering operation is also called loop filtering and filtering operation in a coding loop.

[0026] Based on the relevant description of steps 1) to 5) above, in the embodiment of the present application, a basic operation flowchart of a video encoder is provided. Please refer to FIG. 1a. FIG. 1a is a basic operation flowchart of a video encoder given as an example. Here, in FIG. 1a, a current block is a k-th CU (s shown in FIG. 1a) in a current frame (current image). k Here, k is a positive integer equal to or less than the total number of CUs included in the current frame. k [x,y] represents a pixel point (abbreviated as pixel) whose coordinates are [x,y] in the kth CU, where x represents the horizontal coordinate of the pixel and y represents the vertical coordinate of the pixel. k After performing processing such as motion compensation and intra-prediction on [x,y], a prediction signal is obtained.

number

[0027] A: The output data of the quantization process may be sent to an entropy encoder for entropy encoding to obtain an encoded bitstream (i.e., a video bitstream), which may then be output to a buffer for storage and awaiting transmission.

number

[0028] In some embodiments, mainstream video coding standards such as HEVC, VVC, AVS3, AV1, and AV2 all adopt a block-based hybrid coding framework, which divides original video data into a series of coding blocks and combines video coding methods such as prediction, transformation, and entropy coding to achieve video data compression. Here, motion compensation is a prediction method commonly used in video coding. Motion compensation derives a prediction value of a current block from a previously coded reference block based on the characteristics of redundancy of video content in the time domain or spatial domain. Such prediction methods include inter prediction, intra block copy prediction, intra string copy prediction, etc. In a specific coding implementation, these prediction methods can be used alone or in combination. For a coding block using these prediction methods, it is usually necessary to explicitly or implicitly code one or more two-dimensional displacement vectors in the video bitstream, which indicate the displacement of the current block (or a block co-located with the current block) relative to one or more reference blocks.

[0029] It should be noted that the displacement vector may have different names in different prediction modes and different implementations. In this specification, the following will be described in a unified manner: 1) the displacement vector in inter prediction is called a motion vector (abbreviated as MV), 2) the displacement vector in intra block copy prediction is called a block vector (abbreviated as BV), and 3) the displacement vector in intra string copy prediction is called a string vector (abbreviated as SV). In the following, the related technology of inter prediction will be introduced by taking inter prediction as an example.

[0030] Inter prediction: Inter prediction utilizes the correlation in the time domain of video, and predicts pixels of a current image using pixels of neighboring encoded images, thereby achieving the purpose of effectively removing redundancy in the time domain of video, and efficiently saving bits of encoded residual data. As shown in FIG. 1b, FIG. 1b is a schematic diagram of inter prediction provided in an embodiment of the present application. Here, in FIG. 1b, P is a current frame, Pr is a reference frame, B is a current block, and Br is a reference block of B. B' and B have the same coordinate position in the image (i.e., B' is a block at the same position as B), and the coordinate of Br is (x r ,y r ), and the coordinates of B' are (x,y). The displacement between the current block and its reference block is called the motion vector (MV), i.e., MV=(x r -x,y r -y).

[0031] Considering that neighboring blocks in the time or space domain have strong correlation, MV prediction techniques can be used to further reduce the bits required for encoding MVs. In H.265 / HEVC, inter prediction includes two MV prediction techniques: Merge and Advanced Motion Vector Prediction (AMVP). In Merge mode, an MV candidate list is created for a current prediction unit (PU), in which there are five candidate MVs (and their corresponding reference images). These five candidate MVs are scanned, and the one with the smallest rate distortion cost is selected as the optimal MV. If the encoder and decoder create the candidate list in the same manner, the encoder only needs to transmit the index of the optimal MV in the candidate list. In AV1 and AV2, a technique called Dynamic motion vector prediction (DMVP) is used to predict MVs.

[0032] In order to obtain a better prediction effect, all the current mainstream video coding standards allow the use of multiple reference frames for inter prediction. The AV1 standard and the currently developed AOM next-generation standard AV2 include a prediction mode called compound prediction. In this compound prediction mode, it is allowed to perform inter prediction using two reference frames for a current block, and derive a predicted value of the current block by a weighted combination of inter prediction values, or to derive a predicted value of the current block by a weighted combination of an inter prediction value derived using one reference frame and an intra prediction value derived using a current frame. Here, the current block refers to a coding block being coded (or decoded). In the embodiment of the present application, both the inter prediction value and the intra prediction value are hereinafter referred to as a reference prediction value. The following is a formula for deriving a predicted value of a current block during compound prediction.

[0033] P(x,y)=(w(x,y)·P0(x,y) + (1-w(x,y))·P1(x,y)) / 2 where P(x,y) is the predicted value of the current block, P0(x,y) and P1(x,y) are the two reference predictors corresponding to the current block (x,y), and w(x,y) is the weight used for the first reference predictor P0(x,y).

[0034] Optionally, in video coding, integer calculation is usually used instead of floating-point calculation to reduce the complexity of weighted prediction. The following is a formula for deriving a predicted value of a current block by integer calculation:

[0035] P(x,y)=(w(x,y)×P0(x,y) + (64-w(x,y))×P1(x,y) + 32)>>6 Here, both the weight w(x,y) and the reference predictors P0(x,y), P1(x,y) are integer types, a right shift operation is used instead of division, ">>6" represents a 6-bit right shift, 64 can be divided by a 6-bit right shift, and 32 is an offset added for rounding.

[0036] According to the current video coding standard, a special weighting mode is used in composite prediction, that is, P0 and P1 have equal weighting values, and the weights corresponding to the reference predictive values ​​corresponding to different positions are all set to fixed values. The specific formula is as follows: P(x,y)=(32×P0(x,y) + 32×P1(x,y) + 32)>>6

[0037] Second, video decoding

[0038] On the decoding side, after obtaining a video bitstream for each CU, on the one hand, first perform entropy decoding of the video bitstream to obtain various predictive coding mode information and quantized transform coefficients, and then perform inverse quantization and inverse transform on each transform coefficient to obtain a residual video signal. On the other hand, a prediction signal (hereinafter referred to as a predicted value) corresponding to the CU can be obtained based on the known predictive coding mode information, and a reconstructed video signal is obtained by adding the residual video signal and the prediction signal. This reconstructed video signal can be used to reconstruct a decoded image corresponding to the CU. Finally, a loop filtering operation needs to be performed on the reconstructed video signal to generate a final output signal.

[0039] Based on the above related description, an embodiment of the present application provides a video processing method that can be applied to a video encoder or video compression product that uses composite prediction (or weighted prediction based on multiple reference frames). The general principle of the video processing method is as follows.

[0040] Encoding side: By using composite prediction for a CU included in a video frame, N reference prediction values ​​of the CU (N is an integer greater than 1) are obtained, and a prediction value of the CU can be obtained by adaptively selecting appropriate weights for the CU and performing weighted prediction based on the importance of the N reference prediction values. Weighted prediction refers to performing weighted prediction processing of the N reference prediction values ​​of the CU using adaptively selected weights. Next, a video bitstream is generated by performing video encoding processing based on the prediction value of the CU, and this video bitstream is transmitted to the decoding side.

[0041] Decoding side: When decoding a CU in a video bitstream, it is determined to perform composite prediction of the CU in the video bitstream based on information of the predictive coding mode, and based on the importance of N reference predictive values, an appropriate weight is adaptively selected for the CU to perform weighted prediction, and then, based on the adaptively selected weight, weighted prediction of the N reference predictive values ​​is performed to obtain a predicted value of the CU, and a decoded image corresponding to the CU can be reconstructed using the predicted value of the CU.

[0042] As described above, the current video coding standard specifies that in composite prediction, weighted prediction is performed using equal weight values ​​for reference prediction values ​​derived from different reference blocks. However, in practical applications, reference prediction values ​​derived from different reference blocks may have unequal importance, and using equal weight values ​​cannot embody the importance difference of the reference prediction values. In this case, the use of the conventional standard will affect prediction accuracy. The embodiment of the present application is an improvement of the current video coding standard, in which the importance of reference prediction values ​​derived from different reference blocks is fully taken into consideration in composite prediction, and in composite prediction, appropriate weights are adaptively selected for CUs based on the importance of each reference prediction value to perform weighted prediction, and the weighted prediction method in the video coding standard is extended, which can improve the prediction accuracy of CUs and improve coding performance.

[0043] Next, a video processing system provided in an embodiment of the present application will be described. Please refer to FIG. 2. FIG. 2 is a schematic diagram of the architecture of the video processing system provided in an embodiment of the present application. The video processing system 20 may include an encoding device 201 and a decoding device 202. The encoding device 201 is located on the encoding side, and the decoding device is located on the decoding side. The encoding device 201 may be a terminal or a server. The decoding device 202 may be a terminal or a server. A communication connection can be established between the encoding device 201 and the decoding device 202. Here, the terminal may be, but is not limited to, a smartphone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, an in-vehicle terminal, a smart TV, etc. The server may be an independent physical server, a server cluster or a distributed system consisting of multiple physical servers, or a cloud server that provides base cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0044] (1) About Encoding Device 201

[0045] The encoding device 201 can obtain a video to be encoded, which can be obtained by shooting with a shooting device or by generating with a computer. The shooting device can be a hardware component provided in the encoding device 201. For example, the shooting device can be a normal camera, a stereoscopic camera, a light field camera, etc. provided in a terminal. The shooting device can be a hardware device connected to the encoding device 201, such as a camera connected to a server.

[0046] Here, one video includes one or more video frames, and the encoding device 201 can divide each video frame into one or more CUs and encode each CU. When encoding any CU, composite prediction of the CU being encoded (hereinafter referred to as a current block) is performed to obtain N reference prediction values ​​of the current block, and factors such as the bit rate consumed in the weighted prediction process and the quality loss in encoding the current block are comprehensively considered to determine the importance of each reference prediction value, and further, an appropriate target weight group can be adaptively selected for the current block based on the importance of each reference prediction value. This target weight group may include one or more weight values. Next, the weight values ​​in the target weight group are used to perform weighted prediction processing of the N reference prediction values ​​to obtain a prediction value of the current block. Here, the prediction value of the current block can be understood as a prediction signal corresponding to the current block. The prediction value of the current block can be used to reconstruct a decoded image corresponding to the current block.

[0047] Here, the N reference prediction values ​​of the current block are derived from the N reference blocks of the current block. One reference prediction value corresponds to one reference block. The video frame in which the reference block is located is the reference frame, and the video frame in which the current block is located is the current frame. The positional relationship between the N reference blocks and the current block may include, but is not limited to, any one of the following (1) to (4). (1) The N reference blocks are located in the N reference frames, respectively, and the N reference frames and the current frame belong to different video frames in the video bitstream. (2) The N reference blocks are located in the same reference frame, and the same reference frame and the current frame belong to different video frames in the video bitstream. (3) One or more reference blocks of the N reference blocks are located in the current frame, and the remaining reference blocks of the N reference blocks are located in one or more reference frames, and the one or more reference frames and the current frame belong to different video frames in the video bitstream. (4) Both the N reference blocks and the current block are located in the current frame. As can be seen from this, the prediction modes of the composite prediction in the embodiments of the present application include inter prediction modes (i.e., at least two reference frames are allowed to be used for inter prediction), combined prediction modes (i.e., at least one reference frame is allowed to be used for inter prediction and the current frame is allowed to be used for intra prediction), and intra prediction modes (i.e., the current frame is allowed to be used for intra prediction).

[0048] Corresponding to different prediction modes of the composite prediction, the derivation method of the N reference prediction values ​​of the current block may include any one of the following (1) to (2). (1) The N reference prediction values ​​of the current block are derived by performing inter prediction using the N reference blocks of the current block. In this case, any of the N reference prediction values ​​may be called an inter prediction value. (2) At least one of the N reference prediction values ​​of the current block is derived by performing inter prediction using at least one of the N reference blocks of the current block. This part of the reference prediction value may be called an inter prediction value. The remaining reference prediction values ​​are derived by performing intra prediction using the remaining reference blocks of the N reference blocks. This part of the reference prediction values ​​may be called an intra prediction value.

[0049] Next, the encoding device 201 obtains a video bitstream by performing operations such as transform coding, quantization, and entropy coding on the video based on the predicted values ​​of the CUs included in the video frame, and transmits this video bitstream to the decoding device 202 so that the decoding device 202 performs a decoding process on the video bitstream.

[0050] (2) About the Decryption Device 202

[0051] After receiving the video bitstream transmitted from the encoding device 201, the decoding device 202 can perform a decoding process of the video bitstream to reconstruct a video corresponding to the video bitstream. Specifically, on the one hand, the decoding device 202 may obtain a prediction mode and a quantized transform coefficient of each CU in the video bitstream by performing entropy coding of the video bitstream, obtain N reference prediction values ​​of the current block by performing composite prediction of the current block (i.e., the CU being decoded) according to the prediction mode of the current block, and determine whether the use of adaptive weighted prediction is allowed for the current block.

[0052] If it is determined that the adaptive weighted prediction is allowed to be used for the current block, a target weight list may be determined from one or more weight lists according to the importance of the N reference predictors, and a target weight group for weighted prediction may be determined from the target weight list for the current block. The target weight group includes one or more weight values. Then, the weight values ​​in the target weight group are used to directly perform weighted prediction processing of the N reference predictors to obtain a predicted value for the current block. If it is determined that the adaptive weighted prediction is not allowed to be used for the current block, the weighted prediction processing of the N reference predictors may be performed according to a conventional video coding standard. For example, the weighted prediction processing of each reference predictor is performed using an equal weight value to obtain a predicted value for the current block.

[0053] On the other hand, the decoding device 202 obtains residual signal values ​​of the current block by performing inverse quantization and inverse transformation on the quantized transform coefficients, obtains reconstructed values ​​of the current block by performing overlapping based on the predicted values ​​of the current block and the residual signal values, and then reconstructs a decoded image corresponding to the current block according to the reconstructed values, where the decoded image may be used as a reference image for decoding other CUs, or may be used for video reconstruction.

[0054] In an embodiment of the present application, composite prediction is used for both video encoding and decoding, and in the composite prediction, the importance of reference prediction values ​​derived from different reference blocks is fully taken into consideration. In the composite prediction, weighted prediction is allowed to be performed by adaptively selecting an appropriate weight for a CU based on the importance of each reference prediction value, thereby extending the weighted prediction method in the video encoding standard, improving the prediction accuracy of the CU, and improving the encoding and decoding performance.

[0055] Next, a video processing method provided in an embodiment of the present application will be described. Please refer to FIG. 3. FIG. 3 is a schematic diagram of a flow of the video processing method provided in an embodiment of the present application. This video processing method may be executed by a decoding device in the above-mentioned video processing system. The video processing method in this embodiment may include the following steps S301 to S303.

[0056] In S301, N (N is an integer greater than 1) reference prediction values ​​of the current block are obtained by performing composite prediction of the current block in the video bitstream, where the current block refers to a coding block being decoded in the video bitstream, the N reference prediction values ​​are derived from N reference blocks of the current block, the N reference blocks are coding blocks in the video bitstream that are referenced when decoding the current block, and there is a one-to-one correspondence between the reference prediction values ​​and the reference blocks.

[0057] The video bitstream may include one or more video frames, and each video frame may include one or more coding blocks. When decoding the video bitstream, the decoding device may obtain coding blocks from the video bitstream, and the coding block currently being decoded may be the current block. The above N reference prediction values ​​may be derived from N reference blocks, and one reference prediction value corresponds to one reference block, that is, the N reference prediction values ​​are obtained based on the N reference blocks, and the reference prediction values ​​and the reference blocks correspond one-to-one, and the N reference blocks are coding blocks in the video bitstream that are referenced when the current block is decoded. In the embodiment of the present application, the video frame in which the reference block is located may be a reference frame, and the video frame in which the current block is located is the current frame.

[0058] Here, the positional relationship between the N reference blocks and the current block includes any one of the following (1) to (4). (1) The N reference blocks may be located in N reference frames, respectively, and the N reference frames and the current frame may belong to different video frames in the video bitstream. For example, when N=2, one of the two reference blocks is located in reference frame 1, and the other reference block is located in reference frame 2, and reference frame 1, reference frame 2, and the current frame belong to different video frames in the video bitstream. (2) The N reference blocks may be located in the same reference frame, and the same reference frame and the current frame may belong to different video frames in the video bitstream. For example, when N=2, both of the two reference blocks are located in reference frame 1, and reference frame 1 and the current frame belong to different video frames in the video bitstream. (3) One or more of the N reference blocks are located in the current frame, the remaining of the N reference blocks are located in one or more reference frames, and the one or more reference frames and the current frame belong to different video frames in the video bitstream. For example, when N=4, the four reference blocks are reference block 1, reference block 2, reference block 3, and reference block 4, respectively, and reference block 1 is located in the current frame, and the remaining reference block 2, reference block 3, and reference block 4 are all located in reference frame 1, and the reference frame 1 and the current frame belong to different video frames in the video bitstream. Also, for example, reference block 1 and reference block 2 may be located in the current frame, the remaining reference block 3 is located in reference frame 1, and reference block 4 is located in reference frame 2, and the reference frame 1, reference frame 2, and the current frame belong to different video frames in the video bitstream. (4) All of the N reference blocks and the current block are located in the current frame, and for example, when N=2, all of the two reference blocks are located in the current frame.

[0059] As can be seen from the positional relationships between the N reference blocks and the current block shown in (1) to (4) above, the prediction modes of the composite prediction in the embodiments of the present application include inter prediction mode (i.e., at least two reference frames are allowed to be used for inter prediction), combined prediction mode (i.e., at least one reference frame is allowed to be used for inter prediction and the current frame is allowed to be used for intra prediction), and intra prediction mode (i.e., the current frame is allowed to be used for intra prediction).

[0060] Corresponding to different prediction modes of composite prediction, the derivation method of the N reference prediction values ​​of the current block may include any one of the following (1) to (2). (1) The N reference prediction values ​​of the current block are derived by performing inter prediction using the N reference blocks of the current block. In this case, any of the N reference prediction values ​​may be called an inter prediction value. (2) At least one of the N reference prediction values ​​of the current block is derived by performing inter prediction using at least one of the N reference blocks of the current block. This part of the reference prediction value may be called an inter prediction value. The remaining reference prediction value is derived by performing intra prediction using the remaining reference blocks of the N reference blocks. This part of the reference prediction value may be called an intra prediction value. In this embodiment, the N reference blocks may be from different video frames, specifically, from the current frame in which the current block is located or from different reference frames, thereby being applicable to different composite prediction scenarios and ensuring the performance of encoding and decoding of composite prediction in different scenarios.

[0061] In one embodiment, before performing step S302, the decoding device may first determine whether the current block satisfies the condition for adaptive weighted prediction, i.e., determine the condition for adaptive weighted prediction, and perform step S302 if the current block satisfies the condition for adaptive weighted prediction. By determining whether the current block satisfies the condition for adaptive weighted prediction, the decoding device can adaptively select weights to perform weighted prediction, which can further improve prediction accuracy and improve coding performance.

[0062] Here, when the current block satisfies the conditions for adaptive weighted prediction, at least one of the following a) to j) is included.

[0063] a) a sequence header of a frame sequence to which the current block belongs includes a first indication field, and the first indication field indicates that adaptive weighted prediction is allowed to be used for any of the coding blocks in the frame sequence, where the frame sequence refers to a sequence of multiple video frames in a video bitstream. It should be understood that when the sequence header of the frame sequence includes the first indication field, the first indication field is for indicating that adaptive weighted prediction is allowed to be used for all coding blocks included in the entire frame sequence.

[0064] In one implementation, the first indication field may be represented as seq_acp_flag. The first indication field may indicate whether adaptive weighting prediction is permitted for the coding blocks in the frame sequence according to a value. If the first indication field is a first predetermined value (e.g., 1), it indicates that adaptive weighting prediction is permitted for all coding blocks in the frame sequence, and it may be determined that the current block satisfies the adaptive weighting prediction condition. If the first indication field is a second predetermined value (e.g., 0), it indicates that adaptive weighting prediction is not permitted for all coding blocks in the frame sequence, and it may be determined that the current block does not satisfy the adaptive weighting prediction condition.

[0065] b) a slice header of a current slice to which the current block belongs includes a second indication field, and the second indication field indicates that adaptive weighted prediction is allowed to be used for the coding blocks in the current slice. Here, one video frame can be divided into multiple image slices, and each image slice includes one or more coding blocks. The current slice refers to the image slice to which the current block belongs, i.e., the image slice being decoded. It should be understood that if the slice header of the current slice includes the second indication field, the second indication field can be used to indicate that adaptive weighted prediction is allowed to be used for all coding blocks included in the current slice.

[0066] In one implementation, the second indication field may be represented as slice_acp_flag. The second indication field may indicate whether adaptive weighting prediction is permitted for the coding blocks in the current slice depending on the value. If the second indication field is a first predetermined value (e.g., 1), it indicates that adaptive weighting prediction is permitted for all coding blocks in the current slice, and it may be determined that the current block satisfies the adaptive weighting prediction condition. If the second indication field is a second predetermined value (e.g., 0), it indicates that adaptive weighting prediction is not permitted for all coding blocks in the current slice, and it may be determined that the current block does not satisfy the adaptive weighting prediction condition.

[0067] c) a third indication field is included in the frame header of the current frame in which the current block is located, and the third indication field indicates that adaptive weighted prediction is allowed to be used for the coding blocks in the current frame. It should be understood that, when the frame header of the current frame includes the third indication field, the third indication field can be used to indicate that adaptive weighted prediction is allowed to be used for all coding blocks included in the current frame.

[0068] In one implementation, the third indication field may be represented as pic_acp_flag. The third indication field indicates whether adaptive weighting prediction is permitted for the current frame depending on its value. If the third indication field is a first predetermined value (e.g., 1), it indicates that adaptive weighting prediction is permitted for the coding block in the current frame, and it may be determined that the current block satisfies the adaptive weighting prediction condition. If the third indication field is a second predetermined value (e.g., 0), it indicates that adaptive weighting prediction is not permitted for the coding block in the current frame, and it may be determined that the current block does not satisfy the adaptive weighting prediction condition.

[0069] d) During composite prediction, for a current block, at least two reference frames are used for inter prediction.

[0070] e) During composite prediction, for a current block, at least one reference frame is used for inter prediction, and the current frame is used for intra prediction. For example, if for a current block, reference frame 1 is used for inter prediction, and the current frame is used for intra prediction, it can be determined that the current block satisfies the condition of adaptive weighted prediction.

[0071] f) If the motion type of the current block is a specified motion type, for example, the motion type of the current block is simple_translation (simple translation), then determine that the current block satisfies the conditions for adaptive weighted prediction.

[0072] g) If a predetermined motion vector prediction mode is used for the current block, for example, the predetermined motion vector prediction mode used for the current block is NEAR_NEARMV, it can be determined that the current block satisfies the condition for adaptive weighted prediction.

[0073] In the AV1 and AV2 standards, a technique called dynamic motion vector prediction (DMVP) is used to predict MVs, and MVs can be predicted by spatially adjacent blocks in a current frame or by time-domain adjacent blocks in a reference frame. In the case of single reference inter prediction, each reference frame has its own prediction MV list. In the case of combined inter prediction, a prediction MV group list is constructed from prediction mv lists corresponding to different reference frames, and multiple types of prediction MV modes, such as NEAR_NEARMV, NEAR_NEWMV, NEW_NEAR_MV, NEW_NEWMV, GLOBAL_GLOBALMV, JOINT_NEWMV, etc., are allowed to be used. Here,

[0074] NEAR_NEARMV: indicates that the MVs corresponding to the two reference frames are MVs in the predicted MV group.

[0075] NEAR_NEWMV: The first MV is the first predicted MV in the predicted MV group, and the second MV is derived from the decoded MVD in the video bitstream and the second predicted MV in the predicted MV group. Here, the Motion Vector Difference (MVD) refers to the difference between the current MV and the predicted MV (MVP).

[0076] NEW_NEARMV: The first MV is derived from the decoded MVD in the video bitstream and the first predicted MV in the predicted MV group, and the second MV is the second predicted MV in the predicted MV group.

[0077] NEW_NEWMV: The first MV is derived from the decoded MVD1 in the video bitstream and the first predicted MV in the predicted MV group, and the second MV is derived from the decoded MVD2 in the video bitstream and the second predicted MV in the predicted MV group.

[0078] GLOBAL_GLOBALMV: Derive MV from global motion information at each frame level.

[0079] NEW_NEWMV: Similar to NEW_NEWMV, but the video bitstream contains only one MVD, and the other MVDs are derived from information between reference frames.

[0080] The predetermined motion vector prediction mode in the embodiment of the present application may be one or more of the above NEAR_NEARMV, NEAR_NEWMV, NEW_NEARMV, NEW_NEWMV, GLOBAL_GLOBALMV, and JOINT_MV. Note that the predetermined motion vector prediction mode is not limited to the above MV prediction mode, and in the case of other standards such as H.265 and H.266, the motion vector prediction mode may be determined by combining merge and AMVP.

[0081] h) A predetermined interpolation filter is used for the current block. Video coding usually includes many coding tools, and the coding tools include multiple types of interpolation filters, such as a linear interpolation filter, a cascaded integrator-comb (cic) interpolation filter, and the like. The predetermined interpolation filter in the embodiment of the present application may be any one of the multiple types of interpolation filters. For example, the predetermined interpolation filter may be a linear interpolation filter. If a linear interpolation filter is used for the current block, it may be determined that the current block satisfies the condition of adaptive weighted prediction.

[0082] i) No specific coding tool is used for the current block. For example, no motion vector optimization method based on optical flow is used for the current block. In this case, it can be determined that the current block satisfies the condition of adaptive weighted prediction. Here, video coding standards such as AV2 and H.266 allow the use of a motion vector optimization method based on optical flow, which is a method of subdividing a motion vector based on the derivation of an optical flow equation.

[0083] j) If a reference frame used for a current block during composite prediction satisfies a certain condition, it may be determined that the current block satisfies the condition of adaptive weighted prediction. Here, the certain condition includes one or more of the following (i.e., the certain condition may include one or more of (1) and (2)). (1) An orientation relationship in a video bitstream between a reference frame used for composite prediction and a current frame satisfies a certain relationship. Here, the orientation relationship satisfies the certain relationship includes any one of the following: all of the reference frames used are located before the current frame, all of the reference frames used are located after the current frame, and some of the reference frames used are located before the current frame and the remaining reference frames are located after the current frame. A video bitstream includes multiple video frames, and any video frame corresponds to a frame display time in a video, and the orientation relationship can actually be understood as a front-back order of frame display times. For example, the orientation relationship can be understood as the frame display time of the reference frame being earlier than the frame display time of the current frame when all of the reference frames used are located before the current frame. Any reference frame used is located after the current frame, which can be understood to mean that the frame display time of the reference frame is later than the frame display time of the current frame.

[0084] (2) If the absolute value of the importance difference between the reference predictors corresponding to the reference frames used in the composite prediction is equal to or greater than a predetermined threshold, it may be determined that the current block satisfies the condition of adaptive weighted prediction. Here, the predetermined threshold may be set as necessary. In one implementation, the importance of the reference predictor corresponding to the reference frame may be measured by an importance metric value. In this case, the importance difference between the reference predictors corresponding to the reference frames used in the composite prediction may be determined based on the importance metric value between the reference predictors corresponding to the reference frames used. For example, when N=2, in the composite prediction, the two reference predictors are reference predictor 1 corresponding to reference frame 1 and reference predictor 2 corresponding to reference frame 2, respectively. If the importance metric value of reference prediction value 1 is D0 and the importance metric value of reference prediction value 2 is D1, the importance difference between the reference prediction values ​​corresponding to the two reference frames is the difference D0-D1 between the importance metric value D0 of reference prediction value 1 and the importance metric value D1 of reference prediction value 2. In other words, the absolute value of the importance difference between the two reference prediction values ​​is ΔD=abs(D0-D1), where abs() represents the calculation of the absolute value.

[0085] It should be understood that the above-mentioned conditions of adaptive weighted prediction satisfied by the current block may be used alone or in combination. For example, the condition of adaptive weighted prediction satisfied by the current block may include the case where the sequence header of the frame sequence to which the current block belongs includes a first indication field, and the first indication field indicates that adaptive weighted prediction is permitted to be used for the coding block in the frame sequence, and the motion type of the current block is a specified motion type. Also, for example, the frame header of the current frame to which the current block is located includes a third indication field, and the third indication field indicates that adaptive weighted prediction is permitted to be used for the coding block in the current frame, and a predetermined motion vector prediction mode is used for the current block. The present application is not limited thereto. This embodiment supports the decoding device to flexibly configure each condition item of the conditions of adaptive weighted prediction, and can be applied to different composite prediction scenarios, and ensure the performance of the coding and decoding of composite prediction in different scenarios.

[0086] In S302, a target weight group for weighted prediction is determined for the current block based on the importance of the N reference predictors, where the importance represents the degree of influence of each reference predictor on the decoding performance of the current block.

[0087] The importance of the reference predictor may be determined comprehensively based on factors such as the bit rate consumed in the weighted prediction process and the quality loss in the coding of the current block. For example, if the bit rate consumption increases obviously when a certain reference predictor is used to perform weighted prediction on the current block, this indicates that the reference predictor does not contribute significantly to reducing the bit rate consumption, and its importance is low. Also, if the quality loss increases obviously when a certain reference predictor is used to perform weighted prediction on the current block, this indicates that the reference predictor contributes little to reducing the quality loss, and its importance is low. Based on the importance of the reference predictor, the degree of influence of the reference predictor on the decoding performance of the current block can be determined. The target weight group includes one or more weight values, and these weight values ​​act on each reference predictor in the weighted prediction. If the importance of a certain reference predictor is low, this reference predictor corresponds to a weight value with a small value in the target weight group, and if the importance of a certain reference predictor is high, this reference predictor corresponds to a weight value with a large value in the target weight group. In other words, the target weight group is a group of weights that is selected after considering the influence of each reference predictor on factors such as bitrate consumption and quality loss, so that the total cost of weighted prediction (i.e., bitrate consumption cost, quality loss cost, bitrate consumption and quality loss cost) is small and the encoding / decoding performance is excellent.

[0088] In one embodiment, a video bitstream includes one or more weight lists, each weight list includes one or more weight groups, each weight group includes one or more weight values, the number of weight values ​​in each weight group may be the same or different, the values ​​of the weight values ​​in each weight group may be the same or different, and the order of the weight values ​​in each weight group may be the same or different. The following is an example of a weight list:

[0089] i, Weight List 1 is represented as {2,4,6,8,10,12,14} / 16, which means that Weight List 1 contains seven weight groups, which are Weight Group 1: {2} / 16, Weight Group 2: {4} / 16, Weight Group 3: {6} / 16, …, Weight Group 7: {14} / 16, respectively.

[0090] ii. Weight list 2 is represented as {14,8,4,12,2} / 16, which means that weight list 2 contains five weight groups: weight group 1: {14} / 16, weight group 2: {8} / 16, …, weight group 5: {2} / 16.

[0091] iii. Weight list 3 is represented as {4,{8,8},12} / 16, which means that weight list 3 contains three weight groups: weight group 1: {4} / 16, weight group 2: {8,8} / 16, and weight group 3: {12} / 16.

[0092] As can be seen from the above example, (1) the number of weight groups included in each weight list is permitted to be different; for example, the number of weight groups in weight list 1 is 7, the number of weight groups in weight list 2 is 5, and the number of weight groups in weight list 1 is 3. (2) the number of weight values ​​included in each weight group is permitted to be the same; for example, the number of weight values ​​included in each weight group in weight list 1 is 1. (3) The number of weight values ​​included in each weight group is also permitted to be different; for example, the number of weight values ​​included in weight group 1 in weight list 3 is 1, but the number of weight values ​​included in weight group 2 in weight list 3 is 2. (4) The weight values ​​in each weight group are permitted to be the same; for example, weight group 1 in weight list 1 includes a weight value of 2 / 16, and weight group 1 in weight list 2 also includes a weight value of 2 / 16. (5) The weight values ​​in each weight group are permitted to be different; for example, weight group 1 in weight list 1 includes a weight value of 2 / 16, but weight group 2 in weight list 1 includes a weight value of 4 / 16. As can be understood, the sum of each weight value provided in one weight group should be equal to 1. In the above example, each weight group includes only one weight value, but each weight value is less than 1, so each weight group implicitly includes other weight values. That is, in a specific application, two weight values ​​are actually provided for each weight group, and the sum of the two weight values ​​is 1. For example, weight group 1 in weight list 1 includes only weight value 2 / 16, but in a specific application, this weight group 1 provides a total of two weight values, weight value 2 / 16 and weight value 14 / 16. Also, for example, weight group 3 in weight list 3 includes only weight value 12 / 16, but in a specific application, this weight group 3 provides a total of two weight values, weight value 12 / 16 and weight value 4 / 16. As can be seen from this, in the embodiment of the present application, when the sum of each weight value included in a weight group in a weight list is less than 1, the implicit weight value of the weight group can be obtained by calculation.

[0093] Below is another example of a weight list:

[0094] i,Weight List 4 contains four weight groups, which are weight group 1: {2,14} / 16, weight group 2: {4,12} / 16, weight group 3: {6,10} / 16, and weight group 4: {8,8} / 16, respectively.

[0095] ii. Weight list 5 contains two weight groups, weight group 1: {4, 12} / 16 and weight group 2: {10, 6} / 16, respectively.

[0096] As can be seen from the above example, (1) the sum of all weight values ​​contained in a weight group in a weight list is equal to 1. (2) The order of weight values ​​contained in each weight group is allowed to be different; for example, weight group 3 in weight list 4 and weight group 2 in weight list 5 contain the same weight value, but the order of each weight value is different.

[0097] In one embodiment, step S302 may include steps s31 to s32.

[0098] In s31, a target weight list is determined from one or more weight lists based on the importance of the N reference predictors.

[0099] Here, when the decoding device determines a target weight list from one or more weight lists according to the importance of N reference predictors, the number of weight lists and the importance metric value of each reference predictor may be taken into account to realize the above. Specific implementation forms may include the following methods:

[0100] Method 1: When the number of weight lists in a video bitstream is only one, the decoding device may directly determine this only weight list in the video bitstream as the target weight list.

[0101] Method 2: When there are multiple weight lists in a video bitstream, the decoding device may introduce importance metric values ​​of reference predictors, and determine a target weight list from among the multiple weight lists based on the importance metric values ​​of the N reference predictors.

[0102] (1) Determine a target weight list based on the absolute value of the difference between the importance metric values ​​of the N reference predictors.

[0103] When the number of weight lists in a video bitstream is M+1 (M is a positive integer equal to or greater than 1), the M+1 groups of weight lists can be written as {w_list1, w_list2...w_listM+1}, where one weight list corresponds to one threshold interval, that is, the number of threshold intervals is also M+1. In one implementation, the decoding device may directly set M+1 threshold intervals as needed. In another implementation, the decoding device may obtain M thresholds and perform division of M+1 threshold intervals based on the M thresholds. For example, obtain M thresholds, which are T1, T2,..., TM, respectively, and then the decoding device performs division of M+1 threshold intervals, which are [0, T1], (T1, T2], (T2, T3],..., (TM, +∞), respectively, based on the M thresholds. Here, [0, T1] corresponds to w_list1, and (T1, T2] corresponds to w_list2, and so on may be inferred. T is an integer equal to or greater than 0.

[0104] The decoding device may obtain the importance metric values ​​of the N reference prediction values ​​and calculate the importance difference between the N reference prediction values. The importance difference between any two reference prediction values ​​is measured by the difference between the importance metric values ​​of the two reference prediction values. Next, the decoding device may determine a threshold interval in which the absolute value of the importance difference between the N reference prediction values ​​is located, and determine a weight list corresponding to the threshold interval in which the absolute value of the importance difference between the N reference prediction values ​​is located as a target weight list. Here, assuming that the importance metric values ​​of any two reference prediction values ​​are D0 and D1, respectively, the importance difference between these two reference prediction values ​​is represented as D0-D1, and the absolute value of the importance difference between these two reference prediction values ​​is represented as ΔD=abs(D0-D1). Note that the importance metric value is an index for measuring the degree of importance, but there may be multiple cases for the criterion for measuring. For example, the criterion for measuring may be that the greater the importance metric value, the higher the degree of importance. Also, for example, the measurement criterion may be that the smaller the importance metric value, the higher the degree of importance. The present application is not limited to this measurement criterion.

[0105] In one embodiment, when N=2, that is, only two reference predictors are included, and the importance metric values ​​of the two reference predictors are D0 and D1, respectively, the decoding device directly determines the weight list corresponding to the threshold interval in which ΔD=abs(D0-D1) is located as the target weight list. For example, when the number of weight lists in the video bitstream is 2, the two weight lists are denoted as {w_list1, w_list2}, respectively, where w_list1 is {8,12,14} / 16, and w_list2 is {12,8,4} / 16. If the threshold interval corresponding to w_list1 is [0,1] and the threshold interval corresponding to w_list2 is (1,+∞), then when ΔD=abs(D0-D1) is less than or equal to 1, the threshold interval in which it is located is [0,1], so w_list1 is determined as the target weight list; conversely, when ΔD=abs(D0-D1) is greater than 1, the threshold interval in which it is located is (1,+∞), so w_list2 is determined as the target weight list.

[0106] In another embodiment, when N>2, the decoding device may calculate the absolute value of the importance difference between any two reference predictors, respectively, find the weight list corresponding to the threshold interval in which each absolute value is located, and determine the weight list with the largest number of corresponding values ​​as the target weight list. For example, when N=3 and the importance metric values ​​of the three reference predictors are D0, D1, and D2, respectively, the decoding device may calculate the absolute value of the importance difference between any two reference predictors, i.e., ΔD=abs(D0-D1), ΔD'=abs(D1-D2), and ΔD''=abs(D0-D1), respectively, determine the weight list corresponding to the threshold interval in which ΔD, ΔD', and ΔD'' are located, respectively, and if two or more of the three correspond to the same weight list, determine the same weight list as the target weight list.

[0107] In another embodiment, when N>2, the decoding device may calculate the absolute value of the importance difference between any two reference prediction values, find the maximum value among the absolute values, and determine the weight list corresponding to the threshold interval in which the maximum value is located as the target weight list. For example, in the above example of N=3, after calculating ΔD, ΔD', and ΔD'', the decoding device finds the maximum value among ΔD, ΔD', and ΔD'', and if the maximum value is ΔD', determine the weight list corresponding to the threshold interval in which ΔD' is located as the target weight list.

[0108] In another embodiment, when N>2, the decoding device may calculate the absolute value of the importance difference between any two reference prediction values, find the minimum value among the absolute values, and determine the weight list corresponding to the threshold interval in which the minimum value is located as the target weight list. For example, in the above example of N=3, after calculating ΔD, ΔD', and ΔD'', the decoding device finds the minimum value among ΔD, ΔD', and ΔD'', and if the minimum value is ΔD, determine the weight list corresponding to the threshold interval in which ΔD is located as the target weight list.

[0109] In another embodiment, when N>2, the decoding device may calculate the absolute value of the importance difference between any two reference predicted values, calculate the average value of each absolute value, and determine the weight list corresponding to the threshold interval in which the average value is located as the target weight list. For example, in the above example of N=3, after calculating ΔD, ΔD', and ΔD'', the decoding device calculates the average value of ΔD, ΔD', and ΔD''=(ΔD+ΔD'+ΔD'') / 3, and determine the weight list corresponding to the threshold interval in which the average value is located as the target weight list.

[0110] In addition, in the embodiment of the present application, the target weight list may be determined using other numerical characteristics of the absolute values ​​of the importance differences between the N reference predictors, for example, the maximum, minimum, or average value of the squared absolute values. The present application is not limited thereto. In the present embodiment, the decoding device determines the target weight list according to a threshold interval in which the differences between the importance metric values ​​of the N reference predictors are located. In this way, it is possible to determine weight values ​​suitable for reconstruction of the current block, which is advantageous to improving the prediction accuracy of the current block and to improve the performance of encoding and decoding.

[0111] (2) A target weight list is determined by comparing the magnitudes of the importance metric values ​​of the reference predictors.

[0112] The N reference predictor values ​​of the current block may include a first reference predictor value and a second reference predictor value, and the video bitstream may include a first weight list and a second weight list. The decoding device may compare an importance metric value of the first reference predictor value with an importance metric value of the second reference predictor value, and may determine the first weight list as a target weight list if it determines that the importance metric value of the first reference predictor value is greater than the importance metric value of the second reference predictor value, and may determine the second weight list as a target weight list if it determines that the importance metric value of the first reference predictor value is equal to or less than the importance metric value of the second reference predictor value.

[0113] For example, the importance metric value of the first reference prediction value is D0, the importance metric value of the second reference prediction value is D1, the first weight list is w_list1, and the second weight list is w_list2. The decoding device may compare D0 and D1, and if D0>D1, determine the first weight list w_list1 as the target weight list, and if D0≦D1, determine the first weight list w_list2 as the target weight list.

[0114] Optionally, the weight values ​​in the first weight list and the weight values ​​in the second weight list are reversed, i.e., w_list2[x]=1-w_list1[x], where x represents a weight value in a weight list. For example, w_list1={0.2,0.4}, w_list2[x]=1-w_list1[x], i.e., w_list2[x]={0.8,0.6}. The sum of weight values ​​in the same order in the first weight list and the second weight list is 1. Also, the weight values ​​in the first weight list and the weight values ​​in the second weight list may be set separately.

[0115] In this embodiment, the decoding device determines a corresponding weight list as a target weight list for the current block based on the magnitude relationship between the two reference predicted values, thereby improving the targeting of the target weight list, which is beneficial to improving the prediction accuracy of the current block and improving the encoding / decoding performance.

[0116] (3) A mathematical sign function, the importance metric value of the reference predictor, is utilized to determine the target weight list.

[0117] The N reference predictors of the current block include a first reference predictor and a second reference predictor, and the video bitstream includes a first weight list, a second weight list, and a third weight list. The decoding device may obtain a code value by calling a mathematical code function to process a difference between an importance metric value of the first reference predictor and an importance metric value of the second reference predictor. If the code value is a first predetermined value (e.g., -1), the decoding device determines the first weight list as a target weight list; if the code value is a second predetermined value (e.g., 0), the decoding device determines the second weight list as a target weight list; if the code value is a third predetermined value (e.g., 1), the decoding device determines the third weight list as a target weight list. Here, the first weight list, the second weight list, and the third weight list are different weight lists, or two of the first weight list, the second weight list, and the third weight list are allowed to be the same weight list.

[0118] Here, the importance metric value of the first reference prediction value is D0, the importance metric value of the second reference prediction value is D1, and a data sign function is called to process the difference between D0 and D1 to obtain a sign value, i.e., sign value=sign(D0-D1), where sign() represents the data sign function. The decoding device determines a corresponding weight list from the three weight lists as a target weight list for the current block based on the magnitude relationship between the two reference prediction values, thereby improving the targeting of the target weight list, which is favorable to improving the prediction accuracy of the current block and improving the performance of encoding and decoding.

[0119] (4) The above methods (1), (2), and (3) may be used alone, or the target weight list may be determined by combining the methods (1), (2), and (3). As an implementation form, assuming that there are M+1 weight lists, for any two reference prediction values, the decoding device first determines a candidate weight list according to a threshold interval in which the absolute value of the importance difference between the two reference prediction values ​​is located in the method (1), and then compares the importance metric value of the first reference prediction value with the importance metric value of the second reference prediction value in the method (2). If it is determined that the importance metric value of the first reference prediction value is greater than the importance metric value of the second reference prediction value, it directly determines the candidate weight list as the target weight list, and if it is determined that the importance metric value of the first reference prediction value is equal to or less than the importance metric value of the second reference prediction value, it determines a weight list corresponding to a weight value opposite to that in the candidate weight list as the target weight list.

[0120] For example, the importance metric value of the first reference predictor is D0, the importance metric threshold of the second reference predictor is D1, and the candidate weight list w_list1 can be determined by the manner (1). If D0>D1, the decoding device determines the weight list w_list1 as the target weight list, and if D0≦D1, the decoding device determines the weight list w_list2 as the target weight list, where the weight values ​​in w_list2 and w_list1 are reversed, i.e. w_list2[x]=1−w_list1[x].

[0121] In another implementation, assume that there are 3*(M+1) weight lists in the video bitstream, that is, one threshold interval can correspond to three weight lists. For any two reference predictors, the decoding device may first determine a threshold interval in which the absolute value of the importance difference between the first reference predictor and the second reference predictor is located in manner (1), and determine any of the three weight lists corresponding to this threshold interval as a candidate weight list. That is, the candidate weight list may include a first weight list, a second weight list, and a third weight list. Then, the decoding device calls a mathematical sign function in manner (2) to obtain a sign value by processing the difference between the importance metric value of the first reference predictor and the importance metric value of the second reference predictor. When the code value is a first predetermined value, the decoding device determines the first weight list as the target weight list, when the code value is a second predetermined value, the decoding device determines the second weight list as the target weight list, and when the code value is a third predetermined value, the decoding device determines the third weight list as the target weight list.

[0122] For example, in method (1), three weight lists {w_list1, w_list2, w_list3} are determined as candidate weight lists, and then a mathematical sign function is called to obtain a code value by processing the difference between the importance metric value of the first reference predictor value and the importance metric value of the second reference predictor value. If the code value is -1, the decoding device determines w_list1 as the target weight list, if the code value is 0, the decoding device determines w_list2 as the target weight list, and if the code value is 1, the decoding device determines w_list3 as the target weight list.

[0123] Any one of the N reference predictors is denoted as reference predictor i (i is an integer less than or equal to N), where reference predictor i is derived from reference block i, the video frame in which reference block i is located is reference frame i, and the video frame in which the current block is located is the current frame, and the importance metric value of reference predictor i may be determined by any one of the following methods:

[0124] Method 1: Calculate based on the picture order count (POC) of the current frame in the video bitstream and the picture order count of the reference frame i in the video bitstream. Specifically, the decoding device may calculate the difference between the picture order count of the current frame in the video bitstream and the picture order count of the reference frame i in the video bitstream, and the absolute value of this difference may be the importance metric value of the reference prediction value i. For example, if the picture order count of the current frame in the video bitstream is represented as cur_poc and the picture order count of the reference frame i in the video bitstream is represented as ref_poc, the importance metric value D of the reference prediction value i is D=abs(cur_poc-ref_poc), where abs() represents obtaining the absolute value.

[0125] Method 2: Calculate based on the picture order count of the current frame in the video bitstream, the picture order count of the reference frame i in the video bitstream, and the quality metric Q. Here, the quality metric Q may be determined in various cases. This application is not limited thereto. Specifically, the quality metric Q of the reference frame i may be derived from the quantization information of the current block. For example, the quality metric Q may be set as the base quantization index (base_qindex) of the reference frame i. The base_qindex of any reference frame may be different or the same. In other realizations, the quality metric Q of the reference frame i may be derived from other coding information. For example, the quality metric Q of the reference frame i may be derived from the coding information difference between the coded CU in the reference frame i and the coded CU in the current frame.

[0126] In one implementation, the decoding device may calculate the difference between the picture order count of the current frame in the video bitstream and the picture order count of the reference frame i in the video bitstream, and then use an objective function to determine the importance metric value of the reference predictor i based on the difference, the quality metric Q, and the importance metric value list.

[0127] where the objective function is D=f(cur_poc-ref_poc)+Q, where D represents the importance metric value of reference predictor i, f(x) is an increasing function, where x=cur_poc-ref_poc, where cur_poc represents the picture order count of the current frame in the video bitstream, and ref_poc represents the picture order count of reference frame i in the video bitstream. As shown in Table 1, the above importance metric value list includes the correspondence between f(x) and the reference importance metric values.

[0128] [Table 1] It should be noted that the above function expression of f(x) is merely an example, and in the embodiments of the present application, changes to the function expression of f(x) are permitted; that is, the embodiments of the present application do not limit the specific expression form of f(x).

[0129] In one embodiment, the importance metric value of the reference frame i may be calculated based on an orientation relationship between the reference frame i and the current frame and a quality metric Q. In one implementation, the decoding device may establish a correspondence relationship between the reference orientation relationship and the reference importance metric value. For example, if the reference orientation relationship is that the reference frame is located before the current frame, the reference importance metric value may correspond to a first value, and if the reference orientation relationship is that the reference frame is located after the current frame, the reference importance metric value may correspond to a second value. Then, the decoding device may calculate an importance metric value of the reference prediction value i based on the reference importance metric value corresponding to the reference frame i and the quality metric Q.

[0130] Method 3: Based on the calculation results of Method 1 and Method 2, calculate the importance metric score of reference frame i, and order the importance metric scores of the reference frames corresponding to the N reference prediction values ​​in ascending order, and determine the index of reference frame i in the ordering as the importance metric value of reference prediction value i.

[0131] In one implementation, a first importance metric value of a reference predictor value i can be calculated by the above method 1, and a second importance metric value of the reference predictor value i can be calculated by the above method 2, and then an importance metric score (e.g., score) of the reference frame i is calculated based on the first importance metric value and the second importance metric value, and then the importance metric scores of the reference frames corresponding to the N reference predictors are ordered in ascending order, and the index of the reference frame i in the ordering is determined as the importance metric value of the reference predictor value i.

[0132] For example, the importance metric score of reference frame 1 corresponding to reference prediction value 1 is 20, the importance metric score of reference frame 2 corresponding to reference prediction value 2 is 30, and the importance metric score of reference frame 3 corresponding to reference prediction value 3 is 40. Next, the importance metric scores of the reference frames corresponding to the three reference prediction values ​​are ordered in ascending order, and the ordering result is reference frame 1, reference frame 2, and reference frame 3. Here, since the index of reference frame 1 in the ordering is 1, the importance metric value of reference prediction value 1 is 1, since the index of reference frame 2 in the ordering is 2, the importance metric value of reference prediction value 2 is 2, and since the index of reference frame 3 in the ordering is 3, the importance metric value of reference prediction value 3 is 3.

[0133] Here, the importance metric score of the reference frame i can be calculated based on the first and second importance metric values ​​in the following ways: (1) Obtain the importance metric value of the reference prediction value i by performing a weighted addition of the first and second importance metric values; (2) Obtain the importance metric value of the reference prediction value i by performing an averaging process of the first and second importance metric values.

[0134] Method 4: In order to obtain an accurate importance metric value, the decoding device may obtain the importance metric value of the reference predictor i by adjusting the calculation result of method 1, method 2, or method 3 according to the prediction mode of the reference predictor i. Here, the prediction mode of the reference predictor i includes any one of an inter prediction mode and an intra prediction mode. In one implementation, the importance metric value of the reference predictor i may be obtained by adjusting the calculation result of method 1, method 2, or method 3 according to an adjustment function. The adjustment function may be, for example, D'=g(D)=a*D+b. Here, D' represents the importance metric value of the reference predictor i, D is the calculation result of the above method 1, method 2, or method 3, and a and b may be adaptively set according to the prediction mode. For example, three specific examples are shown below.

[0135] a) When the prediction mode of the reference predictor value i is inter prediction, a=1 and b=0.

[0136] b) When the prediction mode of the reference predictor i is intra prediction, a=2 and b=0.

[0137] c) When the prediction mode of the reference predictor i is intra prediction, a=0 and b=160.

[0138] It should be understood that in practice, the importance metric value of the reference predicted value may be determined using any one of the above methods 1 to 4 as necessary, but the present application is not limited thereto. The decoding device may determine the importance metric value using various different methods, which may ensure the reliability of the importance metric value, and may be advantageous in improving the prediction accuracy of the current block, thereby improving the performance of encoding and decoding.

[0139] In s32, a target weight group for weighted prediction is selected from the target weight list.

[0140] Here, the selection of a target weight group from the target weight list by the decoding device can be divided into the following two cases.

[0141] (1) The number of weight groups contained in the target weight list is equal to 1. In this case, there is no need to decode the index of the target weight group from the video bitstream, and the weight group in the target weight list is directly taken as the target weight group for weighted prediction.

[0142] (2) The number of weight groups included in the target weight list is greater than 1. That is, the target weight list includes a plurality of weight groups. For example, the target weight list is represented as {{2,14},{4,12},{6,10},{8,8}} / 16, and this target weight list includes 4 weight groups. In this case, the video bitstream can indicate the index of the target weight group during the weighted prediction of the current block. For the index of this target weight group, an encoding method that performs binary encoding with a shortened unary code or a multi-symbol entropy encoding method is used. Here, for the shortened unary code, when the maximum value Max of the syntax element to be encoded is known, assuming that the symbol to be encoded is x, when 0 < x < Max, the unary code is used for the binary conversion of x, and when x = Max, the binary sequence obtained by binary-converting x consists of all 1s and has a length of Max. Next, the decoding device needs to decode the index of the target weight group for weighted prediction from the video bitstream and select the target weight group from the target weight list using the index of the target weight group. Here, the index of the target weight group can indicate the position in the target weight list. For example, the target weight list in the above example includes 4 weight groups, the index of the target weight group decoded from the video bitstream by the decoding device is 2, and based on this index of the target weight group, the position in the target weight list, that is, the second position in the target weight list (that is, weight group 2) is determined, and the determined target weight group is {4,12} / 16. The decoding device can determine the target weight group based on the index decoded from the video bitstream, identify the target weight group, improve the processing efficiency of the selection of the target weight group, and improve the decoding efficiency.

[0143] In this embodiment, the decoding device determines a target weight list based on the importance of the N reference predictors, selects a target weight group for weighted prediction from the target weight list, and selects an appropriate weight group through the selection of the weight list to decode and reconstruct the current block, thereby improving the prediction accuracy of the current block and improving the encoding and decoding performance.

[0144] In S303, a weighted prediction process is performed on N reference predictive values ​​based on the weight values ​​in the target weight group to obtain a predictive value of the current block, and the predictive value of the current block is used to reconstruct a decoded image corresponding to the current block.

[0145] The number of weight values ​​actually provided in the target weight group should correspond to the number of reference predictors. For example, the number of reference predictors is N, and the number of weight values ​​actually provided in the target weight group is also N. For example, when N=3, the three reference predictors are reference predictor 1, reference predictor 2, and reference predictor 3, respectively, and the target weight group may include three weight values, which are weight value 1, weight value 2, and weight value 3, respectively. Here, reference predictor 1 corresponds to weight value 1, reference predictor 2 corresponds to weight value 2, and reference predictor 3 corresponds to weight value 3. In the above example, the target weight group may include only two weight values, which are weight value 1 and weight value, respectively. The sum of this weight value 1 and weight value 2 is less than 1. In this case, the implicit weight value 3=1-weight value 1-weight value 2 in the target weight group can also be obtained by calculation.

[0146] In one embodiment, there are two methods for obtaining a predicted value of a current block by performing weighted prediction processing of N reference predictive values ​​based on target weights: (1) and (2).

[0147] (1) The decoding device may obtain a predicted value of the current block by performing a weighted sum process on each of the N reference predictors using the weight values ​​in the target weight group, where the predicted value P(x,y) of the current block is expressed as:

[0148] P(x,y)=(w1·P0(x,y) + w2·P1(x,y) +····+ wn·P N-1 (x,y) / N.

[0149] where P(x,y) is the prediction value of the current block, and P0(x,y), P1(x,y), . . . , P N-1 (x, y) respectively represent N reference predictors, w1 represents a weight value corresponding to the first reference predictor corresponding to the current block, w2 represents a weight value corresponding to the second reference predictor corresponding to the current block, and by analogy, wn represents a weight value corresponding to the Nth reference predictor corresponding to the current block (x, y).

[0150] In one embodiment, when N=2, i.e., when weighted prediction is performed using two reference predictors for the current block, the number of weight values ​​included in the target weight group is 1, i.e., an implicit weight value is included, and the predicted value P(x,y) of the current block is It is also possible that P(x,y) = (w(x,y) · P0(x,y) + (1-w(x,y)) · P1(x,y)) / 2. where P(x,y) is the predicted value of the current block, P0(x,y) and P1(x,y) are the two reference predictors corresponding to the current block (x,y), w(x,y) is the weight value in the target weight group used for the first reference predictor P0(x,y), and 1-w(x,y) is the implicit weight value used for the second reference predictor P1(x,y).

[0151] (2) In video coding, considering the complexity of prediction value calculation, integer type calculation may be realized by using right shift operation instead of division, and in the form of integer type calculation, weighting process is performed on each of N reference prediction values ​​by using the weight value in the target weight group to obtain the prediction value of the current block. In the form of integer type calculation, the complexity of prediction value calculation can be reduced to a certain extent.

[0152] For example, when the number of reference predictors is two (i.e., N=2) and the number of weight values ​​in the target weight group is one, the decoding device may obtain a predicted value of the current block by performing a weighting process on each of the two reference predictors using the weight values ​​in the target weight group in the form of integer calculation. In this case, the predicted value P(x,y) of the current block is: It is also possible that P(x,y) = (w(x,y) × P0(x,y) + (16 - w(x,y)) × P1(x,y) + 8) >> 4. Here, ">>4" represents a 4-bit right shift, that is, the data in the weighting process can be divided by 16 by a 4-bit right shift, and 8 is an offset added for rounding. Above P0(x,y), P1(x,y), w(x,y), and P(x,y) are all integer types, P(x,y) is the predicted value of the current block, P0(x,y) and P1(x,y) are two reference predicted values ​​corresponding to the current block (x,y), and w(x,y) is the weight value (i.e., the weight value in the target weight group) used for the first predicted value P0(x,y). Also, for example, “>>6” represents a 6-bit right shift, that is, the data in the weighting process can be divided by 64 by a 6-bit right shift, and 32 is an offset added for rounding. In this case, the predicted value P(x,y) of the current block is It is also possible that P(x,y) = (w P0(x,y) + (64-w) P1(x,y) + 32) >> 6.

[0153] In this embodiment, the decoding device performs a weighted sum process on each of the N reference prediction values ​​using the weight values ​​in the target weight group, or performs a weighted sum process in the form of integer calculation to obtain a prediction value of the current block, and ensures the accuracy of the prediction value, which is advantageous to improving the decoding performance of the current block.

[0154] In one embodiment, after obtaining the predicted value of the current block, a overlapping process may be performed based on the residual video signal of the current block and the predicted value to obtain a reconstructed value of the current block, and a decoded image corresponding to the current block may be reconstructed based on the reconstructed value of the current block. On the one hand, the decoded image corresponding to the current block can be used as a reference image during weighted prediction of other coding blocks, and on the other hand, the decoded image corresponding to the current block can also be used to reconstruct the current frame of the current block, and finally, a video can be reconstructed based on the reconstructed video frames.

[0155] In an embodiment of the present application, a decoding device performs composite prediction of a current block in a video bitstream to obtain N reference prediction values ​​of the current block (N is an integer greater than 1). The current block refers to an encoding block being decoded in the video bitstream. The decoding device determines a target weight group for weighted prediction for the current block based on the importance of the N reference prediction values, and performs weighted prediction processing of the N reference prediction values ​​based on the target weight group to obtain a prediction value of the current block. The prediction value of the current block is for reconstructing a decoded image corresponding to the current block. Composite prediction is used for decoding the video, and the importance of the reference prediction values ​​is fully taken into account in the composite prediction, and the composite prediction realizes adaptively selecting an appropriate weight for the current block based on the importance of each reference prediction value to perform weighted prediction, thereby improving the prediction accuracy of the current block and improving the performance of encoding and decoding.

[0156] Please refer to Fig. 4. Fig. 4 is a schematic diagram of a flow of another video processing method provided in an embodiment of the present application. This video processing method may be performed by an encoding device in a video processing system. The video processing method in this embodiment may include the following steps S401 to S405.

[0157] In S401, a current block is obtained by performing a division process of a current frame in a video. The current block refers to a coding block in a video that is being coded by a coding device. Here, the video may include one or more video frames. The current frame refers to a video frame that is being coded. The coding device may obtain one or more coding blocks by dividing the current frame in the video. The current block refers to any one of the coding blocks in the current frame that is being coded.

[0158] In S402, N (N is an integer greater than 1) reference prediction values ​​of the current block are obtained by performing composite prediction of the current block, and the N reference prediction values ​​are derived from N reference blocks of the current block, and the N reference blocks are coding blocks in the video that are referenced when coding the current block, and there is a one-to-one correspondence between the reference prediction values ​​and the reference blocks.

[0159] The above N reference prediction values ​​may be derived from the N reference blocks. One reference prediction value corresponds to one reference block. In the embodiment of the present application, the video frame in which the reference block is located may be a reference frame, and the video frame in which the current block is located is the current frame. Here, the positional relationship between the N reference blocks and the current block includes any one of the following (1) to (4). (1) The N reference blocks may be located in the N reference frames, respectively, and the N reference frames and the current frame may belong to different video frames in the video. (2) The N reference blocks may be located in the same reference frame, and the same reference frame and the current frame may belong to different video frames in the video. (3) One or more reference blocks of the N reference blocks are located in the current frame, and the remaining reference blocks of the N reference blocks are located in one or more reference frames, and the one or more reference frames and the current frame belong to different video frames in the video. (4) Both the N reference blocks and the current block are located in the current frame.

[0160] As can be seen from the positional relationships between the N reference blocks and the current block shown in (1) to (4) above, the prediction modes of the composite prediction in the embodiments of the present application include inter prediction mode (i.e., at least two reference frames are allowed to be used for inter prediction), combined prediction mode (i.e., at least one reference frame is allowed to be used for inter prediction and the current frame is allowed to be used for intra prediction), and intra prediction mode (i.e., the current frame is allowed to be used for intra prediction).

[0161] Corresponding to different modes of composite prediction, the encoding device performing composite prediction of the current block to obtain N reference prediction values ​​of the current block may include any one of the following (1) to (2). (1) During the composite prediction of the current block, N reference prediction values ​​of the current block are obtained by performing inter prediction on the current block using N reference blocks. In this case, the N reference prediction values ​​of the current block are derived by performing inter prediction using the N reference blocks of the current block, respectively. (2) During the composite prediction of the current block, inter prediction is performed on the current block using at least one reference prediction value of the N reference prediction values, and intra prediction is performed using the remaining reference blocks of the N reference blocks, thereby obtaining N reference prediction values. Here, some of the N reference prediction values ​​are derived by performing inter prediction using at least one reference block of the N reference blocks of the current block, and the remaining reference prediction values ​​are derived by performing intra prediction using the remaining reference blocks of the N reference blocks.

[0162] In S403, a target weight group for weighted prediction is determined for the current block based on the importance of the N reference predictors, where the importance represents the degree of influence of each reference predictor on the coding performance of the current block.

[0163] The importance of a reference predictor may be determined comprehensively based on factors such as the bit rate consumed in the weighted prediction process and the quality loss in the coding of the current block. For example, if the bit rate consumption increases obviously when a certain reference predictor is used for the current block, this indicates that the reference predictor does not contribute significantly to reducing the bit rate consumption, and its importance is low. Also, for example, if the quality loss increases obviously when a certain reference predictor is used for the current block, this indicates that the reference predictor does not contribute significantly to reducing the quality loss, and its importance is low. A target weight group includes one or more weight values, and these weight values ​​act on each reference predictor in the weighted prediction. If the importance of a certain reference predictor is low, this reference predictor corresponds to a weight value with a small value in the target weight group, and if the importance of a certain reference predictor is high, this reference predictor corresponds to a weight value with a large value in the target weight group. In other words, the target weight group is a group of weights that is selected after considering the influence of each reference predictor on factors such as bitrate consumption and quality loss, so that the total cost of weighted prediction (i.e., bitrate consumption cost, quality loss cost, bitrate consumption and quality loss cost) is small and the encoding / decoding performance is excellent.

[0164] In one implementation, the decoding device uses weight values ​​in the target weight group to perform weighted prediction processing of the N reference predictors, the consumed bit rate is less than a predetermined bit rate threshold, or performs weighted prediction processing of the N reference predictors according to the weight values ​​in the target weight group to make the quality loss in encoding the current block less than a predetermined loss threshold, or performs weighted prediction processing of the N reference predictors according to the weight values ​​in the target weight group to make the quality loss in encoding the current block less than a predetermined loss threshold. Here, the predetermined bit rate threshold and the predetermined loss threshold may be preset according to actual needs, for example, the corresponding bit rate threshold and loss threshold may be statistically analyzed according to past encoding and decoding records to reduce the overall cost of weighted prediction. In this embodiment, the encoding device can constrain the selection of the target weight group based on at least one of a predetermined bit rate threshold and a predetermined loss threshold, and can select an appropriate target weight group according to actual needs, which is favorable to reducing encoding costs and improving encoding and decoding performance.

[0165] In one embodiment, there is one or more weight lists for encoding. This may be understood as one or more weight lists may be used for encoding. Each weight list includes one or more weight groups, each weight group includes one or more weight values, and the number of weight values ​​included in each weight group is allowed to be the same or different, and the value of the weight values ​​included in each weight group is allowed to be the same or different. In this case, the step of determining a target weight group for weighted prediction for the current block based on the importance of the N reference predictors may include steps s41-s42.

[0166] In s41, a target weight list is determined from one or more weight lists based on the importance of the N reference predictors.

[0167] Here, specific implementations of determining a target weight list from one or more weight lists based on the importance of the N reference predictors may include the following methods.

[0168] Method 1: If there is one weight list in the encoding, we can directly use this weight list as the target weight list.

[0169] Method 2: When there are multiple weight lists in encoding, an importance metric value of a reference predictor may be introduced, and a target weight list may be determined from the multiple weight lists based on the importance metric values ​​of the N reference predictors.

[0170] (1) Determine a target weight list based on the absolute value of the difference between the importance metric values ​​of the N reference predictors.

[0171] When the number of weight lists in encoding is M+1 (M is a positive integer equal to or greater than 1), one weight list corresponds to one threshold interval, that is, the number of threshold intervals is also M+1. Next, the encoding device may obtain the importance metric values ​​of N reference prediction values ​​and calculate the importance difference between the N reference prediction values. The importance difference between any two reference prediction values ​​is measured by the difference between the importance metric values ​​of the any two reference prediction values. Specifically, the difference between the importance metric between any two reference prediction values ​​is calculated, and the difference between the importance metric values ​​of the any two reference prediction values ​​is taken as the importance difference between the any two reference prediction values, and then the threshold interval in which the absolute value of the importance difference between the N reference prediction values ​​is located is determined, and the weight list corresponding to the threshold interval in which the absolute value of the importance difference between the N reference prediction values ​​is located may be determined as the target weight list.

[0172] (2) A target weight list is determined by comparing the magnitudes of the importance metric values ​​of the reference predictors.

[0173] The N reference predictor values ​​of the current block may include a first reference predictor value and a second reference predictor value, and the weight list for encoding includes a first weight list and a second weight list. The encoding device may compare the magnitude between the importance metric value of the first reference predictor value and the importance metric value of the second reference predictor value, and determine the first weight list as a target weight list if it determines that the importance metric value of the first reference predictor value is greater than the importance metric value of the second reference predictor value, and determine the second weight list as a target weight list if it determines that the importance metric value of the first reference predictor value is equal to or less than the importance metric value of the second reference predictor value.

[0174] Optionally, the weight values ​​in the first weight list and the weight values ​​in the second weight list are reversed.

[0175] (3) A mathematical sign function, the importance metric value of the reference predictor, is utilized to determine the target weight list.

[0176] The N reference predictor values ​​of the current block include a first reference predictor value and a second reference predictor value, and the weight list in encoding includes a first weight list, a second weight list, and a third weight list. The encoding device may obtain a code value by calling a mathematical code function to process a difference between the importance metric value of the first reference predictor value and the importance metric value of the second reference predictor value. If the code value is a first predetermined value (e.g., -1), the first weight list is determined as a target weight list; if the code value is a second predetermined value (e.g., 0), the second weight list is determined as a target weight list; and if the code value is a third predetermined value (e.g., 1), the third weight list is determined as a target weight list.

[0177] (4) The above methods (1), (2), and (3) may be used alone, or the target weight list may be determined by combining the methods (1), (2), and (3). As one implementation form, assuming that there are M+1 weight lists, the weight list corresponding to the threshold interval in which the absolute value of the importance difference is located can be determined as the candidate weight list in the method (1). Next, the N reference prediction values ​​include a first reference prediction value and a second prediction value, and the importance metric value of the first reference prediction value and the importance metric value of the second reference prediction value are compared in size in the method (2). If it is determined that the importance metric value of the first reference prediction value is greater than the importance metric value of the second reference prediction value, the candidate weight list is directly determined as the target weight list, and if it is determined that the importance metric value of the first reference prediction value is equal to or less than the importance metric value of the second reference prediction value, the weight list corresponding to the inverse weight value of the weight value in the candidate weight list is determined as the target weight list.

[0178] Here, any one of the N reference predictors is represented as reference predictor i (i is an integer less than or equal to N), the reference predictor i is derived from reference block i, the video frame in which reference block i is located is reference frame i, the video frame in which the current block is located is the current frame, and the importance metric value of reference predictor i may be determined by any one of the following methods:

[0179] Method 1: Calculate based on the picture order count (POC) of the current frame in the video and the picture order count of the reference frame i in the video. Specifically, the encoding device may calculate the difference between the picture order count of the current frame in the video and the picture order count of the reference frame i in the video, and the absolute value of this difference may be the importance metric value of the reference prediction value i. For example, if the picture order count of the current frame in the video is represented as cur_poc and the picture order count of the reference frame i in the video is represented as ref_poc, the importance metric value D of the reference prediction value i is D=abs(cur_poc-ref_poc), where abs() represents obtaining the absolute value.

[0180] Method 2: Calculate based on the picture order count of the current frame in the video, the picture order count of the reference frame i in the video, and the quality metric Q. Here, the quality metric Q may be determined in various cases. This application is not limited thereto. Specifically, the quality metric Q of the reference frame i may be derived from the quantization information of the current block. For example, the quality metric Q may be set as the base quantization index (base_qindex) of the reference frame i. The base_qindex of any reference frame may be different or the same. In other implementations, the quality metric Q of the reference frame i may be derived from other coding information. For example, the quality metric Q of the reference frame i may be derived from the coding information difference between the coded CU in the reference frame i and the coded CU in the current frame.

[0181] In one implementation, the encoding device may calculate the difference between the picture order count of a current frame in the video and the picture order count of a reference frame i in the video, and then use an objective function to determine an importance metric value for reference predictor i based on the difference, a quality metric Q, and a list of importance metric values.

[0182] In one embodiment, the importance metric value of the reference frame i may be calculated based on an orientation relationship between the reference frame i and the current frame and a quality metric Q. In one implementation, a correspondence relationship between the reference orientation relationship and the reference importance metric value may be established. For example, if the reference orientation relationship is that the reference frame is located before the current frame, the reference importance metric value may correspond to a first value, and if the reference orientation relationship is that the reference frame is located after the current frame, the reference importance metric value may correspond to a second value. Then, the encoding device may calculate an importance metric value of the reference prediction value i based on the reference importance metric value corresponding to the reference frame i and the quality metric Q.

[0183] Method 3: Based on the calculation results of Method 1 and Method 2, calculate the importance metric score of reference frame i, and order the importance metric scores of the reference frames corresponding to the N reference prediction values ​​in ascending order, and determine the index of reference frame i in the ordering as the importance metric value of reference prediction value i.

[0184] In one implementation, a first importance metric value of a reference predictor value i can be calculated by the above method 1, and a second importance metric value of the reference predictor value i can be calculated by the above method 2, and then an importance metric score (e.g., score) of the reference frame i is calculated based on the first importance metric value and the second importance metric value, and then the importance metric scores of the reference frames corresponding to the N reference predictors are ordered in ascending order, and the index of the reference frame i in the ordering is determined as the importance metric value of the reference predictor value i.

[0185] Here, the importance metric score of the reference frame i can be calculated based on the first and second importance metric values ​​in the following ways: (1) Obtain the importance metric value of the reference prediction value i by performing a weighted addition of the first and second importance metric values; (2) Obtain the importance metric value of the reference prediction value i by performing an averaging process of the first and second importance metric values.

[0186] Method 4: To obtain an accurate importance metric value, the importance metric value of the reference predictor i may be obtained by adjusting the calculation result of method 1, method 2, or method 3 according to the prediction mode of the reference predictor i, where the prediction mode of the reference predictor i includes any one of an inter prediction mode and an intra prediction mode.

[0187] It should be understood that in practice, the importance metric value of the reference predictor value may be determined using any one of the above methods 1 to 4 according to need, but the present application is not limited thereto.

[0188] In s42, a target weight group for weighted prediction is selected from the target weight list.

[0189] Here, the target weight list includes one or more weight groups.

[0190] (1) If the number of weight groups included in the target weight list is equal to 1, the weight groups in the target weight list are directly taken as the target weight groups for weighted prediction.

[0191] (2) When the target weight list contains multiple weight groups, a target weight group for weighted prediction is selected from the target weight list.

[0192] When performing the weighted prediction process, bit rate may be consumed, and the coding block may suffer quality loss during coding. Therefore, the coding device may first obtain coding performance when performing weighted prediction process for each of the N reference predictors using each weight group in the target weight list, and then determine the weight group in the target weight list with the best coding performance as the target weight group. In one implementation, when weighted prediction processing of N reference predictors is performed using weight values ​​in the target weight group, the consumed bit rate is the minimum value of the consumed bit rate corresponding to all weight groups in the target weight list; alternatively, when weighted prediction processing of N reference predictors is performed based on weight values ​​in the target weight group, the quality loss in encoding the corresponding current block is the minimum value of the quality loss corresponding to all weight groups in the target weight list; alternatively, when weighted prediction processing of N reference predictors is performed using weight values ​​in the target weight group, the consumed bit rate is the minimum value of the consumed bit rate corresponding to all weight groups in the target weight list, and when weighted prediction processing of N reference predictors is performed based on weight values ​​in the target weight group, the quality loss in encoding the corresponding current block is the minimum value of the quality loss corresponding to all weight groups in the target weight list. The encoding device can determine the weight group with the best encoding performance as the target weight group. Encoding the current according to the target weight group can obtain the best encoding performance, improving the encoding performance of the video.

[0193] It should be understood that, when the encoding side selects a target weight group for weighted prediction, it needs to try each weight group in turn from the target weight list to obtain a target weight group for weighted prediction. The decoding side does not need to try in turn, and indicates the index of the target weight group to be used in the video bitstream. The decoding side only needs to decode the index of the target weight group from the video bitstream, and find the target weight group in the target weight list according to the index of the target weight group. The encoding device determines the target weight list according to the importance of the N reference predictors, and selects the target weight group for weighted prediction from the target weight list, and can select an appropriate weight group to reconstruct the current block through the selection of the weight list, thereby improving the prediction accuracy of the current block and improving the encoding performance.

[0194] In S404, a weighted prediction process is performed on the N reference predictors based on the weights in the target weight group to obtain a predicted value of the current block.

[0195] Here, for a specific implementation of step S404, refer to the specific implementation of step S303 above.

[0196] In S405, a video bitstream is generated by encoding the video based on the predicted value of the current block. The encoding device may generate the video bitstream by encoding the video based on the predicted value of the current block and the index of the target weight group. Here, for generating the video bitstream by encoding the video based on the predicted value of the current block and the index of the target weight group, refer to the above description of the encoding.

[0197] In an embodiment of the present application, an encoding device obtains a current block, which is a coding block being coded in the video, by performing a division process of a current frame in the video, obtains N reference prediction values ​​of the current block by performing composite prediction of the current block, determines a target weight group for weighted prediction for the current block based on the importance of the N reference prediction values, obtains a prediction value of the current block by performing a weighted prediction process of the N reference prediction values ​​based on the target weight group, and encodes the video based on the prediction value of the current block to generate a video bitstream. Composite prediction is used for coding the video, the importance of the reference prediction value is derived in the composite prediction, and the composite prediction realizes adaptively selecting an appropriate weight for the current block based on the importance of each reference prediction value to perform weighted prediction, thereby improving the prediction accuracy of the current block and improving the performance of coding and decoding.

[0198] Please refer to Fig. 5. Fig. 5 is a schematic diagram of the configuration of a video processing device provided in an embodiment of the present application. The video processing device may be provided in a computer device provided in an embodiment of the present application. The computer device may be a decoding device mentioned in the above method embodiment. The video processing device shown in Fig. 5 may be computer-readable instructions (including program code) executed in the computer device. The video processing device may be used to perform some or all of the steps in the method embodiment shown in Fig. 3. Referring to Fig. 5, the video processing device may include the following units:

[0199] The processing unit 501 performs composite prediction of a current block in a video bitstream to obtain N (N is an integer greater than 1) reference prediction values ​​of the current block, where the current block is a coding block to be decoded in the video bitstream, the N reference prediction values ​​are derived from the N reference blocks of the current block, the N reference blocks are coding blocks in the video bitstream that are referenced when decoding the current block, and there is a one-to-one correspondence between the reference prediction values ​​and the reference blocks.

[0200] The determining unit 502 determines a target weight group for weighted prediction for the current block based on the importance of the N reference predictors, where the target weight group includes one or more weight values, and the importance is for representing the degree of influence of each reference predictor on the decoding performance of the current block.

[0201] The processing unit 501 further performs weighted prediction processing of the N reference predictive values ​​based on the weight values ​​in the target weight group to obtain a predictive value of the current block, and the predictive value of the current block is for reconstructing a decoded image corresponding to the current block.

[0202] In one embodiment, the video frame in which the reference block is located is a reference frame, and the video frame in which the current block is located is the current frame, and the positional relationship between the N reference blocks and the current block includes any one of the following: the N reference blocks are respectively located in the N reference frames, and the N reference frames and the current frame belong to different video frames in the video bitstream; the N reference blocks are located in the same reference frame, and the same reference frame and the current frame belong to different video frames in the video bitstream; one or more of the N reference blocks are located in the current frame, and the remaining reference blocks of the N reference blocks are located in one or more reference frames, and the one or more reference frames and the current frame belong to different video frames in the video bitstream; and both the N reference blocks and the current block are located in the current frame.

[0203] In one embodiment, the processing unit 501 further determines a condition for adaptive weighted prediction, and if the current block satisfies the condition for adaptive weighted prediction, determines a target weight group for weighted prediction for the current block based on the importance of the N reference predictors.

[0204] In one embodiment, if the current block satisfies the condition for adaptive weighted prediction, the following is performed: a sequence header of a frame sequence to which the current block belongs includes a first indication field, and the first indication field indicates that adaptive weighted prediction is allowed for the coding blocks in the frame sequence, where the frame sequence is a sequence of video frames in a video bitstream; a slice header of a current slice to which the current block belongs includes a second indication field, and the second indication field indicates that adaptive weighted prediction is allowed for the coding blocks in the current slice, where the current slice is an image slice to which the current block belongs, and the image slice is divided from a current frame in which the current block is located; a frame header of a current frame to which the current block belongs includes a third indication field, and the third indication field indicates that adaptive weighted prediction is allowed for the coding blocks in the current frame; and during the composite prediction, the current block is selected from the group consisting of the first and second indication fields. and wherein at least two reference frames are used for inter prediction; during composite prediction, for a current block, at least one reference frame is used for inter prediction and a current frame is used for intra prediction; the motion type of the current block is a specified motion type; a predetermined motion vector prediction mode is used for the current block; a predetermined interpolation filter is used for the current block; a specific coding tool is not used for the current block; and a reference frame used for composite prediction for the current block satisfies a specific condition, the specific condition including one or more of: an orientation relationship in the video bitstream between the reference frame used and the current frame satisfies a specific relationship; and an absolute value of a significance difference between the reference prediction values ​​corresponding to the reference frame used is equal to or greater than a specific threshold, and wherein the orientation relationship satisfies the specific relationship includes at least one of the following: any of the reference frames used is located before the current frame; any of the reference frames used is located after the current frame;some of the reference frames used are located before the current frame, and the remaining reference frames are located after the current frame;

[0205] In one embodiment, the video bitstream includes one or more weight lists, each weight list includes one or more weight groups, each weight group includes one or more weight values, the number of weight values ​​in each weight group is allowed to be the same or different, and the values ​​of the weight values ​​included in each weight group are allowed to be the same or different, and the determining unit 502 specifically includes:

[0206] A target weight list may be determined from among the one or more weight lists based on the importance of the N reference predictors, and a target weight group for weighted prediction may be selected from the target weight list.

[0207] In one embodiment, the number of weight lists in the video bitstream is M+1 (M is a positive integer equal to or greater than 1), and one weight list corresponds to one threshold interval, and the determining unit 502 may specifically obtain the importance metric values ​​of N reference predictors, calculate the importance difference between the N reference predictors, determine the threshold interval in which the absolute value of the importance difference between the N reference predictors is located, and determine the weight list corresponding to the threshold interval in which the absolute value of the importance difference between the N reference predictors is located as the target weight list. Here, the importance difference between any two reference predictors is measured by the difference between the importance metric values ​​of the any two reference predictors.

[0208] In one embodiment, the N reference predictors of the current block include a first reference predictor and a second reference predictor, and the video bitstream includes a first weight list and a second weight list, and the determining unit 502 may specifically compare the magnitude between the importance metric value of the first reference predictor and the importance metric value of the second reference predictor, and determine the first weight list as the target weight list if the importance metric value of the first reference predictor is greater than the importance metric value of the second reference predictor, and determine the second weight list as the target weight list if the importance metric value of the first reference predictor is less than or equal to the importance metric value of the second reference predictor. Here, the sum of the weight values ​​of the same order in the first weight list and the second weight list is 1.

[0209] In one embodiment, the N reference predictors of the current block include a first reference predictor and a second reference predictor, and the video bitstream includes a first weight list, a second weight list, and a third weight list, and the determining unit 502 may specifically call a mathematical sign function to obtain a sign value by processing the difference between the importance metric value of the first reference predictor and the importance metric value of the second reference predictor, and determine the first weight list as the target weight list when the sign value is a first predetermined value, determine the second weight list as the target weight list when the sign value is a second predetermined value, and determine the third weight list as the target weight list when the sign value is a third predetermined value. Here, the first weight list, the second weight list, and the third weight list are different weight lists, or two weight lists among the first weight list, the second weight list, and the third weight list are allowed to be the same weight list.

[0210] In one embodiment, the method includes: a method for determining whether a video frame in which a current frame is located is a video frame in which a video bitstream is a current frame; a method for determining whether a video frame in which a current frame is located is a video frame in which a video bitstream is a current frame; and a method for determining whether a video frame in which a current frame is located is a video frame in which a video bitstream is a current frame. In one embodiment, the method includes: a method for determining whether a video frame in which a current frame is located is a video frame in which a video bitstream is a current frame; method 2, calculating an importance metric score for the reference frame i based on the pixel order count and the pixel order count; method 3, calculating an importance metric score for the reference frame i based on the calculation results of method 1 and method 2, ordering the importance metric scores of the reference frames corresponding to the N reference predictors in ascending order, and determining an index of the reference frame i in the ordering as the importance metric value of the reference predictor i; and method 4, obtaining the importance metric value of the reference predictor i by adjusting the calculation result of method 1, method 2, or method 3 based on a prediction mode of the reference predictor i, where the prediction mode of the reference predictor i includes any one of an inter prediction mode and an intra prediction mode.

[0211] In one embodiment, the number of weight groups included in the target weight list is greater than 1, and the determining unit 502 may specifically decode an index of a target weight group for weighted prediction from the video bitstream, and select a target weight group from the target weight list according to the index of the target weight group, where the index of the target weight group is encoded using a shortened unary code binarization encoding scheme or a multi-symbol entropy encoding scheme.

[0212] In one embodiment, the processing unit 501 may specifically use the weight values ​​in the target weight group to perform a weighted sum process on each of the N reference predictors to obtain a predicted value of the current block, or may use the weight values ​​in the target weight group in the form of integer calculation to perform a weighting process on each of the N reference predictors to obtain a predicted value of the current block.

[0213] In an embodiment of the present application, composite prediction is performed on a current block, which is a coding block being decoded in a video bitstream, to obtain N reference prediction values ​​of the current block, and adaptively select a target weight group for weighted prediction for the current block based on the importance of the N reference prediction values, and perform weighted prediction processing of the N reference prediction values ​​based on the target weight group to obtain a prediction value of the current block, the prediction value of the current block being used to reconstruct a decoded image corresponding to the current block. Composite prediction is used for decoding the video, and the importance of the reference prediction values ​​is fully taken into account in the composite prediction, and an appropriate weight is determined for the current block based on the importance of each reference prediction value in the composite prediction to perform weighted prediction, thereby improving the prediction accuracy of the current block and improving the performance of encoding and decoding.

[0214] Please refer to Fig. 6. Fig. 6 is a schematic diagram of the configuration of a video processing device provided in an embodiment of the present application. The video processing device may be provided in a computer device provided in an embodiment of the present application. The computer device may be the encoding device mentioned in the above method embodiment. The video processing device shown in Fig. 6 may be computer-readable instructions (including program code) executed in the computer device. The video processing device may be used to perform some or all of the steps in the method embodiment shown in Fig. 4. Referring to Fig. 6, the video processing device may include the following units:

[0215] The processing unit 601 obtains a current block by performing a segmentation process of a current frame in a video.

[0216] The processing unit 601 further performs composite prediction of the current block to obtain N (N is an integer greater than 1) reference prediction values ​​of the current block, where the N reference prediction values ​​are derived from N reference blocks of the current block, and the N reference blocks are coding blocks in the video that are referenced when coding the current block, and the reference prediction values ​​and the reference blocks have a one-to-one correspondence.

[0217] The determining unit 602 determines a target weight group for weighted prediction for the current block based on the importance of the N reference predictors, where the target weight group includes one or more weight values, and the importance is for representing the degree of influence of each reference predictor on the encoding performance of the current block.

[0218] The processing unit 601 further performs weighted prediction processing of the N reference predictive values ​​based on the weight values ​​in the target weight group to obtain a predictive value of the current block, and the predictive value of the current block is for reconstructing a decoded image corresponding to the current block.

[0219] Processing unit 601 further generates a video bitstream by encoding the video based on the prediction value of the current block.

[0220] In one embodiment, the determining unit 602 may specifically determine a target weight list from one or more weight lists according to the importance of the N reference predictors, and select a target weight group for weighted prediction from the target weight list. Each weight list includes one or more weight groups, each weight group includes one or more weight values, the number of weight values ​​included in each weight group is allowed to be the same or different, and the values ​​of the weight values ​​included in each weight group are allowed to be the same or different.

[0221] In one embodiment, when weighted prediction processing of N reference predictive values ​​is performed using weight values ​​in a target weight group, the bit rate consumed is less than a predetermined bit rate threshold, or weighted prediction processing of N reference predictive values ​​is performed based on weight values ​​in the target weight group, thereby reducing the quality loss in encoding of a current block to less than a predetermined loss threshold, or when weighted prediction processing of N reference predictive values ​​is performed using weight values ​​in a target weight group, the bit rate consumed is less than a predetermined bit rate threshold and weighted prediction processing of N reference predictive values ​​is performed based on weight values ​​in the target weight group, thereby reducing the quality loss in encoding of a current block to less than a predetermined loss threshold.

[0222] In one embodiment, the target weight group is a weight group with the best encoding performance in the target weight list, where the best encoding performance includes: when performing weighted prediction processing of N reference predictors using weight values ​​in the target weight group, the consumed bit rate is the minimum value among the consumed bit rates corresponding to all weight groups in the target weight list; or when performing weighted prediction processing of N reference predictors based on weight values ​​in the target weight group, the quality loss in encoding of the corresponding current block is the minimum value among the quality losses corresponding to all weight groups in the target weight list; or when performing weighted prediction processing of N reference predictors using weight values ​​in the target weight group, the consumed bit rate is the minimum value among the consumed bit rates corresponding to all weight groups in the target weight list, and when performing weighted prediction processing of N reference predictors based on weight values ​​in the target weight group, the quality loss in encoding of the corresponding current block is the minimum value among the quality losses corresponding to all weight groups in the target weight list.

[0223] In the embodiment of the present application, a current block, which is a coding block being coded in the video, is obtained by performing a division process of a current frame in the video, N reference prediction values ​​of the current block are obtained by performing composite prediction of the current block, a target weight group for weighted prediction is determined for the current block based on the importance of the N reference prediction values, a weighted prediction process of the N reference prediction values ​​is performed based on the target weight group to obtain a prediction value of the current block, and a video bitstream is generated by coding the video based on the prediction value of the current block. Composite prediction is used for coding the video, and the importance of the reference prediction values ​​is fully taken into consideration, and in the composite prediction, an appropriate weight is determined for the current block based on the importance of each reference prediction value to perform weighted prediction, thereby improving the prediction accuracy of the current block and improving the performance of coding and decoding.

[0224] Furthermore, in the embodiment of the present application, a schematic diagram of the configuration of a computer device is also provided. For the schematic diagram of the configuration of the computer device, refer to FIG. 7. The computer device may be the above-mentioned encoding device or decoding device. The computer device may include a processor 701, an input device 702, an output device 703, and a memory 704. The above-mentioned processor 701, the input device 702, the output device 703, and the memory 704 are connected via a bus. The memory 704 stores computer-readable instructions, the computer-readable instructions include program instructions, and the processor 701 executes the program instructions stored in the memory 704.

[0225] In an embodiment of the present application, when the computing device is the above-mentioned decoding device, the processor 701 executes executable program code in the memory 704 to perform each step of the video processing method relating to decoding as performed by the above-mentioned decoding device.

[0226] Optionally, when the computing device is the above-mentioned encoding device, in an embodiment of the present application, the processor 701 executes executable program code in the memory 704 to perform each step of the video processing method related to encoding performed by the above-mentioned encoding device.

[0227] In addition, in the embodiment of the present application, a computer-readable storage medium is also provided, in which computer-readable instructions are stored, the computer-readable instructions including program instructions, and when a processor executes the program instructions, the method in the embodiment corresponding to the above-mentioned FIG. 3 and FIG. 4 can be executed. Therefore, further description is omitted here. For technical details not disclosed in the embodiment of the computer-readable storage medium of the present application, please refer to the description of the embodiment of the method of the present application. For example, the program instructions may be arranged to be executed on one computer device, on multiple computer devices at one location, or on multiple computer devices distributed at multiple locations and connected to each other via a communication network.

[0228] According to one aspect of the present application, a computer program product is provided that includes computer-readable instructions. The computer-readable instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer-readable instructions from the computer-readable storage medium, and when the processor executes the computer-readable instructions, the computer device can execute the method in the embodiment corresponding to FIG. 3 and FIG. 4 above. Therefore, further description is omitted here. As can be understood by those skilled in the art, all or part of the flow of the method according to the above embodiment may be realized by instructing related hardware through computer-readable instructions. The computer-readable instructions may be stored in a computer-readable storage medium. When the program is executed, the flow of each method embodiment as described above is executed. Here, the storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), etc.

[0229] The above are only preferred embodiments of the present application, and of course, they do not limit the scope of the present application. As can be understood by those skilled in the art, the equivalent modifications that realize all or part of the above embodiment flow and comply with the claims of the present application still belong to the scope of the present invention.

Claims

1. 1. A computer device implemented video processing method, comprising: performing composite prediction of a current block in a video bitstream to obtain N reference prediction values ​​(N is an integer greater than 1) of the current block, the current block being a coding block to be decoded in the video bitstream, the N reference prediction values ​​being derived from N reference blocks of the current block, the N reference blocks being coding blocks in the video bitstream that are referenced when the current block is decoded, and the reference prediction values ​​and the reference blocks having a one-to-one correspondence; determining a target weight group for weighted prediction for the current block based on the importance of the N reference predictors, the target weight group including one or more weight values, the importance representing the degree of influence of each reference predictor on the decoding performance of the current block; and performing a weighted prediction process of the N reference predictors based on weight values ​​in the target weight group to obtain a predicted value of the current block, the predicted value of the current block being for reconstructing a decoded image corresponding to the current block.

13. A video processing method comprising:

2. a video frame in which a reference block is located is a reference frame, and a video frame in which the current block is located is a current frame; The positional relationship between the N reference blocks and the current block is The N reference blocks are located in N reference frames, respectively, and the N reference frames and the current frame belong to different video frames in the video bitstream; The N reference blocks are located in a same reference frame, and the same reference frame and the current frame belong to different video frames in the video bitstream; one or more of the N reference blocks are located in the current frame, and the remaining reference blocks of the N reference blocks are located in one or more reference frames, and the one or more reference frames and the current frame belong to different video frames in the video bitstream; and wherein the N reference blocks and the current block are located in the current frame.

2. The video processing method of claim 1.

3. determining conditions for adaptive weighted prediction; If the current block satisfies the condition for the adaptive weighted prediction, performing a step of determining a target weight group for weighted prediction for the current block based on the importance of the N reference predictors.

2. The video processing method of claim 1.

4. If the current block satisfies the condition for adaptive weighted prediction, a first indication field is included in a sequence header of a frame sequence to which the current block belongs, and the first indication field indicates that adaptive weighted prediction is allowed for coding blocks in the frame sequence, the frame sequence being a sequence of video frames in the video bitstream; and a slice header of a current slice to which the current block belongs includes a second indication field, and the second indication field indicates that adaptive weighted prediction is allowed for coding blocks in the current slice, the current slice being an image slice to which the current block belongs, the image slice being divided from a current frame in which the current block is located; A third indication field is included in a frame header of a current frame in which the current block is located, and the third indication field indicates that adaptive weighted prediction is allowed for a coding block in the current frame; and During the composite prediction, at least two reference frames are used for inter prediction for the current block; and During the composite prediction, for the current block, at least one reference frame is used for inter prediction and the current frame is used for intra prediction; the motion type of the current block is a specified motion type; A predetermined motion vector prediction mode is used for the current block; A predetermined interpolation filter is used for the current block; and if no specific coding tool is used for the current block; and (c) a reference frame used during the composite prediction for the current block satisfies a specific condition, the specific condition including one or more of: an orientation relationship in the video bitstream between the reference frame used and the current frame satisfies a predetermined relationship; and an absolute value of a significance difference between reference prediction values ​​corresponding to the reference frame used is equal to or greater than a predetermined threshold; The orientation relationship satisfies a predetermined relationship, and includes any one of the following: all of the reference frames used are located before the current frame; all of the reference frames used are located after the current frame; and some of the reference frames used are located before the current frame and the remaining reference frames are located after the current frame.

4. A video processing method according to claim 3.

5. the video bitstream includes one or more weight lists, each weight list includes one or more weight groups, each weight group includes one or more weight values, the number of weight values ​​included in each weight group may be the same or different, and the values ​​of the weight values ​​included in each weight group may be the same or different; The step of determining a target weight group for weighted prediction for the current block based on the importance of the N reference predictors comprises: determining a target weight list from among the one or more weight lists based on the importance of the N reference predictors; selecting a target weight group for weighted prediction from the target weight list; 2. The video processing method of claim 1.

6. The number of weight lists in the video bitstream is M+1 (M is a positive integer equal to or greater than 1), and one weight list corresponds to one threshold interval; determining a target weight list from the one or more weight lists based on the importance of the N reference predictors, obtaining importance metric values ​​of the N reference predictors and calculating an importance difference between the N reference predictors, where the importance difference between any two reference predictors is measured as the difference between the importance metric values ​​of the any two reference predictors; determining a threshold interval in which the absolute value of the importance difference between the N reference predicted values ​​is located; determining a weight list corresponding to a threshold interval in which an absolute value of the importance difference between the N reference predicted values ​​is located as a target weight list; 6. A video processing method according to claim 5.

7. The N reference predictors of the current block include a first reference predictor and a second reference predictor, and the video bitstream includes a first weight list and a second weight list; determining a target weight list from the one or more weight lists based on the importance of the N reference predictors, a step of comparing a magnitude between an importance metric value of the first reference prediction value and an importance metric value of the second reference prediction value; determining the first weight list as a target weight list if the importance metric value of the first reference predictor value is greater than the importance metric value of the second reference predictor value; determining the second weight list as a target weight list if the importance metric value of the first reference predictor value is less than or equal to the importance metric value of the second reference predictor value; the sum of weight values ​​in the same order in the first weight list and the second weight list is 1; 6. A video processing method according to claim 5.

8. The N reference predictors of the current block include a first reference predictor and a second reference predictor, and the video bitstream includes a first weight list, a second weight list, and a third weight list; determining a target weight list from the one or more weight lists based on the importance of the N reference predictors, obtaining a sign value by processing the difference between the importance metric value of the first reference predictor value and the importance metric value of the second reference predictor value by calling a mathematical sign function; determining the first weight list as a target weight list if the code value is a first predetermined value; determining the second weight list as a target weight list if the code value is a second predetermined value; if the code value is a third predetermined value, determining the third weight list as a target weight list; the first weight list, the second weight list, and the third weight list are each a different weight list, or two of the first weight list, the second weight list, and the third weight list are allowed to be the same weight list; 6. A video processing method according to claim 5.

9. Any one of the N reference predictors is represented as a reference predictor i (i is an integer equal to or less than N), the reference predictor i being derived from a reference block i, the video frame in which the reference block i is located is a reference frame i, and the video frame in which the current block is located is a current frame; The importance metric value of the reference prediction value i is a method for performing a calculation based on a picture order count of the current frame in the video bitstream and a picture order count of the reference frame i in the video bitstream; Method 2, performing a calculation based on a picture order count of the reference frame i in the video bitstream, a quality metric Q, and a picture order count of the current frame in the video bitstream, where the quality metric Q is determined based on quantization information of the current block or coding information of the reference frame i; Method 3: Calculate the importance metric score of the reference frame i according to the calculation result of method 1 and the calculation result of method 2, order the importance metric scores of the reference frames corresponding to the N reference predictors in ascending order, and determine the index of the reference frame i in the ordering as the importance metric value of the reference predictor i; and method 4, which obtains an importance metric value for the reference predictor value i by adjusting the calculation result of method 1, method 2, or method 3 based on the prediction mode of the reference predictor value i; The prediction mode of the reference predicted value i includes any one of an inter prediction mode and an intra prediction mode. Video processing method according to any one of claims 6 to 8.

10. the number of weight groups included in the target weight list is greater than 1; The step of selecting a target weight group for weighted prediction from the target weight list includes: decoding an index of a target weight group for weighted prediction from the video bitstream; selecting the target weight group from the target weight list by an index of the target weight group; The index of the target weight group is encoded using a binarization encoding method with a shortened unary code or a multi-symbol entropy encoding method.

6. A video processing method according to claim 5.

11. The step of obtaining a predicted value of the current block by performing a weighted prediction process of the N reference predictors based on the weight values ​​in the target weight group includes: performing a weighted sum operation on each of the N reference predictors using the weights in the target weight group to obtain a prediction of the current block; or performing a weighting process on each of the N reference predictors using a weight value in the target weight group in the form of an integer calculation to obtain a predicted value of the current block; 2. The video processing method of claim 1.

12. 1. A computer device implemented video processing method, comprising: obtaining a current block by performing a segmentation process of a current frame in a video; performing hybrid prediction of the current block to obtain N reference prediction values ​​of the current block (N is an integer greater than 1), the N reference prediction values ​​being derived from N reference blocks of the current block, the N reference blocks being coding blocks in the video that are referenced when coding the current block, and the reference prediction values ​​and the reference blocks are in one-to-one correspondence; determining a target weight group for weighted prediction for the current block based on the importance of the N reference predictors, the target weight group including one or more weight values, the importance being for representing the degree of influence of each reference predictor on the encoding performance of the current block; performing a weighted prediction process of the N reference predictors based on weight values ​​in the target weight group to obtain a predicted value of the current block, the predicted value of the current block being used to reconstruct a decoded image corresponding to the current block; and encoding the video based on the prediction value of the current block to generate a video bitstream.

13. A video processing method comprising:

13. The step of determining a target weight group for weighted prediction for the current block based on the importance of the N reference predictors comprises: determining a target weight list from among one or more weight lists based on the importance of the N reference predictors, each weight list including one or more weight groups, each weight group including one or more weight values, the number of weight values ​​included in each weight group being allowed to be the same or different, and the values ​​of the weight values ​​included in each weight group being allowed to be the same or different; selecting a target weight group for weighted prediction from the target weight list; 13. A video processing method according to claim 12.

14. When performing weighted prediction processing of the N reference predictors using the weight values ​​in the target weight group, the consumed bit rate is less than a predetermined bit rate threshold, or Performing a weighted prediction process of the N reference predictors based on the weight values ​​in the target weight group to reduce the quality loss in encoding the current block below a predetermined loss threshold; or When performing weighted prediction processing of the N reference predictor values ​​using the weight values ​​in the target weight group, a consumed bit rate is smaller than a predetermined bit rate threshold, and by performing weighted prediction processing of the N reference predictor values ​​based on the weight values ​​in the target weight group, a quality loss in encoding of the current block is made smaller than a predetermined loss threshold.

14. A video processing method according to claim 13.

15. The target weight group is a weight group having the best coding performance in the target weight list, The best encoding performance includes: when weighted prediction processing of N reference predictors is performed using weight values ​​in a target weight group, the consumed bit rate is the minimum of the consumed bit rates corresponding to all weight groups in the target weight list; when weighted prediction processing of N reference predictors is performed based on weight values ​​in a target weight group, the quality loss in encoding of the corresponding current block is the minimum of the quality losses corresponding to all weight groups in the target weight list; or when weighted prediction processing of N reference predictors is performed using weight values ​​in a target weight group, the consumed bit rate is the minimum of the consumed bit rates corresponding to all weight groups in the target weight list, and when weighted prediction processing of N reference predictors is performed based on weight values ​​in the target weight group, the quality loss in encoding of the corresponding current block is the minimum of the quality losses corresponding to all weight groups in the target weight list.

14. A video processing method according to claim 13.

16. 1. A video processing device comprising: a processing unit for performing composite prediction of a current block in a video bitstream to obtain N reference prediction values ​​(N is an integer greater than 1) of the current block, the current block being a coding block to be decoded in the video bitstream, the N reference prediction values ​​being derived from N reference blocks of the current block, the N reference blocks being coding blocks in the video bitstream that are referenced when the current block is decoded, and the reference prediction values ​​and the reference blocks having a one-to-one correspondence; A determination unit for determining a target weight group for weighted prediction for the current block based on the importance of the N reference predictors, the target weight group including one or more weight values, and the importance is for representing a degree of influence of each reference predictor on a decoding performance of the current block; The processing unit further performs a weighted prediction process on the N reference predictors based on weight values ​​in the target weight group to obtain a predictive value of the current block, the predictive value of the current block being for reconstructing a decoded image corresponding to the current block.

13. A video processing device comprising:

17. 1. A video processing device comprising: a processing unit for obtaining a current block by performing a division process on a current frame in a video, and for obtaining N reference prediction values ​​of the current block (N is an integer greater than 1) by performing composite prediction on the current block, the N reference prediction values ​​being derived from N reference blocks of the current block, the N reference blocks being coding blocks in the video that are referenced when coding the current block, and the reference prediction values ​​and the reference blocks being in one-to-one correspondence; A determination unit for determining a target weight group for weighted prediction for the current block based on the importance of the N reference predictors, the target weight group including one or more weight values, and the importance representing a degree of influence of each reference predictor on the encoding performance of the current block; The processing unit further performs a weighted prediction process on the N reference predictors based on weight values ​​in the target weight group to obtain a predictive value of the current block, the predictive value of the current block being for reconstructing a decoded image corresponding to the current block; The processing unit further comprises: encoding the video based on the prediction value of the current block to generate a video bitstream.

13. A video processing device comprising:

18. 1. A computer device comprising: A processor suitable for executing computer readable instructions; a computer readable storage medium having stored thereon computer readable instructions which, when executed by the processor, cause the video processing method of any one of claims 1 to 15 to be performed.

1. A computer device comprising:

19. 16. A computer readable storage medium having stored thereon computer readable instructions which, when executed by a processor, cause a video processing method according to any one of claims 1 to 15 to be performed.

20. A computer program product comprising computer readable instructions which, when executed by a processor, causes the implementation of a video processing method according to any one of claims 1 to 15.

Citation Information

Patent Citations

  • Method for encoding moving image and method for decoding moving image

    JP2004007379A

  • Image prediction / encoding device, image prediction / encoding method, image prediction / encoding program, image prediction / decoding device, image prediction / decoding method, and image prediction / decoding program

    JP2008283662A

  • Compound prediction for video coding

    US20180205964A1

  • Systems and methods for generalized multi-hypothesis prediction for video coding

    US20190230350A1

  • Method and apparatus for processing video signal

    US20190246133A1