Video encoding method and apparatus, video decoding method and apparatus, computer readable medium, and electronic device
By decoding the video code stream and sorting the mode list, and selecting a low-cost angle-weighted prediction mode, the problem of high overhead for encoding the index information of the angle-weighted prediction mode in the prior art is solved, and the effect of improving the encoding and decoding performance is achieved.
Patent Information
- Application Number
- PCT/CN2024/131393
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-12
- Filing Date
- 2024-11-11
- Publication Date
- 2025-06-19
AI Technical Summary
In the prior art, the encoding bit overhead of the angle-weighted prediction mode index information is relatively large, resulting in low encoding and decoding performance.
By decoding the video code stream, the pattern index information of the mode list is obtained, and then sorted from low to high according to the cost of the angle-weighted prediction mode, and the corresponding angle-weighted prediction mode is selected to reduce the encoding bit overhead of the index information.
It effectively reduces the encoding bit overhead of the angle-weighted prediction mode index information and improves the video encoding and codec performance.
Smart Images

Figure CN2024131393_19062025_PF_FP_ABST
Abstract
Description
Video encoding and decoding method, device, computer-readable medium, and electronic device
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 12, 2023, with application number 202311706064.9 and invention name “Video Coding and Decoding Method, Device, Computer-Readable Medium and Electronic Device”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of computers and communications technology, and more specifically, to a video encoding and decoding method, apparatus, computer-readable medium, and electronic device. Background Art
[0003] In relevant audio and video standards (such as the second phase of AVS3), Angular Weighted Prediction AWP (Angular Weighted Prediction) and SAWP (Spatial Angular Weighted Prediction) technologies are used. These prediction technologies use weight masks to weight the two prediction blocks to achieve the combination of different parts of the prediction blocks. The coding block using angular weighted prediction technology needs to encode a mode index in the bitstream. The mode index can determine the reference weight configuration and weight prediction angle required to derive the weight matrix. Currently, AVS3 supports 56 angular weighted prediction modes, and encoding these mode indexes in the bitstream has a large bit overhead.
[0004] Summary of the Invention
[0005] The embodiments of the present application provide a video encoding and decoding method, apparatus, computer-readable medium, and electronic device, which can effectively reduce the coding bit overhead of angle-weighted prediction mode index information, thereby improving encoding and decoding performance.
[0006] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by practice of the present application.
[0007] According to one aspect of an embodiment of the present application, a video decoding method is provided, comprising: decoding a video code stream to obtain mode index information for a rearranged mode list; sorting the multiple angle-weighted prediction modes in order from low to high cost according to the current block using the multiple angle-weighted prediction modes to obtain the rearranged mode list; selecting a corresponding angle-weighted prediction mode from the rearranged mode list according to the mode index information, and deriving a weight matrix of the current block according to the selected angle-weighted prediction mode; performing weighted prediction according to the weight matrix to obtain a prediction value corresponding to the current block.
[0008] According to one aspect of an embodiment of the present application, a video encoding method is provided, including: sorting the multiple angle-weighted prediction modes in order from low to high according to the cost of using the multiple angle-weighted prediction modes for the current block to obtain a rearranged mode list; determining mode index information based on the rearranged mode list and the selected angle-weighted prediction mode; deriving a weight matrix of the current block based on the selected angle-weighted prediction mode; performing weighted prediction based on the weight matrix to obtain a prediction value corresponding to the current block, encoding the current block based on the prediction value, and encoding the mode index information in a video code stream.
[0009] According to one aspect of an embodiment of the present application, a video decoding device is provided, including: a decoding unit, configured to decode a video code stream to obtain mode index information for a rearranged mode list; a sorting unit, configured to sort the multiple angle-weighted prediction modes in order from low to high according to the cost of using the multiple angle-weighted prediction modes for the current block, to obtain the rearranged mode list; a selection unit, configured to select the corresponding angle-weighted prediction mode from the rearranged mode list according to the mode index information, and derive the weight matrix of the current block according to the selected angle-weighted prediction mode; a processing unit, configured to perform weighted prediction according to the weight matrix to obtain a prediction value corresponding to the current block.
[0010] According to one aspect of an embodiment of the present application, a video encoding device is provided, including: a sorting unit, configured to sort the multiple angle-weighted prediction modes in order from low to high according to the cost of using the multiple angle-weighted prediction modes for the current block, to obtain a rearranged mode list; a determination unit, configured to determine mode index information according to the rearranged mode list and the selected angle-weighted prediction mode; a calculation unit, configured to derive a weight matrix of the current block according to the selected angle-weighted prediction mode; an encoding unit, configured to perform weighted prediction according to the weight matrix, to obtain a prediction value corresponding to the current block, to encode the current block according to the prediction value, and to encode the mode index information in a video code stream.
[0011] According to one aspect of an embodiment of the present application, a computer-readable medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method described in the above embodiment is implemented.
[0012] According to one aspect of an embodiment of the present application, an electronic device is provided, comprising: one or more processors; and a storage device for storing one or more computer programs, wherein when the one or more computer programs are executed by the one or more processors, the electronic device implements the method described in the above embodiment.
[0013] According to one aspect of an embodiment of the present application, a computer program product is provided, the computer program product including a computer program stored in a computer-readable storage medium. A processor of an electronic device reads and executes the computer program from the computer-readable storage medium, causing the electronic device to perform the methods provided in the various optional embodiments described above.
[0014] In the technical solutions provided in some embodiments of the present application, mode index information for a rearranged mode list is obtained by decoding a video code stream, and then the angle-weighted prediction modes are sorted in order from low to high according to the cost of using multiple angle-weighted prediction modes for the current block to obtain a rearranged mode list. Then, the corresponding angle-weighted prediction mode is selected from the rearranged mode list according to the mode index information, so that the angle-weighted prediction mode with a higher probability of being selected (i.e., a lower cost) has a smaller index value in the rearranged mode list, thereby effectively reducing the coding bit overhead of the angle-weighted prediction mode index information, which is beneficial to improving the encoding and decoding performance.
[0015] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] FIG1 is a schematic diagram showing an exemplary system architecture to which the technical solutions of the embodiments of the present application can be applied;
[0017] FIG2 is a schematic diagram showing the placement of a video encoding device and a video decoding device in a streaming transmission system;
[0018] FIG3 shows a basic flow chart of a video encoder;
[0019] FIG4 shows a schematic diagram of a block division structure in the HEVC standard;
[0020] FIG5 shows a block partition structure diagram in the AVS3 standard;
[0021] FIG6 shows a schematic diagram of angular prediction direction in intra-frame prediction mode;
[0022] FIG7 shows a schematic diagram of intra-frame prediction;
[0023] FIG8 shows a schematic diagram of inter-frame prediction;
[0024] FIG9 shows a schematic diagram of the prediction process of the AWP mode;
[0025] FIG10 shows a schematic diagram of eight weight generation angles;
[0026] FIG11 is a schematic diagram showing seven reference weight prediction positions;
[0027] FIG12 shows a schematic diagram of angle partitioning in an angle-weighted prediction mode;
[0028] FIG13 shows a flowchart of a video decoding method according to an embodiment of the present application;
[0029] FIG14 shows a flowchart of a video encoding method according to an embodiment of the present application;
[0030] FIG15 is a diagram showing an image of a reference weight derivation function in the form of a sigmoid function according to an embodiment of the present application;
[0031] FIG16 is a diagram showing an image of a reference weight derivation function based on a hyperbolic tangent function according to one embodiment of the present application;
[0032] FIG17 is a diagram showing an image of a reference weight derivation function based on a cosine function according to an embodiment of the present application;
[0033] FIG18 shows a block diagram of a video decoding apparatus according to an embodiment of the present application;
[0034] FIG19 shows a block diagram of a video encoding apparatus according to an embodiment of the present application;
[0035] FIG20 shows a schematic structural diagram of a computer system suitable for implementing an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0036] Example embodiments will now be described in a more complete manner with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to these examples; rather, these embodiments are provided to make this application more comprehensive and complete and to fully convey the concepts of the example embodiments to those skilled in the art.
[0037] FIG1 shows a schematic diagram of an exemplary system architecture to which the technical solutions of the embodiments of the present application can be applied.
[0038] 1 , a system architecture 100 includes a plurality of terminal devices that can communicate with each other via, for example, a network 150. For example, the system architecture 100 may include a first terminal device 110 and a second terminal device 120 interconnected via the network 150. In the embodiment of FIG1 , the first terminal device 110 and the second terminal device 120 perform unidirectional data transmission.
[0039] For example, the first terminal device 110 can encode video data (such as a video picture stream captured by the terminal device 110) for transmission to the second terminal device 120 via the network 150. The encoded video data is transmitted in the form of one or more encoded video streams. The second terminal device 120 can receive the encoded video data from the network 150, decode the encoded video data to restore the video data, and display the video picture based on the restored video data.
[0040] In one embodiment of the present application, the system architecture 100 may include a third terminal device 130 and a fourth terminal device 140 for performing bidirectional transmission of encoded video data, such as during a video conference. For bidirectional data transmission, each of the third terminal device 130 and the fourth terminal device 140 may encode video data (e.g., a video image stream captured by the terminal device) for transmission to the other of the third terminal device 130 and the fourth terminal device 140 via a network 150. Each of the third terminal device 130 and the fourth terminal device 140 may also receive the encoded video data transmitted by the other of the third terminal device 130 and the fourth terminal device 140, decode the encoded video data to recover the video data, and display the video image on an accessible display device based on the recovered video data.
[0041] In the embodiment shown in FIG. 1 , the first terminal device 110 , the second terminal device 120 , the third terminal device 130 , and the fourth terminal device 140 may be servers or terminals, but the principles disclosed herein are not limited thereto.
[0042] A server can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. A terminal can be, but is not limited to, a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, intelligent voice interaction device, smartwatch, smart home appliance, vehicle-mounted terminal, aircraft, etc.
[0043] The network 150 shown in FIG1 represents any number of networks, including, for example, wired and / or wireless communication networks, for transmitting encoded video data between the first terminal device 110, the second terminal device 120, the third terminal device 130, and the fourth terminal device 140. The communication network 150 may exchange data using circuit-switched and / or packet-switched channels. The network may include a telecommunications network, a local area network, a wide area network, and / or the Internet. For the purposes of this application, unless otherwise explained below, the architecture and topology of the network 150 may not be relevant to the operations disclosed herein.
[0044] In one embodiment of the present application, FIG2 illustrates the placement of a video encoding device and a video decoding device in a streaming environment. The subject matter disclosed herein is equally applicable to other video-enabled applications, including, for example, video conferencing, digital television, and storing compressed video on digital media such as CDs, DVDs, and memory sticks.
[0045] The streaming system may include an acquisition subsystem 213, which may include a video source 201, such as a digital camera, that creates an uncompressed video picture stream 202. In one embodiment, the video picture stream 202 includes samples captured by the digital camera. The video picture stream 202 is depicted as a thicker line to emphasize the higher data volume of the video picture stream compared to the encoded video data 204 (or the encoded video stream 204). The video picture stream 202 may be processed by an electronic device 220, which includes a video encoding device 203 coupled to the video source 201. The video encoding device 203 may include hardware, software, or a combination of hardware and software to implement or embody various aspects of the disclosed subject matter, as described in greater detail below. The encoded video data 204 (or the encoded video stream 204) is depicted as a thinner line to emphasize the lower data volume of the encoded video data 204 (or the encoded video stream 204), which may be stored on a streaming server 205 for future use. One or more streaming client subsystems, such as client subsystem 206 and client subsystem 208 in FIG2 , can access streaming server 205 to retrieve copies 207 and 209 of encoded video data 204. Client subsystem 206 can include, for example, a video decoding device 210 in electronic device 230. Video decoding device 210 decodes the incoming copy 207 of the encoded video data and generates an output video picture stream 211 that can be presented on a display 212 (e.g., a screen) or another presentation device. In some streaming systems, the encoded video data 204, video data 207, and video data 209 (e.g., video code streams) can be encoded according to certain video encoding / compression standards.
[0046] It should be noted that the electronic device 220 and the electronic device 230 may include other components not shown in the figure. For example, the electronic device 220 may include a video decoding device, and the electronic device 230 may also include a video encoding device.
[0047] In one embodiment of the present application, taking the international video coding standards HEVC (High Efficiency Video Coding), VVC (Versatile Video Coding), and China's national video coding standard AVS as examples, after a video frame image is input, the video frame image will be divided into several non-overlapping processing units according to a block size, and each processing unit will perform similar compression operations. This processing unit is called CTU (Coding Tree Unit), or LCU (Largest Coding Unit). The CTU can be further divided into more refined parts to obtain one or more basic coding units CU (Coding Unit). CU is the most basic element in a coding link.
[0048] In another embodiment, this processing unit can also be called a coding tile (i.e., tile), which is a rectangular area of a multimedia data frame that can be independently decoded and encoded. In the AV1 standard, a coding tile can be further divided into more refined sub-blocks (SBs). The SB is the starting point for block division and can be further divided into multiple sub-blocks. The superblock is then further divided into one or more blocks. Each block is the most basic element in an encoding process. Optionally, an SB can contain several Bs.
[0049] The above division method for video frame images can be called block partition structure. The following introduces some concepts in the encoding process:
[0050] Predictive Coding: Predictive coding includes intra-frame prediction and inter-frame prediction. The original video signal is predicted by a selected reconstructed video signal to produce a residual video signal. The encoder needs to determine which predictive coding mode to use for the current coding unit (or coding block) and inform the decoder. Intra-frame prediction uses the predicted signal from a previously encoded and reconstructed region within the same image; inter-frame prediction uses the predicted signal from a previously encoded image (called a reference image) that is different from the current image.
[0051] Transform & Quantization: After the residual video signal undergoes transformations such as the Discrete Fourier Transform (DFT) and Discrete Cosine Transform (DCT), the signal is converted to the transform domain, where the coefficients are known as transform coefficients. The transform coefficients are then subjected to a lossy quantization operation, which loses some information, making the quantized signal more suitable for compression. In some video coding standards, more than one transform scheme may be available. Therefore, the encoder must select one for the current coding unit (or coding block) and inform the decoder of this selection. The level of quantization is typically determined by the quantization parameter (QP). A larger QP value means that a wider range of coefficients will be quantized to the same output, which typically results in greater distortion and a lower bitrate. Conversely, a smaller QP value means that a smaller range of coefficients will be quantized to the same output, which typically results in less distortion and a higher bitrate.
[0052] Entropy Coding or Statistical Coding: The quantized transform domain signal is statistically compressed and encoded based on the frequency of occurrence of each value, and finally a binary (0 or 1) compressed code stream is output. At the same time, the encoding generates other information, such as the selected coding mode, motion vector data, etc., which also need to be entropy coded to reduce the bit rate. Statistical coding is a lossless coding method that can effectively reduce the bit rate required to express the same signal. Common statistical coding methods include variable length coding (VLC) or context-based binary arithmetic coding (CABAC).
[0053] The context-based binary arithmetic coding (CABAC) process consists of three main steps: binarization, context modeling, and binary arithmetic coding. After binarization of the input syntax elements, the binary data can be encoded using either the normal coding mode or the bypass coding mode. The bypass coding mode eliminates the need to assign a specific probability model to each binary bit. Instead, the input binary bit bin values are directly encoded using a simple bypass encoder, speeding up both encoding and decoding. Generally, different syntax elements are not completely independent, and even the same syntax elements have some memory. Therefore, according to conditional entropy theory, conditional coding using other coded syntax elements can further improve coding performance compared to independent encoding or memoryless coding. This coded symbol information used as a condition is called context. In the normal coding mode, the binary bits of the syntax elements are sequentially fed into the context modeler. The encoder assigns an appropriate probability model to each input binary bit based on the values of previously coded syntax elements or binary bits. This process is known as context modeling. The context model corresponding to the syntax element can be located using ctxIdxInc (context index increment) and ctxIdxStart (context index Start). After the bin value and the assigned probability model are fed into the binary arithmetic encoder for encoding, the context model needs to be updated based on the bin value, which is the adaptive process in encoding.
[0054] Loop Filtering: The changed and quantized signal will be reconstructed through inverse quantization, inverse transformation and prediction compensation operations to obtain a reconstructed image. Compared with the original image, due to the influence of quantization, some information of the reconstructed image is different from the original image, that is, the reconstructed image will produce distortion. Therefore, the reconstructed image can be filtered, such as deblocking filter (DB), SAO (Sample Adaptive Offset) or ALF (Adaptive Loop Filter) and other filters, which can effectively reduce the degree of distortion caused by quantization. Since these filtered reconstructed images will be used as a reference for subsequent encoded images to predict future image signals, the above filtering operation is also called loop filtering, that is, filtering operation within the encoding loop.
[0055] In one embodiment of the present application, FIG3 shows a basic flow chart of a video encoder, in which intra-frame prediction is used as an example for explanation. k[x,y] and predicted image signal Perform difference operation to obtain the residual signal u k [x,y], residual signal u k [x,y] is transformed and quantized to obtain the quantized coefficients. The quantized coefficients are entropy coded to obtain the encoded bit stream, and the reconstructed residual signal u' is obtained by inverse quantization and inverse transformation. k [x,y], predicted image signal and the reconstructed residual signal u' k [x,y] superposition generates image signal Image signal On the one hand, it is input to the intra-frame mode decision module and the intra-frame prediction module for intra-frame prediction processing, and on the other hand, the reconstructed image signal s' is output through loop filtering. k [x,y], reconstructed image signal s' k [x,y] can be used as the reference image for the next frame for motion estimation and motion compensation prediction. Then based on the result of motion compensation prediction s' r [x+m x ,y+m y ] and intra prediction results Get the predicted image signal of the next frame And continue to repeat the above process until the encoding is completed.
[0056] Based on the above encoding process, at the decoding end, after obtaining the compressed bitstream (i.e., bitstream), entropy decoding is performed on each coding unit (or coding block) to obtain various mode information and quantization coefficients. The quantized coefficients are then dequantized and inversely transformed to produce a residual signal. Furthermore, based on the known coding mode information, a prediction signal corresponding to the coding unit (or coding block) can be obtained. The residual signal is then added to the prediction signal to produce a reconstructed signal. The reconstructed signal then undergoes loop filtering and other operations to produce the final output signal.
[0057] Current mainstream video coding standards, such as HEVC (High Efficiency Video Coding), VVC (Versatile Video Coding), AVS3, AV1 (Alliance for Open Media Video 1), and AV2 (Alliance for Open Media Video 2), all employ a block-based hybrid coding framework. Specifically, this involves dividing the raw video data into a series of coding blocks and combining them with prediction, transform, and entropy coding techniques to achieve video data compression.
[0058] In a block-based hybrid coding framework, a video image is divided into several non-overlapping processing units for video compression. This processing unit is called a CTU (Coding Tree Unit). The CTU can be further divided into one or more basic coding units, called CUs (Coding Units). Each CU is the most basic element in the coding process, and each CU can select different coding modes. Figure 4 shows a schematic diagram of the block division structure in the HEVC standard. A CTU can be divided downward using a quadtree.
[0059] The AVS3 standard uses a basic block partitioning structure of QT (Quad-Tree) + BT (Binary-Tree) + EQT (Extended Quad-Tree). Specifically, the representation of the QT+BT+EQT basic block partitioning structure in the bitstream of AVS3 is shown in Figure 5. For a CU, the first step is to determine whether to use QT for partitioning. If QT is used, QT partitioning is performed directly. If not, a further determination is made whether to not partition. If not, the process ends. If partitioning is required, a determination is made whether to use EQT or BT. Furthermore, whether EQT or BT is used requires a determination of whether to partition horizontally or vertically. Block partitioning is a recursive process starting from the LCU and proceeding downwards. During this recursive process, the optimal partitioning method and encoding mode are determined through optimization on the encoder side.
[0060] Intra-frame prediction is a commonly used predictive coding technology. It is based on the correlation between the pixels of the video image in the spatial domain and derives the predicted value of the current coding block from the adjacent coded areas. The second phase of AVS3 adopts the extended intra-frame angle prediction mode (Extended Intra Prediction Mode, abbreviated as EIPM). In the previous generation AVS2, there are 33 intra-frame prediction modes, including 30 angle prediction modes and 3 special prediction modes (Plane prediction mode, DC prediction mode and Bilinear prediction mode). They are encoded using 2 MPMs (Most Probable Modes), and the remaining modes are encoded using 5-bit fixed-length encoding. To support more refined angle prediction, AVS3 expands the angle prediction modes to 62, as shown in Figure 6. The newly added angle prediction modes are numbered 34 to 65.
[0061] When the angle prediction mode is used, the pixel points in the current prediction block use the reference pixel value of the corresponding position on the reference pixel row or column as the prediction value according to the direction corresponding to the angle of the prediction mode. As shown in Figure 7, for the pixel point P in the prediction block, the position of the reference pixel will be determined from the pixel row that has been encoded above according to the prediction angle in the figure, and then the reference pixel value will be used as the prediction value of the pixel point P. It should be noted that not all pixel positions point to reference pixel positions with integer pixel accuracy. For example, the reference pixel position of the pixel point P in Figure 7 is a sub-pixel position between pixels B and C, so the predicted pixel value of this position needs to be obtained by interpolation of the surrounding pixels. In order to improve the efficiency of intra-frame prediction, on-chip memory is usually used to store the reference pixels of intra-frame prediction.
[0062] As shown in Figure 8, inter-frame prediction uses the correlation in the video time domain to predict the pixels of the current image using the pixels of the adjacent coded images to effectively remove the temporal redundancy of the video and effectively save the bits of the coded residual data. Wherein, P represents the current frame, Pr represents the reference frame, B represents the current coded block, and Br represents the reference block of B. The coordinates of B' in the reference frame are the same as the coordinates of B in the current frame. The coordinates of Br are (x r ,y r ), the coordinates of B' are (x, y), and the displacement between the current coding block and its reference block is called the motion vector (MV), where MV = (x r -x,y r -y).
[0063] The second phase of AVS3 also adopted the AWP mode for inter-frame prediction and the SAWP mode for intra-frame prediction. As shown in Figure 9, the AWP mode uses the intra-frame angle prediction idea to derive the weight value of each pixel position: first set the reference weight value of the current block's surrounding positions (whole pixel positions and sub-pixel positions), and then use the angle prediction method to obtain the weight value corresponding to each pixel position, and then use the obtained weight to achieve weighted prediction of two different inter-frame prediction values. SAWP uses a similar method to derive weights to achieve weighted prediction of two intra-frame prediction values. The following is an introduction to the AWP mode:
[0064] The angle-weighted prediction mode supports a minimum block size of 8 and a maximum block size of 64, with a total of 8 angles supported. As shown in Figure 10, these 8 angles have five slope absolute values: {horizontal, vertical, 1, 2, 1 / 2}. As shown in Figure 11, each angle supports 7 reference weight configurations, so for each block, the angle-weighted mode has a total of 56 modes. The reference weight configuration is a distribution function of the reference weight value obtained based on the reference weight index value. The non-strictly monotonically increasing function is assigned based on the 8-equal-division point position of the reference weight effective length, where the reference weight effective length is calculated from the prediction angle and the current block size.
[0065] As shown in Figure 12, the angles supported by the angle-weighted prediction mode are divided into four partitions: angle partition 0, angle partition 1, angle partition 2, and angle partition 3. The formula for deriving pixel-by-pixel weights varies slightly depending on the angle region. Specifically, the bitstream contains the AwpIndex field, which indicates the weighting mode used for the current coding block. The following formula is used to determine the relevant parameters of the angle-weighted mode based on the AwpIndex field.
[0066] stepIndex=(AwpIndex>>3)-3
[0067] modAngNum=AwpIndex%8
[0068] angleAreaIndex=modAngNum>>1
[0069] Wherein, the “>>” in the above formula represents a right shift operation. After obtaining the above parameters, the weight matrix of the angle weighting mode can be derived according to these parameters in the following way:
[0070] First, calculate the effective length vL of the reference weight. The length vL of the reference weight is expressed using 1 / 2 pixel precision. For the first four angles (i.e., when the angle area indicator angleAreaIndex is equal to 0 and 1), the reference weight is 1 column to the left of the current block; for the last four angles (i.e., when the angle area indicator angleAreaIndex is equal to 2 and 3), the reference weight is 1 row above the current block. If the width of the current block is W and the height is H, then the effective length vL is calculated as shown in Table 1:
[0071] Table 1
[0072] In Table 1, "<<" indicates a left shift operation. After calculating the effective length vL of the reference weight, the reference weight Lw is filled in at each sample position x (x ≤ vL) as follows. PictureAwpRefineIndex is the image header index that controls whether the reference weight is adjusted.
[0073] Lw[x]=Clip3(0,8,(x-fP)<<shift)
[0074] shift=PictureAwpRefineIndex? 2:0
[0075] o=PictureAwpRefineIndex? 3:1
[0076] The value of fP is determined according to the following Table 2:
[0077] Table 2
[0078] Secondly, fill the luminance weight matrix according to the reference weight. Let the luminance weight matrix be BwLuma(x,y), then BwLuma(x,y)=Lw[tP].
[0079] For the SAWP mode or the AWP mode in the B frame, the value of tP is determined according to the following Table 3:
[0080] Table 3
[0081] For the AWP mode in P frames, the value of tP is determined according to the following Table 4:
[0082] Table 4
[0083] After obtaining the luminance weight matrix, the chrominance weight matrix BwChroma(x,y) is filled according to the luminance weight matrix.
[0084] For the SAWP mode or the AWP mode in the B frame, the chroma weight matrix is derived according to the following formula:
[0085] BwChroma[x][y]=BwLuma[x<<1][y<<1]
[0086] For the AWP mode in P frames, the chrominance weight matrix is derived according to the following formula:
[0087] BwChroma[x][y]=BwLuma[(x>>2)<<3][(y>>2)<<3]
[0088] Finally, the weighted prediction value pred is calculated based on the derived weight matrix and the predicted value x,y .
[0089] pred x,y =predA x,y *w x,y +predB x,y *(1-w x,y )
[0090] Among them, predA x,y with predB x,y Represents the two predicted values of weighted prediction, w x,y Indicates the weight value at (x, y). When the weight is 0, the predicted value is predB x,y ; When the weight is the maximum value 1, the predicted value is predA x,y .
[0091] In video coding, weights are quantized into integers to reduce floating-point operations. The weight range is set to [0, m]. The value of m is set according to the precision of the weight. For example, if the weight is represented by 3 bits, then m = 8. The weighted prediction value pred x,y The formula can be expressed as:
[0092] pred x,y =(predA x,y *w x,y +predB x,y *(8-w x,y )+4)>>3
[0093] Among them, when the weight is 0, the predicted value is predB x,y ; When the weight is the maximum value m, the predicted value is predA x,y .
[0094] The weight matrix w in AWP and SAWP x,yAccording to the reference weight, the range of the mixed area of the reference weight is L, and its range can be confirmed according to the starting position and the end position (p0, p1). The values of p0 and p1 can be the same or different. i It can be derived according to the following formula:
[0095] or
[0096] It can be seen from the above formula that the reference weight w=f(d) in the mixed area.
[0097] In the specific implementation, for functions whose value range is larger than the weight value range, the clip function can be used to clip the function, that is, w = clip(0,m,f(d)). Alternatively, integer weights can be used to reduce complexity.
[0098] Assume that the position of the dividing boundary is c, and the sample point x on the reference weight is i The distance to the mixed area is d. d can be determined based on the distance between the current sample point and the boundary of the mixed area, that is, d = x i -c.
[0099] Alternatively, d can also be calculated based on the range of the blending area. Assuming that offset is the offset from the center position to the starting position, d can be expressed as: d = x i -p0+offset.
[0100] Wherein, d can be represented by a preset precision integer, for example, d can be represented by, but not limited to, 8-pixel precision, 4-pixel precision, 2-pixel precision, 1-pixel precision, 1 / 2-pixel precision, 1 / 4-pixel precision, or 1 / 64-pixel precision. At the same time, the weight quantized into an integer can be obtained according to the following formula: q =round(m*f(d)); or use the following formula to calculate:
[0101] w q =clip(0,m,round(m*f(d q )))
[0102] Among them, m is the maximum weight, round() is the rounding function, and f(d) is the function derived from the reference weight.
[0103] Since the coding block using the angle-weighted prediction technology needs to encode a mode index in the bitstream, the mode index can determine the reference weight configuration and weight prediction angle required to derive the weight matrix. Currently, AVS3 supports 56 angle-weighted prediction modes, and encoding these mode indexes in the bitstream has a large bit overhead. Based on this, the technical solution of the embodiment of the present application proposes that the angle-weighted prediction modes can be rearranged so that the angle-weighted prediction modes with a higher probability of being selected have smaller index values in the mode list, thereby effectively reducing the coding bit overhead of the angle-weighted prediction modes, thereby improving the video encoding and decoding efficiency.
[0104] The following is a detailed description of the implementation details of the technical solution of the embodiment of the present application:
[0105] FIG13 shows a flowchart of a video decoding method according to an embodiment of the present application. The video decoding method can be performed by a device with computing and processing capabilities, such as a terminal device or a server. Referring to FIG13 , the video decoding method includes at least steps S1310 to S1340, which are described in detail as follows:
[0106] In step S1310 , the video stream is decoded to obtain mode index information for the rearranged mode list.
[0107] In one embodiment of the present application, a video code stream is a code stream obtained by encoding a sequence of video image frames. The video image frame sequence includes a series of images, each of which can be further divided into slices or patches, and the slices can be divided into a series of LCUs (or CTUs), each of which contains several CUs. The video image frame is encoded in blocks. The coding block or current block in the embodiment of the present application can be a CU, or a block smaller than a CU, such as a smaller block obtained by dividing the CU.
[0108] In some optional embodiments, at least one of the following index information obtained by decoding a video bitstream may be used as the mode index information: index information indicating an angle-weighted prediction mode, index information indicating a reference weight configuration, index information indicating a prediction angle, index information indicating a reference weight derivation method, and index information indicating a weight mode for the angle-weighted prediction mode. The rearranged mode list is a mode list obtained by rearranging one or more index information included in the mode index information.
[0109] Alternatively, the mode index information may be index information for indicating the angle-weighted prediction mode. In this case, the rearranged mode list is a mode list obtained by rearranging the index information of the angle-weighted prediction mode.
[0110] Optionally, the pattern index information may be index information for indicating a reference weight configuration and index information for indicating a prediction angle. In this case, the rearranged pattern list may be a pattern list obtained by rearranging one or more of the index information for indicating a reference weight configuration and the index information for indicating a prediction angle.
[0111] For example, the index information used to indicate the reference weight configuration is recorded as step_idx, and the index information used to indicate the prediction angle is recorded as angle_idx. Currently, there are 56 AWP modes, namely 7 weight configurations (step_idx) × 8 weight prediction angles (angle_idx). The original arrangement of these 56 AWP modes is 8 weight prediction angles corresponding to the first weight configuration, 8 weight prediction angles corresponding to the second weight configuration, ..., 8 weight prediction angles corresponding to the seventh weight configuration. Then the mode index information awp_mode_idx = step_idx × angle_num + angle_idx, where angle_num is the number of prediction angles.
[0112] If, when decoding the mode index information, step_idx is obtained without reordering, and the weighted angle index for the reordered mode list is obtained by decoding, it is recorded as angle_idx_org. Then, during the subsequent reordering, the corresponding angle-weighted prediction mode step_idx×angle_idx is calculated for the eight different angle_idx values, and the angle_idx is reordered using template matching. After the reordering, the actual weighted angle index used can be found from the reordered mode list based on angle_idx_org.
[0113] Similarly, the step_idx can be rearranged in the same manner without rearranging the angle_idx. Other combinations of the pattern index information in the above embodiment can also be processed in a similar manner.
[0114] Alternatively, the mode index information may be a combination of index information indicating the reference weight derivation method and index information indicating the angle-weighted prediction mode. In this case, the rearranged mode list may be a mode list obtained by rearranging one or more of the index information indicating the reference weight derivation method and the index information indicating the reference weight derivation method.
[0115] For example, the index information indicating the reference weight derivation method is recorded as blend_idx and the index information indicating the angle weighted prediction mode is recorded as awp_mode_idx. Then a similar method can be used to reorder the angle weighted prediction mode based on awp_mode_idx instead of reordering the blend_idx when decoding.
[0116] In some optional embodiments, other index information can be derived based on the index information indicating the angle-weighted prediction mode. For example, the index information indicating the angle-weighted prediction mode is recorded as awp_mode_idx, the index information indicating the reference weight configuration is recorded as step_idx, and the index information indicating the prediction angle is recorded as angle_idx. Then, other index information can be derived based on awp_mode_idx in the following manner:
[0117] angle_idx=awp_mode_idx%awp_angle_num;
[0118] step_idx=awp_mode_idx / awp_angle_num;
[0119] Among them, awp_angle_num=8.
[0120] Optionally, the pattern index information may be index information for indicating a reference weight derivation method, index information for indicating a reference weight configuration, and index information for indicating a prediction angle. In this case, the rearranged pattern list may be a pattern list obtained by rearranging one or more of the index information for indicating a reference weight derivation method, the index information for indicating a reference weight configuration, and the index information for indicating a prediction angle.
[0121] Optionally, the mode index information may be index information for indicating a weight mode of an angle-weighted prediction mode. In this case, the rearranged mode list may be a mode list obtained by rearranging the index information for indicating a weight mode of an angle-weighted prediction mode. Among them, the index information for indicating the weight mode of an angle-weighted prediction mode may be used to derive index information for indicating an angle-weighted prediction mode and index information for indicating a reference weight derivation method. The index information for indicating an angle-weighted prediction mode may be used to derive index information for indicating a reference weight configuration and index information for indicating a prediction angle. For example, the index information for indicating the weight mode of an angle-weighted prediction mode is recorded as cu_awp_blend_mode_idx, the index information for indicating the angle-weighted prediction mode is recorded as awp_mode_idx, the index information for indicating the reference weight derivation method is recorded as blend_idx, the index information for indicating the reference weight configuration is recorded as step_idx, and the index information for indicating the prediction angle is recorded as angle_idx. Then, other index information may be derived in the following manner:
[0122] awp_mode_idx=cu_awp_blend_mode_idx%awp_mode_num;
[0123] blend_idx=cu_awp_blend_mode_idx / awp_mode_num;
[0124] angle_idx=awp_mode_idx%awp_angle_num;
[0125] step_idx=awp_mode_idx / awp_angle_num;
[0126] Among them, awp_angle_num=8; awp_mode_num=56.
[0127] In step S1320 , the multiple angle weighted prediction modes are sorted according to the order of the costs of using the multiple angle weighted prediction modes for the current block from low to high, to obtain a rearranged mode list.
[0128] Optionally, the angle weighted prediction mode in the embodiment of the present application may be an AWP mode or a SAWP mode.
[0129] In some optional embodiments, in step S1320, an index flag included in the video bitstream is decoded; if the index flag determines that the current block allows the use of the rearranged angle-weighted prediction mode, the multiple angle-weighted prediction modes are sorted in ascending order of cost to obtain a rearranged mode list. In other words, before sorting the angle-weighted prediction modes, the index flag included in the video bitstream may also be decoded; if the index flag determines that the current block allows the use of the rearranged angle-weighted prediction mode, the process of sorting the multiple angle-weighted prediction modes is executed.
[0130] In some optional embodiments, whether the current sequence allows the use of the rearrangement-based angle weighted prediction mode can be determined based on a sequence header flag bit included in the sequence header information. For example, if the value of the sequence header flag bit is 1, it indicates that the current sequence allows the use of the rearrangement-based angle weighted prediction mode; if the value of the sequence header flag bit is 0, it indicates that the current sequence does not allow the use of the rearrangement-based angle weighted prediction mode.
[0131] In some optional embodiments, whether the current image allows the use of the rearrangement-based angle weighted prediction mode can be determined based on an image header flag included in the image header information. For example, if the value of the image header flag is 1, it indicates that the current image allows the use of the rearrangement-based angle weighted prediction mode; if the value of the image header flag is 0, it indicates that the current image does not allow the use of the rearrangement-based angle weighted prediction mode.
[0132] In some optional embodiments, whether the current slice allows the use of the rearrangement-based angle weighted prediction mode can be determined based on a slice header flag included in the slice header information. For example, if the value of the slice header flag is 1, it indicates that the current slice allows the use of the rearrangement-based angle weighted prediction mode; if the value of the slice header flag is 0, it indicates that the current slice does not allow the use of the rearrangement-based angle weighted prediction mode.
[0133] In some optional embodiments, whether the rearrangement-based angle weighted prediction mode is allowed can also be determined based on two or more of the sequence header flag contained in the sequence header information, the image header flag contained in the image header information, and the slice header flag contained in the slice header information.
[0134] For example, it is possible to determine whether a coding block using a weighted prediction mode is allowed to use a rearrangement-based angle weighted prediction mode based on the sequence header flag and the image header flag. Specifically, if the value of the sequence header flag and the value of the image header flag are both 1, it means that the current image is allowed to use a rearrangement-based angle weighted prediction mode; if the value of the sequence header flag is 1 and the value of the image header flag is 0, it means that the current image is not allowed to use a rearrangement-based angle weighted prediction mode; if the value of the sequence header flag is 0, regardless of the value of the image header flag (in fact, there is no need to decode the value of the image header flag), it can be considered that the current sequence is not allowed to use a rearrangement-based angle weighted prediction mode.
[0135] In some optional embodiments, the index flag used to determine whether the current block is allowed to use the rearrangement-based angle weighted prediction mode may include one flag bit or multiple flag bits; the different values of this flag bit or multiple flag bits are used to indicate whether the corresponding coding block is allowed to use the angle weighted prediction mode, and when the angle weighted prediction mode is allowed, whether the rearrangement-based angle weighted prediction mode is allowed.
[0136] For example, the index flag includes a flag, and when the flag value is 0, it indicates that the corresponding coding block is not allowed to use the angle weighted prediction mode; when the flag value is 1, it indicates that the corresponding coding block is allowed to use the angle weighted prediction mode, but is not allowed to use the rearrangement-based angle weighted prediction mode; when the flag value is 2, it indicates that the corresponding coding block is allowed to use the angle weighted prediction mode, and is allowed to use the rearrangement-based angle weighted prediction mode.
[0137] For another example, the index flag contains two flags (denoted as the first flag and the second flag). If the first flag is 0, it is possible to determine that the corresponding coding block does not allow the use of the angle-weighted prediction mode without decoding the second flag; if the first flag is 1 and the second flag is 0, it is determined that the corresponding coding block allows the use of the angle-weighted prediction mode, but does not allow the use of the rearrangement-based angle-weighted prediction mode; if the first flag is 1 and the second flag is 1, it is determined that the corresponding coding block allows the use of the angle-weighted prediction mode and allows the use of the rearrangement-based angle-weighted prediction mode.
[0138] In some optional embodiments, the index flag may be obtained by decoding the video stream when the image frame type is a specified type (e.g., an I-frame image or a B-frame image). Alternatively, the index flag may be obtained by decoding the video stream when the image frame type is not a specified type (e.g., a non-I-frame image).
[0139] In one embodiment of the present application, if it is determined that the current block allows the use of the angle-weighted prediction mode, syntax elements for deriving the weight matrix and syntax elements for determining multiple prediction values can be decoded from the video code stream.
[0140] In some optional embodiments, the syntax elements used to derive the weight matrix include at least one of the following syntax elements: index information for indicating the angle-weighted prediction mode, index information for indicating the reference weight configuration, index information for indicating the prediction angle, index information for indicating the reference weight derivation method, index information for indicating the weight pattern of the angle-weighted prediction mode, and index information for indicating the size of the reference weight mixing area.
[0141] Alternatively, the syntax element used to derive the weight matrix may be index information indicating the angle-weighted prediction mode.
[0142] Optionally, the syntax elements used to derive the weight matrix may be index information for indicating a reference weight derivation method and index information for indicating an angle-weighted prediction mode.
[0143] Optionally, the syntax elements used to derive the weight matrix may be index information for indicating a reference weight derivation method, index information for indicating a reference weight configuration, and index information for indicating a prediction angle.
[0144] Alternatively, the syntax element used to derive the weight matrix may be index information indicating a weight mode of the angle-weighted prediction mode.
[0145] Optionally, the syntax elements used to derive the weight matrix may be index information for indicating a reference weight configuration and index information for indicating a prediction angle.
[0146] In some optional embodiments, the syntax elements for determining multiple prediction values include at least one of the following syntax elements: index information for determining the predicted motion vector of the reference block, syntax elements for determining the correction of the motion vector, and syntax elements for determining the intra-frame prediction mode (the SAWP mode requires the syntax elements for determining the intra-frame prediction mode).
[0147] Optionally, the syntax elements used to determine whether the motion vector needs to be corrected include: index information for indicating whether the motion vector needs to be corrected, index information for indicating the motion vector correction step size, and index information for indicating the motion vector correction direction. The index information for indicating the motion vector correction step size and the index information for indicating the motion vector correction direction are primarily used to derive MVD (Motion Vector Difference), such that motion vector MV = MVP (Motion Vector Predictor) + MVD.
[0148] Optionally, the index information indicating whether the motion vector needs to be corrected and the index information indicating the motion vector difference are used. That is, in this embodiment, the MVD can be directly obtained by decoding the bitstream.
[0149] In some optional embodiments, all or part of the syntax elements used to derive the weight matrix and the syntax elements used to determine the multiple prediction values can be decoded in a specified manner. The following is a detailed description:
[0150] Optionally, some or all of the binary bits of these syntax elements may be decoded using variable-length codes, such as K-order Exponential Golomb codes, truncated unary codes, truncated binary codes, etc.
[0151] Optionally, part or all of the binary bits of these syntax elements may be decoded using a fixed-length code.
[0152] Optionally, different parts of the binary bits of these syntax elements may be decoded using different variable-length codes, for example, the prefix part may be decoded using a truncated unary code, and the suffix part may be decoded using a truncated binary code.
[0153] Optionally, the portion of these syntax elements whose binary bits are smaller than a set threshold may be decoded using a decoding method corresponding to the context-based binary encoding method, and the remaining portion may be decoded using a decoding method corresponding to the bypass encoding method.
[0154] Optionally, some or all of the binary bits of these syntax elements can be decoded using a combination of variable-length codes and fixed-length codes. For example, the prefix portion is decoded using a variable-length code, and the suffix portion is decoded using a fixed-length code. For another example, if the number of decoded angle-weighted prediction modes (less than or equal to the total number of AWP modes) is divided into multiple groups according to a set grouping method, the index information indicating the group number in the mode index information can be decoded using a variable-length code, and the index information indicating the elements within the group in the mode index information can be decoded using a fixed-length code.
[0155] In one embodiment of the present application, before sorting the angle-weighted prediction modes, it is necessary to calculate the cost of using various angle-weighted prediction modes for the current block. Specifically, it is necessary to traverse the various decoded angle-weighted prediction modes, then obtain the current template corresponding to the current block, and obtain the prediction template corresponding to the reference block of the current block. Then, based on the weights derived from each angle-weighted prediction mode, determine the template weight corresponding to the prediction template, and calculate the weighted prediction template based on the template weight and the prediction template. Finally, based on the weighted prediction template and the current template, calculate the cost of each angle-weighted prediction mode.
[0156] In some optional embodiments, the current template corresponding to the current block includes at least one of the following sample points: sample points located in a set row above the current block, sample points located in a set column to the left of the current block, sample points obtained by sampling adjacent sample points of the current block according to a set interval, and sample points located in a set row above the current block and in a set column to the left of the current block.
[0157] In some optional embodiments, the prediction template corresponding to the reference block of the current block includes at least one of the following samples: samples located in a set row above the reference block, samples located in a set column to the left of the reference block, samples located in a set row within the reference block, samples located in a set column within the reference block, samples obtained by sampling adjacent samples of the reference block according to a set interval, samples obtained by sampling samples within the reference block according to a set interval, samples located in a set row above the reference block and in a set column to the left of the reference block, and samples located in a set row above the reference block and in a set column to the left of the reference block.
[0158] It should be noted that the prediction template may correspond to the current template, that is, the position of the prediction template relative to the reference block is the same as the position of the current template relative to the current block. Alternatively, the prediction template and the current template may not correspond.
[0159] In some optional embodiments, after selecting a current template, the samples in the current template may be corrected. Furthermore, after selecting a prediction template, the samples in the prediction template may also be corrected. Optionally, the correction processing of the samples includes one or more of filtering, linear mapping, and nonlinear mapping.
[0160] In some optional embodiments, for an angle-weighted prediction mode whose derived weight is greater than a weight threshold, the template weight is set to 1, and for an angle-weighted prediction mode whose derived weight is less than the weight threshold, the template weight is set to 0; or if the position of the sample point corresponding to the prediction template after projection according to the weighted prediction angle is to the left of the center position of the mixed region, the template weight is set to 0; otherwise (i.e., the position of the sample point corresponding to the prediction template after projection according to the weighted prediction angle is to the right of the center position of the mixed region or overlaps with the center position), the template weight is set to 1. In other words, after deriving the weight corresponding to the prediction template based on each angle-weighted prediction mode, the template weight can also be determined using the derived weight.
[0161] In some optional embodiments, when calculating the cost of the angle-weighted prediction mode, the sum of absolute differences (SAD), or the sum of squares of differences (SSD), or the mean-reduced SAD (MR-SAD) between the weighted prediction template and the current template can be calculated as the cost of the angle-weighted prediction mode.
[0162] In some optional embodiments, when calculating the cost, only samples whose differences are less than a set threshold may be calculated. That is, the sum of absolute errors, the sum of squares of the differences, or the mean reduced sum of absolute errors between the weighted prediction template and the target samples in the current template whose differences are less than the set threshold is calculated.
[0163] In some optional embodiments, when sorting the angle-weighted prediction modes, if the number of angle-weighted prediction modes sorted in descending order of cost reaches a set number, the sorting may be stopped; wherein the set number is less than or equal to the total number of angle-weighted prediction modes. Optionally, the set number may be the number of decoded angle-weighted prediction modes.
[0164] In step S1330, a corresponding angle-weighted prediction mode is selected from the rearranged mode list according to the mode index information, and a weight matrix of the current block is derived according to the selected angle-weighted prediction mode.
[0165] In some optional embodiments, the process of deriving the weight matrix for the current block based on the selected angle-weighted prediction mode may include deriving reference weights for the current block based on the selected angle-weighted prediction mode, and then calculating the weight matrix corresponding to the current block based on the reference weights and the weight prediction angle used by the current block. The specific calculation process can be referred to the description in the aforementioned embodiment and will not be repeated here.
[0166] In step S1340, weighted prediction is performed according to the weight matrix to obtain a prediction value corresponding to the current block.
[0167] The process of performing weighted prediction according to the weight matrix to obtain the predicted value corresponding to the current block can refer to the introduction in the aforementioned embodiment and will not be repeated here.
[0168] FIG14 shows a flowchart of a video encoding method according to an embodiment of the present application. The video encoding method can be performed by a device with computing and processing capabilities, such as a terminal device or a server. Referring to FIG14 , the video encoding method includes at least steps S1410 to S1440, which are described in detail as follows:
[0169] In step S1410 , the multiple angle weighted prediction modes are sorted according to the order of the costs of using the multiple angle weighted prediction modes for the current block from low to high, to obtain a rearranged mode list.
[0170] In step S1420, mode index information is determined according to the rearranged mode list and the selected angle-weighted prediction mode.
[0171] In step S1430 , a weight matrix of the current block is derived according to the selected angle-weighted prediction mode.
[0172] In step S1440, weighted prediction is performed according to the weight matrix to obtain a prediction value corresponding to the current block, the current block is encoded according to the prediction value, and the mode index information is encoded in the video code stream.
[0173] It should be noted that the specific processing process at the encoding end is similar to that at the decoding end and will not be repeated here.
[0174] In summary, the technical solution of the embodiment of the present application enables the angle-weighted prediction mode with a higher probability of being selected (i.e., lower cost) to have a smaller index value in the rearranged mode list, thereby effectively reducing the coding bit overhead of the angle-weighted prediction mode index information, which is beneficial to improving the encoding and decoding performance.
[0175] The following is a detailed description of the implementation details of the technical solution of the embodiment of the present application from the perspective of the decoding end:
[0176] The technical solution of the embodiment of the present application mainly includes the following four processes: Process 1, decoding the code stream, determining whether the coding block uses the angle-weighted prediction mode; Process 2, if the angle-weighted prediction mode is used, decoding the code stream, determining the relevant syntax elements of the angle-weighted prediction mode, including the syntax elements for determining the weight matrix and the syntax elements for determining the prediction values of each part; Process 3, if the rearrangement-based angle-weighted prediction mode is used, deriving the rearranged mode list and the corresponding prediction mode; Process 4, deriving the weight matrix according to the derived prediction mode, and performing weighted prediction.
[0177] The technical solution of the embodiment of the present application is applicable not only to the AWP mode, but also to the SWAP mode. The following takes the AWP mode as an example to describe the above four processes in detail:
[0178] Process 1: Decode the code stream and determine whether the coding block uses the angle-weighted prediction mode.
[0179] The following methods can be used individually or in combination:
[0180] In one embodiment of the present application, the code stream includes a sequence header syntax element indicating whether the current sequence uses the AWP mode and whether the rearrangement-based angle weighted prediction mode is used.
[0181] It should be noted that the rearranged angle-weighted prediction mode refers to the fact that the decoder needs to rearrange the angle-weighted prediction mode. Therefore, it can also be called the decoder's angle-weighted prediction mode (DecodeAWP, abbreviated as DAWP). The following uses the rearranged angle-weighted prediction mode as an example. Of course, the encoder also includes rearrangement operations to maintain consistency with the decoder.
[0182] In some optional embodiments, the bitstream may include two flags, such as the sequence header angle-weighted prediction mode flag seq_awp_flag and the sequence header rearrangement-based angle-weighted prediction mode flag seq_dawp_flag, to indicate whether the current sequence uses the AWP mode. The specific indication method can be shown in Table 5 ("x" in Table 5 indicates that no decoding is required):
[0183] Table 5
[0184] seq_awp_flag can be a binary variable with a value of 1 indicating that the current sequence allows the use of AWP mode; a value of 0 indicates that the use of AWP mode is not permitted. seq_dawp_flag can be a binary variable with a value of 1 indicating that the use of DAWP mode is permitted; a value of 0 indicates that the use of DAWP mode is not permitted. Referring to Table 5, when seq_awp_flag is 0, it can be determined that the current sequence does not allow the use of AWP mode without decoding seq_dawp_flag. When seq_awp_flag is 1 and seq_dawp_flag is 0, it can be determined that the current sequence allows the use of AWP mode but does not allow the use of DAWP mode. When both seq_awp_flag and seq_dawp_flag are 1, it can be determined that the current sequence allows the use of AWP mode and allows the use of DAWP mode.
[0185] As an example, the decoding process of seq_awp_flag and seq_dawp_flag is shown in Table 6:
[0186] Table 6
[0187] As shown in Table 6, first decode seq_awp_flag. If seq_awp_flag indicates that the current sequence allows the use of AWP mode, then decode seq_dawp_flag. If seq_awp_flag indicates that the current sequence does not allow the use of AWP mode, then there is no need to decode seq_dawp_flag. The descriptors of seq_awp_flag and seq_dawp_flag can both be u(1), that is, a 1-bit unsigned integer. The value of SeqAwpFlag is equal to the value of seq_awp_flag. If seq_awp_flag does not exist in the bitstream, the value of SeqAwpFlag is 0.
[0188] In some optional embodiments, there may be only one flag in the bitstream, namely the sequence header angle weighted prediction mode flag seq_awp_flag, to indicate whether the current sequence uses the AWP mode. The specific indication method may be as shown in Table 7:
[0189] Table 7
[0190] seq_awp_flag indicates the mode type of the angle-weighted prediction mode. A value of "0" indicates that AWP is not allowed. A value greater than "0" indicates that AWP can be used and indicates whether DAWP is allowed. The descriptor of seq_awp_flag can be u(2), a 2-bit unsigned integer, or ue(v), an unsigned integer syntax element.
[0191] As shown in Table 7, when seq_awp_flag is 0, it indicates that the current sequence does not allow the use of AWP mode; when seq_awp_flag is 1, it indicates that the current sequence allows the use of AWP mode and does not allow the use of DAWP mode; when seq_awp_flag is 2, it indicates that the current sequence allows the use of AWP mode and allows the use of DAWP mode.
[0192] In one embodiment of the present application, the code stream includes a picture header syntax element indicating whether the current picture uses the AWP mode and whether the rearrangement-based angle weighted prediction mode is used.
[0193] In some optional embodiments, the code stream may include two flags, such as the picture header angle-weighted prediction mode flag pic_awp_flag and the picture header rearrangement-based angle-weighted prediction mode flag pic_dawp_flag, to indicate whether the current picture uses the AWP mode. The specific indication method can be shown in Table 8 ("x" in Table 8 indicates that no decoding is required):
[0194] Table 8
[0195] Among them, pic_awp_flag can be a binary variable, with a value of 1 indicating that the current image allows the use of AWP mode; a value of 0 indicates that the use of AWP mode is not allowed. pic_dawp_flag can be a binary variable, with a value of 1 indicating that the use of DAWP mode is allowed; a value of 0 indicates that the use of DAWP mode is not allowed. Referring to Table 8, when pic_awp_flag is 0, it can be determined that the current image does not allow the use of AWP mode without decoding pic_dawp_flag; when pic_awp_flag is 1 and pic_dawp_flag is 0, it can be determined that the current image allows the use of AWP mode and does not allow the use of DAWP mode; when pic_awp_flag and pic_dawp_flag are both 1, it can be determined that the current image allows the use of AWP mode and allows the use of DAWP mode.
[0196] As an example, the decoding process of pic_awp_flag and pic_dawp_flag is shown in Table 9:
[0197] Table 9
[0198] Referring to Table 9, decode pic_awp_flag first. If pic_awp_flag indicates that the current image allows the use of AWP mode, then decode pic_dawp_flag. If pic_awp_flag indicates that the current image does not allow the use of AWP mode, then decode pic_dawp_flag. The descriptors of pic_awp_flag and pic_dawp_flag can both be u(1), that is, a 1-bit unsigned integer. The value of PicAwpFlag is equal to the value of pic_awp_flag. If pic_awp_flag does not exist in the bitstream, the value of PicAwpFlag is 0.
[0199] In some optional embodiments, there may be only one flag in the code stream, namely the picture header angle weighted prediction mode flag pic_awp_flag, to indicate whether the current picture uses the AWP mode. The specific indication method may be as shown in Table 10:
[0200] Table 10
[0201] pic_awp_flag indicates the mode type of the angle-weighted prediction mode. A value of "0" indicates that AWP is not allowed. A value greater than "0" indicates that AWP can be used and indicates whether DAWP is allowed. The descriptor of pic_awp_flag can be u(2), a 2-bit unsigned integer, or ue(v), an unsigned integer syntax element.
[0202] Referring to Table 10, when pic_awp_flag is 0, it indicates that the current image does not allow the use of AWP mode; when pic_awp_flag is 1, it indicates that the current image allows the use of AWP mode and does not allow the use of DAWP mode; when pic_awp_flag is 2, it indicates that the current image allows the use of AWP mode and allows the use of DAWP mode.
[0203] In some optional embodiments, only one picture header syntax element, pic_dawp_flag, may be present in the codestream to indicate whether the current picture uses the DAWP mode. The specific indication method may be combined with seq_awp_flag as shown in Table 11 ("x" in Table 11 indicates that decoding is not required).
[0204] Table 11
[0205] seq_awp_flag can be a binary variable, with a value of 1 indicating that the AWP mode is allowed, and a value of 0 indicating that the AWP mode is not allowed. pic_dawp_flag can be a binary variable, with a value of 1 indicating that the DAWP mode is allowed, and a value of 0 indicating that the DAWP mode is not allowed.
[0206] Referring to Table 11, when seq_awp_flag is 0, it is possible to determine that the current image does not allow the use of the AWP mode (the entire sequence does not allow the use of the AWP mode) without decoding the pic_dawp_flag; when seq_awp_flag is 1 and pic_dawp_flag is 0, it can be determined that the current image allows the use of the AWP mode but does not allow the use of the DAWP mode; when seq_awp_flag and pic_dawp_flag are both 1, it can be determined that the current image allows the use of the AWP mode and allows the use of the DAWP mode.
[0207] As an example, the decoding process of seq_awp_flag and pic_dawp_flag is shown in Table 12:
[0208] Table 12
[0209] As shown in Table 12, seq_awp_flag is decoded first. If seq_awp_flag indicates that the current sequence allows the use of AWP mode, pic_dawp_flag is decoded. The descriptors for both seq_awp_flag and pic_dawp_flag can be u(1), a 1-bit unsigned integer. The value of SeqAwpFlag is equal to the value of seq_awp_flag. If seq_awp_flag does not exist in the bitstream, the value of SeqAwpFlag is 0.
[0210] In some optional embodiments, only one picture header syntax element, pic_dawp_flag, may be present in the codestream to indicate whether the current picture uses the DAWP mode. The specific indication method may be combined with seq_dawp_flag as shown in Table 13 ("x" in Table 13 indicates that decoding is not required).
[0211] Table 13
[0212] seq_dawp_flag can be a binary variable, with a value of 1 indicating that the DAWP mode is allowed; a value of 0 indicating that the DAWP mode is not allowed. pic_dawp_flag can be a binary variable, with a value of 1 indicating that the DAWP mode is allowed; a value of 0 indicating that the DAWP mode is not allowed.
[0213] As shown in Table 13, when seq_dawp_flag is 0, it is possible to determine that the current image does not allow the use of DAWP mode (the entire sequence does not allow the use of DAWP mode) without decoding pic_dawp_flag; when seq_dawp_flag is 1 and pic_dawp_flag is 0, it can be determined that the current image does not allow the use of DAWP mode; when seq_dawp_flag and pic_dawp_flag are both 1, it can be determined that the current image allows the use of DAWP mode.
[0214] As an example, the decoding process of seq_dawp_flag and pic_dawp_flag is shown in Table 14:
[0215] Table 14
[0216] As shown in Table 14, seq_dawp_flag is decoded first. If seq_dawp_flag indicates that the current sequence allows the use of DAWP mode, pic_dawp_flag is decoded. The descriptors for both seq_dawp_flag and pic_dawp_flag can be u(1), a 1-bit unsigned integer. The value of SeqDawpFlag is equal to the value of seq_dawp_flag. If seq_dawp_flag does not exist in the bitstream, the value of SeqDawpFlag is 0.
[0217] In one embodiment of the present application, the code stream includes a slice header syntax element indicating whether the current slice uses the AWP mode.
[0218] In some optional embodiments, the bitstream may include two flags, such as the slice header angle weighted prediction mode flag slice_awp_flag and the slice header angle weighted prediction mode flag slice_dawp_flag, to indicate whether the current slice uses the AWP mode. The specific indication method may be as shown in Table 15 ("x" in Table 15 indicates that no decoding is required):
[0219] Table 15
[0220] Among them, slice_awp_flag can be a binary variable, with a value of 1 indicating that the current slice allows the use of AWP mode; a value of 0 indicates that the use of AWP mode is not allowed. slice_dawp_flag can be a binary variable, with a value of 1 indicating that the use of DAWP mode is allowed; a value of 0 indicates that the use of DAWP mode is not allowed. Referring to Table 15, when slice_awp_flag is 0, it can be determined that the current slice does not allow the use of AWP mode without decoding slice_dawp_flag; when slice_awp_flag is 1 and slice_dawp_flag is 0, it can be determined that the current slice allows the use of AWP mode and does not allow the use of DAWP mode; when slice_awp_flag and slice_dawp_flag are both 1, it can be determined that the current slice allows the use of AWP mode and allows the use of DAWP mode.
[0221] As an example, the decoding process of slice_awp_flag and slice_dawp_flag is shown in Table 16:
[0222] Table 16
[0223] As shown in Table 16, slice_awp_flag is decoded first. If slice_awp_flag indicates that the current slice allows the use of AWP mode, slice_dawp_flag is decoded. If slice_awp_flag indicates that the current slice does not allow the use of AWP mode, there is no need to decode slice_dawp_flag. The descriptors of slice_awp_flag and slice_dawp_flag can both be u(1), that is, a 1-bit unsigned integer. The value of SliceAwpFlag is equal to the value of slice_awp_flag. If slice_awp_flag does not exist in the bitstream, the value of SliceAwpFlag is 0.
[0224] In some optional embodiments, there may be only one flag in the bitstream, namely the slice header angle weighted prediction mode flag slice_awp_flag, to indicate whether the current slice uses the AWP mode. The specific indication method may be as shown in Table 17:
[0225] Table 17
[0226] The slice_awp_flag descriptor indicates the mode type of the angle-weighted prediction mode. A value of "0" indicates that the AWP is not allowed. A value greater than "0" indicates that the AWP can be used and indicates whether the DAWP is allowed. The descriptor of slice_awp_flag can be u(2), which is a 2-bit unsigned integer, or ue(v), which is an unsigned integer syntax element.
[0227] Referring to Table 17, when slice_awp_flag is 0, it indicates that the current slice does not allow the use of AWP mode; when slice_awp_flag is 1, it indicates that the current slice allows the use of AWP mode and does not allow the use of DAWP mode; when slice_awp_flag is 2, it indicates that the current slice allows the use of AWP mode and allows the use of DAWP mode.
[0228] In some optional embodiments, only one slice header syntax element slice_dawp_flag may be present in the bitstream to indicate whether the current slice uses the DAWP mode. The specific indication method may be combined with pic_awp_flag as shown in Table 18 ("x" in Table 18 indicates that decoding is not required).
[0229] Table 18
[0230] pic_awp_flag can be a binary variable, with a value of 1 indicating that the AWP mode is allowed, and a value of 0 indicating that the AWP mode is not allowed. slice_dawp_flag can be a binary variable, with a value of 1 indicating that the DAWP mode is allowed, and a value of 0 indicating that the DAWP mode is not allowed.
[0231] As shown in Reference Table 18, when pic_awp_flag is 0, it is possible to determine that the current slice does not allow the use of AWP mode (the entire image does not allow the use of AWP mode) without decoding slice_dawp_flag; when pic_awp_flag is 1 and slice_dawp_flag is 0, it can be determined that the current slice allows the use of AWP mode but does not allow the use of DAWP mode; when pic_awp_flag and slice_dawp_flag are both 1, it can be determined that the current slice allows the use of AWP mode and allows the use of DAWP mode.
[0232] As an example, the decoding process of pic_awp_flag and slice_dawp_flag is shown in Table 19:
[0233] Table 19
[0234] Referring to Table 19, pic_awp_flag is decoded first. If pic_awp_flag indicates that the current image allows the use of AWP mode, slice_dawp_flag is decoded. The descriptors for both pic_awp_flag and slice_dawp_flag can be u(1), a 1-bit unsigned integer. The value of PicAwpFlag is equal to the value of pic_awp_flag. If pic_awp_flag does not exist in the bitstream, the value of PicAwpFlag is 0.
[0235] In some optional embodiments, only one slice header syntax element, slice_dawp_flag, may be present in the bitstream to indicate whether the current slice uses the DAWP mode. The specific indication method may be combined with pic_dawp_flag as shown in Table 20 ("x" in Table 20 indicates that no decoding is required).
[0236] Table 20
[0237] pic_dawp_flag can be a binary variable, with a value of 1 indicating that the DAWP mode is allowed; a value of 0 indicating that the DAWP mode is not allowed. slice_dawp_flag can be a binary variable, with a value of 1 indicating that the DAWP mode is allowed; a value of 0 indicating that the DAWP mode is not allowed.
[0238] As shown in Reference Table 20, when pic_dawp_flag is 0, it is possible to determine that the current slice does not allow the use of DAWP mode (the entire image does not allow the use of DAWP mode) without decoding slice_dawp_flag; when pic_dawp_flag is 1 and slice_dawp_flag is 0, it can be determined that the current slice does not allow the use of DAWP mode; when pic_dawp_flag and slice_dawp_flag are both 1, it can be determined that the current slice allows the use of DAWP mode.
[0239] As an example, the decoding process of pic_dawp_flag and slice_dawp_flag is shown in Table 21:
[0240] Table 21
[0241] Referring to Table 21, pic_dawp_flag is decoded first. If pic_dawp_flag indicates that the current image allows the use of DAWP mode, slice_dawp_flag is decoded. The descriptors for both pic_dawp_flag and slice_dawp_flag can be u(1), a 1-bit unsigned integer. The value of PicDawpFlag is equal to the value of pic_dawp_flag. If pic_dawp_flag does not exist in the bitstream, the value of PicDawpFlag is 0.
[0242] In one embodiment of the present application, the decoding of the above-mentioned high-level syntax elements (i.e., sequence header syntax elements, picture header syntax elements, and slice header syntax elements) can be determined based on the picture type. For example, one or more of the above-mentioned high-level syntax elements may be decoded only for specific picture types, or one or more of the above-mentioned high-level syntax elements may be decoded only for non-specific picture types.
[0243] Specifically, for example, the dawp_flag related flag bit is decoded only in I-frame images, or only in B-frame images, or in P-frame and B-frame images. For another example, the dawp_flag related flag bit can be decoded only in non-I-frame images.
[0244] Process 2: If the angle weighted prediction mode is used, the code stream is decoded to determine the relevant syntax elements of the angle weighted prediction mode, including the syntax elements for determining the weight matrix and the syntax elements for determining the prediction values of each part.
[0245] In one embodiment of the present application, the grammatical elements for determining the weight matrix include at least one of the following grammatical elements: index information for indicating the angle-weighted prediction mode, index information for indicating the reference weight configuration, index information for indicating the prediction angle, index information for indicating the reference weight derivation method, index information for indicating the weight mode of the angle-weighted prediction mode, and index information for indicating the size of the reference weight mixing area.
[0246] The index information used to indicate the reference weight derivation method may be index information used to indicate the reference weight derivation function, or index information used to indicate the reference weight derivation function parameters.
[0247] The following lists the reference weight derivation functions that can be used in the embodiments of the present application, where in the following function, d represents the distance between the reference weight sample and the mixed area, which can be the distance between the reference weight sample and the starting position of the mixed area, or the distance between the reference weight sample and the midpoint position of the mixed area, and of course, the distance between the reference weight sample and the end position of the mixed area, etc. Of course, an offset value can also be added to these distances. Optionally, the distance represented by d can be directional, such as if the reference weight sample position is pos, the mixed area is c, and the value of d is pos-c, then a negative value of d indicates the left side of the mixed area, and a positive value of d indicates the right side of the mixed area. Of course, the distance represented by d can also be non-directional, so the absolute value of the distance can be calculated by abs(pos-c).
[0248] In one embodiment of the present application, the reference weight derivation function may use a linear function, such as: w = s*(d+k); wherein s and k are parameter values of the function, which may be preset values or indicated by code stream information.
[0249] In one embodiment of the present application, the reference weight derivation function may use a sigmoid function, such as: Where s and k are function parameter values, which can be preset values or indicated by the bitstream information. For example, when s = 2, the function graph is shown in Figure 15. The weight of the mixed area can be set according to the function value corresponding to [-2, 2].
[0250] In one embodiment of the present application, the reference weight derivation function may use a hyperbolic function, such as the tanh function: w = 0.5*tanh(s*d)+0.5; where s is the function parameter value, which may be a preset value or indicated by bitstream information. For example, when s = 1.3, the function graph is shown in Figure 16, and the weight of the mixed area can be set according to the function value corresponding to [-2, 2].
[0251] In one embodiment of the present application, the reference weight derivation function may use a form based on a trigonometric function, such as a form based on a cosine function: Where s is the parameter value of the function, which can be a preset value or indicated by the bitstream information. For example, when s = 0.5, the function graph is shown in Figure 17. The weight of the mixed area can be set according to the function value corresponding to [-2, 2].
[0252] In one embodiment of the present application, the reference weight derivation function may use an exponential-based function, for example, if d is less than 0, then w=e s1(d-s2) ; If d is greater than or equal to 0, then Where s1 and s2 are function parameter values, which can be preset values or indicated by the bitstream information. For example, if s1 = 2 and s2 = 0.3465735, the weight of the mixed area can be set according to the function value corresponding to [-2, 2].
[0253] In one embodiment of the present application, the reference weight derivation function may use a polynomial-based function, such as a polynomial with a quadratic power.
[0254] In one embodiment of the present application, the reference weight derivation function can use a piecewise function to derive weights, and use different weight derivation functions according to the value of d. For example, when d is less than or equal to 0, a sigmoid function is used, otherwise a linear function is used.
[0255] In one embodiment of the present application, the syntax elements for determining each partial prediction value include an index mvp_idx_n for determining the predicted motion vector of the reference block, and related syntax elements for determining the correction of the motion vector. It should be noted that the "partial prediction value" in this embodiment generally refers to the prediction value of two parts, but in other embodiments of the present application, it can also refer to the prediction value of more parts.
[0256] In one embodiment of the present application, the relevant syntax elements for determining whether to modify a motion vector may include a flag mvr_flag_n indicating whether the motion vector needs to be modified, an index value mvr_step_n indicating the step size of the motion vector modification, and an index value mvr_dir_n indicating the direction of the motion vector modification. mvr_step_n and mvr_dir_n are used to derive mvd, where the motion vector mv = mvp + mvd. In other embodiments of the present application, mvd can also be directly decoded from the bitstream without decoding mvr_step_n and mvr_dir_n.
[0257] In one embodiment of the present application, the syntax elements or parts of the syntax elements in the above embodiment can be decoded using variable-length codes, such as K-order Exponential Golomb codes, truncated unary codes, truncated binary codes, etc. Optionally, different parts of the syntax elements in the above embodiment can be decoded using different variable-length codes, such as using truncated unary codes to decode the prefix part and using truncated binary codes to decode the suffix part.
[0258] In one embodiment of the present application, the syntax elements or parts of the syntax elements in the above embodiments may also be decoded using fixed-length codes.
[0259] In one embodiment of the present application, the syntax elements in the above embodiment may also be decoded using a combination of variable-length codes and fixed-length codes. For example, the prefix portion is decoded using a variable-length code, and the suffix portion is decoded using a fixed-length code.
[0260] In one embodiment of the present application, the portion of the syntax element in the above embodiment whose binary bits are less than a set threshold can be decoded using a decoding method corresponding to the context-based binary encoding method, and the remaining portion can be decoded using a decoding method corresponding to the bypass coding method (Bypass Coding Mode). For example, assuming that the length of the above syntax element is M, the portion whose binary bits are less than the threshold th is decoded using context-based binary encoding, and the remaining portion is decoded using the bypass mode.
[0261] In one embodiment of the present application, if the total number of AWP mode decoding modes awp_sig_mode (less than or equal to 56 AWP modes) is divided into multiple groups, each group having awp_sig_divisor elements, then there are a total of awp_sig_group = Ceil(awp_sig_mode / awp_sig_divisor) groups. Then, variable-length code decoding can be used for group numbers, and fixed-length code decoding can be used for index information of elements within the group.
[0262] Optionally, the total number of modes awp_sig_mode obtained by decoding and the grouping method (such as the number of each group or the number of groups, etc.) can be indicated in a high-level syntax element (such as one or more of a sequence header, a picture header, and a slice header).
[0263] As an example, the number of AWP modes is 56. Assuming that the total number of modes obtained by decoding is 56, the total number of modes is divided into 7 groups, with 8 modes in each group. Then the group number can use a truncated unary code and the element index within the group can use a truncated binary code; or the group number can use a context-based binary code and the element index within the group can use the bypass mode.
[0264] As an example, the number of AWP modes is 56. Assuming that the total number of modes obtained by decoding is 28, the total number of modes is divided into 7 groups, with 4 modes in each group. Then the group number can use truncated unary code and the element index within the group can use truncated binary code; or the group number can use context-based binary encoding and the element index within the group can use bypass mode.
[0265] Process 3: If the rearranged angle-weighted prediction mode is used, a rearranged mode list and corresponding prediction mode are derived.
[0266] In one embodiment of the present application, if it is determined that the rearrangement-based angle weighted prediction mode can be used, the decoding end needs to confirm the availability of the DAWP mode. If it is not available, the AWP mode will not be rearranged. When confirming the availability of the DAWP mode, it is possible to check whether the current template corresponding to the current block (all reconstructed samples around the current block can be used as a template) has been reconstructed (such as checking whether the samples to the left of the current block are available, checking whether the samples above the current block are available, etc.); check whether the size of the current block meets the conditions (assuming that only blocks of a set size can use the DAWP mode); check whether the color components of the current block meet the conditions (assuming that only the luminance component or the chrominance component can use the DAWP mode), etc.
[0267] In one embodiment of the present application, if it is determined that a rearranged angle-weighted prediction mode can be used, and the DAWP mode is determined to be available in the above manner, then the decoded angle-weighted prediction modes can be traversed to calculate the cost corresponding to each angle-weighted prediction mode, and then the rearranged AWP mode list awp_cost_list can be derived based on the cost.
[0268] In some optional embodiments, when calculating the cost corresponding to each angle weighted prediction mode, the following steps need to be performed (the steps may be performed in any order):
[0269] Step a: Get the current template (all reconstructed sample points around the current block can be used as templates).
[0270] Optionally, the current template may be N rows of samples above the current block (may include the upper left and / or upper right), with a width and height of (tw0, th0).
[0271] Optionally, the current template may be M columns of samples to the left (may include the upper left and / or lower left) of the current block, with a width and height of (tw1, th1).
[0272] Optionally, the current template may be sample points obtained by sampling at fixed intervals, for example, sampling at intervals of one row (and / or column).
[0273] Optionally, the current template may use the templates to the left and above at the same time. For example, the current template is the sample points in one row above and one column to the left of the current block, and the width and height of the current template are consistent with those of the current block.
[0274] In some optional embodiments, the above methods for selecting the current template may be arbitrarily combined to obtain the current template.
[0275] In some optional embodiments, the samples of the current template may be subjected to some correction processing, such as filtering, linear mapping, nonlinear mapping, etc.
[0276] Step b: Obtain prediction templates for each part.
[0277] In one embodiment of the present application, it is necessary to determine the position of the reference block according to the motion vector, and then select a prediction template according to the position of the reference block.
[0278] In some optional embodiments, the position of the prediction template may correspond to the position of the current template. For example, if the current template is N rows of samples above the current block (including the upper left and / or upper right), then the prediction template may select N rows of samples above the reference block (including the upper left and / or upper right). If the current template is M columns of samples to the left of the current block (including the upper left and / or lower left), then the prediction template may select M columns of samples to the left of the reference block (including the upper left and / or lower left).
[0279] In some optional embodiments, the position of the prediction template may not correspond to the position of the current template. For example, to reduce complexity, samples in the reference block may be used as the prediction template.
[0280] Specifically, the prediction template may be N rows (may include the upper left and / or upper right) of samples located in the reference block, with a width and height of (tw2, th2).
[0281] Optionally, the prediction template may be M columns (may include the upper left and / or lower left) of samples located in the reference block, with a width and height of (tw3, th3).
[0282] Optionally, the prediction template may also be obtained by sampling at fixed intervals, for example, sampling at intervals of one row (and / or column).
[0283] Optionally, the prediction template can use both the left and top templates. For example, the prediction template is the samples in the top row and left column of the reference block, and the width and height of the prediction template are consistent with those of the reference block. Alternatively, the prediction template can be the samples in the first row and first column of the reference block, and the width and height of the prediction template are consistent with those of the reference block.
[0284] In some optional embodiments, the above methods for selecting a prediction template may be arbitrarily combined to obtain a prediction template.
[0285] In some optional embodiments, the samples of the prediction template may be subjected to some correction processing, such as filtering, linear mapping, nonlinear mapping, etc.
[0286] In some optional embodiments, a motion vector of the prediction template can be calculated based on the relative position of the prediction template and the reference block, and motion compensation can be performed to derive a prediction value. Optionally, if the motion vector is pixel-wise, interpolation filtering can be performed, and different interpolation filters can be used, such as a bilinear compensation interpolation filter.
[0287] In one embodiment of the present application, the process of deriving the predicted value in the AWP mode is as follows: Assume that the motion vector cu_mv of the current block is (mv_x, mv_y), and the upper left corner of the prediction template is located in a row above the reference block and has the same width as the current block, that is, the width and height (tpl_w, tpl_h) is (cu_w, 1). Then the motion vector tpl_mv from the current block to the prediction template is (mv_x, mv_y - 1). Based on the coordinates (cu_x, cu_y) of the current block and tpl_mv, the predicted value of the prediction template can be derived. Here, tpl is the prediction template, tpl_w is the width of the prediction template, tpl_h is the height of the prediction template, and cu_w is the width of the current block.
[0288] In one embodiment of the present application, the process of deriving the predicted value in the SAWP mode is as follows: Assume that the intra prediction mode of the current block is intra_mode, and the upper left corner of the prediction template is located in a row above the reference block and has the same width as the current block, that is, the width and height (tpl_w, tpl_h) is (cu_w, 1). Then, based on the intra prediction mode intra_mode, the intra prediction value of the current template is derived using the reconstructed pixels above and to the left of the current prediction template tpl. Here, tpl_w is the width of the prediction template, tpl_h is the height of the prediction template, and cu_w is the width of the current block.
[0289] Step c, derive the template weight.
[0290] In some alternative embodiments, the weight of the prediction template can be derived according to the weight derivation method in the angular weighted prediction mode, where the relative position relationship between the prediction template and the reference block needs to be considered.
[0291] In some alternative embodiments, a weight threshold can be set, and the template weight greater than or equal to this weight threshold is set to 1, and the template weight less than this weight threshold is set to 0.
[0292] In some alternative embodiments, the reference weight configuration and weight prediction angle corresponding to the prediction template can be obtained, and then the center position cp of the reference weight mixing area is determined accordingly. After that, the position tp of the sample points corresponding to the prediction template after projection according to the weight prediction angle is calculated. If tp < cp (i.e., the projection position is to the left of cp), the template weight of the sample points corresponding to the prediction template is set to 0, otherwise it is set to 1.
[0293] In some alternative embodiments, the template weight can be initialized in advance and does not need to be repeatedly derived.
[0294] Step d, derive the weighted prediction template.
[0295] After calculating the template weight of the prediction template, you can use the template weight to perform a weighted combination of the various prediction templates to obtain your prediction template.
[0296] Step e: Calculate the cost.
[0297] In some optional embodiments, the SAD of the weighted prediction template and the current template, or the SSE of the weighted prediction template and the current template, or the MR-SAD of the weighted prediction template and the current template may be calculated as the cost of the angle weighted prediction mode.
[0298] Optionally, when calculating the cost, only the portion of sample point differences that are smaller than a threshold value may be counted.
[0299] In some optional embodiments, after the costs of the angle-weighted prediction modes are calculated, the angle-weighted prediction modes may be sorted in ascending order of cost to obtain a rearranged AWP mode list awp_cost_list.
[0300] Optionally, when sorting in ascending order of cost, if the number of sorted angle-weighted prediction modes reaches a set number (such as the number of decoded angle-weighted prediction modes), the sorting may be stopped.
[0301] In one embodiment of the present application, after the rearranged AWP mode list awp_cost_list is derived, the mode in the rearranged AWP mode list corresponding to the mode index information obtained by decoding is the mode finally used for weighted prediction, which is recorded as final_awp_mode.
[0302] In step 4, a weight matrix is derived according to the derived prediction model to perform weighted prediction.
[0303] In one embodiment of the present application, after determining the mode final_awp_mode for weighted prediction, a weight matrix may be derived according to final_awp_mode, and then weighted prediction may be performed based on the weight matrix to obtain a predicted value.
[0304] It should be noted that the technical solution of the embodiment of the present application is not only applicable to the reordering processing of the AWP mode, but also applicable to the reordering processing of the SAWP mode.
[0305] The following describes an embodiment of the device of the present application, which can be used to perform the method described in the above embodiment of the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the above method embodiment of the present application.
[0306] FIG18 shows a block diagram of a video decoding apparatus according to an embodiment of the present application. The video decoding apparatus may be provided in a device having a computing and processing function, such as a terminal device or a server.
[0307] 18 , a video decoding apparatus 1800 according to an embodiment of the present application includes: a decoding unit 1802 , a sorting unit 1804 , a selecting unit 1806 , and a processing unit 1808 .
[0308] Among them, the decoding unit 1802 is configured to decode the video code stream to obtain mode index information for the rearranged mode list; the sorting unit 1804 is configured to sort the multiple angle weighted prediction modes in order from low to high according to the cost of using the multiple angle weighted prediction modes for the current block to obtain the rearranged mode list; the selection unit 1806 is configured to select the corresponding angle weighted prediction mode from the rearranged mode list according to the mode index information, and derive the weight matrix of the current block according to the selected angle weighted prediction mode; the processing unit 1808 is configured to perform weighted prediction according to the weight matrix to obtain the prediction value corresponding to the current block.
[0309] In some embodiments of the present application, based on the aforementioned scheme, the sorting unit 1804 is configured as follows: if it is determined according to the index flag that the current block allows the use of the rearranged angle-weighted prediction mode, then the multiple angle-weighted prediction modes are sorted in order from low to high according to the costs of the multiple angle-weighted prediction modes to obtain the rearranged mode list.
[0310] In some embodiments of the present application, based on the aforementioned solution, the index flag includes at least one of the following flags: a sequence header flag included in sequence header information, a picture header flag included in picture header information, and a slice header flag included in slice header information;
[0311] Among them, the sequence header flag is used to indicate whether the current sequence allows the use of the angle weighted prediction mode based on rearrangement; the image header flag is used to indicate whether the current image allows the use of the angle weighted prediction mode based on rearrangement; and the slice header flag is used to indicate whether the current slice allows the use of the angle weighted prediction mode based on rearrangement.
[0312] In some embodiments of the present application, based on the aforementioned scheme, the index flag bit includes one flag bit or multiple flag bits; the different values of the one flag bit or multiple flag bits are used to indicate whether the corresponding coding block is allowed to use the angle-weighted prediction mode, and when the angle-weighted prediction mode is allowed to be used, whether the rearrangement-based angle-weighted prediction mode is allowed to be used.
[0313] In some embodiments of the present application, based on the aforementioned scheme, the decoding unit 1802 is configured as follows: if the video image frame type is a specified type, then the index flag is decoded from the video code stream; or if the video image frame type is a non-specified type, then the index flag is decoded from the video code stream.
[0314] In some embodiments of the present application, based on the aforementioned scheme, the decoding unit 1802 is further configured to: if the current block uses an angle-weighted prediction mode, decode from the video bit stream to obtain syntax elements for deriving a weight matrix and syntax elements for determining multiple prediction values.
[0315] In some embodiments of the present application, based on the aforementioned scheme, the syntax elements for deriving the weight matrix include at least one of the following syntax elements: index information for indicating the angle-weighted prediction mode, index information for indicating the reference weight configuration, index information for indicating the prediction angle, index information for indicating the reference weight derivation method, index information for indicating the weight mode of the angle-weighted prediction mode, and index information for indicating the size of the reference weight mixing area.
[0316] In some embodiments of the present application, based on the aforementioned solution, the syntax elements used to derive the weight matrix include:
[0317] Index information indicating an angle-weighted prediction mode; or
[0318] Index information indicating a reference weight derivation method and index information indicating an angle-weighted prediction mode; or
[0319] Index information indicating a reference weight derivation method, index information indicating a reference weight configuration, and index information indicating a prediction angle; or
[0320] Index information for indicating a weight mode of an angle-weighted prediction mode; or
[0321] Index information indicating a reference weight configuration and index information indicating a prediction angle.
[0322] In some embodiments of the present application, based on the aforementioned solution, the syntax elements for determining the multiple prediction values include at least one of the following syntax elements: index information for determining a predicted motion vector of a reference block, a syntax element for determining a correction to a motion vector, and a syntax element for determining an intra-frame prediction mode;
[0323] The syntax elements for determining whether to modify the motion vector include: index information for indicating whether the motion vector needs to be modified, index information for indicating the motion vector modification step size, and index information for indicating the motion vector modification direction; or
[0324] Index information indicating whether a motion vector needs to be corrected and index information indicating a motion vector difference.
[0325] In some embodiments of the present application, based on the aforementioned solution, part or all of the binary bits of the syntax element are decoded using variable-length codes; or
[0326] Part or all of the binary bits of the syntax element are decoded using a fixed-length code; or
[0327] Different parts of the binary bits of the syntax element are decoded using different variable-length codes; or
[0328] Part or all of the binary bits of the syntax element are decoded using a combination of variable-length code and fixed-length code; or
[0329] The portion of the syntax element whose binary bits are smaller than the set threshold is decoded using a decoding method corresponding to the context-based binary encoding method, and the remaining portion is decoded using a decoding method corresponding to the bypass encoding method.
[0330] In some embodiments of the present application, based on the aforementioned scheme, part or all of the syntax elements are decoded using a combination of variable-length codes and fixed-length codes, including: if the number of modes of the decoded angle-weighted prediction mode is divided into multiple groups according to a set grouping method, then the index information used to indicate the group number in the mode index information is decoded using a variable-length code, and the index information used to indicate the elements within the group in the mode index information is decoded using a fixed-length code.
[0331] In some embodiments of the present application, based on the aforementioned scheme, the video decoding device further includes: an acquisition unit, configured to acquire, for each angle-weighted prediction mode among the multiple angle-weighted prediction modes, a current template corresponding to the current block, and acquire a prediction template corresponding to the reference block of the current block; a weight derivation unit, configured to determine the template weight corresponding to the prediction template based on the weights derived from each angle-weighted prediction mode; a weight calculation unit, configured to calculate a weighted prediction template based on the template weight and the prediction template; and a cost calculation unit, configured to calculate the cost of each angle-weighted prediction mode based on the weighted prediction template and the current template.
[0332] In some embodiments of the present application, based on the aforementioned solution, the current template corresponding to the current block includes at least one of the following samples:
[0333] The sample point located at the set row above the current block;
[0334] The sample point located in the set column to the left of the current block;
[0335] Sample points are obtained by sampling the adjacent sample points of the current block according to the set interval;
[0336] The sample that is set rows above the current block and set columns to the left of the current block.
[0337] In some embodiments of the present application, based on the aforementioned solution, the prediction template corresponding to the reference block of the current block includes at least one of the following samples:
[0338] Sample points located in a set row above the reference block;
[0339] Sample points located in a set column to the left of the reference block;
[0340] Sample points of a set row within the reference block;
[0341] Sample points of a set column within the reference block;
[0342] Sample points are obtained by sampling adjacent sample points of the reference block according to the set interval;
[0343] Sample points are obtained by sampling the sample points in the reference block according to the set interval;
[0344] Sample points located at a set row above the reference block and a set column to the left of the reference block;
[0345] The sample points are set in the row above the reference block and in the column set in the left of the reference block.
[0346] In some embodiments of the present application, based on the aforementioned scheme, the acquisition unit is further configured to perform at least one of the following processing methods: performing correction processing on the samples in the current template, and performing correction processing on the samples in the prediction template; the correction processing includes one or more of filtering processing, linear mapping processing, and nonlinear mapping processing.
[0347] In some embodiments of the present application, based on the aforementioned solution, the weight derivation unit is further configured to: for an angle-weighted prediction mode whose derived weight is greater than a weight threshold, set the template weight to 1; and for an angle-weighted prediction mode whose derived weight is less than the weight threshold, set the template weight to 0; or
[0348] If the position of the sample point corresponding to the prediction template after projection according to the weight prediction angle is located on the left side of the center position of the mixed area, the template weight is set to 0; otherwise, the template weight is set to 1.
[0349] In some embodiments of the present application, based on the aforementioned scheme, the cost calculation unit is configured to: calculate the sum of absolute errors, or the sum of squares of differences, or the average reduced sum of absolute errors between the weighted prediction template and the current template as the cost of the weighted prediction mode of each angle.
[0350] In some embodiments of the present application, based on the aforementioned scheme, the cost calculation unit is configured to: calculate the sum of absolute errors, or the sum of squares of the differences, or the average reduced sum of absolute errors between the weighted prediction template and the target sample points in the current template whose sample point differences are less than a set threshold.
[0351] In some embodiments of the present application, based on the aforementioned scheme, the sorting unit 1804 is configured as follows: if the number of angle-weighted prediction modes sorted in order of the cost from low to high reaches a set number, the sorting is stopped; wherein the set number is less than or equal to the total number of angle-weighted prediction modes.
[0352] In some embodiments of the present application, based on the aforementioned solution, the decoding unit 1802 is configured to: decode from the video bitstream to obtain at least one of the following index information as the mode index information: index information indicating the angle-weighted prediction mode, index information indicating the reference weight configuration, index information indicating the prediction angle, index information indicating the reference weight derivation method, and index information indicating the weight mode of the angle-weighted prediction mode;
[0353] The rearranged pattern list is a pattern list obtained by rearranging one or more index information included in the pattern index information.
[0354] FIG19 shows a block diagram of a video encoding apparatus according to an embodiment of the present application. The video encoding apparatus may be provided in a device having a computing and processing function, such as a terminal device or a server.
[0355] 19 , a video encoding apparatus 1900 according to an embodiment of the present application includes a sorting unit 1902 , a determining unit 1904 , a calculating unit 1906 , and an encoding unit 1908 .
[0356] Among them, the sorting unit 1902 is configured to sort the multiple angle-weighted prediction modes in order from low to high according to the cost of using the multiple angle-weighted prediction modes for the current block, and obtain a rearranged mode list; the determination unit 1904 determines the mode index information based on the rearranged mode list and the selected angle-weighted prediction mode; the calculation unit 1906 is configured to derive the weight matrix of the current block according to the selected angle-weighted prediction mode; the encoding unit 1908 is configured to perform weighted prediction based on the weight matrix to obtain the prediction value corresponding to the current block, encode the current block according to the prediction value, and encode the mode index information in the video code stream.
[0357] FIG20 shows a schematic structural diagram of a computer system of an electronic device suitable for implementing an embodiment of the present application. The electronic device may be a video decoding device or a video encoding device in the aforementioned embodiment.
[0358] It should be noted that the computer system 2000 of the electronic device shown in FIG20 is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present application.
[0359] As shown in Figure 20, the computer system 2000 may include a central processing unit (CPU) 2001, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 2002 or the program loaded from the storage part 2008 into the random access memory (RAM) 2003, such as the method described in the above embodiment. In the RAM 2003, various programs and data required for system operation are also stored. The CPU 2001, ROM 2002 and RAM 2003 are connected to each other via a bus 2004. An input / output (I / O) interface 2005 is also connected to the bus 2004.
[0360] The following components can be connected to the I / O interface 2005: an input section 2006 including a keyboard, a mouse, and the like; an output section 2007 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 2008 including a hard disk and the like; and a communication section 2009 including a network interface card such as a LAN (Local Area Network) card or a modem. The communication section 2009 performs communication processing via a network such as the Internet. A drive 2010 is also connected to the I / O interface 2005 as needed. Removable media 2011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, and the like, are installed in the drive 2010 as needed, so that computer programs read therefrom can be installed into the storage section 2008 as needed.
[0361] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program is used to perform the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 2009, and / or installed from a removable medium 2011. When the computer program is executed by the central processing unit (CPU) 2001, the various functions defined in the system of the present application are performed.
[0362] It should be noted that the computer-readable medium shown in the embodiments of the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a computer program that can be used by or in combination with an instruction execution system, device or device. In the present application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries a computer-readable computer program. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. A computer program embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, or any suitable combination thereof.
[0363] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of the systems, methods, and computer program products according to various embodiments of the present application. Each box in the flowchart or block diagram may represent a module, program segment, or part of a code, and the above-mentioned module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box may also occur in an order different from that marked in the accompanying drawings. For example, two boxes shown in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, as well as the combination of boxes in the block diagram or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or can be implemented using a combination of dedicated hardware and a computer program. The units involved in the embodiments described in the present application can be implemented by software or by hardware, and the units described can also be set in a processor. The names of these units do not constitute a limitation on the units themselves in some cases.
[0364] As another aspect, the present application further provides a computer-readable medium, which may be included in the electronic device described in the above embodiments, or may exist independently without being incorporated into the electronic device. The computer-readable medium carries one or more computer programs, and when the one or more computer programs are executed by the electronic device, the electronic device implements the method described in the above embodiments.
[0365] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiment of the application, the features and functions of two or more modules or units described above can be concretized in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.
[0366] Through the description of the above implementation methods, it is easy for those skilled in the art to understand that the example implementation methods described here can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the implementation method of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes a number of instructions to enable an electronic device to execute a method according to the implementation method of the present application. For example, the electronic device can be a video decoding device, and the video decoding device can execute the video decoding method shown in Figure 13; for another example, the electronic device can be a video encoding device, and the video encoding device can execute the video encoding method shown in Figure 14.
[0367] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of this application and include common knowledge or customary techniques in the art that are not disclosed herein.
[0368] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
Claims
1. A video decoding method, characterized in that: include: Decoding the video code stream to obtain mode index information for the rearranged mode list; Sorting the multiple angle weighted prediction modes according to the order of the costs of using the multiple angle weighted prediction modes for the current block from low to high, to obtain the rearranged mode list; Selecting a corresponding angle weighted prediction mode from the rearranged mode list according to the mode index information, and deriving a weight matrix of the current block according to the selected angle weighted prediction mode; A weighted prediction is performed according to the weight matrix to obtain a prediction value corresponding to the current block.
2. The video decoding method according to claim 1, characterized in that: The method of sorting the multiple angle weighted prediction modes in order of cost from low to high for the current block to obtain the rearranged mode list includes: Decoding the index flag contained in the video bit stream; If it is determined according to the index flag that the current block allows the use of the rearranged angle-weighted prediction mode, the multiple angle-weighted prediction modes are sorted in order of their costs from low to high to obtain the rearranged mode list.
3. The video decoding method according to claim 2, characterized in that: The index flag includes at least one of the following flags: The sequence header flag bit contained in the sequence header information, the image header flag bit contained in the image header information, and the slice header flag bit contained in the slice header information; Among them, the sequence header flag is used to indicate whether the current sequence allows the use of the angle weighted prediction mode based on rearrangement; the image header flag is used to indicate whether the current image allows the use of the angle weighted prediction mode based on rearrangement; and the slice header flag is used to indicate whether the current slice allows the use of the angle weighted prediction mode based on rearrangement.
4. The video decoding method according to claim 2 or 3, characterized in that: The index flag bit includes one flag bit or multiple flag bits; The one flag bit or different values of the multiple flag bits are used to indicate whether the corresponding coding block is allowed to use the angle weighted prediction mode, and when the angle weighted prediction mode is allowed, whether the rearrangement-based angle weighted prediction mode is allowed.
5. The video decoding method according to any one of claims 2 to 4, characterized in that: The index flag bits contained in the decoded video code stream include: If the video image frame type is a specified type, the index flag is obtained by decoding from the video code stream; or If the video image frame type is a non-specified type, the index flag is obtained by decoding the video code stream.
6. The video decoding method according to any one of claims 1 to 5, characterized in that: Also includes: If the current block uses the angle weighted prediction mode, syntax elements for deriving a weight matrix and syntax elements for determining a plurality of prediction values are decoded from the video bitstream.
7. The video decoding method according to claim 6, characterized in that: The syntax element used to derive the weight matrix includes at least one of the following syntax elements: Index information used to indicate the angle-weighted prediction mode, index information used to indicate the reference weight configuration, index information used to indicate the prediction angle, index information used to indicate the reference weight derivation method, index information used to indicate the weight mode of the angle-weighted prediction mode, and index information used to indicate the size of the reference weight mixing area.
8. The video decoding method according to claim 6, characterized in that: The grammatical elements used to derive the weight matrix include: Index information for indicating an angle-weighted prediction mode; or Index information for indicating a reference weight derivation method and index information for indicating an angle weighted prediction mode; or Index information for indicating a reference weight derivation method, index information for indicating a reference weight configuration, and index information for indicating indicating index information of the prediction angle; or Index information for indicating a weight mode of an angle-weighted prediction mode; or Index information indicating a reference weight configuration and index information indicating a prediction angle.
9. The video decoding method according to any one of claims 6 to 8, characterized in that: The syntax elements for determining the plurality of prediction values include at least one of the following syntax elements: index information for determining a prediction motion vector of a reference block, a syntax element for determining a correction of a motion vector, and a syntax element for determining an intra prediction mode; The syntax element used to determine whether the motion vector needs to be corrected includes: index information used to indicate whether the motion vector needs to be corrected, index information used to indicate the motion vector correction step size, and index information used to indicate the motion vector correction direction; or Index information used to indicate whether the motion vector needs to be corrected and index information used to indicate the motion vector difference.
10. The video decoding method according to any one of claims 6 to 9, characterized in that: Part or all of the binary bits of the syntax element are decoded using variable length code; or Part or all of the binary bits of the syntax element are decoded using a fixed-length code; or Different parts of the binary bits of the syntax element are decoded using different variable length codes respectively; or Part or all of the binary bits of the syntax element are decoded by combining variable-length code and fixed-length code; or The portion of the syntax element whose binary bits are smaller than the set threshold is decoded using a decoding method corresponding to the context-based binary encoding method, and the remaining portion is decoded using a decoding method corresponding to the bypass encoding method.
11. The video decoding method according to claim 10, characterized in that: Part or all of the syntax elements are decoded by combining variable-length codes and fixed-length codes, including: If the number of modes of the decoded angle weighted prediction mode is divided into multiple groups according to the set grouping method, the index information used to indicate the group number in the mode index information is decoded using a variable length code, and the index information used to indicate the elements within the group in the mode index information is decoded using a fixed length code.
12. The video decoding method according to any one of claims 1 to 11, characterized in that: Before sorting the multiple angle weighted prediction modes according to the order of the costs of using the multiple angle weighted prediction modes for the current block from low to high, the method further includes: For each of the multiple angle weighted prediction modes, obtaining a current template corresponding to the current block, and obtaining a prediction template corresponding to a reference block of the current block; Determining a template weight corresponding to the prediction template based on the weights derived from the weighted prediction modes of each angle; Calculating a weighted prediction template according to the template weight and the prediction template; The costs of the weighted prediction modes of each angle are calculated according to the weighted prediction template and the current template.
13. The video decoding method according to claim 12, characterized in that: The current template includes at least one of the following samples: The sample points located at the set row above the current block; The sample points located in the set column to the left of the current block; Sample points are obtained by sampling the adjacent sample points of the current block according to the set interval; The sample points are located at the set rows above the current block and at the set columns to the left of the current block.
14. The video decoding method according to claim 12 or 13, characterized in that: The prediction template includes at least one of the following samples: Sample points located in a set row above the reference block; Sample points located in a set column to the left of the reference block; Sample points of a set row within a reference block; Sample points of a set column within a reference block; Sample points are obtained by sampling the sample points adjacent to the reference block according to the set interval; Sample points obtained by sampling the sample points in the reference block according to the set interval; Sample points located at a set row above the reference block and a set column to the left of the reference block; Sample points at the top set row in the reference block and at the left set column in the reference block.
15. The video decoding method according to any one of claims 12 to 14, characterized in that: It also includes at least one of the following processing methods: performing correction processing on the sample points in the current template, and performing correction processing on the sample points in the prediction template; The correction processing includes one or more of filtering processing, linear mapping processing, and nonlinear mapping processing.
16. The video decoding method according to any one of claims 12 to 15, characterized in that: The determining the template weight corresponding to the prediction template based on the weighted prediction modes of each angle includes: For an angle-weighted prediction mode whose derived weight is greater than a weight threshold, the template weight is set to 1, and for an angle-weighted prediction mode whose derived weight is less than the weight threshold, the template weight is set to 0; or If the position of the sample point corresponding to the prediction template after projection according to the weight prediction angle is located on the left side of the center position of the mixed area, the template weight is set to 0; otherwise, the template weight is set to 1.
17. The video decoding method according to any one of claims 12 to 16, characterized in that: The calculating the cost of each angle weighted prediction mode according to the weighted prediction template and the current template includes: The sum of absolute errors, or the sum of squares of differences, or the average reduced sum of absolute errors between the weighted prediction template and the current template is calculated as the cost of each angle weighted prediction mode.
18. The video decoding method according to claim 17, characterized in that: The calculating the absolute error sum, or the square sum of the difference, or the average reduced absolute error sum between the weighted prediction template and the current template comprises: The absolute error sum, or the square sum of the difference, or the average reduced absolute error sum between the weighted prediction template and the target sample points whose sample point difference in the current template is less than a set threshold is calculated.
19. The video decoding method according to any one of claims 1 to 18, characterized in that: The method of sorting the multiple angle weighted prediction modes in order of cost from low to high for the current block to obtain the rearranged mode list includes: If the number of angle weighted prediction modes sorted in the order of the cost from low to high reaches a set number, the sorting is stopped; wherein the set number is less than or equal to the total number of angle weighted prediction modes.
20. The video decoding method according to any one of claims 1 to 19, characterized in that: The decoding of the video code stream to obtain mode index information for the rearranged mode list includes: At least one of the following index information is obtained by decoding the video bitstream as the mode index information: index information for indicating an angle weighted prediction mode, index information for indicating a reference weight configuration, index information for indicating a prediction angle, index information for indicating a reference weight derivation method, and index information for indicating a weight mode of the angle weighted prediction mode; The rearranged pattern list is a pattern list obtained by rearranging one or more index information included in the pattern index information.
21. A video encoding method, characterized in that: include: Sorting the multiple angle weighted prediction modes according to the order of costs of using the multiple angle weighted prediction modes for the current block from low to high to obtain a rearranged mode list; Determining mode index information according to the re-arranged mode list and the selected angle weighted prediction mode; deriving a weight matrix of the current block according to the selected angle weighted prediction mode; A weighted prediction is performed according to the weight matrix to obtain a prediction value corresponding to the current block, the current block is encoded according to the prediction value, and the mode index information is encoded in a video bitstream.
22. A video decoding device, characterized in that: include: A decoding unit configured to decode the video code stream to obtain mode index information for the rearranged mode list; A sorting unit configured to sort the multiple angle weighted prediction modes according to the order of the cost of using the multiple angle weighted prediction modes for the current block from low to high, so as to obtain the rearranged mode list; a selection unit configured to select a corresponding angle weighted prediction mode from the rearranged mode list according to the mode index information, and derive a weight matrix of the current block according to the selected angle weighted prediction mode; The processing unit is configured to perform weighted prediction according to the weight matrix to obtain a prediction value corresponding to the current block.
23. A video encoding device, characterized in that: include: A sorting unit configured to sort the multiple angle weighted prediction modes according to the order of the cost of using the multiple angle weighted prediction modes for the current block from low to high, to obtain a rearranged mode list; a determining unit, determining mode index information according to the re-arranged mode list and the selected angle weighted prediction mode; A calculation unit, configured to derive a weight matrix of the current block according to a selected angle weighted prediction mode; The encoding unit is configured to perform weighted prediction according to the weight matrix to obtain a prediction value corresponding to the current block, encode the current block according to the prediction value, and encode the mode index information in a video bit stream.
24. A computer readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the video decoding method according to any one of claims 1 to 20 or the video encoding method according to claim 21 is implemented.
25. An electronic device, characterized in that: include: one or more processors; A memory for storing one or more computer programs, which, when executed by the one or more processors, enables the electronic device to implement the video decoding method described in any one of claims 1 to 20, or the video encoding method described in claim 21.
26. A method for processing a code stream, characterized in that: A video code stream is stored on a non-transitory computer-readable medium, wherein the video code stream is decoded based on the video decoding method according to any one of claims 1 to 20, or generated according to the video encoding method according to claim 21.
Citation Information
Patent Citations
Intra-frame prediction mode coding and decoding method and device
CN101854551A
Method for coding hybrid image
CN102223541A
System and method for MPEG2 / h.264 high speed digital video transcoding
KR1020080093302A
Wide-angle intra prediction for video coding
US20200112728A1