Video Coding Method, Apparatus, Computer Program, and Electronic Device
By applying a correction value to adjust prediction modes derived from TIMD, the video coding process achieves improved accuracy and performance by addressing the mismatch between template and current area texture characteristics.
Patent Information
- Application Number
- JP2024563818
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-10-21
- Filing Date
- 2023-05-25
- Publication Date
- 2025-07-10
- Estimated Expiration
- 2043-05-25
AI Technical Summary
The TIMD technology assumes that the texture characteristics of the template area and the current prediction target area are the same, but in actual encoding, this assumption is often not valid, leading to suboptimal prediction modes and affecting video coding performance.
Introduce a correction value to adjust the candidate prediction mode derived from TIMD, using correction value indication information to improve the accuracy and self-adaptive ability of the video coding process.
The introduction of a correction value enhances the accuracy and self-adaptive ability of TIMD, resulting in improved video coding performance by better aligning prediction modes with actual texture characteristics.
Smart Images

Figure 2025521399000001_ABST
Abstract
Description
Technical Field
[0001] (Cross-reference to Related Applications) This application claims priority to a Chinese patent application filed with the China National Intellectual Property Administration on October 21, 2022, with application number 202211296332.X and invention title "Video Coding Method, Apparatus, Computer-readable Medium, and Electronic Device", the entire content of which is incorporated herein by reference.
[0002] This application relates to the video coding technology in the field of codec technology. Specifically, it relates to a video coding method, apparatus, computer-readable medium, and electronic device.
Background Art
[0003] The TIMD (Template Based Intra Mode Derivation) technology is introduced into the ECM (Enhanced Compression Model) standard. Similar to the DIMD (Decoder-side Intra Mode Derivation) technology, TIMD can derive a prediction mode by performing the same operations on the template area on both the encoding side and the decoding side, and reconstruct the current CU (Coding Unit) based on the derived prediction mode. Since TIMD does not need to encode the prediction mode, this technology can reduce the size of the code stream.
[0004] The TIMD technology assumes that the texture characteristics of the template area and the current prediction target area are the same. However, in actual encoding, since the texture characteristics of the template area cannot fully represent the texture characteristics of the current area, the prediction mode derived from the template area is not necessarily suitable for the current CU, which further affects the video coding performance.
Summary of the Invention
[0005] The embodiments of the present application provide a video coding method, apparatus, computer-readable medium, and electronic device, which can improve the accuracy and self-adaptive ability of TIMD by introducing a correction value, and further improve the video coding performance.
[0006] Other characteristics and advantages of the present application will become apparent from the following detailed description, or some of them can be learned through the practice of the present application.
[0007] According to an aspect of the embodiments of the present application, a video decoding method is provided. The video decoding method includes steps of obtaining correction value indication information when decoding from a video code stream and using template-based intra-mode derivation (TIMD); obtaining a candidate prediction mode of a current block derived based on TIMD; correcting the candidate prediction mode with the correction value indicated by the correction value indication information to obtain a corrected TIMD prediction mode; and performing a decoding process on the current block based on the corrected TIMD prediction mode.
[0008] According to an aspect of the embodiments of the present application, a video encoding method is provided. The video encoding method includes steps of obtaining a candidate prediction mode of a current block derived based on TIMD; determining a correction value for the candidate prediction mode based on rate distortion cost when correction processing needs to be performed on the candidate prediction mode of the current block; generating correction value indication information based on the correction value; and adding the correction value indication information to a video code stream.
[0009] According to one aspect of the embodiments of the present application, a video decoding device is provided. The video decoding device includes a decoding unit configured to obtain correction value indication information when decoding from a video code stream and using template-based intra-mode derivation (TIMD), an acquisition unit configured to acquire a candidate prediction mode of a current block derived based on TIMD, a correction unit configured to correct the candidate prediction mode with a correction value indicated by the correction value indication information to obtain a corrected TIMD prediction mode, and a processing unit configured to perform a decoding process on the current block based on the corrected TIMD prediction mode.
[0010] According to one aspect of the embodiments of the present application, a video encoding device is provided. The video encoding device includes an acquisition unit configured to acquire a candidate prediction mode of a current block derived based on TIMD, a determination unit configured to determine a correction value for the candidate prediction mode based on rate distortion cost when it is determined that correction processing needs to be performed on the candidate prediction mode of the current block, a generation unit configured to generate correction value indication information based on the correction value, and an addition unit configured to add the correction value indication information to a video code stream.
[0011] According to one aspect of the embodiments of the present application, a computer-readable medium is provided. A computer program is stored in the computer-readable medium. When the computer program is executed by a processor, the video encoding method or video decoding method described in the above embodiments is realized.
[0012] According to one aspect of the embodiments of the present application, an electronic device is provided. The electronic device includes one or more processors and a storage device for storing one or more computer programs. When the one or more computer programs are executed by the one or more processors, the electronic device realizes the video encoding method or video decoding method described in the above embodiments.
[0013] According to one aspect of the embodiments of the present application, a computer program product is provided, the computer program product includes a computer program, and the computer program is stored in a computer-readable storage medium. The processor of the electronic device reads and executes the computer program from the computer-readable storage medium, and causes the electronic device to execute the video encoding method or the video decoding method provided in the above various selectable embodiments.
[0014] According to one aspect of the embodiments of the present application, a code stream is provided, and the code stream is a code stream according to the method described in the above first aspect or a code stream generated by the method described in the above second aspect.
[0015] In the technical solutions provided by some embodiments of the present application, when decoding from a video code stream to obtain correction value indication information when using TIMD, the candidate prediction mode obtained based on TIMD is corrected by the correction value indicated by the correction value indication information, and then based on the corrected TIMD prediction mode, decoding processing is performed on the current block, so that by introducing the correction value, the accuracy and self-adaptive ability of TIMD can be improved, and furthermore, the codec performance of the video can be improved.
[0016] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory and do not limit the present application.
Brief Description of the Drawings
[0017]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Best Mode for Carrying Out the Invention
[0018] Now, exemplary embodiments will be described more comprehensively with reference to the drawings. However, the exemplary embodiments may be implemented in various forms and should not be understood as being limited to these examples. On the contrary, the disclosure of these embodiments aims to make the present application more comprehensive and complete, and to convey the technical idea of the exemplary embodiments to those skilled in the art in a comprehensive manner.
[0019] Furthermore, the features, structures, or characteristics described in the present application may be combined in one or more embodiments in any appropriate manner. Since there are many specific details in the following description, the embodiments of the present application can be fully understood. However, those skilled in the art should note that when implementing the technical solution of the present application, it may not be necessary to use all the detailed features in the embodiments, and one or more specific details may be omitted, or other methods, elements, devices, steps, etc. may be adopted.
[0020] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or may be implemented by one or more hardware modules or integrated circuits, or may be implemented by different networks and / or processor devices and / or microcontroller devices.
[0021] The flowcharts shown in the drawings are only illustrative explanations and do not necessarily include all contents and operations / steps, nor are they required to be executed in the described order. For example, some operations / steps may be decomposed, while some operations / steps may be merged or partially merged, so the actual execution order may be changed according to the actual situation.
[0022] It should be noted that the "plurality" mentioned in this specification refers to two or more. "And / or" describes the relationship of related objects and indicates that three types of relationships can exist. For example, A and / or B can indicate three situations: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally indicates that the related objects before and after are in an "or" relationship.
[0023] The solutions provided by the embodiments of the present application can be applied to the field of digital video encoding technologies, including, but not limited to, the fields of image codec, video codec, hardware video codec, dedicated circuit video codec, and real-time video codec. Further, the solutions provided by the embodiments of the present application may be combined with video and video encoding standards (AVS: Audio Video coding Standard), the second-generation AVS standard (AVS2), or the third-generation AVS standard (AVS3). For example, it includes, but is not limited to, the H.264 / Audio Video coding (AVC) standard, the H.265 / High Efficiency Video Coding (HEVC) standard, and the H.266 / Versatile Video Coding (VVC) standard. Also, the solutions provided by the embodiments of the present application may be used to perform lossy compression on images or lossless compression on images. Here, the lossless compression may be visually lossless compression or mathematically lossless compression.
[0024] FIG. 1 shows a schematic diagram of an exemplary system architecture to which the technical solution of the embodiment of the present application can be applied.
[0025] As shown in FIG. 1, the system architecture 100 includes a plurality of terminal devices, and the terminal devices can communicate with each other via, for example, the network 150. For example, the system architecture 100 can include a first terminal device 110 and a second terminal device 120 interconnected via the network 150. In the embodiment of FIG. 1, the first terminal device 110 and the second terminal device 120 perform unidirectional data transmission.
[0026] For example, the first terminal device 110 can encode video data (e.g., a video image stream collected by the terminal device 110) and transmit it to the second terminal device 120 via the network 150. The encoded video data is transmitted in the form of one or more encoded video code streams. The second terminal device 120 can receive the encoded video data from the network 150, decode the encoded video data to restore the video data, and display a video image based on the restored video data.
[0027] In one embodiment of the present application, the system architecture 100 can include a third terminal device 130 and a fourth terminal device 140 that perform bidirectional transmission of encoded video data. The bidirectional transmission can occur, for example, during a video conference. For the bidirectional data transmission, each terminal device in the third terminal device 130 and the fourth terminal device 140 can encode video data (e.g., a video image stream collected by the terminal device) and transmit it to the other terminal device among the third terminal device 130 and the fourth terminal device 140 via the network 150. Each terminal device in the third terminal device 130 and the fourth terminal device 140 can receive the encoded video data transmitted by the other terminal device among the third terminal device 130 and the fourth terminal device 140, decode the encoded video data to restore the video data, and display a video image on an accessible display device based on the restored video data.
[0028] In the embodiment of FIG. 1, the first terminal device 110, the second terminal device 120, the third terminal device 130, and the fourth terminal device 140 may be servers or terminals. The server may be an independent physical server, or may be a server cluster or a distributed system composed of multiple physical servers, or may be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The terminal may be a smartphone, a tablet, a laptop, a desktop computer, a smart speaker, a smart watch, a smart voice interaction device, smart home appliances, an in-vehicle terminal, an aircraft, etc., but is not limited thereto.
[0029] Network 150 represents any number of networks that transmit encoded video data between the first terminal device 110, the second terminal device 120, the third terminal device 130, and the fourth terminal device 140, and includes, for example, wired and / or wireless communication networks. The communication network 150 can exchange data through a circuit-switched channel and / or a packet-switched channel. The network can include a telecommunications network, a local area network, a wide area network, and / or the Internet. To achieve the objectives of the present application, unless otherwise specified below, the architecture and topology of network 150 may not be important for the operations disclosed in the present application.
[0030] In one embodiment of the present application, FIG. 2 shows an arrangement method of a video encoding device and a video decoding device in a streaming transmission environment. The subject matter disclosed in the present application is similarly applicable to other video-supporting applications, including, for example, compressed video stored in digital media such as video conferences, digital TVs (televisions), CDs, DVDs, memory sticks, and the like.
[0031] The streaming transmission system can include a collection subsystem 213, and the collection subsystem 213 can include a video source 201 such as a digital camera. The video source creates an uncompressed video image stream 202. In an embodiment, the video image stream 202 includes samples taken by a digital camera. Compared with the encoded video data 204 (or encoded video code stream 204), the video image stream 202 is drawn with a thick line to emphasize the high-data-volume video image stream. The video image stream 202 can be processed by an electronic device 220, and the electronic device 220 includes a video encoding device 203 coupled to the video source 201. The video encoding device 203 can include hardware, software, or a combination of hardware and software to implement or carry out each aspect of the subject matter disclosed in more detail below. Compared with the video image stream 202, the encoded video data 204 (or encoded video code stream 204) is drawn with a thin line to emphasize the relatively low-data-volume encoded video data 204 (or encoded video code stream 204), which can be stored in the streaming transmission server 205 for future use. One or more streaming transmission client subsystems, such as client subsystem 206 and client subsystem 208 in FIG. 2, can access the streaming transmission server 205 to retrieve copies 207 and 209 of the encoded video data 204. The client subsystem 206 can include, for example, a video decoding device 210 within an electronic device 230. The video decoding device 210 decodes the received copy 207 of the encoded video data and generates an output video image stream 211 that can be displayed on a display 212 (e.g., a display screen) or another display device. In some streaming transmission systems, the encoded video data 204, video data 207, and video data 209 (e.g., video code stream) can be encoded based on some specific video encoding / compression standard.
[0032] It should be noted that the electronic device 220 and the electronic device 230 may include other components not shown in the figure. For example, the electronic device 220 can include a video decoder, and the electronic device 230 can further include a video encoder.
[0033] In one embodiment of the present application, taking the international video coding standards HEVC (High Efficiency Video Coding), VVC (Versatile Video Coding), and the Chinese national video coding standard AVS (Audio Video coding Standard) as examples, after inputting one video frame image, based on the size of one block, the video frame image is divided into several non-overlapping processing units, and the same compression operation is performed for each processing unit. This processing unit is called a CTU or an LCU (Largest Coding Unit). For the CTU, it can be further subdivided to continue with finer segmentation, and one or more basic coding units (CUs: Coding Unit) can be obtained. The CU is the most basic element of the coding link.
[0034] Hereinafter, some concepts when encoding the CU are introduced.
[0035] Predictive Coding: Predictive coding includes methods such as intra-frame prediction and inter-frame prediction. The original video signal undergoes prediction of the selected reconstructed video signal and then a residual video signal is obtained. The encoding side needs to determine which predictive coding mode to select for the current CU and notify the decoding side. Here, intra-frame prediction means that the predicted signal is from the encoded and reconstructed area within the same image, and inter-frame prediction means that the predicted signal is from an encoded other image (referred to as a reference image) different from the current image.
[0036] Transform & Quantization: After the residual video signal undergoes transformation operations such as DFT (Discrete Fourier Transform) and DCT (Discrete Cosine Transform), the signal is transformed into the transform domain and is called the transform coefficient. A non-invertible quantization operation is further performed on the transform coefficient to discard certain information and make the quantized signal advantageous for compression representation. In some video coding standards, since there can be multiple selectable transformation methods, correspondingly, the encoding side also needs to select one of them for the current CU and notify the decoding side. The degree of quantization fineness is usually determined by the quantization parameter (abbreviated as QP). When the QP value is large, it indicates that coefficients in a larger value range are quantized to the same output, so usually it results in a larger distortion and a relatively low code rate. Conversely, when the QP value is small, it indicates that coefficients in a relatively small value range are quantized to the same output, so usually it results in a relatively small distortion and at the same time corresponds to a relatively high code rate.
[0037] Entropy Coding or Statistical Coding: The quantized transform domain signal performs statistical compression coding based on the occurrence frequency of each value and finally outputs a binary (0 or 1) compressed code stream. At the same time, other information such as the selected coding mode and motion vector data is also generated by the coding, and entropy coding is necessary to reduce the code rate. Statistical coding is a reversible coding method and can effectively reduce the code rate required to represent the same signal. Common statistical coding methods include variable length coding (abbreviated as VLC) or context-based binary arithmetic coding (abbreviated as CABAC).
[0038] The context-based binary arithmetic coding (CABAC) process mainly includes three steps: binarization, context modeling, and binary arithmetic coding. After performing the binarization process on the input grammar element, binary data can be encoded in the general coding mode and the bypass coding mode. In the bypass coding mode, it is not necessary to assign a specific probability model to each binary bit. The input binary bit bin value is directly encoded by a simple bypass encoder, which speeds up the entire encoding and decoding process. Generally, different grammar elements are not completely independent of each other, and the same grammar element itself also has a certain memory property. Therefore, according to the conditional entropy theory, by performing conditional coding using other encoded grammar elements, the coding performance can be further improved compared to independent coding or memoryless coding. The encoded code information used as a condition is called a context. In the general coding mode, the binary bits of the grammar element enter the context model in order, and the encoder assigns an appropriate probability model to each input binary bit based on the value of the previously encoded grammar element or binary bit. This process is the context modeling. The context model corresponding to the grammar element can be located by ctxIdxInc (context index increment) and ctxIdxStart (context index start). After transmitting the bin value and the assigned probability model to the binary arithmetic encoder for encoding, it is necessary to update the context model based on the bin value, which is the self-adaptive process in encoding.
[0039] Loop Filtering: The signal that has undergone transformation and quantization is used to obtain the reconstructed image through inverse quantization, inverse transformation, and prediction compensation operations. Compared with the original image, due to the impact of quantization, some information in the reconstructed image is different from that of the original image, that is, distortion occurs in the reconstructed image. Therefore, a filtering operation can be performed on the reconstructed image to effectively reduce the degree of distortion caused by quantization. These filtered reconstructed images are used to predict future image signals as references for subsequent encoded images. Therefore, the above filtering operation is also called loop filtering, that is, the filtering operation within the encoding loop.
[0040] In one embodiment of the present application, FIG. 3 shows a basic flowchart of a video encoder. In this process, intra-frame prediction is taken as an example for explanation. Here, the original image signal s k [x, y] and the predicted image signal s ^ k [x, y] are subjected to a difference operation to obtain the residual signal u k [x, y]. After performing transformation and quantization processing on the residual signal u k [x, y], quantization coefficients are obtained. The quantization coefficients are used to obtain the encoded bitstream through entropy encoding, while the reconstructed residual signal u k ’ [x, y] is obtained through inverse quantization and inverse transformation processing. The predicted image signal s ^ k [x, y] and the reconstructed residual signal u k ’ [x, y] are superimposed to generate the reconstructed image signal s * k [x, y]. The reconstructed image signal s * k [x, y] is input into the intra-frame mode decision module and the intra-frame prediction module to perform intra-frame prediction processing, while filtering processing is performed through loop filtering. The filtered image signal s ’ k [x, y] is output. The filtered image signal s ’k Motion estimation and motion compensation prediction can be performed with [x, y] as the reference image for the next frame. Next, the result s of the motion compensation prediction ’ r [x + m x , y + m y and the in-frame prediction result f(s * k [x, y]) are used to obtain the predicted image signal s of the next frame ^ k [x, y]. The above process is continuously repeated until the encoding is completed.
[0041] In the above encoding process, loop filtering is one of the core modules of video encoding and can effectively remove multiple encoding distortions. The latest generation of international video encoding standard VVC supports four different types of loop filters: the deblocking filter (abbreviated as DF), the sample adaptive offset (abbreviated as SAO), the adaptive loop filter (abbreviated as ALF), and the cross-component adaptive loop filter (CC-ALF).
[0042] To explore the next-generation compression standard, JVET (Joint Video Experts Team) established the latest ECM reference platform and added the TIMD technology. Similar to the DIMD technology, TIMD can derive the prediction mode and reconstruct the current CU based on the derived prediction mode by performing the same operation on the template area on both the encoding side and the decoding side. Since TIMD does not need to encode the prediction mode, this technology can reduce the size of the code stream.
[0043] The TIMD technology derives the intra - frame angle mode using the reconstructed pixels in the template area adjacent to the current CU. As shown in Figure 4, the area above and to the left of the current CU to be predicted is the template area, and the width of the template area is determined by the size of the current CU. If the width of the current CU is greater than 8, the width of the left template area is 4 pixels; otherwise, the width of the left template area is 2 pixels. If the height of the current CU is greater than 8, the height of the upper template area is 4 pixels; otherwise, the height of the upper template area is 2 pixels.
[0044] TIMD assumes that the texture characteristics of the template area and the current prediction target area match. Since the pixels in the template area have been reconstructed, by traversing the MPM (Most Probable Mode) list, calculating the predicted pixels in the template area, and obtaining the SATD (Sum of Absolute Transformed Difference) between the predicted pixels and the reconstructed pixels, the optimal mode can be selected, and this mode is used as the intra - frame prediction mode for the prediction target area. On the decoder side, the intra - frame prediction mode can be derived by the same derivation method, thereby significantly reducing the number of bits for mode coding.
[0045] Specifically, for the current CU, TIMD derives the intra - frame prediction mode from within the MPM list, maps the intra - frame prediction mode to a finer TIMD prediction mode, calculates the SATD between the predicted pixels and the reconstructed pixels in the template area, and then, the prediction mode M1 with the smallest SATD cost (SATD_cost), and the prediction mode M2 with the second - smallest SATD_cost are selected, and it is determined whether to apply weighted fusion based on the SATD_cost of M1 and M2. That is, SATD_cost(M1)<2×SATD_cost(M2) In this case, M1 and M2 are used to predict the current CU, and the predicted values respectively obtained by predicting with M1 and M2 are weighted and fused as the final predicted value of the current CU. Here, The weight weight1 of M1 = SATD_cost(M1) / (SATD_cost(M1)+SATD_cost(M2)), The weight weight2 of M2 = SATD_cost(M2) / (SATD_cost(M1)+SATD_cost(M2)), and otherwise, only M1 is used to predict the current CU to obtain the predicted value of the current CU. The TIMD technology can also be used in combination with multiple other modes, such as the ISP (Intra Sub-Partitions) mode, the MRL (Multiple Reference Line) mode, the CIIP (Combined Inter and Intra Prediction) mode, the CIIP-TM (Combined Inter and Intra Prediction-Template Matching) mode, the GPM (Geometric Partitioning Mode) mode, etc.
[0046] However, in actual encoding, since the texture characteristics of the template region cannot fully represent the texture characteristics of the current region, the prediction mode derived for the template region is not necessarily suitable for the current CU, thus affecting the codec performance of the video. Therefore, the technical solution of the embodiment of the present application further applies one correction value based on the prediction mode derived by TIMD to more accurately predict the current CU, further improving the accuracy and self-adaptive ability of TIMD, and improving the compression and codec performance of the video.
[0047] In the following, the implementation details of the technical solution of the embodiment of the present application will be described in detail.
[0048] FIG. 5 shows a flowchart of a video decoding method according to an embodiment of the present application. The video decoding method may be executed by a device having a computing function, for example, it may be executed by a terminal device or a server. Referring to FIG. 5, the video decoding method includes at least steps S510 to S540, and the detailed introduction is as follows.
[0049] In step S510, obtain correction value instruction information when decoding from a video code stream and using TIMD.
[0050] In some selectable embodiments, the correction value instruction information includes a first flag bit for indicating whether to use a correction value to correct a candidate prediction mode for the current block, and a correction value instruction flag bit. For example, the first flag bit may be timd_delta_flag. When the value of timd_delta_flag is 1 (for example only), it indicates that it is necessary to use a correction value to correct the candidate prediction mode for the current block. When the value of timd_delta_flag is 0, it indicates that it is not necessary to use a correction value to correct the candidate prediction mode for the current block.
[0051] Optionally, when the value of the first flag bit obtained by decoding from the video code stream is used to indicate that a correction value is to be used to correct the candidate prediction mode for the current block, for example, when the value of timd_delta_flag obtained by decoding is 1, decode the correction value instruction flag bit to obtain the correction value.
[0052] Optionally, when the value of the first flag bit obtained by the decoder by decoding from the video code stream indicates that a correction value is not to be used to correct the candidate prediction mode for the current block, for example, when the value of timd_delta_flag obtained by decoding is 0, it is not necessary to decode the correction value instruction flag bit. In this case, the encoder also does not need to encode the correction value instruction flag bit in the code stream.
[0053] In some selectable embodiments, the correction value indication flag bits can include a second flag bit and at least one third flag bit. The value of the second flag bit is used to indicate the sign of the correction value, and at least one third flag bit is used to indicate the absolute value of the correction value. For example, the second flag bit can be timd_delta_sign. If the value of timd_delta_sign is 1, it represents that the correction value is a positive value (it may also indicate a negative value). If the value of timd_delta_sign is 0, it represents that the correction value is a negative value (it may also indicate a positive value).
[0054] In some selectable embodiments, the value of each third flag bit is used to indicate whether the absolute value of the correction value corresponds to a level in a set of set numerical values. For example, if the set of set numerical values is {3, 6, 9}, the third flag bit can be timd_delta_first_level. If timd_delta_first_level is 1, it represents that the absolute value of the correction value is 3. If timd_delta_first_level is 0, another third flag bit timd_delta_second_level is introduced.
[0055] If timd_delta_second_level is 1, the absolute value of the correction value is 6. If timd_delta_second_level is 0, the absolute value of the correction value is 9. Optionally, If timd_delta_first_level is 1, since the absolute value of the correction value is already known, there is no need to decode other third flag bits. In this case, the encoding side also does not need to encode other third flag bits in the code stream.
[0056] In some selectable embodiments, the value of at least one third flag bit can be used cooperatively to indicate selecting a corresponding level value from a set of set values as the absolute value of the correction value. For example, when the set of set values is {3, 6, 9}, at least one third flag bit may be timd_delta_level, when timd_delta_level is 00, it represents that the absolute value of the correction value is 3, when timd_delta_level is 01, it represents that the absolute value of the correction value is 6, when timd_delta_level is 10, it represents that the absolute value of the correction value is 9.
[0057] In some selectable embodiments, the correction value indication flag bit can also include correction value index information, and the correction value index information is used to indicate selecting a corresponding value from a set of set values as the correction value. For example, when the set of set values is {-9, -6, -3, 3, 6, 9} and the correction value index information is timd_delta_index, when timd_delta_index is 000, 001, 010, 011, 100, 101, it represents that the index of the correction value is 0, 1, 2, 3, 4, 5 respectively, that is, corresponding to the values -9, -6, -3, 3, 6, 9 respectively.
[0058] Of course, in other alternative embodiments, at least one third flag bit can also indicate the absolute value of the correction value in other ways.
[0059] For example, the number of at least one third flag bit can be used to indicate the absolute value of the correction value. Specifically, among this at least one third flag bit, when the values of the flag bits other than the last flag bit are the first numerical value and the value of the last flag bit among this at least one third flag bit is the second numerical value, the number of this at least one third flag bit can be used to indicate the absolute value of the correction value. The first numerical value is 0 and the second numerical value is 1. Or, the first numerical value is 1 and the second numerical value is 0.
[0060] In some selectable embodiments, the correction value indication information in the foregoing embodiments may be for the current block. In other embodiments of the present application, a higher-layer syntax element for indicating whether a plurality of blocks need to modify the candidate prediction mode may be introduced. In this case, a specified flag bit can be obtained by decoding from the syntax elements of the video code stream, and based on the specified flag bit, it can be determined whether the current block needs to modify the candidate prediction mode. Optionally, the syntax element includes at least one of a sequence parameter set (SPS), a picture parameter set (PPS), a flag bit in a picture header, and a flag bit in a slice header.
[0061] For example, when the specified flag bit in the sequence parameter set SPS indicates that the image sequence needs to modify the candidate prediction mode, the decoding side does not need to decode the PPS, the flag bit in the picture header, and the flag bit in the slice header, and the encoding side does not need to encode either.
[0062] When the specified flag bit in the sequence parameter set SPS indicates that the picture sequence does not need to modify the candidate prediction mode, the decoding side can decode the PPS. When the specified flag bit in the PPS indicates that the current picture needs to modify the candidate prediction mode, the decoding side does not need to decode the flag bits in the picture header and the slice header, and the encoding side does not need to encode either.
[0063] When the specified flag bits in the sequence parameter set SPS and PPS indicate that the picture sequence does not need to modify the candidate prediction mode, the decoding side can decode the flag bits in the picture header. When the specified flag bit in the picture header indicates that the current picture needs to modify the candidate prediction mode, the decoding side does not need to decode the flag bits in the slice header, and the encoding side does not need to encode either.
[0064] Continuing to refer to FIG. 5, in step S520, the candidate prediction mode of the current block derived based on TIMD is obtained.
[0065] In one embodiment of the present application, the video code stream is a code stream obtained by encoding a video image frame sequence. Here, the video image frame sequence includes a series of images, each image can be further divided into slices, and the slices can also be divided into a series of LCUs (or CTUs), and an LCU contains several CUs. When encoding a video image frame, encoding processing is performed in units of blocks. In some new video encoding standards, for example, in the H.264 standard, there are macroblocks (MBs), and a macroblock can be further divided into a plurality of prediction blocks that can be used for predictive encoding. In the HEVC standard, using basic concepts such as encoding units CU, prediction units (PU), and transform units (TU), a plurality of block units are functionally divided and represented in a new tree-based structure. For example, a CU can be divided into smaller CUs according to a quadtree, and the smaller CUs can be further divided to form a quadtree structure. The current block in the embodiment of the present application may be a CU, or a block smaller than a CU, for example, a smaller block obtained by dividing a CU.
[0066] Optionally, the candidate prediction mode of the current block derived based on TIMD can be executed using the technical solution of the foregoing embodiment, and the candidate prediction mode may be the prediction mode with the smallest SATD cost (SATD_cost) and the prediction mode with the second smallest SATD_cost.
[0067] In step S530, the candidate prediction mode is corrected by the correction value indicated by the correction value indication information to obtain the corrected TIMD prediction mode.
[0068] In some selectable embodiments, the process of correcting the candidate prediction mode by the correction value indicated by the correction value indication information may be to add the correction value and the candidate prediction mode to obtain the corrected TIMD prediction mode.
[0069] As an option, only one of the candidate prediction modes derived by TIMD may be corrected, or all of the multiple candidate prediction modes derived by TIMD may be corrected. When all of the multiple candidate prediction modes of the current block derived by TIMD are corrected, the sets of correction values corresponding to these multiple candidate prediction modes may be different, and the correction value is determined from the set of correction values based on the correction value indication information. Of course, the sets of correction values corresponding to these multiple candidate prediction modes may be the same.
[0070] In some selectable embodiments, when the candidate prediction mode that needs to be corrected is the target prediction mode, based on the correction value, the predefined prediction mode corresponding to the correction value can be used as the corrected TIMD prediction mode. As an option, the target prediction mode may be a planar mode, a DC mode, or other non-angle prediction modes.
[0071] For example, when the set of correction values can be {-9, -6, -3, 3, 6, 9}, the predefined prediction modes corresponding to these correction values respectively can be a certain prediction mode in the set of {DC mode (mode index 0) / planar mode (mode index 1), TIMD horizontal mode (mode index 34), TIMD diagonal mode (mode index 66), TIMD vertical mode (mode index 98), TIMD inverse diagonal mode 1 (mode index 130), TIMD inverse diagonal mode 2 (mode index 2)}.
[0072] In some selectable embodiments, the set of correction values corresponding to the candidate prediction mode may be determined based on at least one of the candidate prediction mode of the current block derived based on TIMD, the size of the current block, the most probable mode MPM set, the candidate prediction mode derived based on the in-frame mode derivation DIMD on the decoding side, and the magnitude of the cost of TIMD, and the correction value is determined from the set of correction values based on the correction value indication information. As an option, the cost of TIMD may be an SATD cost or the like.
[0073] In some selectable embodiments, the modified TIMD prediction mode may be different from at least one of the candidate prediction modes derived based on TIMD, the prediction modes within the MPM set, the prediction modes within the non-MPM set, the candidate prediction modes derived based on DIMD, and the prediction modes after modifying the candidate prediction modes derived based on DIMD.
[0074] Optionally, the prediction mode after modifying the candidate prediction mode derived based on DIMD can be obtained by modifying it in the same modification manner as TIMD, that is, modifying the candidate prediction mode derived based on DIMD with the modification value indicated in the code stream.
[0075] In one embodiment of the present application, when performing decoding processing in cooperation with TIMD and other modes, the sets of modification values corresponding to different other modes are the same or different.
[0076] For example, when other modes include the multi-reference line MRL mode, modification processing of the TIMD candidate prediction mode is performed on some or all of the reference lines of the MRL mode. Here, each reference line of the MRL mode corresponds to a different set of modification values, or multiple reference lines of the MRL mode correspond to the same set of modification values.
[0077] In some selectable embodiments, when the modified TIMD prediction mode is not within the TIMD prediction angle mode set, the modified TIMD prediction mode based on the candidate prediction mode and the modification value can be adjusted to be within the TIMD prediction angle mode set.
[0078] Specifically, when the correction value is a positive number, the difference between the sum of the candidate prediction mode and the correction value and the first set value can be calculated. Next, the remainder of the difference and the second set value is calculated, and the adjusted TIMD prediction mode is determined based on the remainder. Here, the first set value is a positive integer, for example, it may be 2, and the second set value is determined based on the maximum value of the TIMD prediction angle mode. For example, the second set value is offset + 3, and offset is the maximum value of the TIMD prediction angle mode - 5 (the numerical value is only an example).
[0079] When the correction value is a negative number, the sum of the difference between the candidate prediction mode and the correction value and the third set value is calculated, the remainder of the sum and the second set value is calculated, and the adjusted TIMD prediction mode is determined based on the remainder. Here, the third set value and the second set value are determined based on the maximum value of the TIMD prediction angle mode. For example, the third set value is offset + 1, the second set value is offset + 3, and offset is the maximum value of the TIMD prediction angle mode - 5 (the numerical value is only an example).
[0080] Continuing to refer to FIG. 5, in step S540, decoding processing is performed on the current block based on the corrected TIMD prediction mode.
[0081] In one embodiment of the present application, when correcting the candidate prediction mode, when correcting the first candidate prediction mode derived based on TIMD by the correction value indicated by the correction value indication information to obtain the corrected first candidate prediction mode, performing decoding processing on the current block based on the corrected TIMD prediction mode in step S540 may be to use the predicted value of the corrected first prediction mode for the current block as the predicted value of the current block.
[0082] Alternatively, performing decoding processing on the current block based on the corrected TIMD prediction mode in step S540 may be to fuse the predicted value of the corrected first prediction mode for the current block and the predicted value of the set prediction mode for the current block to obtain the predicted value of the current block.
[0083] As an option, the setting prediction mode includes at least one of a non-angle prediction mode (e.g., Planner mode), a second candidate prediction mode derived based on TIMD, and a prediction mode after modifying the second candidate prediction mode derived based on TIMD.
[0084] As an option, the same correction method as the first candidate prediction mode can be used when correcting the second candidate prediction mode, where the same correction value or a different correction value from the first candidate prediction mode is used to correct the second candidate prediction mode.
[0085] In one embodiment of the present application, the process of fusing the prediction value of the corrected first prediction mode for the current block and the prediction value of the setting prediction mode for the current block may be to perform weighted addition on the prediction value of the corrected first prediction mode for the current block and the prediction value of the setting prediction mode for the current block according to a predetermined weight value to obtain the prediction value of the current block.
[0086] As an option, a predetermined weight value can be determined based on at least one of the candidate prediction mode of the current block derived based on TIMD, the size of the current block, the absolute value of the conversion residual and the SATD cost between the first prediction mode before correction and the setting prediction mode before correction, and the absolute value of the conversion residual and the SATD cost between the first prediction mode after correction and the setting prediction mode after correction.
[0087] For example, when determining the weight based on the SATD cost between the first prediction mode before correction and the setting prediction mode before correction, the weight weight1 of the corrected first prediction mode M1 = SATD_cost(M1) / (SATD_cost(M3)+SATD_cost(M3)), The weight weight3 of the set prediction mode M3 is SATD_cost(M3) / (SATD_cost(M1)+SATD_cost(M3)). Here, SATD_cost(M1) represents the SATD cost of the first prediction mode before correction, and SATD_cost(M3) represents the SATD cost of the set prediction mode before correction.
[0088] When determining the weight based on the SATD cost between the corrected first prediction mode and the corrected set prediction mode, the weight weight1’ of the corrected first prediction mode M1 is SATD_cost(M1’) / (SATD_cost(M3’)+SATD_cost)(M3’)), the weight weight3’ of the set prediction mode M3 is SATD_cost(M3’) / (SATD_cost(M1’)+SATD_cost(M3’)). Here, SATD_cost(M1’) represents the SATD cost of the corrected first prediction mode, and SATD_cost(M3’) represents the SATD cost of the corrected set prediction mode.
[0089] FIG. 6 shows a flowchart of a video encoding method according to an embodiment of the present application. The video encoding method may be executed by a device having a calculation processing function. For example, it may be executed by a terminal device or a server. Referring to FIG. 6, the video encoding method includes at least steps S610 to S640, and the detailed introduction is as follows.
[0090] In step S610, obtain the candidate prediction mode of the current block derived based on TIMD.
[0091] In step S620, when it is necessary to perform correction processing on the candidate prediction mode of the current block, determine the correction value for the candidate prediction mode based on the rate-distortion cost.
[0092] In step S630, generate correction value indication information based on the correction value.
[0093] In step S640, correction value indication information is added to the video code stream.
[0094] It should be noted that the processing process on the encoding side is the same as that on the decoding side. For example, the candidate prediction mode of the current block derived based on TIMD is the same as that on the decoding side. When determining the correction value of the candidate prediction mode, based on RDO (Rate-Distortion Optimization), the correction value corresponding to the smallest RDO can be determined. Since the process of generating correction value indication information based on the correction value is the same as that on the decoding side obtaining the correction value indication information by decoding from the code stream, it will not be repeatedly described.
[0095] Therefore, the technical solution of the embodiment of the present application can mainly improve the accuracy and self-adaptive ability of TIMD by introducing a correction value, and further improve the codec performance of the video.
[0096] In one embodiment of the present application, when making a certain correction to the prediction mode derived by the TIMD mode, for the candidate prediction mode (for example, the candidate prediction mode M1 derived by the first TIMD) derived by a certain TIMD mode, the numerical set of its correction value M_delta may be a predefined numerical set, for example, M_delta ∈ {-9, -6, -3, 3, 6, 9}, or it may be a numerically set dynamically adjusted during encoding and decoding. The numerical set of the correction value may be defined by the upper layer grammar. For example, the numerical set of the correction value is specified in SPS, PPS, and Slice header. On the encoding side, whether to use the correction value and which correction value to use are determined by RDO, thereby ensuring that the prediction mode corresponding to the minimum RDO is used.
[0097] In some selectable embodiments, whether to use a correction value can be indicated by introducing a flag bit timd_delta_flag. For example, when timd_delta_flag is 0, it means that the TIMD derivation mode of the current block does not use any correction value; otherwise, the TIMD derivation mode needs to use the correction value together.
[0098] In one embodiment of the present application, when timd_delta_flag is 1, a flag bit timd_delta_sign can be introduced to represent the sign bit of M_delta. When timd_delta_sign is 1, M_delta is a positive value (or a negative value); when timd_delta_sign is 0, M_delta is a negative value (or a positive value). Further, one or more flag bits can be introduced to represent the magnitude of the correction value. Taking M_delta ∈ {-9, -6, -3, 3, 6, 9} as an example, when the flag bit timd_delta_first_level is introduced, if timd_delta_first_level is 1, it means that the absolute value of M_delta is 3; if timd_delta_first_level is 0, the flag bit timd_delta_second_level is introduced. When timd_delta_second_level is 1, the absolute value of M_delta is 6; when timd_delta_second_level is 0, the absolute value of M_delta is 9.
[0099] In one embodiment of the present application, when timd_delta_flag is 1, a flag bit timd_delta_sign can be introduced to represent the sign bit of M_delta. When timd_delta_sign is 1, M_delta is a positive value (or a negative value), and when timd_delta_sign is 0, M_delta is a negative value (or a positive value). Furthermore, one or more flag bits can be introduced to represent the size of the correction value. Taking M_delta ∈ {-9, -6, -3, 3, 6, 9} as an example, a flag bit timd_delta_level can be introduced. When timd_delta_level is 00, it represents that the absolute value of M_delta is 3. When timd_delta_level is 01, the absolute value of M_delta is 6. When timd_delta_level is 10, the absolute value of M_delta is 9.
[0100] In one embodiment of the present application, when timd_delta_flag is 1, one or more flag bits can be introduced to represent the correction value index. Taking M_delta ∈ {-9, -6, -3, 3, 6, 9} as an example, a flag bit timd_delta_index is introduced. When timd_delta_index is 000, 001, 010, 011, 100, 101, it represents that the index of M_delta is 0, 1, 2, 3, 4, 5 respectively.
[0101] In one embodiment of the present application, when one of the candidate prediction modes derived by TIMD (for example, M1) exceeds the TIMD prediction angle mode set after being corrected, the corrected mode is mapped to a mode within the TIMD prediction angle mode set. For example, as shown in FIG. 7, if the TIMD prediction angle mode set in the ECM is {2, 3, …, 130}, and mode 0 and mode 1 represent the DC mode and the Planar mode respectively, the corrected TIMD mode can be remapped to a mode within the TIMD prediction angle mode set {2, 3, …, 130}.
[0102] Specifically, when the correction value M_delta is a positive number and exceeds the TIMD predicted angle mode set after correction, the corrected TIMD predicted angle mode is mapped to ((M1 - 1 + (M_delta - 1)) % mod) + 2. When the correction value M_delta is a negative number and exceeds the TIMD predicted angle mode set after correction, the corrected TIMD predicted angle mode is mapped to ((M1 + offset - (M_delta - 1)) % mod). Here, % is the operation of taking the remainder, offset is the maximum value of the TIMD predicted angle mode - 5, and mod is offset + 3.
[0103] In one embodiment of the present application, when one of the candidate prediction modes derived by TIMD (for example, M1) is a planar mode, a DC mode, or another non - angle prediction mode, the corrected prediction mode can be mapped to a certain special TIMD prediction mode. For example, when M_delta ∈ {-9, -6, -3, 3, 6, 9}, the corrected prediction mode is mapped to one of the prediction modes in the set {DC mode (mode index 0) / planar mode (mode index 1), TIMD horizontal mode (mode index 34), TIMD diagonal mode (mode index 66), TIMD vertical mode (mode index 98), TIMD inverse diagonal mode 1 (mode index 130), TIMD inverse diagonal mode 2 (mode index 2)}.
[0104] In one embodiment of the present application, the candidate prediction modes derived by different TIMDs may have different correction value sets. For example, the candidate prediction mode M1 derived by the first TIMD and the candidate prediction mode M2 derived by the second TIMD may have different correction value sets.
[0105] In one embodiment of the present application, the set of correction values is determined by the encoded or decoded information, and includes, but is not limited to, one or more candidate prediction modes derived by TIMD, block size, MPM mode set, one or more modes derived by DIMD, cost size (e.g., SATD_cost) matching the template of TIMD, etc.
[0106] In one embodiment of the present application, after the candidate prediction mode derived by a certain TIMD is corrected by the correction value, it cannot include one or more modes derived by TIMD.
[0107] In one embodiment of the present application, after the candidate prediction mode derived by a certain TIMD is corrected by the correction value, it cannot include the modes within the MPM mode set.
[0108] In one embodiment of the present application, after the candidate prediction mode derived by a certain TIMD is corrected by the correction value, it cannot include the modes within the non-MPM mode set.
[0109] In one embodiment of the present application, after the candidate prediction mode derived by a certain TIMD (e.g., the candidate prediction mode M1 derived by the first TIMD) is corrected, the current block can be predicted by the corrected prediction mode to obtain a predicted value. Next, multiple methods of weighted fusion can be used for the predicted value of the current block by another TIMD-derived mode (e.g., the candidate prediction mode M2 derived by the second TIMD) to obtain the final predicted value of the current block.
[0110] In one embodiment of the present application, after correcting the candidate prediction mode (e.g., the candidate prediction mode M1 derived by the first TIMD) derived by a certain TIMD, the current block can be predicted by the corrected prediction mode to obtain a predicted value. Then, multiple methods of weighted fusion can be used for the predicted value of the current block by another TIMD-derived mode (e.g., the candidate prediction mode M2 derived by the second TIMD) to obtain the final predicted value of the current block.
[0111] As an option, after modifying a candidate prediction mode (for example, candidate prediction mode M1 derived by a certain TIMD, such as the first TIMD), the current block can be predicted by the modified prediction mode to obtain a predicted value, and then directly use it as the final prediction result.
[0112] As an option, after modifying a candidate prediction mode (for example, candidate prediction mode M1 derived by a certain TIMD, such as the first TIMD), the current block can be predicted by the modified prediction mode to obtain a predicted value, and then weighted fusion is performed on the predicted value of the current block by a non-angle mode such as the Planar mode, thereby obtaining the final prediction result.
[0113] As an option, after modifying a candidate prediction mode (for example, candidate prediction mode M1 derived by a certain TIMD, such as the first TIMD), the current block can be predicted by the modified prediction mode to obtain a predicted value, and then weighted fusion is performed on the predicted value of the current block by a candidate prediction mode derived by another TIMD (for example, prediction mode M2 derived by the second TIMD), thereby obtaining the final prediction result.
[0114] As an option, after modifying a candidate prediction mode (for example, candidate prediction mode M1 derived by a certain TIMD, such as the first TIMD), the current block can be predicted by the modified prediction mode to obtain a predicted value, and then another candidate prediction mode derived by a different TIMD (for example, candidate prediction mode M2 derived by the second TIMD) is modified using the same correction value M_delta. Next, weighted fusion is performed on the predicted value of the current block by the modified M1 and M2, thereby obtaining the final prediction result.
[0115] As an option, after modifying a candidate prediction mode (for example, candidate prediction mode M1 derived by a certain TIMD, such as the candidate prediction mode derived by the first TIMD), the current block can be predicted by the modified prediction mode to obtain a predicted value. Next, a different correction value M_delta is used to correct the candidate prediction mode (for example, candidate prediction mode M2 derived by the second TIMD) derived by another TIMD. Next, weighted fusion is performed on the predicted values of the current block by the modified M1 and M2, thereby obtaining the final prediction result.
[0116] In some selectable embodiments, when encoding and transmitting the correction value M_delta used for the candidate prediction mode (for example, candidate prediction mode M2 derived by the second TIMD) derived by another TIMD, the method of the foregoing embodiment can be used to transmit the flag bit information alone.
[0117] In one embodiment of the present application, the weights for weighted fusion are determined by the encoded or decoded information, and include, but are not limited to, one or more modes derived by TIMD, block size, etc.
[0118] As an option, the weights for weighted fusion can be obtained by calculating based on the SATD_cost of the candidate prediction mode before correction, and can also be obtained by calculating based on the SATD_cost of the candidate prediction mode after correction.
[0119] For example, when modifying all of the candidate prediction mode M1 derived by the first TIMD and the candidate prediction mode M2 derived by the second TIMD, the current block is predicted by each of the modified prediction modes M1 and M2 to obtain predicted values, and further weighted fusion is performed on the predicted values of the modified M1 and M2 with respect to the predicted value of the current block to obtain the final prediction result.
[0120] Weights can be determined based on the SATD cost between the candidate prediction mode M1 before correction and the candidate prediction mode M2 before correction. That is, the weight weight1 of the candidate prediction mode M1 after correction = SATD_cost(M1) / (SATD_cost(M2) + SATD_cost(M2)), and the weight weight2 of the candidate prediction mode M2 after correction = SATD_cost(M2) / (SATD_cost(M1) + SATD_cost(M2)). Here, SATD_cost(M1) represents the SATD cost of the candidate prediction mode M1 before correction, and SATD_cost(M2) represents the SATD cost of the candidate prediction mode M2 before correction.
[0121] Weights can also be determined based on the SATD cost between the candidate prediction mode M1 after correction and the candidate prediction mode M2 after correction. That is, the weight weight1’ of the candidate prediction mode M1 after correction = SATD_cost(M1’) / (SATD_cost(M2’) + SATD_cost(M2’)), and the weight weight2’ of the candidate prediction mode M2 after correction = SATD_cost(M2’) / (SATD_cost(M1’) + SATD_cost(M2’)). Here, SATD_cost(M1’) represents the SATD cost of the candidate prediction mode M1 after correction, and SATD_cost(M2’) represents the SATD cost of the candidate prediction mode M2 after correction.
[0122] In one embodiment of the present application, the correction value of TIMD can be applied to one or more other modes used in cooperation with TIMD. Optionally, when TIMD is used in cooperation with other modes, TIMD can design a unique set of correction values for each mode used in cooperation with TIMD, or can also design the same set of correction values for multiple other modes used in cooperation with TIMD.
[0123] Specifically, when TIMD is used in cooperation with the MRL mode, a unique set of correction values can be designed for each reference line of each MRL, or the same set of correction values can also be designed for the reference lines of multiple MRLs.
[0124] Alternatively, when the TIMD is used in cooperation with the MRL mode, the correction mode of the TIMD can be applied to some or all of the reference lines. For example, the correction mode of the TIMD is applied only to the most adjacent MRL reference line.
[0125] In one embodiment of the present application, one or more flag bits are introduced into upper-layer syntax elements (e.g., SPS, PPS, Picture header, Slice header), and the flag bits are used to indicate whether to introduce a correction value based on the TIMD or DIMD, or the TIMD and DIMD modes. These upper-layer syntax elements combine the flag bits of the current block in the foregoing embodiments to determine whether the current block needs to correct the candidate prediction mode derived by the TIMD or DIMD.
[0126] The embodiments of the apparatus of the present application will be described below, which can be used to execute the method described in the foregoing embodiments of the present application. For details not disclosed in the embodiments of the apparatus of the present application, please refer to the embodiments of the method of the present application above.
[0127] FIG. 8 shows a block diagram of a video decoding apparatus according to an embodiment of the present application. The video decoding apparatus may be provided in a device having a computing processing function, for example, may be provided in a terminal device or a server.
[0128] Referring to FIG. 8, a video decoding apparatus 800 according to an embodiment of the present application includes a decoding unit 802, an acquisition unit 804, a correction unit 806, and a processing unit 808.
[0129] Here, the decoding unit 802 is configured to obtain correction value indication information when decoding from a video code stream and using template-based intra-mode derivation (TIMD). The acquisition unit 804 is configured to acquire a candidate prediction mode of the current block derived based on TIMD. The correction unit 806 is configured to correct the candidate prediction mode with the correction value indicated by the correction value indication information to obtain a corrected TIMD prediction mode. The processing unit 808 is configured to perform decoding processing on the current block based on the corrected TIMD prediction mode.
[0130] In some embodiments of the present application, based on the foregoing solution, the correction value indication information includes a first flag bit for indicating whether to use a correction value to correct the candidate prediction mode for the current block, and a correction value indication flag bit. The decoding unit 802 is configured to decode the correction value indication flag bit and obtain the correction value when the value of the first flag bit obtained by decoding from the video code stream indicates that a correction value is to be used to correct the candidate prediction mode for the current block.
[0131] In some embodiments of the present application, based on the foregoing solution, the correction value indication flag bit includes a second flag bit and at least one third flag bit. The value of the second flag bit is used to indicate the sign of the correction value, and at least one third flag bit is used to indicate the absolute value of the correction value.
[0132] In some embodiments of the present application, based on the foregoing solution, the value of each third flag bit is used to indicate whether the absolute value of the correction value corresponds to a level in a set of preset values.
[0133] In some embodiments of the present application, based on the foregoing solution, the values of at least one third flag bit are used in cooperation to indicate selecting a value of a corresponding level from a set of preset values as the absolute value of the correction value.
[0134] In some embodiments of the present application, based on the foregoing solution, the correction value indication flag bit includes correction value index information, and the correction value index information is used to indicate selecting a corresponding value from the set of set values as the correction value.
[0135] In some embodiments of the present application, based on the foregoing solution, when the corrected TIMD prediction mode is not within the TIMD prediction angle mode set, the video decoding device further includes an adjustment unit configured to adjust the corrected TIMD prediction mode based on the candidate prediction mode and the correction value so that it is within the TIMD prediction angle mode set.
[0136] In some embodiments of the present application, based on the foregoing solution, when the correction value is a positive number, the adjustment unit is configured to calculate the difference between the sum of the candidate prediction mode and the correction value and a first set value, calculate the remainder of the difference and a second set value, and determine the adjusted TIMD prediction mode based on the remainder. The first set value is a positive integer, and the second set value is determined based on the maximum value of the TIMD prediction angle mode.
[0137] In some embodiments of the present application, based on the foregoing solution, when the correction value is a negative number, the adjustment unit is configured to calculate the sum of the difference between the candidate prediction mode and the correction value and a third set value, calculate the remainder of the sum and the second set value, and determine the adjusted TIMD prediction mode based on the remainder. The third set value and the second set value are determined based on the maximum value of the TIMD prediction angle mode.
[0138] In some embodiments of the present application, based on the foregoing solution, when the candidate prediction mode is the target prediction mode, the correction unit 806 is configured to use the predefined prediction mode corresponding to the correction value based on the correction value as the corrected TIMD prediction mode.
[0139] In some embodiments of the present application, based on the foregoing solution, when the candidate prediction mode of the current block derived based on TIMD includes a plurality of candidate prediction modes, the sets of correction values corresponding to the plurality of candidate prediction modes are different, and the correction value is determined from the set of correction values based on the correction value indication information.
[0140] In some embodiments of the present application, based on the foregoing solution, the set of correction values corresponding to the candidate prediction mode is determined based on at least one of the candidate prediction mode of the current block derived based on TIMD, the size of the current block, the most probable mode (MPM) set, the candidate prediction mode derived based on the in-frame mode derivation (DIMD) on the decoding side, and the magnitude of the cost of TIMD, and the correction value is determined from the set of correction values based on the correction value indication information.
[0141] In some embodiments of the present application, based on the foregoing solution, the corrected TIMD prediction mode is different from at least one of the candidate prediction mode derived based on TIMD, the prediction mode within the MPM set, the prediction mode outside the MPM set, the candidate prediction mode derived based on DIMD, and the prediction mode after correcting the candidate prediction mode derived based on DIMD.
[0142] In some embodiments of the present application, based on the foregoing solution, the correction unit 806 is configured to correct the first candidate prediction mode derived based on TIMD with the correction value indicated by the correction value indication information to obtain the corrected first prediction mode.
[0143] The processing unit 808 is configured to use the prediction value of the corrected first prediction mode for the current block as the prediction value of the current block, or fuse the prediction value of the corrected first prediction mode for the current block and the prediction value of the set prediction mode for the current block to obtain the prediction value of the current block.
[0144] In some embodiments of the present application, based on the foregoing solution, the set prediction mode includes at least one of a non-angle prediction mode, a second candidate prediction mode derived based on TIMD, and a prediction mode after correcting the second candidate prediction mode derived based on TIMD.
[0145] Here, the same correction value or a different correction value as that of the first candidate prediction mode is used to correct the second candidate prediction mode.
[0146] In some embodiments of the present application, based on the foregoing solution, the processing unit 808 is configured to perform weighted addition on the predicted value of the corrected first prediction mode for the current block and the predicted value of the set prediction mode for the current block according to a predetermined weight value to obtain the predicted value of the current block.
[0147] In some embodiments of the present application, based on the foregoing solution, a predetermined weight value is determined based on at least one of the candidate prediction mode of the current block derived based on TIMD, the size of the current block, the absolute value of the conversion residual and the SATD cost between the first prediction mode before correction and the set prediction mode before correction, and the absolute value of the conversion residual and the SATD cost between the first prediction mode after correction and the set prediction mode after correction.
[0148] In some embodiments of the present application, based on the foregoing solution, when performing decoding processing in cooperation with other modes using TIMD, the set of correction values corresponding to different other modes may be the same or different, and the correction value is determined from the set of correction values based on the correction value indication information.
[0149] In some embodiments of the present application, based on the foregoing solution, the other mode includes a multi-reference line MRL mode, and the correction unit 806 is further configured to perform correction processing of the TIMD candidate prediction mode on some or all of the reference lines of the MRL mode. Each reference line of the MRL mode corresponds to a different set of correction values, or multiple reference lines of the MRL mode correspond to the same set of correction values.
[0150] In some embodiments of the present application, based on the foregoing solution, the decoding unit 802 is further configured to obtain a specified flag bit by decoding from the syntax elements of the video code stream, and based on the specified flag bit, determine whether the current block needs to modify the candidate prediction mode. The syntax elements include at least one of a sequence parameter set SPS, a picture parameter set PPS, a flag bit in a picture header, and a flag bit in a slice header.
[0151] FIG. 9 shows a block diagram of a video encoding device according to an embodiment of the present application. The video encoding device may be provided in a device with a computing processing function, for example, it may be provided in a terminal device or a server.
[0152] Referring to FIG. 9, a video encoding device 900 according to an embodiment of the present application includes an acquisition unit 902, a determination unit 904, a generation unit 906, and an addition unit 908.
[0153] Here, the acquisition unit 902 is configured to acquire the candidate prediction mode of the current block derived based on TIMD. The determination unit 904 is configured to determine a correction value for the candidate prediction mode based on the rate-distortion cost when it is determined that correction processing needs to be performed on the candidate prediction mode of the current block. The generation unit 906 is configured to generate correction value indication information based on the correction value. The addition unit 908 is configured to add the correction value indication information to the video code stream.
[0154] FIG. 10 shows a schematic structural diagram of a computer system suitable for implementing the electronic device of the embodiment of the present application.
[0155] It should be noted that the computer system 1000 of the electronic device shown in FIG. 10 is only an example and does not limit the functions and usage scopes of the embodiments of the present application.
[0156] As shown in FIG. 10, the computer system 1000 includes a central processing unit (CPU) 1001, which can execute various appropriate operations and processes based on a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage unit 1008 into a random access memory (RAM) 1003. For example, it can execute the method described in the above embodiments. Various programs and data necessary for system operations are further stored in the RAM 1003. The CPU 1001, ROM 1002, and RAM 1003 are interconnected via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0157] An input unit 1006 including a keyboard, a mouse, etc., an output unit 1007 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc., a storage unit 1008 including a hard disk, etc., and a communication unit 1009 including a local area network (LAN) card, a network interface card such as a modem, etc. are connected to the I / O interface 1005. The communication unit 1009 performs communication processing via a network such as the Internet. A driver 1010 is also connected to the I / O interface 1005 as needed. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed in the driver 1010 as needed, facilitating the installation of a computer program read therefrom into the storage unit 1008 as needed.
[0158] In particular, according to the embodiments of the present application, the above-described process described with reference to the flowchart can be implemented as a computer software program. For example, the embodiments of the present application include a computer program product, the computer program product includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication unit 1009 and / or installed from the removable medium 1011. When the computer program is executed by the central processing unit (CPU) 1001, various functions limited to the system of the present application are executed.
[0159] It should be noted that the computer-readable medium shown in the embodiments of the present application may be a computer-readable signal medium, or a computer-readable storage medium, or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of the computer-readable storage medium may include, but are not limited to, an electrical connection with one or more wires, a portable computer magnetic disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program is used by or in combination with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier, and a computer-readable computer program is carried thereon. Such a propagated data signal may take various forms, including, but not limited to, an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer-readable signal medium may be any computer-readable medium other than the computer-readable storage medium, and the computer-readable medium can transmit, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The computer program included in the computer-readable medium may be transmitted by any suitable medium, including, but not limited to, wireless, wired, or any suitable combination of the above.
[0160] Flowcharts and block diagrams in the drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Here, each block in a flowchart or block diagram can represent a module, a program segment, or a portion of code, and the above-mentioned module, program segment, or portion of code includes one or more executable instructions for implementing a defined logical function. It should also be noted that in some alternative implementations, the functions shown within a block may be executed in an order different from that shown in the drawings. For example, two consecutively displayed blocks may actually be executed substantially in parallel, or in a reverse order depending on the functions involved. Each block in a block diagram or flowchart diagram, and combinations of blocks in a block diagram or flowchart diagram, may be implemented by a system based on dedicated hardware for performing the defined functions or operations, or may be implemented by a combination of dedicated hardware and a computer program.
[0161] The units according to the embodiments of the present application may be implemented in software, may be implemented in hardware, and the described units may be provided in a processor. Here, the names of these units do not constitute a limitation to the unit itself in some cases.
[0162] In another aspect, the present application further provides a computer-readable medium, which may be included in the electronic device described in the above embodiments, or may exist independently without being incorporated into the electronic device. The above computer-readable medium carries one or more computer programs, and when the above one or more computer programs are executed by the electronic device, the electronic device is caused to implement the method described in the above embodiments.
[0163] In another aspect, the present application further provides a computer program product or a computer program, the computer program product or the computer program including computer instructions stored in a computer-readable storage medium. At this time, when the processor reads the computer instructions from the computer-readable storage medium and executes the computer instructions, the encoding method or the decoding method provided in various selectable manners as described above is realized.
[0164] In another aspect, the present application further provides a code stream, which may be a code stream decoded by the decoding method provided by the present application or a code stream generated by the encoding method provided by the present application.
[0165] In another aspect, the present application further provides a codec system including the encoder and the decoder as described above.
[0166] It should be noted that although some modules or units of the devices for operation execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above may be embodied in one module or unit. Conversely, the features and functions of one module or unit described above may further be embodied by being divided and implemented by a plurality of modules or units.
[0167] From the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein may be implemented by software or by a combination of software and the necessary hardware. Therefore, the technical solution according to the embodiments of the present application can be embodied in the form of a software product, and the software product may be stored in a non-volatile storage medium (which may be a CD-ROM, a USB disk, a mobile hard disk, etc.) or on a network, and includes several instructions for causing a computing device (which may be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.
[0168] Those skilled in the art can easily conceive of other implementation solutions of the present application after considering the specification and practicing the embodiments disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, and these variations, uses, or adaptations comply with the general principles of the present application and include known knowledge or conventional technical means in the technical field not disclosed by the present application.
[0169] It should be understood that the present application is not limited to the exact structures already described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is limited only by the appended claims.
Claims
**Claim 1** A video decoding method executed by a video decoding apparatus, comprising: obtaining correction value indication information when decoding from a video code stream and using template-based intra-mode derivation (TIMD); obtaining a candidate prediction mode of a current block derived based on TIMD; correcting the candidate prediction mode by a correction value indicated by the correction value indication information to obtain a corrected TIMD prediction mode; and performing a decoding process on the current block based on the corrected TIMD prediction mode. **Claim 2** The correction value indication information includes a first flag bit for indicating whether to use a correction value to correct the candidate prediction mode for the current block, and a correction value indication flag bit, wherein the step of obtaining correction value indication information when decoding from the video code stream and using template-based intra-mode derivation (TIMD) includes: when the value of the first flag bit obtained by decoding from the video code stream indicates that a correction value is to be used to correct the candidate prediction mode for the current block, decoding the correction value indication flag bit to obtain the correction value. The video decoding method according to claim 1. **Claim 3** The correction value indication flag bit includes a second flag bit and at least one third flag bit, wherein the value of the second flag bit is used to indicate the sign of the correction value, and the at least one third flag bit is used to indicate the absolute value of the correction value. The video decoding method according to claim 2. **Claim 4** The value of each third flag bit is used to indicate whether the absolute value of the correction value corresponds to a level in a set of preset values. The video decoding method according to claim 3. **Claim 5** The values of the at least one third flag bit are used in cooperation to indicate selecting a value of a corresponding level from a set of preset values as the absolute value of the correction value. The video decoding method according to claim 3. **Claim 6** The correction value indication flag bit includes correction value index information, and the correction value index information is used to indicate selecting a corresponding value from a set of preset values as the correction value. The video decoding method according to claim 2.
7. When the corrected TIMD prediction mode is not within the TIMD prediction angle mode set, further including the step of adjusting the corrected TIMD prediction mode based on the candidate prediction mode and the correction value so as to be within the TIMD prediction angle mode set, characterized in that The video decoding method according to claim 1.
8. The step of adjusting the corrected TIMD prediction mode based on the candidate prediction mode and the correction value so as to be within the TIMD prediction angle mode set is When the correction value is a positive number, calculating the difference between the sum of the candidate prediction mode and the correction value and a first set value, calculating the remainder of the difference and a second set value, and determining the adjusted TIMD prediction mode based on the remainder, wherein the first set value is a positive integer and the second set value is determined based on the maximum value of the TIMD prediction angle mode, characterized in that The video decoding method according to claim 7.
9. The step of adjusting the corrected TIMD prediction mode based on the candidate prediction mode and the correction value so as to be within the TIMD prediction angle mode set is When the correction value is a negative number, calculating the sum of the difference between the candidate prediction mode and the correction value and a third set value, calculating the remainder of the sum and the second set value, and determining the adjusted TIMD prediction mode based on the remainder, wherein the third set value and the second set value are determined based on the maximum value of the TIMD prediction angle mode, characterized in that The video decoding method according to claim 7.
10. The step of correcting the candidate prediction mode with the correction value indicated by the correction value indication information to obtain the corrected TIMD prediction mode is When the candidate prediction mode is the target prediction mode, including the step of setting the predefined prediction mode corresponding to the correction value as the corrected TIMD prediction mode based on the correction value, characterized in that The video decoding method according to claim 1.
11. When the candidate prediction mode of the current block derived based on TIMD includes a plurality of candidate prediction modes, the correction value sets corresponding to the plurality of candidate prediction modes are different, and the correction value is determined from the correction value set based on the correction value indication information, characterized in that The video decoding method according to claim 1.
12. Further comprising the step of determining a set of correction values corresponding to the candidate prediction mode based on at least one of the candidate prediction mode of the current block derived based on TIMD, the size of the current block, the most probable mode MPM set, the candidate prediction mode derived based on the in-frame mode derivation (DIMD) on the decoding side, and the magnitude of the cost of TIMD, wherein the correction value is determined from the set of correction values based on the correction value indication information. The video decoding method according to claim 1.
13. The corrected TIMD prediction mode is different from at least one of the candidate prediction mode derived based on TIMD, the prediction mode within the MPM set, the prediction mode within the non-MPM set, the candidate prediction mode derived based on DIMD, and the prediction mode after correcting the candidate prediction mode derived based on DIMD. The video decoding method according to claim 1.
14. The step of correcting the candidate prediction mode with the correction value indicated by the correction value indication information to obtain a corrected TIMD prediction mode includes the step of correcting the first candidate prediction mode derived based on TIMD with the correction value indicated by the correction value indication information to obtain a corrected first prediction mode. The step of performing decoding processing on the current block based on the corrected TIMD prediction mode includes using the predicted value of the corrected first prediction mode for the current block as the predicted value of the current block, or fusing the predicted value of the corrected first prediction mode for the current block and the predicted value of the set prediction mode for the current block to obtain the predicted value of the current block. The video decoding method according to claim 1.
15. The set prediction mode is including at least one of a non-angular prediction mode, a second candidate prediction mode derived based on TIMD, and a prediction mode after correcting the second candidate prediction mode derived based on TIMD. using the same correction value or a different correction value as the first candidate prediction mode to correct the second candidate prediction mode. The video decoding method according to claim 14.
16. The step of fusing the predicted value of the corrected first prediction mode for the current block and the predicted value of the set prediction mode for the current block to obtain the predicted value of the current block includes: Performing weighted addition on the predicted value of the corrected first prediction mode for the current block and the predicted value of the set prediction mode for the current block according to a predetermined weight value to obtain the predicted value of the current block, characterized in that The video decoding method according to claim 14.
17. The candidate prediction mode of the current block derived based on TIMD, The size of the current block, The absolute value of the conversion residual and the SATD cost between the first prediction mode before correction and the set prediction mode before correction, and The absolute value of the conversion residual and the SATD cost between the first prediction mode after correction and the set prediction mode after correction, Further including the step of determining the predetermined weight value based on at least one of them, characterized in that The video decoding method according to claim 16.
18. When performing decoding processing in cooperation with other modes for TIMD, the set of correction values corresponding to different other modes may be the same or different, and the correction value is determined from the set of correction values based on the correction value indication information, characterized in that The video decoding method according to claim 1.
19. The other mode includes the multi-reference line MRL mode, and the video decoding method further includes: Performing correction processing on the TIMD candidate prediction mode for some or all of the reference lines of the MRL mode, Each reference line of the MRL mode corresponds to a different set of correction values, or a plurality of reference lines of the MRL mode correspond to the same set of correction values, characterized in that The video decoding method according to claim 18.
20. The video decoding method further includes: Obtaining a specified flag bit by decoding from the syntax elements of the video code stream, and determining whether the current block needs to correct the candidate prediction mode based on the specified flag bit, The syntax elements include at least one of a sequence parameter set (SPS), a picture parameter set (PPS), a flag bit in a picture header, and a flag bit in a slice header, characterized in that The video decoding method according to claim 1.
21. A video encoding method executed by a video encoding device, comprising: obtaining a candidate prediction mode of a current block derived based on TIMD; when correction processing needs to be performed on the candidate prediction mode of the current block, determining a correction value for the candidate prediction mode based on rate distortion cost; generating correction value indication information based on the correction value; adding the correction value indication information to a video code stream.
22. A video decoding device, comprising: a decoding unit configured to obtain correction value indication information when decoding from a video code stream and using template-based intra-mode derivation (TIMD); an obtaining unit configured to obtain a candidate prediction mode of a current block derived based on TIMD; a correction unit configured to correct the candidate prediction mode with the correction value indicated by the correction value indication information to obtain a corrected TIMD prediction mode; a processing unit configured to perform decoding processing on the current block based on the corrected TIMD prediction mode.
23. A video encoding device, comprising: an obtaining unit configured to obtain a candidate prediction mode of a current block derived based on TIMD; a determining unit configured to determine a correction value for the candidate prediction mode based on rate distortion cost when it is determined that correction processing needs to be performed on the candidate prediction mode of the current block; a generating unit configured to generate correction value indication information based on the correction value; an adding unit configured to add the correction value indication information to a video code stream.
24. An electronic device, comprising: one or more processors; and a memory for storing one or more computer programs, wherein when the one or more computer programs are executed by the one or more processors, the electronic device realizes the video decoding method according to any one of Claims 1 to 20 or the video encoding method according to Claim 21.
25. A computer program that causes an electronic device to execute the video decoding method according to any one of claims 1 to 20, or to implement the video encoding method according to claim 21.
26. A code stream that is decoded based on the video decoding method according to any one of claims 1 to 20, or that is generated based on the video encoding method according to claim 21.
Citation Information
Patent Citations
Method and apparatus for encoding or decoding video signals
JP2011515060A
Implicit and semi-implicit intra-mode signal transmission method and apparatus in video encoders and decoders
JP2012517736A
Image encoding / decoding method and apparatus, and recording medium storing bitstream
JP2025502798A
Method and apparatus for encoding and decoding video signal
US20180324418A1
Method and system for decoder-side intra mode derivation for block-based video coding
US20190166370A1